A Spanish answering service for a US trade line takes calls from Mexican, Central American, Caribbean, South American and US-born bilingual callers at once. This one is not offered, and this page is the measurement that explains why — along with what a Spanish emergency vocabulary gets wrong when it is built by people who did not grow up speaking it.
The safety gate's Spanish half was put through five regional panels plus two adversarial corpora, merged into 198 sentences. 13 of 198 agreed with intended behaviour at the start. 194 do now. Spanish hazards passing in silence went from 110 to zero; ordinary Spanish calls wrongly hung up on went from 42 to zero. That is measured, not reviewed — no native speaker has signed it off, so Spanish stays off.
On Call Bell has not answered a phone call yet. The safety gate, the triage rules, the call flow and the evidence trail are written and tested offline; the telephony is the next piece of work. Nothing on this page describes how fast or how reliably a live line answers, because no live line exists to measure. There is nothing to sign up for and no demo number. Everything below is either a decided price or a number you can count in the code.
The single most-agreed finding, raised independently by all five regions and both adversarial corpora: the Spanish hazard-probe vocabulary contained sonando, suena, humo, fuego, chispas, silbido, chispeando, chiflido, pitando, pita, quemandose, llamas, incendio, pitido. Every entry is a sight or a sound. The probe question asks about a smell.
So a caller who opened with "no" and then described the smell in any words the top tier did not carry was scored a clean denial and the call booked. Verified across regions: no, nomás huele feo (Mexico), no, huele a huevo podrido (Central America), no pero hay una peste ahí (Caribbean), no, a gas no, pero hay un olor rarísimo (South America), no, but there's a smell (US-born bilingual).
That is a structural defect, not a vocabulary opinion, and it is the class of finding this exercise is genuinely good for: it was verified by executing the shipped gate, and a native reviewer cannot overturn it.
The merged corpus is not a specification. Some of its labels are stricter than the shipped policy, and where the two disagree the shipped design wins if it can say why. It asserts properties that must not degrade, with the numbers written down so a change that quietly trades one for another fails a test instead of becoming a slightly different number nobody reads.
| Split | What it contains | Sentences |
|---|---|---|
| Miss corpus | Hazards a Spanish-speaking caller would report that the gate let through in silence | 70 |
| False-positive corpus | Ordinary Spanish plumbing calls the gate reacted to and should not have | 58 |
| Proposed additions | Regional phrasings the panels argued should be recognised | 70 |
| Total | Agreeing with intended behaviour: 13 at the start, 194 now | 198 |
Strict adjacency, and Spanish speakers insert words. The gas rule required huele next to a gas. Any word between broke it: huele como a gas, huele bien feo a gas, nomás huele un poquito a gas all downgraded from a stop to a question.
Candela — a whole region could not report a fire. In Venezuela and coastal Colombia, se cogió candela el calentador is how you say the water heater caught fire. It passed in silence, and as a probe answer no, pero hay candela was scored a clean denial.
The commonest carbon monoxide tell was in the wrong grammatical form. The rule required the noun phrase dolor de cabeza. What people say is the verb: nos duele la cabeza a todos en la casa — everybody in the house has one, and headache is the first symptom on the CDC's list — and that passed.
Inconsciente is hospital Spanish. Four of the five regions said the same sentence in the same words: nobody says it at 11pm. What a frightened caller says is no despierta, se privó, le falta el aire, está tirado en el piso — and the feminine desmayada, which was missing while desmayado was present.
The most emphatic denial in the language stopped the call. Claro que no was listed as a clean denial in one place and matched as an affirmation in another, so a caller saying "of course not" got hung up on. The written brief said it was handled; the code disagreed; there was no test case.
This matters more than the numbers, and stating it is the reason Spanish is still off. The five regional panels were model personas, not five people. Thirty of the panel's load-bearing claims were re-run against the shipped gate before anything was written, and every one reproduced — but that only validates the structural findings.
Frequency is guesswork. Whether botando gas beats fugando gas in Miami, or chillando beats pitando in Phoenix, decides which additions are worth their false-positive cost. A model's sense of frequency comes from text distribution, which over-weights written and Peninsular Spanish — the exact bias that produced the original list's problems.
Written prose is not speech. Every sentence in the corpus is imagined speech. What speech recognition hands a gate for a frightened caller talking over a crying child is an empirical question, and it drives at least a third of the vocabulary decisions here. It needs recorded audio, not a reviewer.
Agreement between the panels is weaker evidence than it looks. All five saw the same brief and the same file. "All five regions said X" is one observation with five restatements. Their disagreements are the encouraging part, because they were specific and opposite-signed — one region warning another off an addition, one reporting a word's meaning inverting between countries.
Before Spanish is offered as a line language, a human native speaker has to reject the over-broad additions, decide which denials count as clean, listen to the stop script — the one that tells a caller to get outside and then dial 911, which is what the CPSC advises — in the real voice, and read the false-positive half harder than the miss half. Until then the honest position is that the measurement improved and the product has not shipped.
Three questions, and they work on any vendor selling bilingual answering.
Which Spanish? A single-dialect vocabulary is invisible until a caller from the wrong country needs it. Regadera is a shower in Mexico and a watering can nearly everywhere else; a corpus written in one dialect is how one-region gaps stay hidden.
Was the false-positive half measured? Every region here wrote a longer, more confident miss list than false-positive list, because misses are easier to imagine than the ordinary calls a word list wrongly eats. The false-positive corpus found 58 wrongly-handled ordinary calls. That ratio, not the miss count, is what decides whether a shop keeps Spanish switched on.
Has a native speaker signed it off, or has a model? If the honest answer is a model, the honest product decision is the one on this page. The English half of the same exercise is published as its own ledger, and the design that makes both measurable is described on the AI answering page.
No. Spanish is not offered as a line language. The gate's Spanish half went from 13 of 198 panel sentences to 194, but the panels were model personas rather than people, and no human native speaker has signed the vocabulary off.
That the hazard probe asked about a smell while its vocabulary contained only sights and sounds, so a caller who said "no" and then described the smell was scored a clean denial. It also found strict word adjacency that Spanish speakers break, a whole region unable to report a fire because candela was missing, and the commonest carbon monoxide phrasing present only as a noun.
Spanish hazards passing in silence went from 110 to zero and ordinary Spanish calls wrongly hung up on went from 42 to zero, with agreement rising from 13 of 198 to 194. The remaining disagreements are tier arguments that each land on a single clarifying question rather than a hang-up.
Because a panel of model personas cannot tell you how often a phrasing is really used, cannot produce what speech recognition emits for a frightened caller, and converges on safe answers about register. A human native speaker has to reject the over-broad additions and listen to the stop script before it ships.
Yes. Ask which Spanish the vocabulary was written in, whether the false-positive half was measured as hard as the miss half, and whether a native speaker or a model signed it off. A single-dialect vocabulary is invisible until a caller from the wrong country needs it.
If you run a shop and something on this page is wrong — a rule that would wake you for nothing, a phrase the word list should catch and does not — write to hello@oncallbell.com. A person reads it. That is more useful to us than a signup, and there is no signup yet anyway.
Last checked against the code on 2026-09-06