Languages
Where we are, per capability.
Every language above is paired with English, because that is how it is spoken in an exam room. There is no Gujarati-only mode and there is no English-only mode; the model decodes the mixture.
A vendor claiming uniform production support across a dozen Indic languages has either not measured or is not telling you. Six of the eleven below are beta or in training.
| Language | Code-switched recognition All | Clinical entity linking Scribe | Speech synthesis Front Desk · Follow-Up | Red-flag lexicon Front Desk · Follow-Up |
|---|---|---|---|---|
| Hindi हिन्दी | Production — we would put a pilot on it | Production — we would put a pilot on it | Production — we would put a pilot on it | Production — we would put a pilot on it |
| Gujarati ગુજરાતી | Production — we would put a pilot on it | Production — we would put a pilot on it | Production — we would put a pilot on it | Production — we would put a pilot on it |
| Telugu తెలుగు | Production — we would put a pilot on it | Production — we would put a pilot on it | Beta — works, fails visibly, we want your corrections | Production — we would put a pilot on it |
| Tamil தமிழ் | Production — we would put a pilot on it | Beta — works, fails visibly, we want your corrections | Beta — works, fails visibly, we want your corrections | Production — we would put a pilot on it |
| Marathi मराठी | Beta — works, fails visibly, we want your corrections | Beta — works, fails visibly, we want your corrections | Beta — works, fails visibly, we want your corrections | Beta — works, fails visibly, we want your corrections |
| Punjabi ਪੰਜਾਬੀ | Beta — works, fails visibly, we want your corrections | Beta — works, fails visibly, we want your corrections | Beta — works, fails visibly, we want your corrections | Beta — works, fails visibly, we want your corrections |
| Bengali বাংলা | Beta — works, fails visibly, we want your corrections | Beta — works, fails visibly, we want your corrections | In training — not shipped, listed so you know where we are going | Beta — works, fails visibly, we want your corrections |
| Kannada ಕನ್ನಡ | Beta — works, fails visibly, we want your corrections | In training — not shipped, listed so you know where we are going | In training — not shipped, listed so you know where we are going | Beta — works, fails visibly, we want your corrections |
| Malayalam മലയാളം | Beta — works, fails visibly, we want your corrections | In training — not shipped, listed so you know where we are going | In training — not shipped, listed so you know where we are going | Beta — works, fails visibly, we want your corrections |
| Urdu اُردُو | In training — not shipped, listed so you know where we are going | In training — not shipped, listed so you know where we are going | In training — not shipped, listed so you know where we are going | In training — not shipped, listed so you know where we are going |
| Spanish Español | In training — not shipped, listed so you know where we are going | In training — not shipped, listed so you know where we are going | In training — not shipped, listed so you know where we are going | In training — not shipped, listed so you know where we are going |
On the red-flag lexicon
The red-flag lexicon is built per language by linguists in colloquial register, not translated from a list of clinical terms. “छाती में भारी लग रहा है” has to fire as reliably as “chest pain”. Each language is reviewed by a native-speaking clinician before that tier moves to production.
Published accuracy
We have not published accuracy numbers yet.
The eval harness is built and the numbers are not ready. Publishing a figure we cannot stratify by language and code-mix band would be worse than publishing nothing, because the aggregate number is the one that hides the failure. When these tables have real values in them, they will name the eval set and the date.
If you are evaluating us, ask for the current run. We will send it whether or not it flatters us.
How we will measure it
Published in advance, so that when the numbers arrive you can check they were produced the way we said they would be.
- Stratified by code-mix ratio, not averaged
- 0%, 10–30%, 30–60% and 60%+ English within the utterance. The middle bands are where English-first models collapse, and they are also the most common register in a real exam room. An aggregate number conceals exactly the thing you are buying.
- Entity-level F1, not just word error rate
- WER is table stakes and the least informative number here. Medication, dosage, symptom, duration and negation are scored separately, because those are the errors that reach a patient.
- Numeric and dosage exact-match, reported separately
- Spoken quantities crossing a code-switch boundary are where this product could hurt someone. That number gets its own row and is never folded into an average.
- Stratified by speaker age band
- Models tuned on young urban speakers degrade badly on a 74-year-old with a regional vocabulary from the 1950s — and that speaker is disproportionately this product's patient.
- Negation is evaluated as its own component
- “No chest pain” and “chest pain” differ by one token and by everything.
- Latency at p50, p95 and p99
- p99 is what the angry customer experienced. Reporting only the median is a way of not answering.
Pilot
Book a pilot call.
Thirty minutes. We will run the demo on the languages your panel actually speaks, show you the eval methodology, and tell you plainly which parts are not ready. If we are not a fit we would rather find out on this call.
One reply from a person. We do not run a nurture sequence and we do not share your address.