Code-switch detection is weak even on clean audio. On Code-Switch F1 — a labeling score for how well a model tags which language each word is, not whether it heard it right — ElevenLabs' Scribe v2 manages 0.83 and Sarvam's Saaras v3 just 0.58. And because borrowed words are 10–30% of a mixed-language sentence, that small labeling slip compounds: an acoustically perfect transcript comes back at 15% Word Error Rate.
Nothing was misheard. The error is in the grading — not the model, not the dataset. Here's how it forms, and how we removed it.
The words we borrow
Over half the world is bilingual, and bilinguals mix. Hindi–English "Hinglish," Spanish–English "Spanglish," Arabic–French across the Maghreb, Tagalog–English "Taglish," Swahili–English "Sheng" — everyday speech in most of the world braids two languages into one sentence. The tools that grade speech AI were built as if it didn't.
A Hindi speaker says maine usko call kiya ("I called him"). That "call" is a nativized loanword — a foreign word absorbed into Hindi, said with Hindi sounds — and it can be written two valid ways: call or कॉल. IIT's COMI-LINGUA, the largest hand-checked Hindi–English code-mixed dataset, records both scripts side by side. A test that assumes a single correct spelling is measuring the wrong thing.
कॉल ≠ call
Word Error Rate compares the model's transcript to the SOP reference word by word, and has no idea if two spellings can be the same word.
Take that four-word sentence. Our baseline SOP is to write each word in the script of its own language. So the reference keeps the English loanword in English letters:
मैंने उसको call किया
The model hears it perfectly — and writes that same word in Hindi letters:
मैंने उसको कॉल किया
Same word, same sound. But कॉल ≠ call, so the scoreboard counts a wrong word — one error in a four-word line. 25%, out of thin air. This is not a speech-recognition failure — the model mishearing the audio. The transcript is right, the measurement invented the error.
Which still compounds into 15%
It isn't a fluke, because the trigger is in every sentence. WER = (Substitutions + Deletions + Insertions) / total reference words. Loanwords are 10–30% of a mixed-language turn, each a coin-toss to disagree with the reference. That's how a clean transcript compounds to 15% WER.
The fix — the measurement. Not the model.
A normalization pipeline with a judgement layer that only ever edits the model's output, never the reference.
Step zero — tag the language. Token-level language identification: every word labeled Hindi, English, or other. Without it you can't enforce a script convention. COMI-LINGUA formalises exactly this task, and records both scripts side by side.
Then four layers, running on the model's output only — never the reference:
1. Lexical Mapping — a bidirectional lookup between Latin-script tokens and their native-script renderings, plus observed spelling variants, so call and कॉल resolve to one word. This is the only language-specific piece, and it does most of the work: on its own it takes 15.0% to 3.3%.
2. Confidence Tiers — bucket each token by decoder confidence. The low-confidence ones cluster at switch boundaries; route them to a second pass instead of accepting them at face value.
3. Collision Guard — stop genuinely distinct native words from being flattened into an English candidate on phonetic similarity alone. सर is not always sir.
4. SOP Rules — our baseline standard is write each word in the script of its own language. SOP Rules lets a client override it. Same audio, different declared convention, and the number stays trustworthy for whoever's asking.
Those last three take 3.3% to 3.0%.
Swap the lexicon and the same stack runs on Spanglish, Taglish, Sheng. The conflict is always the same: two valid spellings of one spoken word.
What we didn't fix
We didn't improve Code-Switch F1 either. The model labels language exactly as before. What changed is that a script difference between two valid spellings no longer counts as a misheard word.
That this is artifact-removal and not gaming rests on two things: the baseline reference commits to one canonical form per word, and the mapping is declared up front and applied to every output identically — never tuned per example.
A number you can trust
Score against a benchmark that assumes one right spelling and you don't just miscount — you teach the next model, and the next dataset, to treat the way people talk as a defect.
Which is why code-switched Indian speech needs its own tests from scratch, as the COSTA work had to be built for Bengali, Hindi, Marathi, and Telugu. If you'd rather not build it, hand us the project — the normalized output comes back with the evaluation already resolved.
