ASR Models Have Different Failure Geometries
Blogs
ASR model comparisonSeptember 28, 2026Divyajot Singh6 min read

ASR Models Have Different Failure Geometries

We tend to think of speech recognizers as points on a single axis. One model has 18% WER, another has 21%, and the first is therefore a somewhat better version of the second.

The raw results from BRIDGE ASR 2.0 suggest a different picture. Across current ASR systems evaluated on conversational speech, changes in model size, generation and serving tier frequently do not reduce errors uniformly. They change which errors a system makes, which languages it struggles with, and sometimes even which recordings it finds difficult.

For an engineer, that distinction matters. A three-point WER improvement caused by fewer substitutions is not equivalent to the same improvement caused by eliminating catastrophic deletions.

image
image

A more useful representation of an ASR system is therefore an error vector:

M=(S,D,I,CS,P,…)

where substitutions, deletions, insertions, code-switch behaviour and phonetic error jointly describe the system.

Looking at BRIDGE this way exposes several behaviours that disappear inside a leaderboard.

Gemini Pro does not simply look like a better Flash

On 311 Indic files transcribed by both Gemini 2.5 Flash and Gemini 2.5 Pro, Pro improves average language-weighted WER from roughly 34.3% to 25.4%.

But the decomposition is more interesting than the 8.9-point improvement.

Flash averages roughly 18.3% substitution error, compared with 23.3% for Pro. Pro wins because deletion falls from about 7.9% to 4.0%, while insertion collapses from roughly 9.6% to 2.4%.

So Pro is not uniformly better at recognizing words. In fact, it substitutes more of them. What changes dramatically is its tendency to omit speech or generate content that should not be present.

Code-switching moves just as sharply. Flash retains around 55% of expected switched tokens, while Pro reaches roughly 79% recall. Their code-switch F1 differs by almost 30 points.

There is another clue. Once average language difficulty is removed, the file-level error correlation between Flash and Pro is only about 0.20. They frequently do not find the same utterances difficult.

That makes it hard to think of Flash as simply a cheaper Pro for speech workloads. The tier change appears to shift the recognizer into a different failure regime.

Instead of describing an upgrade as:

ΔM=ΔWER

it may be more useful to report:

ΔM=(ΔS,ΔD,ΔI,ΔCS,…)

For production systems, that vector is often more informative than the scalar improvement.

Scaling is surprisingly non-monotonic across languages

GPT-4o Transcribe and GPT-4o Mini Transcribe show a different anomaly.

Across 298 matched Indic files, the full model is about 2.7 percentage points worse in average language-weighted WER than Mini. That average, however, hides extreme language-dependent reversals.

Mini beats the full model by roughly 37 points on Nepali, 24 on Marathi and 15 on Kannada. Yet full GPT-4o is approximately 25 points better on Urdu and 8 points better on Tamil.

This is not a simple "small model beats large model" story. The direction of the effect changes with language.

Whatever differs between the two transcription systems appears to interact strongly with linguistic context. That could involve training mixture, post-training, language conditioning or decoding, but the available data cannot tell us which mechanism is responsible.

The behavioural result is still important: greater model capacity does not produce a monotonic multilingual ASR improvement.

A natural follow-up experiment would hold the acoustic environment fixed while varying language. Measuring script choice, language identification, deletion rate and code-switch retention could help locate whether these reversals originate primarily in acoustic recognition or in the linguistic prior imposed during decoding.

BRIDGE does not answer that question. It tells us where to look.

Saaras V4 improves without changing what it finds difficult

Saaras V3 to V4 shows a very different kind of model improvement from Gemini.

Across 216 recordings shared by both versions, V4 reduces language-weighted WER from roughly 17.0% to 16.4% and performs better on about 62% of matched files. Most of that gain comes from lower substitution and insertion error, while deletion rises slightly and code-switch F1 remains almost unchanged.

The more interesting result is what happens when we compare which recordings the two systems find difficult.

After removing each model's average language difficulty, the correlation between V3 and V4 file-level error is approximately:

R≈0.97

That is remarkably high. Files that are difficult for V3 are almost exactly the files that remain difficult for V4.

The improvement is also concentrated rather than uniform. V4 reduces lwWER by roughly 3.9 points on Urdu and 3.3 on Odia, with smaller gains on Nepali, Assamese and Hindi. At the same time, it regresses slightly on Tamil, Kannada, Marathi, Telugu and Malayalam.

This gives us two useful forms of ASR progress.

One is refinement: roughly the same failure surface, but lower error across parts of it. The other is regime change: a new model develops a substantially different set of strengths and weaknesses.

Saaras V3 to V4 looks strongly like the first. Gemini 2.5 Flash to Pro looks much closer to the second.

WER improvement alone makes both transitions look similar: the newer system performs better. The file-level structure tells us something more useful. In Saaras V4, progress appears to come from improving performance within an existing robustness regime rather than learning a fundamentally different notion of which speech is difficult.

Silence exposes the recognizer, not just the segmenter

Long silence produces another useful example.

The straightforward explanation is that long gaps trigger bad endpointing or segmentation. That is certainly plausible, but decomposing the resulting errors shows that systems react differently to the same condition.

For Scribe v2, controlling for language, call duration, overlap and conversational density, calls containing gaps longer than 20 seconds show roughly a 7.9-point increase in language-weighted WER. Surprisingly, much of that increase appears in substitutions rather than deletions.

GPT-4o shows a different signature, with long-gap degradation much more strongly associated with deletion errors. Other systems show increased insertion behaviour.

So "silence causes segmentation failures" is too coarse.

Silence is better treated as a perturbation. Different architectures and serving pipelines react differently after their temporal context is interrupted.

That suggests a simple controlled experiment: take identical recordings and synthetically insert 0, 5, 10, 20 and 40 seconds of silence at fixed locations, then measure

S(t),D(t),I(t)

as functions of the inserted gap.

If one model begins deleting, another substituting and another inserting, we have learned considerably more than we would from saying that WER increased by seven points.

image
image

From leaderboards to failure maps

There is one important methodological caveat. The BRIDGE evaluation matrix is not perfectly rectangular because different providers support different language sets. Unsupported languages should not be interpreted in the same way as poor performance within a model's intended operating envelope.

More broadly, though, the raw results suggest that the interesting question is no longer simply where a recognizer sits on a leaderboard.

Gemini's tier change reorganizes the composition of error. GPT-4o's scale change reverses direction depending on language. Saaras improves while retaining almost exactly the same map of difficult recordings. Long silence pushes different systems toward different error types.

Those behaviours are not details around the benchmark score. They are the behaviour of the models.

For researchers, the more useful question is how training, post-training, decoding and serving choices reshape this failure geometry.

For engineers, the question is even simpler:

Not which ASR model fails least, but which one fails in ways your application can tolerate.

Want to challenge BRIDGE ASR 2.0? We would love to hear from you.

Explore the full BRIDGE ASR 2.0 Report here: Click here!

Explore the methodology here: Click here!

Explore the datasets here: Click here!

More Articles