BRIDGE ASR 2.0

00
ASR Models Tested
00
Languages
00
Cohort Dimensions

Every ASR benchmark stops at the score.
We start there.

Physical AI can see. Hearing the real world is next.
The only global independent ASR benchmark evaluating 23 models for real-world deployment across a 6-metric stack.

Context

Why this Benchmark exists?

From speaker recruitment to evaluation pipeline — the decisions that make BRIDGE reproducible, auditable, and resistant to benchmark gaming.

Methodology

How the Benchmark was built ?

A dataset built for how people actually talk, and a metric stack built to show exactly where a model breaks, not just that it did. The scaffolding is simple to state: genuine, spontaneous dual-speaker conversations, tagged across seven cohort dimensions, normalised through the same three-step cleanup before anything is scored, then run identically across all 23 models.

6 Metric Evaluation Stack
WER — Word Error RateCER — Character Error RateSemSim — Semantic SimilarityCS F1 — Code-Switch F1PIER — Phoneme-Informed ERWIL — Word Information Lost
(A harder, more realistic dataset)

Voice AI's requirements have shifted fast from clean, single-speaker dictation toward noisy, overlapping, real conversation and the dataset reflects where that need is now. Every recording is a genuine two-person conversation with real overlapping speech and cross-talk, the exact failure mode that breaks streaming ASR and speaker separation in production, captured across a wide range of everyday acoustic environments rather than a single clean studio.

(Normalised and tagged before a single score runs)

The scaffolding behind every score is the same regardless of model or language: genuine, spontaneous dual-speaker conversations, tagged across seven cohort dimensions, normalised through an identical three-step text cleanup before anything is scored, then run against every provider on the exact same set. That consistency is what makes a result attributable to a specific condition, a language, a noise level, a speaker overlap rather than lost inside a single average rank.

(Core accuracy, and where the errors actually are)

WER and CER anchor the stack, alongside MER and WIL — two bounded variants that stay stable even when a model hallucinates heavily, where plain WER can exceed 100% and break model-to-model comparison. Every error is also classified as a substitution, deletion, or insertion: high substitution means the model is misrecognising speech, high deletion means it's dropping speech outright, usually a segmentation issue — and high insertion means it's inventing words that were never spoken. A single WER collapses all three into one number and hides which is actually happening.

(Script fairness, meaning, and code-switching)

Standard WER counts a correctly recognised English word written in local script as a full error against a Latin-script reference — lwWER and script_penalty separate that script mismatch from a real transcription error. SemanticSim and LevenshteinSim ask a different question again: did the meaning survive even where the exact wording didn't, and how close is the literal string. And because Indic conversation routinely drops English words into a native-language sentence, PIER, CS Recall, CS Precision, and CS F1 measure specifically whether that code-switching was handled — or missed, or fabricated — correctly.

Results

Model Leaderboard

All 23 models across 6 metrics — WER, CER, SemSim, CS F1, PIER, and WIL — scored on a dataset you can inspect yourself: the evaluation set is live here. Filter by language to see how each model performs per-language. Filter by metric to change what the bars represent.

What is Humyn Score?
All languages · Humyn Score (HumynScore)
Higher is better
Humyn Score
0.00
0.15
0.29
0.44
0.59
0.73
0.88
  • ElevenLabs Scribe v2
    0.876
  • Gemini 3 Flash (Preview)
    0.808
  • Gemini 3 Pro (Preview)
    0.798
  • Sarvam saaras v3
    0.787
  • Gemini 2.5 Pro
    0.780
  • Soniox stt-async-v4
    0.775
  • Sarvam saarika v2.5
    0.767
  • Google Chirp 3
    0.748
  • AWS Transcribe
    0.741
  • Gemini 3.5 Transcribe
    0.710
  • Azure (Conv. Transcriber)
    0.706
  • Deepgram Nova-3
    0.688
  • Gemini 2.5 Flash
    0.652
  • Microsoft MAI Transcribe 1.5
    0.647
  • Speechmatics
    0.624
  • Gnani Vachana v3
    0.609
  • Speechmatics Melia
    0.443
  • Gladia v2
    0.350
  • OpenAI GPT-4o mini
    0.245
  • OpenAI GPT-4o
    0.234
  • AssemblyAI Universal
    0.197
  • AssemblyAI Universal 3 Pro
    0.156
  • AssemblyAI Universal 2
    0.156

Bars show each model's Humyn Score for the selected scope, computed across the collected dual-speaker conversations 10–15 minutes long, captured at 44–48 kHz.

FINDINGS #01

One model leads by a clear margin.

ElevenLabs Scribe v2 leads with 10.99% lwWER, compared with 16.32% for Soniox stt-async-v4 — a 5.3pp gap and about a third fewer word errors. It ranks first in 14 of the 18 languages, one more than it wins on WER. Brand name does not guarantee performance. The bottom six Indic models range from 70.77% to 88.03% lwWER.

RESEARCH SUMMARY

Key Findings

A single WER number can only ever say one model beat another. It can't say where either of them is quietly failing, why the number moved, or what to actually do about it — and that's the real job of this section. Every relationship below is a gap: a place where the headline metric hides something, whether that's a metric worth double-checking before you trust it, a model worth routing around, or just a genuinely surprising fact about how these systems break. Read in order, they build on each other, and the picture that comes out the other end is less “here is the winner” and more “here is exactly where every model — including the winner — still has a hole, and what that hole is made of.”

Multi-Dimensional Analysis

Cohort Performance Analysis

Choose a cohort dimension, a metric (WER, CER, SemSim, CS F1, lwWER, PIER, or WIL), and a model to see how performance shifts across conditions. All three filters work together — any combination is valid.

Region · Word Error Rate (WER)
Same StateCross State
ElevenLabs Scribe v2
overall 14.4% · Δ ±1.9pp
Same State
13.5%
Cross State
15.4%
Soniox stt-async-v4
overall 20.2% · Δ ±6.7pp
Same State
23.6%
Cross State
16.9%
Gemini 3 Flash (Preview)
overall 23.4% · Δ ±8.6pp
Same State
27.7%
Cross State
19.1%
Sarvam saaras v3
overall 24.3% · Δ ±2.0pp
Same State
25.3%
Cross State
23.3%
Deepgram Nova-3
overall 25.8% · Δ ±14.6pp
Same State
33.1%
Cross State
18.5%
Google Chirp 3
overall 27.6% · Δ ±0.7pp
Same State
27.3%
Cross State
28.0%
Gemini 3 Pro (Preview)
overall 28.1% · Δ ±1.9pp
Same State
27.1%
Cross State
29.0%
Microsoft MAI Transcribe 1.5
overall 28.1% · Δ ±28.8pp
Same State
42.5%
Cross State
13.7%
Sarvam saarika v2.5
overall 29.4% · Δ ±5.3pp
Same State
26.8%
Cross State
32.1%
AWS Transcribe
overall 30.5% · Δ ±2.4pp
Same State
29.3%
Cross State
31.7%
Gemini 3.5 Transcribe
overall 30.6% · Δ ±1.2pp
Same State
31.2%
Cross State
30.0%
Azure (Conv. Transcriber)
overall 32.6% · Δ ±4.8pp
Same State
35.0%
Cross State
30.2%
Gemini 2.5 Flash
overall 35.1% · Δ ±5.6pp
Same State
32.3%
Cross State
37.9%
Speechmatics
overall 36.4% · Δ ±7.3pp
Same State
40.1%
Cross State
32.7%
Gnani Vachana v3
overall 39.2% · Δ ±5.3pp
Same State
41.8%
Cross State
36.5%
Gemini 2.5 Pro
overall 39.5% · Δ ±23.2pp
Same State
27.9%
Cross State
51.1%
Gladia v2
overall 57.5% · Δ ±38.1pp
Same State
76.5%
Cross State
38.5%
Speechmatics Melia
overall 59.3% · Δ ±1.8pp
Same State
58.4%
Cross State
60.2%
OpenAI GPT-4o
overall 77.2% · Δ ±20.3pp
Same State
87.4%
Cross State
67.1%
OpenAI GPT-4o mini
overall 78.0% · Δ ±12.1pp
Same State
84.1%
Cross State
72.0%
AssemblyAI Universal
overall 79.7% · Δ ±13.2pp
Same State
86.2%
Cross State
73.1%
AssemblyAI Universal 2
overall 80.3% · Δ ±22.9pp
Same State
91.7%
Cross State
68.8%
AssemblyAI Universal 3 Pro
overall 80.3% · Δ ±22.9pp
Same State
91.8%
Cross State
68.9%

Bars show mean Word Error Rate per cohort category for each model. Δ = spread across the displayed categories. Categories with fewer than 2 audio samples for that model are excluded.

6
FINDINGS #01

The usual assumption is that fast conversations are harder. Our data says the bigger problem is long silence.

After controlling for language and model, calls with gaps longer than 20 seconds lose an average of 5.05pp of lwWER, compared with just 1.33pp for rapid turn-taking. The effect shows up across almost every viable model: chirp_3 loses 13.22pp to long silences versus 6.69pp to speed; Gemini 2.5 Pro loses 10.87pp versus 2.16pp; even Scribe v2 loses 7.50pp versus 0.79pp. Controlling for language matters. Long pauses are concentrated in Marwari, which is already the hardest language in the set. Without that control, the apparent silence penalty rises to 16.33pp — overstating the effect of silence itself. The likely culprit is segmentation, not speech recognition.

Multi-Dimensional Analysis

The Hidden Quality Gap

CS F1 measures whether English vocabulary embedded in Indic speech is preserved — not dropped, not transliterated. A model that turns "data backup" into "डेटा बैकअप" scores 0 on CS F1. Invisible to WER. Fatal for downstream applications.

CS F1 Score — All Models (higher = better)
  • ElevenLabs Scribe v2
    0.897
  • Gemini 3 Pro (Preview)
    0.891
  • Gemini 3 Flash (Preview)
    0.886
  • Gemini 2.5 Pro
    0.842
  • Google Chirp 3
    0.806
  • Sarvam saarika v2.5
    0.798
  • Sarvam saaras v3
    0.794
  • AWS Transcribe
    0.780
  • Speechmatics
    0.774
  • Gemini 3.5 Transcribe
    0.773
  • Soniox stt-async-v4
    0.758
  • Azure (Conv. Transcriber)
    0.750
  • Deepgram Nova-3
    0.744
  • Microsoft MAI Transcribe 1.5
    0.711
  • Gnani Vachana v3
    0.703
  • Speechmatics Melia
    0.682
  • Gemini 2.5 Flash
    0.542
  • Gladia v2
    0.441
  • AssemblyAI Universal 3 Pro
    0.417
  • AssemblyAI Universal 2
    0.414
  • OpenAI GPT-4o
    0.343
  • OpenAI GPT-4o mini
    0.340
  • AssemblyAI Universal
    0.157
Strong code-switching — CS F1 ≥ 0.7

These models preserve English vocabulary in Indic speech. Both are viable for enterprise applications where English terminology appears in native-language conversation.

15 models in this bucket · ElevenLabs Scribe v2, Gemini 3 Pro (Preview), Gemini 3 Flash (Preview)
Partial — CS F1 0.2–0.7

These models handle some code-switching. Performance varies by language and English density — verify on your specific use case before shipping.

7 models in this bucket · Speechmatics Melia, Gemini 2.5 Flash, Gladia v2
Zero code-switching — CS F1 < 0.2

These models systematically transliterate or drop English tokens. Not suitable for code-mixed enterprise Indic applications.

1 model in this bucket · AssemblyAI Universal
FINDINGS #01

“Indic models drop English” is true for some models, but misleading for others.

Two nearly independent measures separate the failures: Script penalty: the English survives, but is written in the local script. saarika v2.5, saaras v3, AWS, chirp_3, and Soniox show 6–8pp penalties while maintaining ~90–95% code-switch precision. Code-switch recall: the English disappears entirely. Gemini 2.5 Flash and melia-1 show low recall (~51–52%) despite much smaller script penalties. The first problem is recoverable with transliteration-aware matching. The second is not: there is no token to recover. CS F1 collapses these into one number, making a fixable script error look identical to a fundamental recognition failure.

Access dataset & benchmark

Dataset access
& citation

ASR benchmarks weren't built for the languages you're working on. BRIDGE was.  Get access to the new 6 metric benchmark to evaluate your model.

Audio files, golden transcripts, speaker metadata, cohort labels, and evaluation scripts are available under the IndicBench dataset card. Additional languages and an overlap-focused corpus are in preparation.

If you use this benchmark in your research, please cite the following.

citation.bib
@misc{humynlabs_bridge_v2_2026,
title = {BRIDGE: State of Conversational ASR Across the Global South},
author = {HumynLabs Research Team},
year = {2026},
month = {September},
note = {Independent benchmark evaluating 23 commercial ASR APIs on dual-speaker conversational audio across 21 languages — Indic (18 languages, 22+ Indian states), Latin American Spanish (Argentinian, Peruvian, Venezuelan), Brazilian Portuguese, and Vietnamese — scored on a 6-metric stack across 7 cohort dimensions},
howpublished = {HumynLabs}
}