
BRIDGE ASR 2.0
Every ASR benchmark stops at the score.
We start there.
Physical AI can see. Hearing the real world is next.
The only global independent ASR benchmark evaluating 23 models for real-world deployment across a 6-metric stack.
How the Benchmark was built ?
A dataset built for how people actually talk, and a metric stack built to show exactly where a model breaks, not just that it did. The scaffolding is simple to state: genuine, spontaneous dual-speaker conversations, tagged across seven cohort dimensions, normalised through the same three-step cleanup before anything is scored, then run identically across all 23 models.
Voice AI's requirements have shifted fast from clean, single-speaker dictation toward noisy, overlapping, real conversation and the dataset reflects where that need is now. Every recording is a genuine two-person conversation with real overlapping speech and cross-talk, the exact failure mode that breaks streaming ASR and speaker separation in production, captured across a wide range of everyday acoustic environments rather than a single clean studio.
The scaffolding behind every score is the same regardless of model or language: genuine, spontaneous dual-speaker conversations, tagged across seven cohort dimensions, normalised through an identical three-step text cleanup before anything is scored, then run against every provider on the exact same set. That consistency is what makes a result attributable to a specific condition, a language, a noise level, a speaker overlap rather than lost inside a single average rank.
WER and CER anchor the stack, alongside MER and WIL — two bounded variants that stay stable even when a model hallucinates heavily, where plain WER can exceed 100% and break model-to-model comparison. Every error is also classified as a substitution, deletion, or insertion: high substitution means the model is misrecognising speech, high deletion means it's dropping speech outright, usually a segmentation issue — and high insertion means it's inventing words that were never spoken. A single WER collapses all three into one number and hides which is actually happening.
Standard WER counts a correctly recognised English word written in local script as a full error against a Latin-script reference — lwWER and script_penalty separate that script mismatch from a real transcription error. SemanticSim and LevenshteinSim ask a different question again: did the meaning survive even where the exact wording didn't, and how close is the literal string. And because Indic conversation routinely drops English words into a native-language sentence, PIER, CS Recall, CS Precision, and CS F1 measure specifically whether that code-switching was handled — or missed, or fabricated — correctly.
Model Leaderboard
All 23 models across 6 metrics — WER, CER, SemSim, CS F1, PIER, and WIL — scored on a dataset you can inspect yourself: the evaluation set is live here. Filter by language to see how each model performs per-language. Filter by metric to change what the bars represent.
Humyn Score
- ElevenLabs Scribe v20.876
- Gemini 3 Flash (Preview)0.808
- Gemini 3 Pro (Preview)0.798
- Sarvam saaras v30.787
- Gemini 2.5 Pro0.780
- Soniox stt-async-v40.775
- Sarvam saarika v2.50.767
- Google Chirp 30.748
- AWS Transcribe0.741
- Gemini 3.5 Transcribe0.710
- Azure (Conv. Transcriber)0.706
- Deepgram Nova-30.688
- Gemini 2.5 Flash0.652
- Microsoft MAI Transcribe 1.50.647
- Speechmatics0.624
- Gnani Vachana v30.609
- Speechmatics Melia0.443
- Gladia v20.350
- OpenAI GPT-4o mini0.245
- OpenAI GPT-4o0.234
- AssemblyAI Universal0.197
- AssemblyAI Universal 3 Pro0.156
- AssemblyAI Universal 20.156
Bars show each model's Humyn Score for the selected scope, computed across the collected dual-speaker conversations 10–15 minutes long, captured at 44–48 kHz.
One model leads by a clear margin.
ElevenLabs Scribe v2 leads with 10.99% lwWER, compared with 16.32% for Soniox stt-async-v4 — a 5.3pp gap and about a third fewer word errors. It ranks first in 14 of the 18 languages, one more than it wins on WER. Brand name does not guarantee performance. The bottom six Indic models range from 70.77% to 88.03% lwWER.
Key Findings
A single WER number can only ever say one model beat another. It can't say where either of them is quietly failing, why the number moved, or what to actually do about it — and that's the real job of this section. Every relationship below is a gap: a place where the headline metric hides something, whether that's a metric worth double-checking before you trust it, a model worth routing around, or just a genuinely surprising fact about how these systems break. Read in order, they build on each other, and the picture that comes out the other end is less “here is the winner” and more “here is exactly where every model — including the winner — still has a hole, and what that hole is made of.”
Cohort Performance Analysis
Choose a cohort dimension, a metric (WER, CER, SemSim, CS F1, lwWER, PIER, or WIL), and a model to see how performance shifts across conditions. All three filters work together — any combination is valid.
Bars show mean Word Error Rate per cohort category for each model. Δ = spread across the displayed categories. Categories with fewer than 2 audio samples for that model are excluded.
The usual assumption is that fast conversations are harder. Our data says the bigger problem is long silence.
After controlling for language and model, calls with gaps longer than 20 seconds lose an average of 5.05pp of lwWER, compared with just 1.33pp for rapid turn-taking. The effect shows up across almost every viable model: chirp_3 loses 13.22pp to long silences versus 6.69pp to speed; Gemini 2.5 Pro loses 10.87pp versus 2.16pp; even Scribe v2 loses 7.50pp versus 0.79pp. Controlling for language matters. Long pauses are concentrated in Marwari, which is already the hardest language in the set. Without that control, the apparent silence penalty rises to 16.33pp — overstating the effect of silence itself. The likely culprit is segmentation, not speech recognition.
Access dataset & benchmark
Dataset access
& citation
ASR benchmarks weren't built for the languages you're working on. BRIDGE was. Get access to the new 6 metric benchmark to evaluate your model.
The sample BRIDGE corpus — Indic + Latin American Spanish + Brazilian Portuguese + Vietnamese — including audio files, golden transcripts, speaker metadata, cohort labels, and evaluation scripts is available on below.
Audio files, golden transcripts, speaker metadata, cohort labels, and evaluation scripts are available under the IndicBench dataset card. Additional languages and an overlap-focused corpus are in preparation.
If you use this benchmark in your research, please cite the following.
@misc{humynlabs_bridge_v2_2026,title = {BRIDGE: State of Conversational ASR Across the Global South},author = {HumynLabs Research Team},year = {2026},month = {September},note = {Independent benchmark evaluating 23 commercial ASR APIs on dual-speaker conversational audio across 21 languages — Indic (18 languages, 22+ Indian states), Latin American Spanish (Argentinian, Peruvian, Venezuelan), Brazilian Portuguese, and Vietnamese — scored on a 6-metric stack across 7 cohort dimensions},howpublished = {HumynLabs}}
