50,000 hours of Speech Data for the 10 Languages Driving the Next Billion AI Users
Case Studies
Audio

50,000 hours of Speech Data for the 10 Languages Driving the Next Billion AI Users

Representative speech data at scale is one of the hardest problems in AI development, and it's harder still when the languages you need aren't English, Mandarin, or Spanish.

For underserved language markets, the gap isn't just about volume. Publicly available corpora for languages like Odia, Assamese, or Tagalog are sparse, narrow in demographic range, and rarely reflect the acoustic conditions of real-world deployments. What little exists tends to be read speech, controlled, clean, and largely useless for training models that need to handle the way people actually talk.

A frontier AI lab developing ASR capabilities for customer support in travel-tech ran directly into this problem. The target languages, Bengali, Tamil, Telugu, Odia, Assamese, Korean, Tagalog, Arabic, Vietnamese, and Portuguese, each presented their own data scarcity challenges. The timeline was four months. The target was 50,000+ hours of collected and fully annotated audio.

Project Overview

CategoryDetails
ObjectiveASR model training for Travel-Tech Customer Support
Volume50,000 Hours
LanguagesBengali, Tamil, Telugu, Odia, Assamese, Korean, Tagalog, Arabic, Vietnamese, Portuguese
Audio TypeUncompressed .wav, ≥16 kHz, Single and Multi-speaker Speech
Transcription & Annotation95%+ Accuracy annotated with speaker diarisation, word-level timestamps, disfluency markers, emotion tags, code-switching labels, paralinguistic annotations

The Data Problem Is Harder Than It Looks

The instinct is often to treat underrepresented language data as a procurement problem: find speakers, record audio, transcribe. But that framing misses what actually makes speech data useful for model training.

Real-world customer support audio is full of phenomena that clean corpora systematically exclude. Speakers produce disfluencies, filled pauses like um and uh, false starts, repetitions, and self-corrections, that are not aberrations but core features of spontaneous speech. Models trained without them fail in deployment, because deployment is where disfluencies live.

Then there's code-switching. Across South and Southeast Asian markets, speakers routinely shift between their native language and English, sometimes mid-sentence, sometimes mid-word. A Tamil speaker troubleshooting a booking might say something like "Ithu cancel pannunga, I already paid." without pause or announcement. Any ASR system deployed in these markets that can't handle intra-sentential code-switching isn't fit for purpose. Yet most available corpora are monolingual by design, because code-switching makes transcription harder and annotation more expensive.

Beyond lexical content, spoken language carries prosodic and paralinguistic information, pitch, rhythm, speaking rate, and non-verbal sounds like laughter, sighs, and breathing, that matters enormously for downstream tasks like sentiment analysis and intent detection. This information is almost never captured in standard transcription pipelines.

Collecting representative data means solving all of this simultaneously, across more than ten languages, with consistent quality standards throughout.

Phase One: Audio Collection

The collection scope required both single-speaker and multi-speaker recordings across scripted and naturalistic formats. Scripted content was designed around the domains relevant to travel-tech customer support, booking changes, refunds, itinerary queries, as well as adjacent domains, including healthcare, legal, finance, and e-commerce to broaden model generalization.

Naturalistic multi-speaker recordings introduced the conversational phenomena that scripted audio can't replicate: overlapping speech, interruptions, backchannels, and organic code-switching. These aren't edge cases in customer support, they're the norm.

Speaker diversity was structured across accent, age, gender, and regional variety for each language. The acoustic variation between a Bengali speaker from Dhaka and one from Kolkata, or between an Arabic speaker from Egypt and one from the Gulf, is significant enough to require explicit representation if a model is expected to generalize across those populations.

The output: 50,000+ hours of raw audio across all target languages, balanced across speaker types and domains.

Phase Two: Transcription and Annotation

Outcome

At the scale this project required, speaker recruitment is itself a hard problem. Getting 50,000+ hours of demographically diverse audio across ten languages in four months isn't achievable through standard recruitment pipelines; the throughput isn't there, and neither is the geographic reach.

The distribution model that made this possible was community-led. Across India, Korea, Vietnam, the Philippines, Egypt, and Saudi Arabia, we activated 7,000+ local communities, neighborhood networks, regional associations, and professional groups to reach speakers where they already are. Community leaders functioned as a critical node in the distribution network: people with existing trust relationships and on-ground presence in their communities, who could verify participants, coordinate logistics, and sustain engagement at a pace that centralized recruitment cannot replicate.

This structure matters beyond speed. Community-embedded recruitment produces speaker populations that are genuinely representative of how language is used in a given region, including the dialectal variation, code-switching patterns, and demographic spread that targeted sampling from a distance tends to flatten out.

The completed corpus, 50,000+ hours of collected and annotated audio across six-plus underserved language markets, reflects that breadth. The annotation depth, particularly emotion tagging and disfluency markers, extended the dataset's utility beyond ASR training into customer sentiment analysis. The travel-tech applications the lab was originally building for now operate on data that captures not just what customers say, but the conversational dynamics of how they say it.

The core constraint hasn't changed. Obtaining large-scale, linguistically representative speech data for underserved languages remains difficult, by structure, not by accident. The corpora don't exist because creating them requires investment, operational reach, and the kind of community infrastructure that isn't uniformly distributed. That's the problem this project was built to solve.