
Representative speech data at scale is one of the hardest problems in AI development, and it's harder still when the languages you need aren't English, Mandarin, or Spanish.
For underserved language markets, the gap isn't just about volume. Publicly available corpora for languages like Odia, Assamese, or Tagalog are sparse, narrow in demographic range, and rarely reflect the acoustic conditions of real-world deployments. What little exists tends to be read speech, controlled, clean, and largely useless for training models that need to handle the way people actually talk.
A frontier AI lab developing ASR capabilities for customer support in travel-tech ran directly into this problem. The target languages, Bengali, Tamil, Telugu, Odia, Assamese, Korean, Tagalog, Arabic, Vietnamese, and Portuguese, each presented their own data scarcity challenges. The timeline was four months. The target was 50,000+ hours of collected and fully annotated audio.
Project Overview
| Category | Details |
|---|---|
| Objective | ASR model training for Travel-Tech Customer Support |
| Volume | 50,000 Hours |
| Languages | Bengali, Tamil, Telugu, Odia, Assamese, Korean, Tagalog, Arabic, Vietnamese, Portuguese |
| Audio Type | Uncompressed .wav, ≥16 kHz, Single and Multi-speaker Speech |
| Transcription & Annotation | 95%+ Accuracy annotated with speaker diarisation, word-level timestamps, disfluency markers, emotion tags, code-switching labels, paralinguistic annotations |
The Data Problem Is Harder Than It Looks
The instinct is often to treat underrepresented language data as a procurement problem: find speakers, record audio, transcribe. But that framing misses what actually makes speech data useful for model training.
Real-world customer support audio is full of phenomena that clean corpora systematically exclude. Speakers produce disfluencies, filled pauses like um and uh, false starts, repetitions, and self-corrections, that are not aberrations but core features of spontaneous speech. Models trained without them fail in deployment, because deployment is where disfluencies live.
Then there's code-switching. Across South and Southeast Asian markets, speakers routinely shift between their native language and English, sometimes mid-sentence, sometimes mid-word. A Tamil speaker troubleshooting a booking might say something like "Ithu cancel pannunga, I already paid." without pause or announcement. Any ASR system deployed in these markets that can't handle intra-sentential code-switching isn't fit for purpose. Yet most available corpora are monolingual by design, because code-switching makes transcription harder and annotation more expensive.
Beyond lexical content, spoken language carries prosodic and paralinguistic information, pitch, rhythm, speaking rate, and non-verbal sounds like laughter, sighs, and breathing, that matters enormously for downstream tasks like sentiment analysis and intent detection. This information is almost never captured in standard transcription pipelines.
Collecting representative data means solving all of this simultaneously, across more than ten languages, with consistent quality standards throughout.
Phase One: Audio Collection
The collection scope required both single-speaker and multi-speaker recordings across scripted and naturalistic formats. Scripted content was designed around the domains relevant to travel-tech customer support, booking changes, refunds, itinerary queries, as well as adjacent domains, including healthcare, legal, finance, and e-commerce to broaden model generalization.
Naturalistic multi-speaker recordings introduced the conversational phenomena that scripted audio can't replicate: overlapping speech, interruptions, backchannels, and organic code-switching. These aren't edge cases in customer support, they're the norm.
Speaker diversity was structured across accent, age, gender, and regional variety for each language. The acoustic variation between a Bengali speaker from Dhaka and one from Kolkata, or between an Arabic speaker from Egypt and one from the Gulf, is significant enough to require explicit representation if a model is expected to generalize across those populations.
The output: 50,000+ hours of raw audio across all target languages, balanced across speaker types and domains.
Phase Two: Transcription and Annotation
- 1. Transcription at 95%+ accuracy across ten languages, several of which have limited annotator pools and complex orthographic conventions, is a substantial operational challenge. But accuracy was only the starting point. Each recording was annotated for:
- 2. Speaker diarisation, speaker identification, and turn-taking labels for multi-speaker files, including handling of overlapping speech segments.
- 3. Word-level timestamps, precise utterance segmentation enabling forced alignment, and fine-grained model training.
- 4. Disfluency and non-verbal annotation, explicit markers for filled pauses, false starts, repairs, laughter, sighs, and breathing, preserving the paralinguistic content of the audio rather than cleaning it out.
- 5. Emotion and sentiment tagging, per-turn labels for affective states including frustration, satisfaction, and confusion, structured for downstream sentiment analysis applications.
- 6. Structured metadata, a standardized schema covering domain, speaker demographics, and audio quality applied consistently across all languages, enabling controlled filtering and stratified training splits.
- 7. Code-switching instances were transcribed and tagged within the annotation pipeline rather than normalized away, preserving them as first-class data rather than noise.
Outcome
At the scale this project required, speaker recruitment is itself a hard problem. Getting 50,000+ hours of demographically diverse audio across ten languages in four months isn't achievable through standard recruitment pipelines; the throughput isn't there, and neither is the geographic reach.
The distribution model that made this possible was community-led. Across India, Korea, Vietnam, the Philippines, Egypt, and Saudi Arabia, we activated 7,000+ local communities, neighborhood networks, regional associations, and professional groups to reach speakers where they already are. Community leaders functioned as a critical node in the distribution network: people with existing trust relationships and on-ground presence in their communities, who could verify participants, coordinate logistics, and sustain engagement at a pace that centralized recruitment cannot replicate.
This structure matters beyond speed. Community-embedded recruitment produces speaker populations that are genuinely representative of how language is used in a given region, including the dialectal variation, code-switching patterns, and demographic spread that targeted sampling from a distance tends to flatten out.
The completed corpus, 50,000+ hours of collected and annotated audio across six-plus underserved language markets, reflects that breadth. The annotation depth, particularly emotion tagging and disfluency markers, extended the dataset's utility beyond ASR training into customer sentiment analysis. The travel-tech applications the lab was originally building for now operate on data that captures not just what customers say, but the conversational dynamics of how they say it.
The core constraint hasn't changed. Obtaining large-scale, linguistically representative speech data for underserved languages remains difficult, by structure, not by accident. The corpora don't exist because creating them requires investment, operational reach, and the kind of community infrastructure that isn't uniformly distributed. That's the problem this project was built to solve.