
Voice AI in financial services struggles for predictable reasons: inconsistent normalisation, non-representative voices, or licensing that blocks deployment. Model architectures are mature. The bottleneck is acquiring studio-grade, consent-cleared speech that reflects how users actually talk about money.
We built a 1,000+ hour studio TTS corpus across 12 Indian languages, designed specifically for BFSI deployments. The dataset integrates normalisation-first scripting, production-grade acoustic capture, and structured code-switching - ensuring models trained on it inherit deterministic rendering behaviour, demographic realism, and deployment-safe licensing.
Project Summary
| Category | Specification |
|---|---|
| Objective | Studio-quality TTS corpus for Indian BFSI voice AI |
| Volume | 1,000+ hours scripted speech |
| Languages | Hindi, English, Bengali, Kannada, Tamil, Telugu, Punjabi, Marathi, Gujarati, Assamese, Odia, Malayalam |
| Audio Spec | 48 kHz, 24-bit, mono WAV; studio-controlled capture |
| Normalisation Coverage | Native Utterances for Indian BFSI Voice Models: Currency (₹ lakh/crore), dates, times, OTPs, phone numbers, IFSC, UPI IDs, PAN, PIN codes |
| Licensing | Worldwide & Perpetual for Commercial Use and Derivative Voice Generation |
Voice Sourcing – not Model Training – was the Critical Path
Financial voice agents operate in a narrow acoustic and linguistic regime. Speech must be studio-clean, normalisation-consistent, and delivered in voices that users trust. That requirement collapses the available speaker pool. Each selected speaker had to satisfy three constraints simultaneously:
Native multilingual fluency with natural code-switching Financial conversations in India routinely shift between Hindi and English, or Tamil and English, at sub-sentence granularity. Script fidelity required speakers capable of switching languages without acoustic or prosodic discontinuity.
Consistent effect over long-duration scripted capture: Financial interactions demand warmth and clarity - OTP confirmations, grievance responses, payment reminders, and account notifications - delivered consistently across tens of recording hours.
Explicit consent for derivative voice generation Unlike ASR, TTS training produces reusable voice identities. This required informed licensing agreements covering synthesis, deployment, and derivative generation - constraints not addressed in standard voice recording workflows.
The corpus had to reflect India’s true linguistic topology - not isolated monolingual speech, but structured code-switching across regional language families:
North India: Hindi, Punjabi, Assamese, Odia - Hindi-anchored switching
West India: Marathi, Gujarati, Bengali - Hindi and English integration
South India: Tamil, Telugu, Kannada, Malayalam - English-integrated bilingual usage
Capturing this distribution was essential for downstream model generalisation.
How Humyn Sourced Deployment-Ready TTS Corpus Across 12 Languages
The completed corpus delivered 1,000+ hours of studio-quality speech optimised for production TTS training. Key Properties of the porject includes:
- Acoustic fidelity: Studio-controlled, 48 kHz capture suitable for high-quality neural synthesis
- Normalisation determinism: All numeric, currency, and identifier forms resolved pre-recording
- Code-switching realism: Structured bilingual variants reflecting production linguistic distributions
- Speaker representativeness: Demographically and linguistically aligned with BFSI user populations
- Consent-safe licensing: Full derivative voice generation rights, worldwide and perpetual
Domain balancing ensured production alignment: BFSI interactions were prioritised to maximize normalization coverage, while adjacent conversational domains ensured prosodic and syntactic.
The dataset can move directly into training pipelines without additional normalisation passes, acoustic filtering, or licensing remediation.