
Speech Recognition for BFSI: The Accuracy, Compliance, and Language Requirements That Actually Matter
Summary:
A collections agent at a mid-sized NBFC in Pune reads out a loan repayment disclosure to a customer in Marathi, switches to English for the repayment schedule, and confirms the next EMI date in Hindi. The call lasts six minutes. By regulatory requirement, every word of that interaction needs to be accurately recorded.
The STT system transcribing the call gets the Marathi segment mostly right. It drops two English terms from the repayment schedule. It misrecognises the EMI date. The transcript goes into the CRM. The compliance team reviews a sample and flags nothing because the errors look like minor transcription noise.
Three months later, a customer dispute lands on the desk of the grievance officer. The call record does not accurately reflect what the agent said about the repayment amount. The NBFC has no clean audit trail.
This is not a technology failure. It is a requirements failure. The BFSI enterprise chose an STT system built for general transcription and assumed it would meet compliance-grade requirements. The two are not the same product.
Why BFSI Is the Hardest Environment for Indian STT
Speech recognition for banking and financial services in India sits at the intersection of three overlapping requirements that most generic STT systems are not designed to handle simultaneously.
The first is accuracy on financial vocabulary. Amounts, policy numbers, loan identifiers, product names, and regulatory acronyms like NPA, EMI, DPD, KYC, and OTP must be transcribed correctly. These terms carry legal and financial weight. A substitution error that changes "teen hazaar" to "tees hazaar" or misrecognises "thirty days past due" is not an accuracy statistic. It is a data integrity failure in a regulated record.
The second is multilingual coverage across real call conditions. Indian BFSI contact centers serve customers across 12 major languages. Agents code-switch between Hindi and English, or Tamil and English, within single sentences. The audio comes in on 8kHz telephony from mobile connections in tier-2 and tier-3 cities. This is the hardest audio profile for any STT system, and it is the default condition in Indian BFSI voice operations.
The third is regulatory compliance. RBI guidelines, IRDAI requirements for insurance, DPDP Act obligations for personal data, and TRAI rules for outbound telephony all create specific obligations around how call audio and transcripts are handled, stored, and protected. An STT system that delivers acceptable accuracy but fails on data residency or PII handling is not compliant regardless of its WER numbers.
These three requirements must be evaluated together, not in sequence.
The WER Threshold That BFSI Actually Needs
Word error rate is the primary accuracy metric for STT evaluation, but the acceptable WER threshold for BFSI is lower than for most other enterprise use cases. Understanding why requires thinking about what the transcript is used for.
In a general contact center analytics deployment, a 12% WER means one in eight words is wrong. If errors distribute randomly across filler words and content words, the transcript is still useful for topic classification, sentiment analysis, and agent performance review. High-stakes words will usually be transcribed correctly enough for the analytics to function.
In a BFSI compliance deployment, random error distribution does not hold. Errors cluster on the vocabulary that matters most: amounts, dates, product identifiers, and consent confirmations. These are precisely the terms that appear infrequently in training data and are therefore most likely to be out-of-vocabulary for a general-purpose model. A system with 10% overall WER that systematically misrecognises financial terminology is not a 10% WER system for compliance purposes. It is a system that fails on the words the regulator cares about.
The practical WER threshold for compliance-grade BFSI transcription in India is below 8% on production telephony audio for primary languages, with domain-specific accuracy on financial vocabulary as a separate, stricter evaluation criterion.
For a complete methodology on how to calculate WER and test domain-specific accuracy before committing to a vendor, see How to Benchmark a Speech-to-Text API on Indian Languages Before You Sign Anything.
The Multilingual Challenge in BFSI Contact Centers
Indian BFSI operations span a language distribution that no single monolingual STT system can cover. A bank with national operations handles customer interactions in Hindi, Tamil, Telugu, Malayalam, Kannada, Marathi, Gujarati, Bengali, Punjabi, Odia, Assamese, and English, often within the same contact center.
The language challenge has three dimensions that are specific to BFSI.
Code-switching on financial vocabulary. BFSI agents code-switch more than agents in most other sectors because financial and regulatory terminology in India is predominantly English, even when the conversation is in a regional language. A Tamil-speaking agent explaining a policy exclusion will use Tamil grammatical structure with English terminology: "Aapka claim process aagidu, but pre-existing condition la oru waiting period iruku." The STT system must handle Tamil phonology, English vocabulary, and the specific borrowing patterns of Tamil-English financial speech simultaneously.
For a full treatment of why this breaks most STT systems and what proper code-switching support looks like architecturally, see Why Hinglish Breaks Most STT APIs: The Code-Switching Problem in Indian Voice AI.
Rural and semi-urban audio from field agents. Collections and microfinance operations generate a significant proportion of calls from field agents in rural locations, calling on mobile networks with variable quality. These calls carry 8kHz audio with background noise, GSM codec distortion, and dialectal variation from speakers whose first language may be Bhojpuri, Awadhi, or another regional variant rather than standard Hindi.
WER on this audio profile is substantially higher than on clean urban contact center audio. A system that shows acceptable WER on Kathbath Clean benchmarks may degrade to 30%+ WER on rural collections calls from Bihar or UP. For a detailed analysis of this degradation and the benchmark data that shows it, see our latest blogs.
Diarization for multi-party calls. BFSI calls often involve three-party interactions: agent, customer, and occasionally a supervisor or verification system. Compliance review requires knowing who said what, not just what was said. Speaker diarization, the separation of transcripts by speaker, is not optional in a compliance context. It is a prerequisite for the transcript to be a usable regulatory record. For more on diarization as an evaluation criterion, see Speaker Diarization, Confidence Scores, Latency: The STT Features Enterprises Overlook Until It's Too Late.
Gnani STT API supports all 12 Indian languages relevant to BFSI operations: Hindi, Tamil, Telugu, Malayalam, Kannada, Odia, Marathi, Punjabi, Gujarati, Bengali, Assamese, and English, with training on production telephony audio from Indian financial services deployments.
RBI and Regulatory Compliance: What STT Deployments Must Address
The regulatory landscape for voice AI in Indian BFSI is not static, and compliance requirements around call recording and transcription are becoming more specific. Here are the four regulatory dimensions that any STT deployment in Indian BFSI must address before going live.
RBI guidelines on call recording. The Reserve Bank of India requires regulated entities to maintain records of customer interactions for specified retention periods. For banks and NBFCs, this means call recordings and, increasingly, transcripts must be stored in a tamper-evident format and be accessible for audit. The STT system's output becomes part of the regulated record. Accuracy requirements are therefore not purely operational. They are compliance requirements.
DPDP Act obligations. India's Digital Personal Data Protection Act classifies call recordings and transcripts containing personal financial information as personal data. Voice recordings may also contain biometric data in the form of voiceprints, which are classified as sensitive personal data requiring explicit consent. Any STT deployment that processes call audio from Indian customers must have a consent framework, data minimisation practices, and defined retention and deletion policies for both audio and transcript data. Vendors who store audio or transcripts on offshore infrastructure create additional data residency complexity under DPDP.
IRDAI requirements for insurance. Insurance contact centers face specific IRDAI requirements around disclosure of policy terms, exclusions, and claim procedures. These disclosures must be made in a language the customer understands and must be accurately recorded. STT accuracy on disclosure language, which often involves formal vocabulary in regional languages, is a compliance-grade requirement, not a nice-to-have.
TRAI rules for outbound calls. Outbound voice AI deployments in collections and lead nurturing must comply with TRAI's regulations on automated calls, including disclosure that the caller is an automated system, DND (Do Not Disturb) list compliance, and call timing restrictions. The voice bot's STT layer must accurately capture customer responses, including consent expressions and opt-out requests, because these responses may be needed as evidence of compliance.
Two things worth noting on regulation: these requirements evolve, and this post does not constitute legal or compliance advice. Production deployment decisions should involve your Legal and Compliance teams alongside the technical evaluation.
Data Residency and On-Premises Deployment
For many Indian BFSI enterprises, sending call audio to a third-party cloud infrastructure is not acceptable regardless of STT accuracy. Data residency requirements, internal security policies, and RBI guidance on cloud usage for regulated entities collectively push a significant proportion of BFSI STT deployments toward on-premises or private cloud architectures.
On-premises STT deployment means the transcription model runs on the enterprise's own infrastructure. Audio never leaves the perimeter. Transcripts are generated and stored within the regulated environment. This eliminates data sovereignty concerns but introduces operational complexity: the enterprise takes on model hosting, scaling, and maintenance.
For BFSI buyers evaluating STT vendors, the relevant questions are: does the vendor support on-premises deployment? What are the hardware requirements? How are model updates handled? Can custom vocabulary and fine-tuning be applied to the on-premises model? What SLA applies to on-premises deployments versus cloud API?
Not every STT vendor supports on-premises deployment. Of those that do, not all support the same level of customisation in an on-premises configuration that they offer in their cloud API. This is a non-negotiable evaluation criterion for many BFSI buyers and should be established before investing time in accuracy evaluation.
PII Redaction and Transcript Security
Call transcripts in BFSI contain some of the most sensitive personal data in any enterprise environment: customer names, Aadhaar numbers, PAN numbers, account numbers, loan identifiers, policy numbers, and financial amounts. In collections calls, they also contain information about a customer's financial distress and payment behaviour.
PII redaction, the automatic identification and masking of personally identifiable information in transcripts before storage or downstream processing, is a feature requirement for any BFSI STT deployment. But PII redaction accuracy is a separate evaluation criterion from transcription accuracy, and the two do not always correlate. A vendor with strong WER numbers may have weak PII redaction coverage that misses account numbers formatted in Indian conventions or Aadhaar numbers spoken in Hindi.
When evaluating PII redaction, test with Indian-format identifiers: twelve-digit Aadhaar numbers spoken digit by digit or in pairs, PAN formats like "ABCDE1234F," mobile numbers in ten-digit format, and loan account numbers in your specific CRM format. Do not rely on demo data that uses clearly formatted English-language identifiers.
Voiceprint data, which is generated when STT systems use speaker identification, is classified as biometric sensitive personal data under the DPDP Act. If your deployment involves speaker identification or verification rather than just diarization, your consent framework must explicitly cover voiceprint collection and storage.
Real-Time vs Batch: Which Mode for BFSI
The choice between real-time streaming transcription and batch transcription maps differently to BFSI use cases than to general contact center deployments.
Batch transcription is sufficient and preferable for the majority of BFSI STT use cases: post-call quality monitoring, compliance review, dispute resolution, agent coaching, and speech analytics. Batch produces higher accuracy than streaming on equivalent audio, which matters more in BFSI than in sectors where compliance-grade transcripts are not required.
Real-time transcription is necessary for live agent assist, real-time compliance alerts (for example, flagging when a mandatory disclosure has not been made mid-call), and voice bot deployments where the bot is conducting the conversation. For these use cases, P95 latency under 200ms is the production-ready threshold. Gnani STT API achieves P95 latency under 200ms in streaming mode.
For a detailed breakdown of which contact center use cases require each mode and the accuracy trade-offs involved, see Real-Time vs Batch Transcription: Which STT Mode Does Your Contact Center Actually Need.
What to Validate Before Going Live in BFSI
A BFSI STT evaluation cannot be compressed into a demo and a benchmark sheet. Here is the minimum set of validations before a production commitment.
Accuracy on your audio, not vendor benchmarks. Run a benchmark evaluation on your own production call recordings using ground truth transcripts annotated by native speakers of each language in scope. Calculate WER per language, per audio condition. Test specifically on financial vocabulary. For the complete methodology, see How to Benchmark a Speech-to-Text API on Indian Languages Before You Sign Anything.
Code-switching accuracy. Test specifically on calls where agents switch between the regional language and English within sentences. This is the default communication style in most Indian BFSI contact centers and the condition most likely to produce accuracy failures in general-purpose STT systems.
Diarization accuracy on two and three-party calls. Run the vendor's diarization on a sample of your actual call recordings and manually verify speaker attribution accuracy. Do not accept diarization accuracy claims based on clean two-party test sets if your production calls include three-party interactions or overlapping speech.
PII redaction coverage on Indian-format identifiers. Test with real Indian PII formats: Aadhaar, PAN, account numbers, mobile numbers in your specific call formats. Verify that redaction is applied correctly before transcripts reach any downstream system.
Data residency and deployment model confirmation. Confirm in writing where audio is processed and stored, and whether on-premises deployment is available if required by your security or compliance policy.
Regulatory alignment review. Before go-live, have your Legal and Compliance teams review the deployment architecture, data handling practices, consent framework, and retention policies against current RBI, IRDAI, DPDP, and TRAI requirements applicable to your specific entity type and use cases.
Frequently Asked Questions
- What WER is acceptable for BFSI compliance transcription? Below 8% WER on production telephony audio for primary languages is the target for compliance-grade transcription. More importantly, domain-specific accuracy on financial vocabulary, amounts, identifiers, regulatory acronyms, and consent language, must be evaluated separately from overall WER. A system can show 8% overall WER while systematically failing on the terms that matter most.
- Does speech recognition for banking need to support all Indian languages? It depends on your geographic footprint. National BFSI operations typically need Hindi, English, and the four major Dravidian languages at minimum. Regional operations may need only two or three languages. The relevant requirement is coverage for your actual customer base, not all Indian languages as an abstract standard. Gnani STT API supports 12 Indian languages: Hindi, Tamil, Telugu, Malayalam, Kannada, Odia, Marathi, Punjabi, Gujarati, Bengali, Assamese, and English.
- Is cloud-based STT compliant with RBI guidelines? Cloud deployment for regulated financial data is permissible under RBI guidelines subject to specific conditions around data localisation, vendor due diligence, and audit rights. Many BFSI enterprises prefer private cloud or on-premises deployment for call audio and transcripts due to the sensitivity of the data. Confirm the specific requirements applicable to your entity type with your Legal and Compliance teams.
- How does the DPDP Act affect call transcription in banking? Call transcripts containing customer personal data are classified as personal data under DPDP. This requires a lawful basis for processing, typically legitimate interest for compliance purposes or explicit consent. Voiceprints generated by speaker identification features are sensitive personal data requiring explicit consent. Retention periods must be defined and enforced. Cross-border transfer of personal data is restricted under DPDP and may prohibit use of offshore cloud infrastructure for processing Indian customer call data.
- What is the difference between diarization and speaker identification in BFSI? Diarization segments a transcript by speaker, answering "who spoke when" without identifying who the speakers are. Speaker identification matches speaker segments to known identities. For compliance transcription, diarization is typically sufficient: the record shows agent speech versus customer speech. Speaker identification using voiceprints is a separate, more sensitive capability that requires explicit consent and biometric data handling compliance under DPDP.


