How to Write an STT API RFP: The Questions Your Speech Recognition Vendor Doesn't Want You to Ask

August 12, 2026
14
mins read
Chris Wilson
Content Creator

Summarise with

Be Updated
Get weekly update from Gnani
Thank You! Your submission has been received.
Oops! Something went wrong while submitting the form.

Most enterprise STT RFPs are written by people who have never deployed speech recognition at scale. They are assembled from procurement templates, vendor-supplied capability matrices, and a feature checklist that describes what the business wants the system to do without specifying what it needs the system to handle.

The result is a document that vendors can respond to with polished answers and impressive benchmark numbers, none of which predict production performance on Indian contact center audio. The vendor wins the RFP. The deployment underperforms. The procurement team learns what the right questions were six months after they needed to ask them.

This guide is those questions. It is structured as an RFP framework you can adapt directly, with the reasoning behind each question so you understand what a strong answer looks like versus a deflection.

Why Standard STT RFPs Fail Indian Enterprise Buyers

A standard enterprise software RFP asks vendors to describe their capabilities, list supported languages, provide reference customers, and submit pricing. For most software categories, this works reasonably well. For Indian language STT, it fails in a specific way.

The gap between "we support Hindi" and "we handle Hindi-English code-switched telephony audio from rural UP at 8kHz with GSM codec compression" is not a marketing distinction. It is a 15 to 20 percentage point WER difference in production. A vendor that checks the Hindi box on a capability matrix is telling you nothing useful about whether their system will work on your audio.

The questions in this guide are designed to force vendors out of capability matrix mode and into production evidence mode. Every question has a correct type of answer — specific, measurable, verifiable — and a deflection answer that sounds acceptable but provides no real information. Learning to distinguish between the two is the skill this guide is trying to give you.

For a full methodology on how to independently verify vendor claims through your own benchmark evaluation, see How to Benchmark a Speech-to-Text API on Indian Languages Before You Sign Anything.

Speech to text 2,000 free STT minutes. Gnani Prisma v2.5 · Trained on 1M hours of audio · Best in class multi-lingual, multi-speaker Start building

Section 1: Accuracy and Benchmarks

These are the questions that should open every STT RFP. They establish whether the vendor can provide evidence, not just claims.

Q1. What is your Word Error Rate on Hindi Kathbath Clean and Hindi Gramvaani, measured using Standard WER?

Why this question matters: Kathbath Clean is the standard controlled benchmark for Indian language STT. Gramvaani is rural Hindi telephony audio that represents real Indian contact center conditions. Asking for both forces the vendor to show you the degradation gap between clean and noisy conditions. A vendor who only provides Kathbath numbers has not tested on real telephony audio. A vendor who provides both allows you to calculate the degradation delta, which is a better predictor of production performance than either number alone.

Strong answer: specific WER numbers for both datasets with methodology described (Standard WER, same test split, same audio preprocessing).

Deflection: "Our system achieves industry-leading accuracy on Indian languages." No numbers, no dataset, no methodology.

For Standard WER context: on Kathbath Clean (Hindi), Gnani STT API achieves 7.6% WER. On Gramvaani (rural Hindi telephony), Gnani STT API achieves 22.0% WER — the lowest degradation among providers tested on equivalent audio.

Q2. What is your WER specifically on Hindi-English code-mixed telephony audio in noisy conditions?

Why this question matters: code-switching between Hindi and English within sentences is the default communication style in most Indian BFSI and telecom contact centers. WER on monolingual Hindi and WER on code-mixed Hindi-English are different numbers, measured on different test conditions. A vendor that has not specifically tested for code-switching will not have an answer to this question.

Strong answer: a specific WER number on intra-sentential code-mixed audio, with the test set described.

Deflection: "We support both Hindi and English." This is not an answer to the question.

For reference: on Hindi-English code-mixed telephony audio in noisy conditions, Gnani STT API achieves 9% WER. Global cloud providers typically show 14 to 16% WER on the same audio. For a full analysis of why code-switching is architecturally hard, see Why Hinglish Breaks Most STT APIs: The Code-Switching Problem in Indian Voice AI.

Q3. Provide WER numbers for each of the following languages on Kathbath Clean and Kathbath Noisy: Tamil, Telugu, Kannada, Malayalam, Marathi, Gujarati, Bengali, Punjabi, Odia, Assamese.

Why this question matters: vendors routinely claim support for 20+ Indian languages while having meaningful training data for only three or four. Asking for WER on specific datasets for each language forces the vendor to demonstrate actual performance rather than list coverage. Languages like Odia, Assamese, and Punjabi are where the gaps between claimed and actual performance are largest.

Strong answer: a WER table with numbers for each language on both datasets.

Deflection: "We support all major Indian languages." A support claim without accuracy evidence is not procurement-grade information.

Q4. What datasets were used to train your acoustic model for Indian languages, and what proportion of your training data is telephony audio versus clean studio or read speech?

Why this question matters: training data composition is the primary predictor of production performance on Indian telephony audio. A model trained predominantly on clean studio recordings will degrade sharply on 8kHz GSM telephony. A vendor that cannot answer this question in specific terms either does not know what their model was trained on or knows and does not want to tell you.

Strong answer: specific description of training data sources, audio quality distribution, and how telephony audio is represented.

Deflection: "We train on diverse, high-quality data across Indian languages."

Section 2: Audio Environment and Noise Robustness

Q5. What is your WER on 8kHz audio versus 16kHz audio for Hindi and your other supported Indian languages?

Why this question matters: most Indian contact center calls arrive at 8kHz. Most published benchmarks use 16kHz or higher audio. The WER gap between these two sample rates is typically 5 to 15 percentage points. A vendor who cannot provide 8kHz-specific WER numbers has not tested on the audio format your calls will actually arrive in.

Strong answer: WER numbers for both sample rates on equivalent content.

Deflection: "We support 8kHz audio input." Support is not accuracy. These are different claims.

Q6. What is your degradation delta — the WER increase from clean to noisy audio — for Hindi at 15dB SNR and below?

Why this question matters: the degradation curve, not clean WER alone, determines how the system performs across the full range of audio quality in your operation. A system with good clean WER but sharp degradation will fail on exactly the calls that matter most: rural collections, field agent calls, mobile-heavy outbound campaigns. For a full treatment of how noise degrades STT accuracy on Indian telephony audio, see Speech Recognition in Noisy and Rural India: Why Your STT Fails Where It Matters Most.

Strong answer: WER at multiple SNR thresholds with the degradation delta calculated.

Deflection: "Our system is noise-robust." This is a marketing claim, not a measurement.

Q7. Will you run your system on a test set we provide, drawn from our production call recordings, and provide WER results before contract signing?

Why this question matters: this is the single most important question in any STT RFP. A vendor confident in their production performance will agree to this. A vendor who declines, deflects, or insists on using only their own test data is telling you something critical about the confidence they have in their system on your audio.

Strong answer: yes, with a clear process for receiving test audio, protecting data, and returning results.

Deflection: any version of "our published benchmarks are representative" without agreeing to test on your audio.

Section 3: Language and Linguistic Capability

Q8. How does your system handle intra-sentential code-switching — specifically, a speaker who switches from Hindi to English mid-sentence without any acoustic boundary?

Why this question matters: inter-sentential code-switching (switching between sentences) and intra-sentential code-switching (switching within a sentence) require different architectural solutions. Most vendors handle the former adequately and struggle with the latter. The answer to this question reveals whether the vendor has actually engineered for real Indian speech or has added Hindi as an afterthought to an English-first architecture.

Strong answer: a description of frame-level or joint multilingual language identification, with performance evidence on intra-sentential test cases.

Deflection: "We support multilingual audio." Or: "We detect the language automatically." Neither addresses intra-sentential switching.

Q9. What dialectal and regional variations of Hindi are represented in your training data? Specifically: Bhojpuri-influenced, Rajasthani-influenced, and Awadhi-influenced speech.

Why this question matters: standard Hindi training data over-represents urban, educated Delhi and Mumbai speakers. Field collections calls from eastern UP, Bihar, and Rajasthan involve dialect features that models not trained on these varieties will mishandle. If your deployment covers these geographies, this question is not academic.

Strong answer: specific mention of regional variety coverage with evidence of training data from those regions.

Deflection: "We train on diverse Hindi data." Diversity without specificity is not an answer.

Q10. What is your mechanism for custom vocabulary injection, and how does it affect WER on domain-specific terms?

Why this question matters: product names, regulatory acronyms (NPA, EMI, DPD, KYC, OTP), and financial identifiers are typically out-of-vocabulary for general-purpose STT models. Custom vocabulary injection is the mechanism that corrects this. The relevant question is not whether the feature exists but whether it actually improves WER on your specific domain terms and by how much.

Strong answer: description of the injection mechanism (word list, phonetic spelling, weighted boosting), evidence of WER improvement on domain vocabulary with a test case.

Deflection: "We support custom vocabulary." Feature existence without accuracy impact is not sufficient.

Section 4: Diarization, Latency, and Production Metrics

Q11. What is your speaker diarization error rate on two-party Indian contact center calls, and how does it change on three-party calls?

Why this question matters: quality monitoring, compliance review, and agent coaching all depend on accurate speaker attribution. Diarization error rates on clean English two-party calls, which is what most vendors benchmark, do not predict performance on Indian contact center audio with code-switching, variable acoustic profiles between VoIP agent and mobile customer, and overlapping speech. For a full treatment of diarization as a production metric, see Speaker Diarization, Confidence Scores, and Latency: The STT Metrics Buyers Ignore Until It's Too Late.

Strong answer: a specific diarization error rate on Indian telephony audio, separately for two-party and three-party calls.

Deflection: "Our diarization supports multiple speakers." Not an answer to the question.

Q12. What is your P95 latency for streaming transcription under 500 concurrent connections?

Why this question matters: single-request latency tests are meaningless for production contact center deployments. P95 latency under realistic concurrency is the metric that determines whether live agent assist and real-time compliance monitoring actually work. Gnani STT API achieves P95 latency under 200ms in streaming mode. For a full breakdown of latency requirements by use case, see Real-Time vs Batch Transcription: Which STT Mode Does Your Contact Center Actually Need.

Strong answer: P95 latency measured under a specified concurrency level, with the test conditions described.

Deflection: average latency under a single request, or latency claims without concurrency specification.

Q13. Are your confidence scores calibrated? Provide WER by confidence score bin (above 0.9, 0.7 to 0.9, 0.5 to 0.7, below 0.5) on your Indian language test sets.

Why this question matters: uncalibrated confidence scores make human-in-the-loop routing workflows unreliable. A confidence score that does not predict accuracy cannot be used to route low-certainty transcripts for review. Most vendors do not proactively test confidence calibration. Asking for WER by confidence bin immediately reveals whether the scores convey real information.

Strong answer: a table of WER by confidence bin showing monotonically increasing WER as confidence decreases.

Deflection: "We provide confidence scores with every transcript." Existence of the feature, not evidence of its calibration.

Section 5: Data, Security, and Compliance

Q14. Where is audio data processed and stored? Is India-region or on-premises deployment available?

Why this question matters: for BFSI, insurance, and government deployments, data residency is not optional. RBI guidance on cloud usage, DPDP Act obligations for personal data, and internal security policies collectively restrict where call audio can be sent for processing. A vendor that processes audio only on offshore infrastructure may be disqualified before accuracy is even evaluated. For a full breakdown of regulatory requirements for STT in Indian BFSI, see Speech Recognition for BFSI: The Accuracy, Compliance, and Language Requirements That Matter.

Strong answer: specific confirmation of India-region cloud availability or on-premises deployment, with architecture documentation.

Deflection: "We are SOC 2 certified." Certification is not data residency.

Q15. What is your PII redaction capability, and can you demonstrate it on Indian-format identifiers: 12-digit Aadhaar numbers spoken digit by digit, PAN format, and 10-digit mobile numbers?

Why this question matters: standard PII redaction is trained on US and European identifier formats. Indian identifiers have different structures, different spoken formats, and different positional patterns in call audio. A vendor claiming PII redaction should demonstrate it on the actual formats your calls will contain.

Strong answer: a demo or test output showing correct redaction of Indian-format PII.

Deflection: "We support PII redaction for names, credit cards, and phone numbers." Without Indian-format evidence, this is not a verified capability for your deployment.

Q16. What is your SLA for API uptime, and what happens to in-flight transcription requests during maintenance windows?

Why this question matters: for contact centers running real-time transcription on live calls, API downtime is not a background event. It is a live operational failure. The SLA should specify uptime percentage, maintenance window notification, and the behaviour of in-flight requests. Batch deployments have more tolerance for planned downtime, but real-time deployments need both high uptime and graceful degradation during incidents.

Strong answer: specific uptime SLA (99.9% or higher), maintenance window policy, and defined failover behaviour.

Deflection: "We have enterprise-grade reliability."

Section 6: Commercial and Contractual

Q17. Is your pricing based on audio duration, API calls, or concurrent connections, and how does it scale at 30,000 calls per day?

Why this question matters: STT pricing models vary significantly and the economics shift at scale. A per-call model that looks reasonable at 5,000 calls per day may become uncompetitive at 30,000. Understanding the pricing model and running the numbers at your actual volume is a basic procurement requirement that many teams skip until after the contract is signed.

Strong answer: explicit pricing model with a worked example at your actual call volume.

Deflection: "Contact our sales team for enterprise pricing."

Q18. What does your model improvement cycle look like, and how are accuracy improvements deployed to existing customers?

Why this question matters: STT models improve over time, but improvements are only valuable if they reach your deployment. A vendor who regularly updates their models and automatically improves your transcription accuracy is a different commercial proposition than one whose model is static unless you renegotiate. For on-premises deployments, the update process is especially important.

Strong answer: description of model update cadence, how updates are communicated, and whether existing customers benefit automatically or must recontract.

Deflection: "We continuously improve our models."

How to Use These Questions

Not every question applies to every deployment. Run this framework through your use case filter first.

Post-call analytics and quality monitoring: questions 1, 2, 3, 4, 5, 6, 7, 10, 11, 13, 14, 15, 17 are mandatory.

Real-time agent assist and voice bots: add questions 12 and 16 as hard filters. A vendor that cannot meet P95 latency under 200ms at your concurrency level is disqualified regardless of accuracy.

BFSI and regulated deployments: questions 14 and 15 are non-negotiable. Data residency and PII redaction should be evaluated before accuracy, because a vendor that fails these requirements cannot be deployed regardless of WER.

Multilingual deployments across regional languages: questions 3, 9, and 10 carry the most weight. Accuracy gaps are widest for lower-resource languages like Odia, Assamese, and Punjabi, and vendors who cannot provide specific WER evidence for these languages have not meaningfully tested them.

Any vendor who declines to answer question 7 — running their system on your audio — should be removed from consideration before the rest of the evaluation proceeds. Production evidence on your audio is the only reliable signal. Everything else is a proxy.

Frequently Asked Questions

  1. What should an STT API RFP include for Indian language deployments? At minimum: WER on Kathbath Clean and Gramvaani by language, WER on code-mixed audio, degradation delta from clean to noisy conditions, diarization error rate on Indian telephony, P95 latency under realistic concurrency, confidence score calibration evidence, data residency confirmation, PII redaction demonstration on Indian-format identifiers, and an agreement to benchmark on your own production audio before contract signing.
  2. How do I know if a vendor's WER claims are reliable? Ask for the dataset name, the test split used, the WER calculation method (Standard WER or LLM WER), and whether the test set was provided by the vendor or a third party. Then request permission to run the same evaluation on your own audio. A vendor with reliable claims will agree to independent verification. A vendor who declines is relying on you not to check.
  3. What is the most important question to ask an STT vendor? Question 7: will you run your system on our production audio before contract signing? Every other question can be answered in writing with polished claims. This one requires production evidence. It is the question that most directly predicts whether the vendor's system will work in your environment.
  4. How long should an STT vendor evaluation take? Four to six weeks for a complete evaluation across three to four vendors. Building the test set and creating ground truth transcripts is the longest phase. Running the actual benchmark once the test set is ready takes two to three days per vendor. Do not compress the evaluation under commercial pressure. A six-week evaluation is cheaper than a six-month failed deployment.
  5. Does Gnani STT API support on-premises deployment for BFSI? Yes. Gnani STT API supports on-premises deployment for regulated environments where data residency requirements prevent audio from being processed on external cloud infrastructure. It covers all 12 supported languages: Hindi, Tamil, Telugu, Malayalam, Kannada, Odia, Marathi, Punjabi, Gujarati, Bengali, Assamese, and English.