> For clean Markdown of any page, append `.md` to the page URL. > For a complete documentation index, see https://docs.sarvam.ai/llms.txt. > For full documentation content in one file, see https://docs.sarvam.ai/llms-full.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.sarvam.ai/_mcp/server. # Speech to Text ## Docs - [Speech-to-Text APIs](https://docs.sarvam.ai/api/api-guides-tutorials/speech-to-text/overview.md): Complete overview of Sarvam AI Speech-to-Text APIs including real-time, batch, and streaming options. Process audio with the Saaras model for high-accuracy transcription. - [Which Speech-to-Text API to Use](https://docs.sarvam.ai/api/api-guides-tutorials/speech-to-text/which-api-to-use.md): Compare Sarvam's Speech-to-Text APIs, REST, WebSocket, and Batch (plus their Speech-to-Text-Translate variants), and pick the right one for your audio length, latency, and feature needs. - [Speech-to-Text Rest API](https://docs.sarvam.ai/api/api-guides-tutorials/speech-to-text/rest-api.md): Process short audio files synchronously with immediate response. Instant transcription and translation for quick audio processing with multiple format support. - [Batch Speech-to-Text API](https://docs.sarvam.ai/api/api-guides-tutorials/speech-to-text/batch-api.md): Process large audio files using synchronous or asynchronous methods. Handle up to 2-hour recordings with speaker diarization, timestamps, and advanced transcription features. - [Realtime Streaming Speech-to-Text API](https://docs.sarvam.ai/api/api-guides-tutorials/speech-to-text/realtime-streaming.md): True partial transcripts, live mid-stream reconfiguration, and millisecond-based VAD tuning with saaras:v3-realtime and saaras:v4, Sarvam's next-generation WebSocket streaming API for voice agents and live transcription. - [Best Practices for Speech-to-Text](https://docs.sarvam.ai/api/api-guides-tutorials/speech-to-text/best-practices.md): Connection handling, VAD tuning, and production recommendations for Sarvam AI's Speech-to-Text API. - [FAQs](https://docs.sarvam.ai/api/speech-to-text/faq.md): Frequently asked questions about Sarvam AI speech-to-text services. Get answers about models, pricing, language support, audio formats, and implementation best practices. - [How to select output mode](https://docs.sarvam.ai/api/api-guides-tutorials/speech-to-text/how-to/select-output-mode.md): Choose the right output mode for your speech-to-text use case with Saaras v3 and v4. - [How to specify language codes](https://docs.sarvam.ai/api/api-guides-tutorials/speech-to-text/how-to/specify-language-codes.md): Use BCP-47 language codes for accurate speech-to-text transcription with Saaras v3 and v4. - [How to enable speaker diarization](https://docs.sarvam.ai/api/api-guides-tutorials/speech-to-text/how-to/enable-speaker-diarization.md): Identify and distinguish between multiple speakers in audio using the Batch API. - [Keyterm Prompting](https://docs.sarvam.ai/api/api-guides-tutorials/speech-to-text/how-to/keyterms.md): Use keyterms to help Saaras v4 recognize important names, places, brands, and technical terms. - [IVR & Contact Center Voice Agents](https://docs.sarvam.ai/api/api-guides-tutorials/speech-to-text/use-cases/ivr-contact-center.md): Architecture, model configuration, and latency targets for building telephony-based IVR and contact-center voice agents with Sarvam AI. - [BFSI Voice Bots](https://docs.sarvam.ai/api/api-guides-tutorials/speech-to-text/use-cases/bfsi-voice-bots.md): Architecture, model configuration, and compliance guardrails for building banking, financial services, and insurance voice agents with Sarvam AI. - [EdTech Voice Tutors](https://docs.sarvam.ai/api/api-guides-tutorials/speech-to-text/use-cases/edtech.md): Architecture, model configuration, and pitfalls for building voice-based tutoring agents with Sarvam AI. - [Government Services Voice Agents](https://docs.sarvam.ai/api/api-guides-tutorials/speech-to-text/use-cases/government-services.md): Architecture, model configuration, and pitfalls for building citizen-facing government scheme and services voice agents with Sarvam AI. - [Agri & Rural Voice Agents](https://docs.sarvam.ai/api/api-guides-tutorials/speech-to-text/use-cases/agri-rural.md): Architecture, model configuration, and pitfalls for building farmer-advisory and rural voice agents with Sarvam AI, tuned for low-bandwidth, feature-phone-first users. - [Streaming Speech-to-Text API](https://docs.sarvam.ai/api/api-guides-tutorials/speech-to-text/streaming-api.md): Real-time audio transcription and translation with WebSocket connections. Low-latency streaming for live applications with instant results and interactive features.