> For clean Markdown of any page, append `.md` to the page URL. > For a complete documentation index, see https://docs.sarvam.ai/llms.txt. > For full documentation content in one file, see https://docs.sarvam.ai/llms-full.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.sarvam.ai/_mcp/server. # Text to Speech ## Docs - [Text-to-Speech Overview](https://docs.sarvam.ai/api/api-guides-tutorials/text-to-speech/overview.md): Complete overview of Sarvam AI Text-to-Speech APIs using Bulbul v4 Flash (latest) and Bulbul v3. Convert text to natural speech with real-time and streaming options for Indian languages. - [Voices](https://docs.sarvam.ai/api/api-guides-tutorials/text-to-speech/voices.md): Listen to every Sarvam AI text-to-speech voice before you choose. Bulbul v4 Flash is the latest model; this page lists the v4 catalog plus Bulbul v3 audio previews. - [Which Text-to-Speech API to Use](https://docs.sarvam.ai/api/api-guides-tutorials/text-to-speech/which-api-to-use.md): Compare Sarvam's Text-to-Speech APIs, REST, HTTP streaming, and WebSocket, and pick the right one for your latency, text length, and interactivity needs. - [Text-to-Speech Rest API](https://docs.sarvam.ai/api/api-guides-tutorials/text-to-speech/rest-api.md): Real-time conversion of text into speech using customizable voices. Instant audio generation with multiple voice options and various audio formats for Indian languages. - [HTTP Streaming API](https://docs.sarvam.ai/api/api-guides-tutorials/text-to-speech/streaming-api/http-stream.md): Stream TTS audio over a single HTTP POST request. No WebSocket setup, no connection management, just POST text and pipe the audio response. - [Streaming Text-to-Speech API](https://docs.sarvam.ai/api/api-guides-tutorials/text-to-speech/streaming-api/web-socket.md): Real-time conversion of text into speech using WebSocket connections. Efficient streaming for long texts with progressive audio generation and low-latency playback. - [Pronunciation Dictionary](https://docs.sarvam.ai/api/api-guides-tutorials/text-to-speech/pronunciation-dictionary.md): Teach Bulbul v3 how to say specific words, brand names, abbreviations, regional terms, exactly the way you want, across all 11 supported languages. - [Best Practice Guide for Bulbul v4 Flash](https://docs.sarvam.ai/api/api-guides-tutorials/text-to-speech/best-practice-guide-for-bulbul-v-4-flash.md): Integration guide for Sarvam Bulbul v4 Flash — endpoints, parameters, input formatting, and expert voice picks by use case. - [Best Practices for Bulbul v3](https://docs.sarvam.ai/api/api-guides-tutorials/text-to-speech/bulbul-v3-best-practices.md): Speaker selection, API mode, pace, and temperature tuning for the stable Bulbul v3 TTS model (short-name speakers). - [How to Set the Language](https://docs.sarvam.ai/api/api-guides-tutorials/text-to-speech/how-to/set-the-language.md): Defines the language for text normalization before speech synthesis. - [How to change the speaker voice](https://docs.sarvam.ai/api/api-guides-tutorials/text-to-speech/how-to/change-the-speaker-voice.md): Learn how to choose specific voices for text-to-speech output using the speaker parameter. Explore Bulbul v3's 30+ natural-sounding voices for different languages and use cases. - [How to adjust the pitch (tone)](https://docs.sarvam.ai/api/api-guides-tutorials/text-to-speech/how-to/adjust-the-tone.md): Control the tone of the synthesized speech (bulbul:v2 only). - [How to adjust the pace (speed)](https://docs.sarvam.ai/api/api-guides-tutorials/text-to-speech/how-to/adjust-the-speed.md): Controls the speed at which the speech is delivered. - [How to adjust the loudness (volume)](https://docs.sarvam.ai/api/api-guides-tutorials/text-to-speech/how-to/adjust-the-loudness.md): Controls the volume level of the generated audio (bulbul:v2 only). - [How to set the sample rate (audio quality)](https://docs.sarvam.ai/api/api-guides-tutorials/text-to-speech/how-to/set-the-sample-rate.md): Controls the audio quality and size of the generated output. - [How to enable text preprocessing](https://docs.sarvam.ai/api/api-guides-tutorials/text-to-speech/how-to/enable-text-preprocessing.md): improves pronunciation (bulbul:v2 only). - [How to set the audio format for output using `output_audio_codec`](https://docs.sarvam.ai/api/api-guides-tutorials/text-to-speech/how-to/set-audio-format-for-output.md): Choose the audio format for TTS streaming output. - [How to set `output_audio_bitrate`](https://docs.sarvam.ai/api/api-guides-tutorials/text-to-speech/how-to/set-bitrate-for-output.md): Control the quality and size of the synthesized audio output. - [How to set maximum length for sentence splitting using `max_chunk_length`](https://docs.sarvam.ai/api/api-guides-tutorials/text-to-speech/how-to/set-maximum-length-for-sentence-splitting.md): Control how long each sentence chunk can be when splitting text for streaming TTS. - [How to set buffer size to start processing in Streaming TTS with `min_buffer_size`](https://docs.sarvam.ai/api/api-guides-tutorials/text-to-speech/how-to/set-buffer-size-to-start-processing.md): Define when the TTS engine should start processing text from the buffer. - [IVR & Contact Center Voice Agents](https://docs.sarvam.ai/api/api-guides-tutorials/speech-to-text/use-cases/ivr-contact-center.md): Architecture, model configuration, and latency targets for building telephony-based IVR and contact-center voice agents with Sarvam AI. - [BFSI Voice Bots](https://docs.sarvam.ai/api/api-guides-tutorials/speech-to-text/use-cases/bfsi-voice-bots.md): Architecture, model configuration, and compliance guardrails for building banking, financial services, and insurance voice agents with Sarvam AI. - [EdTech Voice Tutors](https://docs.sarvam.ai/api/api-guides-tutorials/speech-to-text/use-cases/edtech.md): Architecture, model configuration, and pitfalls for building voice-based tutoring agents with Sarvam AI. - [Government Services Voice Agents](https://docs.sarvam.ai/api/api-guides-tutorials/speech-to-text/use-cases/government-services.md): Architecture, model configuration, and pitfalls for building citizen-facing government scheme and services voice agents with Sarvam AI. - [Agri & Rural Voice Agents](https://docs.sarvam.ai/api/api-guides-tutorials/speech-to-text/use-cases/agri-rural.md): Architecture, model configuration, and pitfalls for building farmer-advisory and rural voice agents with Sarvam AI, tuned for low-bandwidth, feature-phone-first users.