REST Stream
Converts the input text into a streamed spoken audio response using a cloned voice.
Step 1: Create a voice with POST /voices/create and save the returned voice_id.
Step 2: Pass that voice_id with text and language_code to this endpoint. The response is a binary audio stream (not JSON), so you can begin playback as soon as the first sentence is ready.
Accepts multipart/form-data.
Base URL: https://api.sarvam.ai.
Auth: send your key in the api-subscription-key header (recommended); Authorization: Bearer <key> is also accepted.
Streaming codecs: mp3 (default), wav, linear16, mulaw, alaw. flac, aac, and opus are not supported on this endpoint because each sentence would be a complete container that cannot be concatenated.
Quality checks (QC/VAD) and duration bounds (min_audio_duration / max_audio_duration) are not accepted — the stream path optimizes for time-to-first-byte.
Billing: charged per character of text against the text_to_speech_voice_cloning API. See the pricing page.
Authentication
Request
Required cloned voice ID (svc-{uuid}). Create one with POST /voices/create, or use a voice from the same organization's Content Studio Voice Library.
Text to synthesize in the cloned voice. Native scripts and code-mixed text are supported. Maximum 1000 characters; longer text is rejected with a 400. Split longer input client-side and synthesize it in parts.
BCP-47 language code for the synthesized output. The cloned voice can speak any supported language, regardless of the reference clip's language (cross-lingual cloning).
Available Options:
as-IN: Assamesebn-IN: Bengalien-IN: English (Indian)gu-IN: Gujaratihi-IN: Hindikn-IN: Kannadaml-IN: Malayalammr-IN: Marathiod-IN: Odiapa-IN: Punjabita-IN: Tamilte-IN: Telugu
Speech speed multiplier. 1.0 is natural speed; 0.5 to 2.0 is accepted. Speed is adjusted without changing pitch.
Audio codec for the streamed output. Default is mp3.
mp3: Self-framing; clients can decode and play as chunks arrive. Recommended.wav: One RIFF header (unknown length), then raw PCM chunks.linear16: Raw 16-bit PCM (no header).mulaw/alaw: 8-bit telephony codecs for IVR pipelines.
Bitrate for lossy codecs (mp3). Default is 128k. Options: 32k, 64k, 96k, 128k, 192k.
Enable pre-TTS text normalization: numbers, currencies and dates are converted to their spoken form before synthesis (e.g. 9840950950 becomes the spoken digit sequence). Defaults to the server-side setting.
Response
Success. Returns a streamed audio response in the requested format (e.g., audio/mpeg for MP3, audio/wav for WAV).