Skip to navigation

REST Stream

View as Markdown

Converts the input text into a streamed spoken audio response using a cloned voice.

Step 1: Create a voice with POST /voices/create and save the returned voice_id.

Step 2: Pass that voice_id with text and language_code to this endpoint. The response is a binary audio stream (not JSON), so you can begin playback as soon as the first sentence is ready.

Accepts multipart/form-data.

Base URL: https://api.sarvam.ai. Auth: send your key in the api-subscription-key header (recommended); Authorization: Bearer <key> is also accepted.

Streaming codecs: mp3 (default), wav, linear16, mulaw, alaw. flac, aac, and opus are not supported on this endpoint because each sentence would be a complete container that cannot be concatenated.

Quality checks (QC/VAD) and duration bounds (min_audio_duration / max_audio_duration) are not accepted — the stream path optimizes for time-to-first-byte.

Billing: charged per character of text against the text_to_speech_voice_cloning API. See the pricing page.

Authentication

api-subscription-keystring
API Key authentication via header

Request

This endpoint expects a multipart form.
voice_idstringRequired

Required cloned voice ID (svc-{uuid}). Create one with POST /voices/create, or use a voice from the same organization's Content Studio Voice Library.

textstringRequired1-1000 characters

Text to synthesize in the cloned voice. Native scripts and code-mixed text are supported. Maximum 1000 characters; longer text is rejected with a 400. Split longer input client-side and synthesize it in parts.

language_codeenumRequired

BCP-47 language code for the synthesized output. The cloned voice can speak any supported language, regardless of the reference clip's language (cross-lingual cloning).

Available Options:

  • as-IN: Assamese
  • bn-IN: Bengali
  • en-IN: English (Indian)
  • gu-IN: Gujarati
  • hi-IN: Hindi
  • kn-IN: Kannada
  • ml-IN: Malayalam
  • mr-IN: Marathi
  • od-IN: Odia
  • pa-IN: Punjabi
  • ta-IN: Tamil
  • te-IN: Telugu
pacedouble or nullOptional0.5-2

Speech speed multiplier. 1.0 is natural speed; 0.5 to 2.0 is accepted. Speed is adjusted without changing pitch.

output_audio_codecenum or nullOptionalDefaults to mp3

Audio codec for the streamed output. Default is mp3.

  • mp3: Self-framing; clients can decode and play as chunks arrive. Recommended.
  • wav: One RIFF header (unknown length), then raw PCM chunks.
  • linear16: Raw 16-bit PCM (no header).
  • mulaw / alaw: 8-bit telephony codecs for IVR pipelines.
Allowed values:
output_audio_bitrateenum or nullOptionalDefaults to 128k

Bitrate for lossy codecs (mp3). Default is 128k. Options: 32k, 64k, 96k, 128k, 192k.

Allowed values:
speech_sample_rateenum or nullOptional
Output sample rate in Hz. Supported values are 8000, 16000, 22050, 24000, 32000, 44100, 48000. If omitted, the API returns 24000 Hz.
enable_text_normalizationboolean or nullOptional

Enable pre-TTS text normalization: numbers, currencies and dates are converted to their spoken form before synthesis (e.g. 9840950950 becomes the spoken digit sequence). Defaults to the server-side setting.

Response

Success. Returns a streamed audio response in the requested format (e.g., audio/mpeg for MP3, audio/wav for WAV).

Errors

400
Bad Request Error
402
Payment Required Error
403
Forbidden Error
422
Unprocessable Entity Error
429
Too Many Requests Error
500
Internal Server Error
502
Bad Gateway Error
503
Service Unavailable Error
504
Gateway Timeout Error