Skip to navigation

HTTP Streaming API

View as Markdown

POST /voices/clone/stream. Send text and a saved voice_id, get a binary audio stream back. The response starts arriving as soon as the first sentence is ready, so you can begin playback without waiting for the full file.

No WebSocket handshake, no config messages, no connection lifecycle. One HTTP request, one streamed response.

Common use cases:

  • Backend audio generation: Pipe audio directly to a file, S3, or a downstream service
  • API proxying: Forward the stream to your frontend or mobile client as-is
  • Serverless / edge: Works in any environment that supports HTTP, no WebSocket runtime needed

HTTP Stream vs WebSocket: When to Use Which

Both give you streaming audio from a cloned voice. The difference is how much control you need.

HTTP StreamWebSocket
ProtocolSingle POST requestPersistent bidirectional connection
SetupZero: it’s a normal HTTP callHandshake + config message before first text
Endpoint/voices/clone/stream/voices/clone/ws
Text inputOne text payload per request (max 1000 characters)Send multiple texts on the same connection (max 2500 characters per message)
Audio outputBinary stream (play/save directly)Base64-encoded chunks (decode each one)
Connection reuseNew connection per requestOne connection, many conversions
Best forOne-shot generation, server-side pipelines, simple integrationsVoice agents, interactive apps, multi-turn conversations

Use HTTP Stream when:

  • You have a complete text and just need audio back
  • You’re generating audio server-side (batch jobs, API endpoints, CI pipelines)
  • Your runtime doesn’t support WebSocket
  • You want the simplest possible integration — curl -N works out of the box

Use WebSocket when:

  • You’re building a conversational agent that streams text incrementally (e.g., from an LLM)
  • You need to send multiple texts without reconnecting
  • Low time-to-first-byte on successive utterances matters (connection is already warm)

See the WebSocket streaming guide for the socket protocol.


Code Examples

Create a voice with POST /voices/create first, then pass the returned voice_id.

curl -N -X POST "https://api.sarvam.ai/voices/clone/stream" \
-H "api-subscription-key: $SARVAM_API_KEY" \
-F "voice_id=svc-your-voice-id" \
-F "text=नमस्ते। यह एक स्ट्रीमिंग टेस्ट है। ऑडियो तुरंत मिलना शुरू हो जाना चाहिए।" \
-F "language_code=hi-IN" \
-F "output_audio_codec=mp3" \
--output output.mp3

Runnable samples are also on the REST Stream endpoint in the API Reference.


Request fields

Same multipart contract as POST /voices/clone, with these differences:

FieldStreaming behavior
output_audio_codecDefaults to mp3 (not wav). Allowed: mp3, wav, linear16, mulaw, alaw. flac / aac / opus are not supported (422).
output_audio_bitrateOptional. Bitrate for lossy codecs (32k, 64k, 96k, 128k, 192k). Default 128k.
min_audio_duration / max_audio_durationNot accepted — duration bounds do not apply to a live stream.
enable_qc / enable_vadNot accepted. Streaming optimizes for time-to-first-byte; QC/VAD used on single-sentence REST calls are bypassed here.

Codec guidance

CodecMIME typeNotes
mp3 (recommended)audio/mpegSelf-framing; clients can decode and play as chunks arrive. Default.
wavaudio/wavOne RIFF header (unknown length), then raw PCM chunks.
linear16audio/pcmRaw PCM only (no header).
mulaw / alawaudio/mulaw / audio/alawTelephony-style encodings.

How a client should handle the stream

  1. Do not parse JSON. A successful response body is raw audio bytes.
  2. Check status before reading the body. Errors before the first audio byte return normal HTTP error JSON (400, 402, 403, 422, 429, 503).
  3. After the first byte, the stream is binary. Mid-stream failures end the connection; there is no in-band error envelope.
  4. Progressive playback: pipe chunks into a decoder or a media player that accepts incremental MP3/PCM.
  5. Disconnect: closing the client connection cancels the stream; generation work already done may still be billed.

Streaming vs non-streaming

POST /voices/clonePOST /voices/clone/stream
ResponseJSON + base64 audioBinary audio (Content-Type per codec)
Default codecwavmp3
When you get audioAfter full synthesisAfter first sentence is ready
Best forShort clips, simple integrationsLonger text, progressive playback

Next steps