> For clean Markdown of any page, append `.md` to the page URL. > For a complete documentation index, see https://docs.sarvam.ai/llms.txt. > For full documentation content in one file, see https://docs.sarvam.ai/llms-full.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.sarvam.ai/_mcp/server. # HTTP Streaming API > Stream cloned-voice audio over a single HTTP POST. No WebSocket setup — POST text and a voice_id, pipe the binary audio response. `POST /voices/clone/stream`. Send text and a saved `voice_id`, get a binary audio stream back. The response starts arriving as soon as the first sentence is ready, so you can begin playback without waiting for the full file. No WebSocket handshake, no config messages, no connection lifecycle. One HTTP request, one streamed response. **Common use cases:** * **Backend audio generation**: Pipe audio directly to a file, S3, or a downstream service * **API proxying**: Forward the stream to your frontend or mobile client as-is * **Serverless / edge**: Works in any environment that supports HTTP, no WebSocket runtime needed --- ## HTTP Stream vs WebSocket: When to Use Which Both give you streaming audio from a cloned voice. The difference is how much control you need. | | HTTP Stream | WebSocket | | -------------------- | --------------------------------------------------------------- | ---------------------------------------------------------------------------- | | **Protocol** | Single `POST` request | Persistent bidirectional connection | | **Setup** | Zero: it's a normal HTTP call | Handshake + config message before first text | | **Endpoint** | `/voices/clone/stream` | `/voices/clone/ws` | | **Text input** | One text payload per request (max 1000 characters) | Send multiple texts on the same connection (max 2500 characters per message) | | **Audio output** | Binary stream (play/save directly) | Base64-encoded chunks (decode each one) | | **Connection reuse** | New connection per request | One connection, many conversions | | **Best for** | One-shot generation, server-side pipelines, simple integrations | Voice agents, interactive apps, multi-turn conversations | **Use HTTP Stream when:** * You have a complete text and just need audio back * You're generating audio server-side (batch jobs, API endpoints, CI pipelines) * Your runtime doesn't support WebSocket * You want the simplest possible integration — `curl -N` works out of the box **Use WebSocket when:** * You're building a conversational agent that streams text incrementally (e.g., from an LLM) * You need to send multiple texts without reconnecting * Low time-to-first-byte on successive utterances matters (connection is already warm) See the [WebSocket streaming guide](/api/api-guides-tutorials/voice-cloning/how-to/websocket-streaming) for the socket protocol. --- ## Code Examples Create a voice with `POST /voices/create` first, then pass the returned `voice_id`. #### cURL ```bash curl -N -X POST "https://api.sarvam.ai/voices/clone/stream" \ -H "api-subscription-key: $SARVAM_API_KEY" \ -F "voice_id=svc-your-voice-id" \ -F "text=नमस्ते। यह एक स्ट्रीमिंग टेस्ट है। ऑडियो तुरंत मिलना शुरू हो जाना चाहिए।" \ -F "language_code=hi-IN" \ -F "output_audio_codec=mp3" \ --output output.mp3 ``` #### Python ```python import httpx API_KEY = "YOUR_SARVAM_API_KEY" URL = "https://api.sarvam.ai/voices/clone/stream" with httpx.stream( "POST", URL, headers={"api-subscription-key": API_KEY}, data={ "text": "नमस्ते। यह एक स्ट्रीमिंग टेस्ट है। ऑडियो तुरंत मिलना शुरू हो जाना चाहिए।", "language_code": "hi-IN", "voice_id": "svc-your-voice-id", "output_audio_codec": "mp3", }, timeout=120.0, ) as response: response.raise_for_status() # fails on 4xx/5xx before any audio with open("output.mp3", "wb") as f: for chunk in response.iter_bytes(4096): f.write(chunk) # or feed to a live player ``` Runnable samples are also on the [REST Stream endpoint](/api-reference/voice-cloning/clone-stream) in the API Reference. --- ## Request fields Same multipart contract as `POST /voices/clone`, with these differences: | Field | Streaming behavior | | ------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------- | | `output_audio_codec` | Defaults to **`mp3`** (not `wav`). Allowed: `mp3`, `wav`, `linear16`, `mulaw`, `alaw`. `flac` / `aac` / `opus` are **not** supported (422). | | `output_audio_bitrate` | Optional. Bitrate for lossy codecs (`32k`, `64k`, `96k`, `128k`, `192k`). Default `128k`. | | `min_audio_duration` / `max_audio_duration` | **Not accepted** — duration bounds do not apply to a live stream. | | `enable_qc` / `enable_vad` | Not accepted. Streaming optimizes for time-to-first-byte; QC/VAD used on single-sentence REST calls are bypassed here. | --- ## Codec guidance | Codec | MIME type | Notes | | ----------------------- | ---------------------------- | -------------------------------------------------------------------- | | **`mp3`** (recommended) | `audio/mpeg` | Self-framing; clients can decode and play as chunks arrive. Default. | | `wav` | `audio/wav` | One RIFF header (unknown length), then raw PCM chunks. | | `linear16` | `audio/pcm` | Raw PCM only (no header). | | `mulaw` / `alaw` | `audio/mulaw` / `audio/alaw` | Telephony-style encodings. | --- ## How a client should handle the stream 1. **Do not parse JSON.** A successful response body is raw audio bytes. 2. **Check status before reading the body.** Errors before the first audio byte return normal HTTP error JSON (`400`, `402`, `403`, `422`, `429`, `503`). 3. **After the first byte, the stream is binary.** Mid-stream failures end the connection; there is no in-band error envelope. 4. **Progressive playback:** pipe chunks into a decoder or a media player that accepts incremental MP3/PCM. 5. **Disconnect:** closing the client connection cancels the stream; generation work already done may still be billed. --- ## Streaming vs non-streaming | | `POST /voices/clone` | `POST /voices/clone/stream` | | ------------------ | -------------------------------- | --------------------------------------- | | Response | JSON + base64 audio | Binary audio (`Content-Type` per codec) | | Default codec | `wav` | `mp3` | | When you get audio | After full synthesis | After first sentence is ready | | Best for | Short clips, simple integrations | Longer text, progressive playback | --- ## Next steps #### [WebSocket streaming](/api/api-guides-tutorials/voice-cloning/how-to/websocket-streaming) Persistent connection for multi-turn and incremental text. #### [Choose audio formats](/api/api-guides-tutorials/voice-cloning/how-to/choose-audio-formats) Codecs and sample rates for REST and streaming. #### [API Reference — REST Stream](/api-reference/voice-cloning/clone-stream) Full request and response schema. > Stream cloned-voice audio over a single HTTP POST. No WebSocket setup — POST text and a voice_id, pipe the binary audio response.