> For clean Markdown of any page, append `.md` to the page URL. > For a complete documentation index, see https://docs.sarvam.ai/llms.txt. > For full documentation content in one file, see https://docs.sarvam.ai/llms-full.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.sarvam.ai/_mcp/server. # WebSocket Streaming API > Stream cloned-voice audio over a persistent WebSocket. Send config once with a voice_id, then text and flush; receive base64 audio frames. `GET /voices/clone/ws` (WebSocket upgrade). Connect once, send a `config` message with a saved `voice_id`, then stream `text` and `flush`. The server returns base64-encoded `audio` frames and a `final` event when synthesis completes. The wire protocol matches Bulbul TTS WebSocket, so a client that already speaks that protocol can retarget with a different path and config payload. **Common use cases:** * **Conversational agents**: Stream TTS responses in real time for voice assistants that use a cloned voice * **Incremental LLM output**: Feed tokens as they arrive; flush when the utterance ends * **Multi-turn sessions**: Reuse one connection across many utterances For one-shot generation without a persistent socket, prefer [HTTP streaming](/api/api-guides-tutorials/voice-cloning/how-to/http-streaming). --- ## Connection ``` wss://api.sarvam.ai/voices/clone/ws ``` **Auth:** send your key in the `api-subscription-key` header, or as a WebSocket subprotocol `api-subscription-key.`. Optional query parameter: `send_completion_event=true` (default) to receive a `final` event when synthesis completes. --- ## Message flow 1. Connect and authenticate. 2. Send `{ "type": "config", "data": { ... } }` once. Must include `target_language_code` and `voice_id`. 3. Send one or more `{ "type": "text", "data": { "text": "..." } }` messages (max 2500 characters each). 4. Send `{ "type": "flush" }` to force remaining buffered text through synthesis. 5. Receive `{ "type": "audio", "data": { "audio": "", "content_type": "...", "request_id": "..." } }` frames. 6. Receive `{ "type": "event", "data": { "event_type": "final" } }` when the utterance is done. 7. Optionally send `{ "type": "ping" }` to keep long-lived connections alive. > **Warning** > > Config uses **`target_language_code`**, not `language_code`. Extra fields are ignored, so sending the REST field name silently does nothing. --- ## Config fields | Field | Required | Description | | ---------------------- | -------- | ------------------------------------------------------------------ | | `target_language_code` | Yes | BCP-47 output language (e.g. `hi-IN`). | | `voice_id` | Yes | Saved cloned voice ID (`svc-{uuid}`) from `POST /voices/create`. | | `output_audio_codec` | No | Default `mp3`. Allowed: `mp3`, `wav`, `linear16`, `mulaw`, `alaw`. | | `output_audio_bitrate` | No | Default `128k`. Options: `32k`, `64k`, `96k`, `128k`, `192k`. | | `speech_sample_rate` | No | Output sample rate in Hz. | | `pace` | No | Speech speed 0.5–2.0. | | `enable_pretts` | No | Pre-TTS text normalization. | | `min_buffer_size` | No | Minimum text buffer before dispatching (30–200). | | `max_chunk_length` | No | Maximum chunk length while splitting (50–500). | Not accepted: `min_audio_duration`, `max_audio_duration`, `enable_qc`, `enable_vad`, `enable_cached_responses`. --- ## Code example ```python import asyncio import base64 import json import websockets API_KEY = "YOUR_SARVAM_API_KEY" URI = "wss://api.sarvam.ai/voices/clone/ws" async def clone_stream(): async with websockets.connect( URI, additional_headers={"api-subscription-key": API_KEY}, subprotocols=[f"api-subscription-key.{API_KEY}"], ) as ws: await ws.send( json.dumps( { "type": "config", "data": { "target_language_code": "hi-IN", "voice_id": "svc-your-voice-id", "output_audio_codec": "mp3", }, } ) ) await ws.send( json.dumps( { "type": "text", "data": { "text": "नमस्ते। यह एक वेबसॉकेट स्ट्रीमिंग टेस्ट है।" }, } ) ) await ws.send(json.dumps({"type": "flush"})) chunks = [] async for raw in ws: msg = json.loads(raw) if msg["type"] == "audio": chunks.append(base64.b64decode(msg["data"]["audio"])) elif msg["type"] == "event" and msg["data"].get("event_type") == "final": break elif msg["type"] == "error": raise RuntimeError(msg["data"]["message"]) with open("output.mp3", "wb") as f: f.write(b"".join(chunks)) asyncio.run(clone_stream()) ``` Full message schemas are on the [WebSocket endpoint](/api-reference/voice-cloning/clone-ws) in the API Reference. --- ## Error handling * Validation and auth failures before or during the session arrive as `{ "type": "error", "data": { "message": "...", "code": ... } }` frames. * After the first audio frame, a mid-stream model failure may close the connection; there is no separate binary error envelope. * Keep the socket alive with periodic `ping` messages on long-lived sessions. --- ## Next steps #### [HTTP streaming](/api/api-guides-tutorials/voice-cloning/how-to/http-streaming) One POST, binary stream — simpler for one-shot generation. #### [API Reference — WebSocket](/api-reference/voice-cloning/clone-ws) Full message schemas and parameters. > Stream cloned-voice audio over a persistent WebSocket. Send config once with a voice_id, then text and flush; receive base64 audio frames.