Skip to navigation

WebSocket Streaming API

View as Markdown

GET /voices/clone/ws (WebSocket upgrade). Connect once, send a config message with a saved voice_id, then stream text and flush. The server returns base64-encoded audio frames and a final event when synthesis completes.

The wire protocol matches Bulbul TTS WebSocket, so a client that already speaks that protocol can retarget with a different path and config payload.

Common use cases:

  • Conversational agents: Stream TTS responses in real time for voice assistants that use a cloned voice
  • Incremental LLM output: Feed tokens as they arrive; flush when the utterance ends
  • Multi-turn sessions: Reuse one connection across many utterances

For one-shot generation without a persistent socket, prefer HTTP streaming.


Connection

wss://api.sarvam.ai/voices/clone/ws

Auth: send your key in the api-subscription-key header, or as a WebSocket subprotocol api-subscription-key.<your-key>.

Optional query parameter: send_completion_event=true (default) to receive a final event when synthesis completes.


Message flow

  1. Connect and authenticate.
  2. Send { "type": "config", "data": { ... } } once. Must include target_language_code and voice_id.
  3. Send one or more { "type": "text", "data": { "text": "..." } } messages (max 2500 characters each).
  4. Send { "type": "flush" } to force remaining buffered text through synthesis.
  5. Receive { "type": "audio", "data": { "audio": "<base64>", "content_type": "...", "request_id": "..." } } frames.
  6. Receive { "type": "event", "data": { "event_type": "final" } } when the utterance is done.
  7. Optionally send { "type": "ping" } to keep long-lived connections alive.

Config uses target_language_code, not language_code. Extra fields are ignored, so sending the REST field name silently does nothing.


Config fields

FieldRequiredDescription
target_language_codeYesBCP-47 output language (e.g. hi-IN).
voice_idYesSaved cloned voice ID (svc-{uuid}) from POST /voices/create.
output_audio_codecNoDefault mp3. Allowed: mp3, wav, linear16, mulaw, alaw.
output_audio_bitrateNoDefault 128k. Options: 32k, 64k, 96k, 128k, 192k.
speech_sample_rateNoOutput sample rate in Hz.
paceNoSpeech speed 0.5–2.0.
enable_prettsNoPre-TTS text normalization.
min_buffer_sizeNoMinimum text buffer before dispatching (30–200).
max_chunk_lengthNoMaximum chunk length while splitting (50–500).

Not accepted: min_audio_duration, max_audio_duration, enable_qc, enable_vad, enable_cached_responses.


Code example

import asyncio
import base64
import json
import websockets
API_KEY = "YOUR_SARVAM_API_KEY"
URI = "wss://api.sarvam.ai/voices/clone/ws"
async def clone_stream():
async with websockets.connect(
URI,
additional_headers={"api-subscription-key": API_KEY},
subprotocols=[f"api-subscription-key.{API_KEY}"],
) as ws:
await ws.send(
json.dumps(
{
"type": "config",
"data": {
"target_language_code": "hi-IN",
"voice_id": "svc-your-voice-id",
"output_audio_codec": "mp3",
},
}
)
)
await ws.send(
json.dumps(
{
"type": "text",
"data": {
"text": "नमस्ते। यह एक वेबसॉकेट स्ट्रीमिंग टेस्ट है।"
},
}
)
)
await ws.send(json.dumps({"type": "flush"}))
chunks = []
async for raw in ws:
msg = json.loads(raw)
if msg["type"] == "audio":
chunks.append(base64.b64decode(msg["data"]["audio"]))
elif msg["type"] == "event" and msg["data"].get("event_type") == "final":
break
elif msg["type"] == "error":
raise RuntimeError(msg["data"]["message"])
with open("output.mp3", "wb") as f:
f.write(b"".join(chunks))
asyncio.run(clone_stream())

Full message schemas are on the WebSocket endpoint in the API Reference.


Error handling

  • Validation and auth failures before or during the session arrive as { "type": "error", "data": { "message": "...", "code": ... } } frames.
  • After the first audio frame, a mid-stream model failure may close the connection; there is no separate binary error envelope.
  • Keep the socket alive with periodic ping messages on long-lived sessions.

Next steps