WebSocket Streaming API
GET /voices/clone/ws (WebSocket upgrade). Connect once, send a config message with a saved voice_id, then stream text and flush. The server returns base64-encoded audio frames and a final event when synthesis completes.
The wire protocol matches Bulbul TTS WebSocket, so a client that already speaks that protocol can retarget with a different path and config payload.
Common use cases:
- Conversational agents: Stream TTS responses in real time for voice assistants that use a cloned voice
- Incremental LLM output: Feed tokens as they arrive; flush when the utterance ends
- Multi-turn sessions: Reuse one connection across many utterances
For one-shot generation without a persistent socket, prefer HTTP streaming.
Connection
Auth: send your key in the api-subscription-key header, or as a WebSocket subprotocol api-subscription-key.<your-key>.
Optional query parameter: send_completion_event=true (default) to receive a final event when synthesis completes.
Message flow
- Connect and authenticate.
- Send
{ "type": "config", "data": { ... } }once. Must includetarget_language_codeandvoice_id. - Send one or more
{ "type": "text", "data": { "text": "..." } }messages (max 2500 characters each). - Send
{ "type": "flush" }to force remaining buffered text through synthesis. - Receive
{ "type": "audio", "data": { "audio": "<base64>", "content_type": "...", "request_id": "..." } }frames. - Receive
{ "type": "event", "data": { "event_type": "final" } }when the utterance is done. - Optionally send
{ "type": "ping" }to keep long-lived connections alive.
Config uses target_language_code, not language_code. Extra fields are ignored, so sending the REST field name silently does nothing.
Config fields
Not accepted: min_audio_duration, max_audio_duration, enable_qc, enable_vad, enable_cached_responses.
Code example
Full message schemas are on the WebSocket endpoint in the API Reference.
Error handling
- Validation and auth failures before or during the session arrive as
{ "type": "error", "data": { "message": "...", "code": ... } }frames. - After the first audio frame, a mid-stream model failure may close the connection; there is no separate binary error envelope.
- Keep the socket alive with periodic
pingmessages on long-lived sessions.