Skip to navigation

WebSocket

View as Markdown

WebSocket channel for real-time voice-cloning synthesis.

Protocol matches Bulbul TTS WebSocket: send config once with a saved voice_id, then incremental text and flush messages. The server streams base64 audio frames and a final event.

Note: This API Reference page is provided for informational purposes only. The Try It playground may not provide the best experience for streaming audio. For optimal streaming performance, please use the SDK or implement your own WebSocket client.

Auth: send your key in the api-subscription-key header, or as a WebSocket subprotocol api-subscription-key.<key>.

Streaming codecs: mp3 (default), wav, linear16, mulaw, alaw. flac, aac, and opus are not supported.

Quality checks (QC/VAD) and duration bounds are not accepted on this path.

Handshake

WSS
wss://api.sarvam.ai/voices/clone/ws

Headers

Api-Subscription-KeystringRequired
API subscription key for authentication

Query parameters

send_completion_eventenumOptionalDefaults to true
Enable completion event notifications when synthesis finishes. When set to true, an event message will be sent when the final audio chunk has been generated.
Allowed values:

Send

Voice Cloning Configure ConnectionobjectRequired
Send initial configuration for voice-cloning streaming
OR
Voice Cloning Send TextobjectRequired
Send text chunk for voice-cloning speech synthesis
OR
Voice Cloning Flush SignalobjectRequired
Send signal to end text streaming for voice cloning
OR
Voice Cloning Ping SignalobjectRequired
Send ping signal to keep the voice-cloning WebSocket connection alive

Receive

Voice Cloning Audio OutputobjectRequired
Receive audio chunks from the voice-cloning WebSocket
OR
Voice Cloning Event NotificationobjectRequired
Receive completion event notifications from the voice-cloning WebSocket (if send_completion_event is enabled)
OR
Voice Cloning Error ResponseobjectRequired
Receive error messages from the voice-cloning WebSocket