> For clean Markdown of any page, append `.md` to the page URL. > For a complete documentation index, see https://docs.sarvam.ai/llms.txt. > For full documentation content in one file, see https://docs.sarvam.ai/llms-full.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.sarvam.ai/_mcp/server. # WebSocket GET /voices/clone/ws WebSocket channel for real-time voice-cloning synthesis. Protocol matches Bulbul TTS WebSocket: send `config` once with a saved `voice_id`, then incremental `text` and `flush` messages. The server streams base64 `audio` frames and a `final` event. **Note:** This API Reference page is provided for informational purposes only. The Try It playground may not provide the best experience for streaming audio. For optimal streaming performance, please use the SDK or implement your own WebSocket client. **Auth:** send your key in the `api-subscription-key` header, or as a WebSocket subprotocol `api-subscription-key.`. **Streaming codecs:** `mp3` (default), `wav`, `linear16`, `mulaw`, `alaw`. `flac`, `aac`, and `opus` are not supported. Quality checks (QC/VAD) and duration bounds are not accepted on this path. Reference: https://docs.sarvam.ai/api-reference/voice-cloning/clone-ws ## AsyncAPI Specification ```yaml asyncapi: 2.6.0 info: title: voiceCloningStreaming version: subpackage_voiceCloningStreaming.voiceCloningStreaming description: > WebSocket channel for real-time voice-cloning synthesis. Protocol matches Bulbul TTS WebSocket: send `config` once with a saved `voice_id`, then incremental `text` and `flush` messages. The server streams base64 `audio` frames and a `final` event. **Note:** This API Reference page is provided for informational purposes only. The Try It playground may not provide the best experience for streaming audio. For optimal streaming performance, please use the SDK or implement your own WebSocket client. **Auth:** send your key in the `api-subscription-key` header, or as a WebSocket subprotocol `api-subscription-key.`. **Streaming codecs:** `mp3` (default), `wav`, `linear16`, `mulaw`, `alaw`. `flac`, `aac`, and `opus` are not supported. Quality checks (QC/VAD) and duration bounds are not accepted on this path. channels: /voices/clone/ws: description: > WebSocket channel for real-time voice-cloning synthesis. Protocol matches Bulbul TTS WebSocket: send `config` once with a saved `voice_id`, then incremental `text` and `flush` messages. The server streams base64 `audio` frames and a `final` event. **Note:** This API Reference page is provided for informational purposes only. The Try It playground may not provide the best experience for streaming audio. For optimal streaming performance, please use the SDK or implement your own WebSocket client. **Auth:** send your key in the `api-subscription-key` header, or as a WebSocket subprotocol `api-subscription-key.`. **Streaming codecs:** `mp3` (default), `wav`, `linear16`, `mulaw`, `alaw`. `flac`, `aac`, and `opus` are not supported. Quality checks (QC/VAD) and duration bounds are not accepted on this path. bindings: ws: query: type: object properties: send_completion_event: $ref: '#/components/schemas/voiceCloningStreaming_send_completion_event' default: 'true' headers: type: object properties: Api-Subscription-Key: type: string publish: operationId: subpackage_voiceCloningStreaming.voiceCloningStreaming-publish summary: Server messages message: oneOf: - $ref: >- #/components/messages/subpackage_voiceCloningStreaming.voiceCloningStreaming-server-0-Voice Cloning Audio Output - $ref: >- #/components/messages/subpackage_voiceCloningStreaming.voiceCloningStreaming-server-1-Voice Cloning Event Notification - $ref: >- #/components/messages/subpackage_voiceCloningStreaming.voiceCloningStreaming-server-2-Voice Cloning Error Response subscribe: operationId: subpackage_voiceCloningStreaming.voiceCloningStreaming-subscribe summary: Client messages message: oneOf: - $ref: >- #/components/messages/subpackage_voiceCloningStreaming.voiceCloningStreaming-client-0-Voice Cloning Configure Connection - $ref: >- #/components/messages/subpackage_voiceCloningStreaming.voiceCloningStreaming-client-1-Voice Cloning Send Text - $ref: >- #/components/messages/subpackage_voiceCloningStreaming.voiceCloningStreaming-client-2-Voice Cloning Flush Signal - $ref: >- #/components/messages/subpackage_voiceCloningStreaming.voiceCloningStreaming-client-3-Voice Cloning Ping Signal servers: Production: url: wss://api.sarvam.ai/ protocol: wss x-default: true components: messages: subpackage_voiceCloningStreaming.voiceCloningStreaming-server-0-Voice Cloning Audio Output: name: Voice Cloning Audio Output title: Voice Cloning Audio Output description: Receive audio chunks from the voice-cloning WebSocket payload: $ref: '#/components/schemas/voiceCloningStreaming_VoiceCloningAudioOutput' subpackage_voiceCloningStreaming.voiceCloningStreaming-server-1-Voice Cloning Event Notification: name: Voice Cloning Event Notification title: Voice Cloning Event Notification description: >- Receive completion event notifications from the voice-cloning WebSocket (if send_completion_event is enabled) payload: $ref: '#/components/schemas/voiceCloningStreaming_VoiceCloningEventResponse' subpackage_voiceCloningStreaming.voiceCloningStreaming-server-2-Voice Cloning Error Response: name: Voice Cloning Error Response title: Voice Cloning Error Response description: Receive error messages from the voice-cloning WebSocket payload: $ref: '#/components/schemas/voiceCloningStreaming_VoiceCloningErrorResponse' subpackage_voiceCloningStreaming.voiceCloningStreaming-client-0-Voice Cloning Configure Connection: name: Voice Cloning Configure Connection title: Voice Cloning Configure Connection description: Send initial configuration for voice-cloning streaming payload: $ref: >- #/components/schemas/voiceCloningStreaming_VoiceCloningConfigureConnection subpackage_voiceCloningStreaming.voiceCloningStreaming-client-1-Voice Cloning Send Text: name: Voice Cloning Send Text title: Voice Cloning Send Text description: Send text chunk for voice-cloning speech synthesis payload: $ref: '#/components/schemas/voiceCloningStreaming_VoiceCloningSendText' subpackage_voiceCloningStreaming.voiceCloningStreaming-client-2-Voice Cloning Flush Signal: name: Voice Cloning Flush Signal title: Voice Cloning Flush Signal description: Send signal to end text streaming for voice cloning payload: $ref: '#/components/schemas/voiceCloningStreaming_VoiceCloningFlushSignal' subpackage_voiceCloningStreaming.voiceCloningStreaming-client-3-Voice Cloning Ping Signal: name: Voice Cloning Ping Signal title: Voice Cloning Ping Signal description: Send ping signal to keep the voice-cloning WebSocket connection alive payload: $ref: '#/components/schemas/voiceCloningStreaming_VoiceCloningPingSignal' schemas: voiceCloningStreaming_send_completion_event: type: string enum: - 'true' - 'false' default: 'true' description: >- Enable completion event notifications when synthesis finishes. When set to true, an event message will be sent when the final audio chunk has been generated. title: voiceCloningStreaming_send_completion_event ChannelsVoiceCloningStreamingMessagesVoiceCloningAudioOutputType: type: string enum: - audio title: ChannelsVoiceCloningStreamingMessagesVoiceCloningAudioOutputType ChannelsVoiceCloningStreamingMessagesVoiceCloningAudioOutputData: type: object properties: content_type: type: string description: MIME type of the audio content (e.g., 'audio/mp3', 'audio/wav') audio: type: string format: base64 description: Base64-encoded audio data ready for playback or download request_id: type: string description: Unique identifier for the request required: - content_type - audio title: ChannelsVoiceCloningStreamingMessagesVoiceCloningAudioOutputData voiceCloningStreaming_VoiceCloningAudioOutput: type: object properties: type: $ref: >- #/components/schemas/ChannelsVoiceCloningStreamingMessagesVoiceCloningAudioOutputType data: $ref: >- #/components/schemas/ChannelsVoiceCloningStreamingMessagesVoiceCloningAudioOutputData required: - type - data title: voiceCloningStreaming_VoiceCloningAudioOutput ChannelsVoiceCloningStreamingMessagesVoiceCloningEventResponseType: type: string enum: - event description: Message type identifier for events title: ChannelsVoiceCloningStreamingMessagesVoiceCloningEventResponseType ChannelsVoiceCloningStreamingMessagesVoiceCloningEventResponseDataEventType: type: string enum: - final description: Type of event that occurred title: >- ChannelsVoiceCloningStreamingMessagesVoiceCloningEventResponseDataEventType ChannelsVoiceCloningStreamingMessagesVoiceCloningEventResponseData: type: object properties: event_type: $ref: >- #/components/schemas/ChannelsVoiceCloningStreamingMessagesVoiceCloningEventResponseDataEventType description: Type of event that occurred message: type: string description: Human-readable description of the event timestamp: type: string format: date-time description: ISO 8601 timestamp when the event occurred required: - event_type title: ChannelsVoiceCloningStreamingMessagesVoiceCloningEventResponseData voiceCloningStreaming_VoiceCloningEventResponse: type: object properties: type: $ref: >- #/components/schemas/ChannelsVoiceCloningStreamingMessagesVoiceCloningEventResponseType description: Message type identifier for events data: $ref: >- #/components/schemas/ChannelsVoiceCloningStreamingMessagesVoiceCloningEventResponseData required: - type - data description: >- Event notification message sent when specific events occur during TTS processing title: voiceCloningStreaming_VoiceCloningEventResponse ChannelsVoiceCloningStreamingMessagesVoiceCloningErrorResponseType: type: string enum: - error title: ChannelsVoiceCloningStreamingMessagesVoiceCloningErrorResponseType ChannelsVoiceCloningStreamingMessagesVoiceCloningErrorResponseData: type: object properties: message: type: string code: type: integer description: Optional error code for programmatic error handling details: type: object additionalProperties: description: Any type description: Additional error details and context information request_id: type: string description: Unique identifier for the request required: - message title: ChannelsVoiceCloningStreamingMessagesVoiceCloningErrorResponseData voiceCloningStreaming_VoiceCloningErrorResponse: type: object properties: type: $ref: >- #/components/schemas/ChannelsVoiceCloningStreamingMessagesVoiceCloningErrorResponseType data: $ref: >- #/components/schemas/ChannelsVoiceCloningStreamingMessagesVoiceCloningErrorResponseData required: - type - data title: voiceCloningStreaming_VoiceCloningErrorResponse ChannelsVoiceCloningStreamingMessagesVoiceCloningConfigureConnectionType: type: string enum: - config title: ChannelsVoiceCloningStreamingMessagesVoiceCloningConfigureConnectionType ChannelsVoiceCloningStreamingMessagesVoiceCloningConfigureConnectionDataTargetLanguageCode: type: string enum: - as-IN - bn-IN - en-IN - gu-IN - hi-IN - kn-IN - ml-IN - mr-IN - od-IN - pa-IN - ta-IN - te-IN description: > BCP-47 language code for the synthesized output. The cloned voice can speak any supported language, regardless of the reference clip's language. title: >- ChannelsVoiceCloningStreamingMessagesVoiceCloningConfigureConnectionDataTargetLanguageCode ChannelsVoiceCloningStreamingMessagesVoiceCloningConfigureConnectionDataSpeechSampleRate: type: string enum: - '8000' - '16000' - '22050' - '24000' - '32000' - '44100' - '48000' description: Output sample rate in Hz. title: >- ChannelsVoiceCloningStreamingMessagesVoiceCloningConfigureConnectionDataSpeechSampleRate ChannelsVoiceCloningStreamingMessagesVoiceCloningConfigureConnectionDataOutputAudioCodec: type: string enum: - mp3 - wav - linear16 - mulaw - alaw default: mp3 description: | Audio codec for streamed frames. Default is `mp3`. Allowed: `mp3`, `wav`, `linear16`, `mulaw`, `alaw`. title: >- ChannelsVoiceCloningStreamingMessagesVoiceCloningConfigureConnectionDataOutputAudioCodec ChannelsVoiceCloningStreamingMessagesVoiceCloningConfigureConnectionDataOutputAudioBitrate: type: string enum: - 32k - 64k - 96k - 128k - 192k default: 128k description: Bitrate for lossy codecs. Default is `128k`. title: >- ChannelsVoiceCloningStreamingMessagesVoiceCloningConfigureConnectionDataOutputAudioBitrate ChannelsVoiceCloningStreamingMessagesVoiceCloningConfigureConnectionData: type: object properties: target_language_code: $ref: >- #/components/schemas/ChannelsVoiceCloningStreamingMessagesVoiceCloningConfigureConnectionDataTargetLanguageCode description: > BCP-47 language code for the synthesized output. The cloned voice can speak any supported language, regardless of the reference clip's language. voice_id: type: string description: | Cloned voice ID (`svc-{uuid}`) from your voice library. Required. Create one with `POST /voices/create`. pace: type: number format: double minimum: 0.5 maximum: 2 description: Speech speed (0.5–2.0). 1.0 = natural speed. speech_sample_rate: $ref: >- #/components/schemas/ChannelsVoiceCloningStreamingMessagesVoiceCloningConfigureConnectionDataSpeechSampleRate description: Output sample rate in Hz. enable_pretts: type: boolean description: > Pre-TTS text normalization: numbers, currencies and dates converted to spoken form before synthesis. Defaults to the server-side setting. output_audio_codec: $ref: >- #/components/schemas/ChannelsVoiceCloningStreamingMessagesVoiceCloningConfigureConnectionDataOutputAudioCodec default: mp3 description: | Audio codec for streamed frames. Default is `mp3`. Allowed: `mp3`, `wav`, `linear16`, `mulaw`, `alaw`. output_audio_bitrate: $ref: >- #/components/schemas/ChannelsVoiceCloningStreamingMessagesVoiceCloningConfigureConnectionDataOutputAudioBitrate default: 128k description: Bitrate for lossy codecs. Default is `128k`. min_buffer_size: type: integer minimum: 30 maximum: 200 default: 50 description: >- Minimum character length that triggers buffer flushing for synthesis. max_chunk_length: type: integer minimum: 50 maximum: 500 default: 100 description: Maximum length for sentence splitting. required: - target_language_code - voice_id title: ChannelsVoiceCloningStreamingMessagesVoiceCloningConfigureConnectionData voiceCloningStreaming_VoiceCloningConfigureConnection: type: object properties: type: $ref: >- #/components/schemas/ChannelsVoiceCloningStreamingMessagesVoiceCloningConfigureConnectionType data: $ref: >- #/components/schemas/ChannelsVoiceCloningStreamingMessagesVoiceCloningConfigureConnectionData required: - type - data description: > Configuration message required as the first message after establishing the WebSocket connection. Pass a saved `voice_id` from `POST /voices/create` (or Content Studio) and the target language. Duration bounds, QC, and VAD are not accepted on this path. title: voiceCloningStreaming_VoiceCloningConfigureConnection ChannelsVoiceCloningStreamingMessagesVoiceCloningSendTextType: type: string enum: - text title: ChannelsVoiceCloningStreamingMessagesVoiceCloningSendTextType ChannelsVoiceCloningStreamingMessagesVoiceCloningSendTextData: type: object properties: text: type: string minLength: 1 maxLength: 2500 description: Text to synthesise. At most 2500 characters per event. required: - text title: ChannelsVoiceCloningStreamingMessagesVoiceCloningSendTextData voiceCloningStreaming_VoiceCloningSendText: type: object properties: type: $ref: >- #/components/schemas/ChannelsVoiceCloningStreamingMessagesVoiceCloningSendTextType data: $ref: >- #/components/schemas/ChannelsVoiceCloningStreamingMessagesVoiceCloningSendTextData required: - type - data title: voiceCloningStreaming_VoiceCloningSendText ChannelsVoiceCloningStreamingMessagesVoiceCloningFlushSignalType: type: string enum: - flush default: flush title: ChannelsVoiceCloningStreamingMessagesVoiceCloningFlushSignalType voiceCloningStreaming_VoiceCloningFlushSignal: type: object properties: type: $ref: >- #/components/schemas/ChannelsVoiceCloningStreamingMessagesVoiceCloningFlushSignalType required: - type description: > Forces the text buffer to process immediately, regardless of the min_buffer_size threshold. Use this when you need to process remaining text that hasn't reached the minimum buffer size. title: voiceCloningStreaming_VoiceCloningFlushSignal ChannelsVoiceCloningStreamingMessagesVoiceCloningPingSignalType: type: string enum: - ping default: ping title: ChannelsVoiceCloningStreamingMessagesVoiceCloningPingSignalType voiceCloningStreaming_VoiceCloningPingSignal: type: object properties: type: $ref: >- #/components/schemas/ChannelsVoiceCloningStreamingMessagesVoiceCloningPingSignalType required: - type description: > Send ping signal to keep the WebSocket connection alive. The connection automatically closes after one minute of inactivity. title: voiceCloningStreaming_VoiceCloningPingSignal ```