> For clean Markdown of any page, append `.md` to the page URL.
> For a complete documentation index, see https://docs.sarvam.ai/llms.txt.
> For full documentation content in one file, see https://docs.sarvam.ai/llms-full.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.sarvam.ai/_mcp/server.

# WebSocket

GET /voices/clone/ws

WebSocket channel for real-time voice-cloning synthesis.

Protocol matches Bulbul TTS WebSocket: send `config` once with a saved
`voice_id`, then incremental `text` and `flush` messages. The server
streams base64 `audio` frames and a `final` event.

**Note:** This API Reference page is provided for informational purposes only.
The Try It playground may not provide the best experience for streaming audio.
For optimal streaming performance, please use the SDK or implement your own WebSocket client.

**Auth:** send your key in the `api-subscription-key` header, or as a
WebSocket subprotocol `api-subscription-key.<key>`.

**Streaming codecs:** `mp3` (default), `wav`, `linear16`, `mulaw`, `alaw`.
`flac`, `aac`, and `opus` are not supported.

Quality checks (QC/VAD) and duration bounds are not accepted on this path.

Reference: https://docs.sarvam.ai/api-reference/voice-cloning/clone-ws

## AsyncAPI Specification

```yaml
asyncapi: 2.6.0
info:
  title: voiceCloningStreaming
  version: subpackage_voiceCloningStreaming.voiceCloningStreaming
  description: >
    WebSocket channel for real-time voice-cloning synthesis.


    Protocol matches Bulbul TTS WebSocket: send `config` once with a saved

    `voice_id`, then incremental `text` and `flush` messages. The server

    streams base64 `audio` frames and a `final` event.


    **Note:** This API Reference page is provided for informational purposes
    only.

    The Try It playground may not provide the best experience for streaming
    audio.

    For optimal streaming performance, please use the SDK or implement your own
    WebSocket client.


    **Auth:** send your key in the `api-subscription-key` header, or as a

    WebSocket subprotocol `api-subscription-key.<key>`.


    **Streaming codecs:** `mp3` (default), `wav`, `linear16`, `mulaw`, `alaw`.

    `flac`, `aac`, and `opus` are not supported.


    Quality checks (QC/VAD) and duration bounds are not accepted on this path.
channels:
  /voices/clone/ws:
    description: >
      WebSocket channel for real-time voice-cloning synthesis.


      Protocol matches Bulbul TTS WebSocket: send `config` once with a saved

      `voice_id`, then incremental `text` and `flush` messages. The server

      streams base64 `audio` frames and a `final` event.


      **Note:** This API Reference page is provided for informational purposes
      only.

      The Try It playground may not provide the best experience for streaming
      audio.

      For optimal streaming performance, please use the SDK or implement your
      own WebSocket client.


      **Auth:** send your key in the `api-subscription-key` header, or as a

      WebSocket subprotocol `api-subscription-key.<key>`.


      **Streaming codecs:** `mp3` (default), `wav`, `linear16`, `mulaw`, `alaw`.

      `flac`, `aac`, and `opus` are not supported.


      Quality checks (QC/VAD) and duration bounds are not accepted on this path.
    bindings:
      ws:
        query:
          type: object
          properties:
            send_completion_event:
              $ref: '#/components/schemas/voiceCloningStreaming_send_completion_event'
              default: 'true'
        headers:
          type: object
          properties:
            Api-Subscription-Key:
              type: string
    publish:
      operationId: subpackage_voiceCloningStreaming.voiceCloningStreaming-publish
      summary: Server messages
      message:
        oneOf:
          - $ref: >-
              #/components/messages/subpackage_voiceCloningStreaming.voiceCloningStreaming-server-0-Voice
              Cloning Audio Output
          - $ref: >-
              #/components/messages/subpackage_voiceCloningStreaming.voiceCloningStreaming-server-1-Voice
              Cloning Event Notification
          - $ref: >-
              #/components/messages/subpackage_voiceCloningStreaming.voiceCloningStreaming-server-2-Voice
              Cloning Error Response
    subscribe:
      operationId: subpackage_voiceCloningStreaming.voiceCloningStreaming-subscribe
      summary: Client messages
      message:
        oneOf:
          - $ref: >-
              #/components/messages/subpackage_voiceCloningStreaming.voiceCloningStreaming-client-0-Voice
              Cloning Configure Connection
          - $ref: >-
              #/components/messages/subpackage_voiceCloningStreaming.voiceCloningStreaming-client-1-Voice
              Cloning Send Text
          - $ref: >-
              #/components/messages/subpackage_voiceCloningStreaming.voiceCloningStreaming-client-2-Voice
              Cloning Flush Signal
          - $ref: >-
              #/components/messages/subpackage_voiceCloningStreaming.voiceCloningStreaming-client-3-Voice
              Cloning Ping Signal
servers:
  Production:
    url: wss://api.sarvam.ai/
    protocol: wss
    x-default: true
components:
  messages:
    subpackage_voiceCloningStreaming.voiceCloningStreaming-server-0-Voice Cloning Audio Output:
      name: Voice Cloning Audio Output
      title: Voice Cloning Audio Output
      description: Receive audio chunks from the voice-cloning WebSocket
      payload:
        $ref: '#/components/schemas/voiceCloningStreaming_VoiceCloningAudioOutput'
    subpackage_voiceCloningStreaming.voiceCloningStreaming-server-1-Voice Cloning Event Notification:
      name: Voice Cloning Event Notification
      title: Voice Cloning Event Notification
      description: >-
        Receive completion event notifications from the voice-cloning WebSocket
        (if send_completion_event is enabled)
      payload:
        $ref: '#/components/schemas/voiceCloningStreaming_VoiceCloningEventResponse'
    subpackage_voiceCloningStreaming.voiceCloningStreaming-server-2-Voice Cloning Error Response:
      name: Voice Cloning Error Response
      title: Voice Cloning Error Response
      description: Receive error messages from the voice-cloning WebSocket
      payload:
        $ref: '#/components/schemas/voiceCloningStreaming_VoiceCloningErrorResponse'
    subpackage_voiceCloningStreaming.voiceCloningStreaming-client-0-Voice Cloning Configure Connection:
      name: Voice Cloning Configure Connection
      title: Voice Cloning Configure Connection
      description: Send initial configuration for voice-cloning streaming
      payload:
        $ref: >-
          #/components/schemas/voiceCloningStreaming_VoiceCloningConfigureConnection
    subpackage_voiceCloningStreaming.voiceCloningStreaming-client-1-Voice Cloning Send Text:
      name: Voice Cloning Send Text
      title: Voice Cloning Send Text
      description: Send text chunk for voice-cloning speech synthesis
      payload:
        $ref: '#/components/schemas/voiceCloningStreaming_VoiceCloningSendText'
    subpackage_voiceCloningStreaming.voiceCloningStreaming-client-2-Voice Cloning Flush Signal:
      name: Voice Cloning Flush Signal
      title: Voice Cloning Flush Signal
      description: Send signal to end text streaming for voice cloning
      payload:
        $ref: '#/components/schemas/voiceCloningStreaming_VoiceCloningFlushSignal'
    subpackage_voiceCloningStreaming.voiceCloningStreaming-client-3-Voice Cloning Ping Signal:
      name: Voice Cloning Ping Signal
      title: Voice Cloning Ping Signal
      description: Send ping signal to keep the voice-cloning WebSocket connection alive
      payload:
        $ref: '#/components/schemas/voiceCloningStreaming_VoiceCloningPingSignal'
  schemas:
    voiceCloningStreaming_send_completion_event:
      type: string
      enum:
        - 'true'
        - 'false'
      default: 'true'
      description: >-
        Enable completion event notifications when synthesis finishes. When set
        to true, an event message will be sent when the final audio chunk has
        been generated.
      title: voiceCloningStreaming_send_completion_event
    ChannelsVoiceCloningStreamingMessagesVoiceCloningAudioOutputType:
      type: string
      enum:
        - audio
      title: ChannelsVoiceCloningStreamingMessagesVoiceCloningAudioOutputType
    ChannelsVoiceCloningStreamingMessagesVoiceCloningAudioOutputData:
      type: object
      properties:
        content_type:
          type: string
          description: MIME type of the audio content (e.g., 'audio/mp3', 'audio/wav')
        audio:
          type: string
          format: base64
          description: Base64-encoded audio data ready for playback or download
        request_id:
          type: string
          description: Unique identifier for the request
      required:
        - content_type
        - audio
      title: ChannelsVoiceCloningStreamingMessagesVoiceCloningAudioOutputData
    voiceCloningStreaming_VoiceCloningAudioOutput:
      type: object
      properties:
        type:
          $ref: >-
            #/components/schemas/ChannelsVoiceCloningStreamingMessagesVoiceCloningAudioOutputType
        data:
          $ref: >-
            #/components/schemas/ChannelsVoiceCloningStreamingMessagesVoiceCloningAudioOutputData
      required:
        - type
        - data
      title: voiceCloningStreaming_VoiceCloningAudioOutput
    ChannelsVoiceCloningStreamingMessagesVoiceCloningEventResponseType:
      type: string
      enum:
        - event
      description: Message type identifier for events
      title: ChannelsVoiceCloningStreamingMessagesVoiceCloningEventResponseType
    ChannelsVoiceCloningStreamingMessagesVoiceCloningEventResponseDataEventType:
      type: string
      enum:
        - final
      description: Type of event that occurred
      title: >-
        ChannelsVoiceCloningStreamingMessagesVoiceCloningEventResponseDataEventType
    ChannelsVoiceCloningStreamingMessagesVoiceCloningEventResponseData:
      type: object
      properties:
        event_type:
          $ref: >-
            #/components/schemas/ChannelsVoiceCloningStreamingMessagesVoiceCloningEventResponseDataEventType
          description: Type of event that occurred
        message:
          type: string
          description: Human-readable description of the event
        timestamp:
          type: string
          format: date-time
          description: ISO 8601 timestamp when the event occurred
      required:
        - event_type
      title: ChannelsVoiceCloningStreamingMessagesVoiceCloningEventResponseData
    voiceCloningStreaming_VoiceCloningEventResponse:
      type: object
      properties:
        type:
          $ref: >-
            #/components/schemas/ChannelsVoiceCloningStreamingMessagesVoiceCloningEventResponseType
          description: Message type identifier for events
        data:
          $ref: >-
            #/components/schemas/ChannelsVoiceCloningStreamingMessagesVoiceCloningEventResponseData
      required:
        - type
        - data
      description: >-
        Event notification message sent when specific events occur during TTS
        processing
      title: voiceCloningStreaming_VoiceCloningEventResponse
    ChannelsVoiceCloningStreamingMessagesVoiceCloningErrorResponseType:
      type: string
      enum:
        - error
      title: ChannelsVoiceCloningStreamingMessagesVoiceCloningErrorResponseType
    ChannelsVoiceCloningStreamingMessagesVoiceCloningErrorResponseData:
      type: object
      properties:
        message:
          type: string
        code:
          type: integer
          description: Optional error code for programmatic error handling
        details:
          type: object
          additionalProperties:
            description: Any type
          description: Additional error details and context information
        request_id:
          type: string
          description: Unique identifier for the request
      required:
        - message
      title: ChannelsVoiceCloningStreamingMessagesVoiceCloningErrorResponseData
    voiceCloningStreaming_VoiceCloningErrorResponse:
      type: object
      properties:
        type:
          $ref: >-
            #/components/schemas/ChannelsVoiceCloningStreamingMessagesVoiceCloningErrorResponseType
        data:
          $ref: >-
            #/components/schemas/ChannelsVoiceCloningStreamingMessagesVoiceCloningErrorResponseData
      required:
        - type
        - data
      title: voiceCloningStreaming_VoiceCloningErrorResponse
    ChannelsVoiceCloningStreamingMessagesVoiceCloningConfigureConnectionType:
      type: string
      enum:
        - config
      title: ChannelsVoiceCloningStreamingMessagesVoiceCloningConfigureConnectionType
    ChannelsVoiceCloningStreamingMessagesVoiceCloningConfigureConnectionDataTargetLanguageCode:
      type: string
      enum:
        - as-IN
        - bn-IN
        - en-IN
        - gu-IN
        - hi-IN
        - kn-IN
        - ml-IN
        - mr-IN
        - od-IN
        - pa-IN
        - ta-IN
        - te-IN
      description: >
        BCP-47 language code for the synthesized output. The cloned voice can
        speak any

        supported language, regardless of the reference clip's language.
      title: >-
        ChannelsVoiceCloningStreamingMessagesVoiceCloningConfigureConnectionDataTargetLanguageCode
    ChannelsVoiceCloningStreamingMessagesVoiceCloningConfigureConnectionDataSpeechSampleRate:
      type: string
      enum:
        - '8000'
        - '16000'
        - '22050'
        - '24000'
        - '32000'
        - '44100'
        - '48000'
      description: Output sample rate in Hz.
      title: >-
        ChannelsVoiceCloningStreamingMessagesVoiceCloningConfigureConnectionDataSpeechSampleRate
    ChannelsVoiceCloningStreamingMessagesVoiceCloningConfigureConnectionDataOutputAudioCodec:
      type: string
      enum:
        - mp3
        - wav
        - linear16
        - mulaw
        - alaw
      default: mp3
      description: |
        Audio codec for streamed frames. Default is `mp3`.
        Allowed: `mp3`, `wav`, `linear16`, `mulaw`, `alaw`.
      title: >-
        ChannelsVoiceCloningStreamingMessagesVoiceCloningConfigureConnectionDataOutputAudioCodec
    ChannelsVoiceCloningStreamingMessagesVoiceCloningConfigureConnectionDataOutputAudioBitrate:
      type: string
      enum:
        - 32k
        - 64k
        - 96k
        - 128k
        - 192k
      default: 128k
      description: Bitrate for lossy codecs. Default is `128k`.
      title: >-
        ChannelsVoiceCloningStreamingMessagesVoiceCloningConfigureConnectionDataOutputAudioBitrate
    ChannelsVoiceCloningStreamingMessagesVoiceCloningConfigureConnectionData:
      type: object
      properties:
        target_language_code:
          $ref: >-
            #/components/schemas/ChannelsVoiceCloningStreamingMessagesVoiceCloningConfigureConnectionDataTargetLanguageCode
          description: >
            BCP-47 language code for the synthesized output. The cloned voice
            can speak any

            supported language, regardless of the reference clip's language.
        voice_id:
          type: string
          description: |
            Cloned voice ID (`svc-{uuid}`) from your voice library. Required.
            Create one with `POST /voices/create`.
        pace:
          type: number
          format: double
          minimum: 0.5
          maximum: 2
          description: Speech speed (0.5–2.0). 1.0 = natural speed.
        speech_sample_rate:
          $ref: >-
            #/components/schemas/ChannelsVoiceCloningStreamingMessagesVoiceCloningConfigureConnectionDataSpeechSampleRate
          description: Output sample rate in Hz.
        enable_pretts:
          type: boolean
          description: >
            Pre-TTS text normalization: numbers, currencies and dates converted
            to spoken

            form before synthesis. Defaults to the server-side setting.
        output_audio_codec:
          $ref: >-
            #/components/schemas/ChannelsVoiceCloningStreamingMessagesVoiceCloningConfigureConnectionDataOutputAudioCodec
          default: mp3
          description: |
            Audio codec for streamed frames. Default is `mp3`.
            Allowed: `mp3`, `wav`, `linear16`, `mulaw`, `alaw`.
        output_audio_bitrate:
          $ref: >-
            #/components/schemas/ChannelsVoiceCloningStreamingMessagesVoiceCloningConfigureConnectionDataOutputAudioBitrate
          default: 128k
          description: Bitrate for lossy codecs. Default is `128k`.
        min_buffer_size:
          type: integer
          minimum: 30
          maximum: 200
          default: 50
          description: >-
            Minimum character length that triggers buffer flushing for
            synthesis.
        max_chunk_length:
          type: integer
          minimum: 50
          maximum: 500
          default: 100
          description: Maximum length for sentence splitting.
      required:
        - target_language_code
        - voice_id
      title: ChannelsVoiceCloningStreamingMessagesVoiceCloningConfigureConnectionData
    voiceCloningStreaming_VoiceCloningConfigureConnection:
      type: object
      properties:
        type:
          $ref: >-
            #/components/schemas/ChannelsVoiceCloningStreamingMessagesVoiceCloningConfigureConnectionType
        data:
          $ref: >-
            #/components/schemas/ChannelsVoiceCloningStreamingMessagesVoiceCloningConfigureConnectionData
      required:
        - type
        - data
      description: >
        Configuration message required as the first message after establishing
        the WebSocket connection.

        Pass a saved `voice_id` from `POST /voices/create` (or Content Studio)
        and the target language.

        Duration bounds, QC, and VAD are not accepted on this path.
      title: voiceCloningStreaming_VoiceCloningConfigureConnection
    ChannelsVoiceCloningStreamingMessagesVoiceCloningSendTextType:
      type: string
      enum:
        - text
      title: ChannelsVoiceCloningStreamingMessagesVoiceCloningSendTextType
    ChannelsVoiceCloningStreamingMessagesVoiceCloningSendTextData:
      type: object
      properties:
        text:
          type: string
          minLength: 1
          maxLength: 2500
          description: Text to synthesise. At most 2500 characters per event.
      required:
        - text
      title: ChannelsVoiceCloningStreamingMessagesVoiceCloningSendTextData
    voiceCloningStreaming_VoiceCloningSendText:
      type: object
      properties:
        type:
          $ref: >-
            #/components/schemas/ChannelsVoiceCloningStreamingMessagesVoiceCloningSendTextType
        data:
          $ref: >-
            #/components/schemas/ChannelsVoiceCloningStreamingMessagesVoiceCloningSendTextData
      required:
        - type
        - data
      title: voiceCloningStreaming_VoiceCloningSendText
    ChannelsVoiceCloningStreamingMessagesVoiceCloningFlushSignalType:
      type: string
      enum:
        - flush
      default: flush
      title: ChannelsVoiceCloningStreamingMessagesVoiceCloningFlushSignalType
    voiceCloningStreaming_VoiceCloningFlushSignal:
      type: object
      properties:
        type:
          $ref: >-
            #/components/schemas/ChannelsVoiceCloningStreamingMessagesVoiceCloningFlushSignalType
      required:
        - type
      description: >
        Forces the text buffer to process immediately, regardless of the
        min_buffer_size threshold. 

        Use this when you need to process remaining text that hasn't reached the
        minimum buffer size.
      title: voiceCloningStreaming_VoiceCloningFlushSignal
    ChannelsVoiceCloningStreamingMessagesVoiceCloningPingSignalType:
      type: string
      enum:
        - ping
      default: ping
      title: ChannelsVoiceCloningStreamingMessagesVoiceCloningPingSignalType
    voiceCloningStreaming_VoiceCloningPingSignal:
      type: object
      properties:
        type:
          $ref: >-
            #/components/schemas/ChannelsVoiceCloningStreamingMessagesVoiceCloningPingSignalType
      required:
        - type
      description: >
        Send ping signal to keep the WebSocket connection alive. The connection
        automatically 

        closes after one minute of inactivity.
      title: voiceCloningStreaming_VoiceCloningPingSignal

```