Skip to navigation

WebSocket

View as Markdown

WebSocket channel for real-time speech to text streaming.

Note: This API Reference page is provided for informational purposes only. The Try It playground may not provide the best experience for streaming audio. For optimal streaming performance, please use the SDK or implement your own WebSocket client.

Handshake

WSS
wss://api.sarvam.ai/speech-to-text/ws

Headers

Api-Subscription-KeystringRequired
API subscription key for authentication

Query parameters

language-codeenumRequired

Specifies the language of the input audio in BCP-47 format.

Available Options:

  • unknown (default): Use when the language is not known; the API will auto-detect.
  • hi-IN: Hindi
  • bn-IN: Bengali
  • gu-IN: Gujarati
  • kn-IN: Kannada
  • ml-IN: Malayalam
  • mr-IN: Marathi
  • od-IN: Odia
  • pa-IN: Punjabi
  • ta-IN: Tamil
  • te-IN: Telugu
  • en-IN: English
  • as-IN: Assamese
  • ur-IN: Urdu
  • ne-IN: Nepali
  • kok-IN: Konkani
  • ks-IN: Kashmiri
  • sd-IN: Sindhi
  • sa-IN: Sanskrit
  • sat-IN: Santali
  • mni-IN: Manipuri
  • brx-IN: Bodo
  • mai-IN: Maithili
  • doi-IN: Dogri
modelenumOptionalDefaults to saaras:v4

Specifies the model to use for speech-to-text conversion.

  • saaras:v4 (default, recommended, latest): Flexible output formats across all modes (transcribe, translate, verbatim, translit, codemix), supporting Global + Indian English and 22 Indic languages.

  • saaras:v3: State-of-the-art model with flexible output formats. Supports multiple modes via the mode parameter: transcribe, translate, verbatim, translit, codemix.

Allowed values:
modeenumOptionalDefaults to transcribe

Mode of operation. Only applicable when using saaras:v3 or saaras:v4 models.

Example audio: 'मेरा फोन नंबर है 9840950950'

  • transcribe (default): Standard transcription in the original language with proper formatting and number normalization.

    • Output: मेरा फोन नंबर है 9840950950
  • translate: Translates speech from any supported Indic language to English.

    • Output: My phone number is 9840950950
  • verbatim: Exact word-for-word transcription without normalization, preserving filler words and spoken numbers as-is.

    • Output: मेरा फोन नंबर है नौ आठ चार zero नौ पांच zero नौ पांच zero
  • translit: Romanization - Transliterates speech to Latin/Roman script only.

    • Output: mera phone number hai 9840950950
  • codemix: Code-mixed text with English words in English and Indic words in native script.

    • Output: मेरा phone number है 9840950950
Allowed values:
keytermsstringOptional

JSON-encoded array of up to 50 domain-specific terms (names, places, brands, technical terms) to bias recognition toward, e.g. keyterms=["Sarvam","New Delhi","Vistaar"]. Each keyterm can contain up to 64 characters; put phrases such as New Delhi in one list item and do not send comma-separated terms in one string. Keyterms bias recognition — they do not guarantee that a term will appear in the transcript. Only supported with model="saaras:v4". Do not use the older keyterm or hotwords fields.

sample_rateenumOptionalDefaults to 16000
Audio sample rate for the WebSocket connection. When specified as a connection parameter, only 16kHz and 8kHz are supported. 8kHz is only available via this connection parameter. If not specified, defaults to 16kHz.
Allowed values:
high_vad_sensitivityenumOptional

Enable high VAD (Voice Activity Detection) sensitivity

Allowed values:
positive_speech_thresholdstringOptionalDefaults to 0.7

VAD probability threshold (0.0–1.0) above which a frame is considered speech. Overrides the server default when provided.

negative_speech_thresholdstringOptionalDefaults to 0.45

VAD probability threshold (0.0–1.0) below which a frame is considered silence. Overrides the server default (or the high_vad_sensitivity preset) when provided.

min_speech_framesstringOptionalDefaults to 2
Minimum number of consecutive speech frames required to start a speech segment. Overrides the server default when provided.
first_turn_min_speech_framesstringOptionalDefaults to 8
Minimum speech frames required specifically for the first user turn. Overrides the server default when provided.
negative_frames_countstringOptionalDefaults to 18

Number of negative (silence) frames needed within the window to end a speech segment. Overrides the server default (or the high_vad_sensitivity preset) when provided.

negative_frames_windowstringOptionalDefaults to 24

Sliding window size (in frames) over which negative frames are counted. Overrides the server default (or the high_vad_sensitivity preset) when provided.

start_speech_volume_thresholdstringOptional

Volume level (dB) below which audio is considered too quiet to be speech. When not provided, no volume-based filtering is applied.

interrupt_min_speech_framesstringOptionalDefaults to 2

Minimum speech frames required to register a barge-in / interruption. Overrides the server default when provided.

pre_speech_pad_framesstringOptionalDefaults to 9
Number of audio frames to prepend before the detected speech onset, ensuring the beginning of speech is not clipped. Overrides the server default when provided.
num_initial_ignored_framesstringOptionalDefaults to 0
Number of leading audio frames to skip entirely at connection start. Useful for discarding connection setup noise. Overrides the server default when provided.
vad_signalsenumOptional
Enable VAD signals in response
Allowed values:
flush_signalenumOptional
Signal to flush the audio buffer and finalize transcription
Allowed values:
input_audio_codecenumOptional

Audio codec/format of the input stream. Use this when sending raw PCM audio. Supported values: wav, pcm_s16le, pcm_l16, pcm_raw.

Allowed values:

Send

Audio Transcription MessageobjectRequired
Send audio data for real-time speech to text streaming
OR
Speech Flush SignalobjectRequired
Send signal to flush audio buffer and finalize transcription

Receive

TranscriptionobjectRequired
Receive real-time transcription results from the WebSocket