WebSocket
WebSocket channel for real-time speech to text streaming.
Note: This API Reference page is provided for informational purposes only. The Try It playground may not provide the best experience for streaming audio. For optimal streaming performance, please use the SDK or implement your own WebSocket client.
Handshake
Headers
Query parameters
Specifies the language of the input audio in BCP-47 format.
Available Options:
unknown(default): Use when the language is not known; the API will auto-detect.hi-IN: Hindibn-IN: Bengaligu-IN: Gujaratikn-IN: Kannadaml-IN: Malayalammr-IN: Marathiod-IN: Odiapa-IN: Punjabita-IN: Tamilte-IN: Teluguen-IN: Englishas-IN: Assameseur-IN: Urdune-IN: Nepalikok-IN: Konkaniks-IN: Kashmirisd-IN: Sindhisa-IN: Sanskritsat-IN: Santalimni-IN: Manipuribrx-IN: Bodomai-IN: Maithilidoi-IN: Dogri
Specifies the model to use for speech-to-text conversion.
-
saaras:v4 (default, recommended, latest): Flexible output formats across all modes (transcribe, translate, verbatim, translit, codemix), supporting Global + Indian English and 22 Indic languages.
-
saaras:v3: State-of-the-art model with flexible output formats. Supports multiple modes via the
modeparameter: transcribe, translate, verbatim, translit, codemix.
Mode of operation. Only applicable when using saaras:v3 or saaras:v4 models.
Example audio: 'मेरा फोन नंबर है 9840950950'
-
transcribe (default): Standard transcription in the original language with proper formatting and number normalization.
- Output:
मेरा फोन नंबर है 9840950950
- Output:
-
translate: Translates speech from any supported Indic language to English.
- Output:
My phone number is 9840950950
- Output:
-
verbatim: Exact word-for-word transcription without normalization, preserving filler words and spoken numbers as-is.
- Output:
मेरा फोन नंबर है नौ आठ चार zero नौ पांच zero नौ पांच zero
- Output:
-
translit: Romanization - Transliterates speech to Latin/Roman script only.
- Output:
mera phone number hai 9840950950
- Output:
-
codemix: Code-mixed text with English words in English and Indic words in native script.
- Output:
मेरा phone number है 9840950950
- Output:
JSON-encoded array of up to 50 domain-specific terms (names, places,
brands, technical terms) to bias recognition toward, e.g.
keyterms=["Sarvam","New Delhi","Vistaar"]. Each keyterm can contain
up to 64 characters; put phrases such as New Delhi in one list item
and do not send comma-separated terms in one string. Keyterms bias
recognition — they do not guarantee that a term will appear in the
transcript. Only supported with model="saaras:v4". Do not use
the older keyterm or hotwords fields.
Enable high VAD (Voice Activity Detection) sensitivity
VAD probability threshold (0.0–1.0) above which a frame is considered speech. Overrides the server default when provided.
VAD probability threshold (0.0–1.0) below which a frame is considered silence. Overrides the server default (or the high_vad_sensitivity preset) when provided.
Number of negative (silence) frames needed within the window to end a speech segment. Overrides the server default (or the high_vad_sensitivity preset) when provided.
Sliding window size (in frames) over which negative frames are counted. Overrides the server default (or the high_vad_sensitivity preset) when provided.
Volume level (dB) below which audio is considered too quiet to be speech. When not provided, no volume-based filtering is applied.
Minimum speech frames required to register a barge-in / interruption. Overrides the server default when provided.
Audio codec/format of the input stream. Use this when sending raw PCM audio. Supported values: wav, pcm_s16le, pcm_l16, pcm_raw.