Realtime Streaming
WebSocket channel for real-time speech to text streaming powered by the
saaras:v3-realtime (default) or saaras:v4 model.
Authentication: Pass your API subscription key via the
API-SUBSCRIPTION-KEY header, or (in browsers) via the WebSocket
subprotocol api-subscription-key.<key> — the server echoes the
subprotocol back.
Note: This API Reference page is provided for informational purposes only. The Try It playground may not provide the best experience for streaming audio. For optimal streaming performance, please use the SDK or implement your own WebSocket client.
Handshake
Headers
Query parameters
BCP-47 language code of the input audio. Required.
Supported languages (saaras:v3-realtime and saaras:v4):
auto: Adaptive automatic language detectionen-IN: Englishhi-IN: Hindibn-IN: Bengalikn-IN: Kannadaml-IN: Malayalammr-IN: Marathior-IN: Odiapa-IN: Punjabita-IN: Tamilte-IN: Telugugu-IN: Gujaratias-IN: Assameseur-IN: Urdune-IN: Nepalikok-IN: Konkaniks-IN: Kashmirisd-IN: Sindhisa-IN: Sanskritsat-IN: Santalimni-IN: Manipuribrx-IN: Bodomai-IN: Maithilidoi-IN: Dogri
Speech-to-text model for this endpoint. saaras:v3-realtime (default) or saaras:v4.
Controls audio chunking and latency behaviour.
- fast: Partials at low latency.
- balanced (default): Partials with better accuracy.
- simulated: No partials.
Task applied on the final transcript, not on partial transcripts. Streaming partials are always straight transcription regardless of this value.
- transcribe (default): Standard transcription in the source language script.
- translate: Translation to English.
- verbatim: Exact transcription in the original script (includes disfluencies).
- translit: Transliteration to Latin script.
- codemix: Code-mixed output (native + English).
Optional context/terminology hint to bias decoding. Applied on the final transcript.
JSON-encoded array of up to 50 domain-specific terms (names, places,
brands, technical terms) to bias recognition toward, e.g.
keyterms=["Sarvam","New Delhi","Vistaar"]. Each keyterm can contain
up to 64 characters. Applied to the transcripts for the whole
session; only supported with model="saaras:v4". Do not use the
older keyterm or hotwords fields.
Turn detection strategy.
- vad (default): Voice Activity Detection automatically detects turn boundaries.
- manual: The client delimits turns with
speech_start/speech_endevents.
Audio encoding of the input stream. Mono only.
Audio sample rate in Hz. Must match the actual audio.
Accepted values:
800016000(default)
Any other value closes the connection with code 4000.
When true, adds segment-level start_s / end_s audio-time offsets
to transcript.final.
Speech-detection sensitivity (0.0–1.0; higher = needs louder/clearer
speech). Only applies when endpointing=vad.
Audio (ms, ≥0) included before the detected speech onset. Only applies
when endpointing=vad.
Silence (ms, ≥0) after speech that marks end-of-turn. Only applies when
endpointing=vad.
Minimum speech duration (ms, ≥0) for a valid utterance. Only applies
when endpointing=vad.