WebSocket
WebSocket channel for real-time speech to text streaming with English translation.
Note: This API Reference page is provided for informational purposes only. The Try It playground may not provide the best experience for streaming audio. For optimal streaming performance, please use the SDK or implement your own WebSocket client.
Handshake
Headers
Query parameters
Model to be used for speech to text translation.
- saaras:v2.5 (default): Translation model that translates audio from any spoken Indic language to English.
- Example: Hindi audio → English text output
For the latest model (saaras:v3), use the /speech-to-text endpoint with mode="translate".
Enable high VAD (Voice Activity Detection) sensitivity
VAD probability threshold (0.0–1.0) above which a frame is considered speech. Overrides the server default when provided.
VAD probability threshold (0.0–1.0) below which a frame is considered silence. Overrides the server default (or the high_vad_sensitivity preset) when provided.
Number of negative (silence) frames needed within the window to end a speech segment. Overrides the server default (or the high_vad_sensitivity preset) when provided.
Sliding window size (in frames) over which negative frames are counted. Overrides the server default (or the high_vad_sensitivity preset) when provided.
Volume level (dB) below which audio is considered too quiet to be speech. When not provided, no volume-based filtering is applied.
Minimum speech frames required to register a barge-in / interruption. Overrides the server default when provided.
Audio codec/format of the input stream. Use this when sending raw PCM audio. Supported values: wav, pcm_s16le, pcm_l16, pcm_raw.