Skip to navigation

Realtime Streaming

View as Markdown

WebSocket channel for real-time speech to text streaming powered by the saaras:v3-realtime (default) or saaras:v4 model.

Authentication: Pass your API subscription key via the API-SUBSCRIPTION-KEY header, or (in browsers) via the WebSocket subprotocol api-subscription-key.<key> — the server echoes the subprotocol back.

Note: This API Reference page is provided for informational purposes only. The Try It playground may not provide the best experience for streaming audio. For optimal streaming performance, please use the SDK or implement your own WebSocket client.

Handshake

WSS
wss://api.sarvam.ai/speech-to-text-realtime/ws

Headers

Api-Subscription-KeystringRequired
API subscription key for authentication

Query parameters

language_codeenumRequired

BCP-47 language code of the input audio. Required.

Supported languages (saaras:v3-realtime and saaras:v4):

  • auto: Adaptive automatic language detection
  • en-IN: English
  • hi-IN: Hindi
  • bn-IN: Bengali
  • kn-IN: Kannada
  • ml-IN: Malayalam
  • mr-IN: Marathi
  • or-IN: Odia
  • pa-IN: Punjabi
  • ta-IN: Tamil
  • te-IN: Telugu
  • gu-IN: Gujarati
  • as-IN: Assamese
  • ur-IN: Urdu
  • ne-IN: Nepali
  • kok-IN: Konkani
  • ks-IN: Kashmiri
  • sd-IN: Sindhi
  • sa-IN: Sanskrit
  • sat-IN: Santali
  • mni-IN: Manipuri
  • brx-IN: Bodo
  • mai-IN: Maithili
  • doi-IN: Dogri
modelenumOptionalDefaults to saaras:v3-realtime

Speech-to-text model for this endpoint. saaras:v3-realtime (default) or saaras:v4.

Allowed values:
stream_typeenumOptionalDefaults to balanced

Controls audio chunking and latency behaviour.

  • fast: Partials at low latency.
  • balanced (default): Partials with better accuracy.
  • simulated: No partials.
Allowed values:
modeenumOptionalDefaults to transcribe

Task applied on the final transcript, not on partial transcripts. Streaming partials are always straight transcription regardless of this value.

  • transcribe (default): Standard transcription in the source language script.
  • translate: Translation to English.
  • verbatim: Exact transcription in the original script (includes disfluencies).
  • translit: Transliteration to Latin script.
  • codemix: Code-mixed output (native + English).
Allowed values:
promptstringOptional

Optional context/terminology hint to bias decoding. Applied on the final transcript.

keytermsstringOptional

JSON-encoded array of up to 50 domain-specific terms (names, places, brands, technical terms) to bias recognition toward, e.g. keyterms=["Sarvam","New Delhi","Vistaar"]. Each keyterm can contain up to 64 characters. Applied to the transcripts for the whole session; only supported with model="saaras:v4". Do not use the older keyterm or hotwords fields.

endpointingenumOptionalDefaults to vad

Turn detection strategy.

  • vad (default): Voice Activity Detection automatically detects turn boundaries.
  • manual: The client delimits turns with speech_start / speech_end events.
Allowed values:
encodingenumOptionalDefaults to linear16

Audio encoding of the input stream. Mono only.

Allowed values:
sample_rateenumOptionalDefaults to 16000

Audio sample rate in Hz. Must match the actual audio.

Accepted values:

  • 8000
  • 16000 (default)

Any other value closes the connection with code 4000.

Allowed values:
return_timestampsenumOptionalDefaults to false

When true, adds segment-level start_s / end_s audio-time offsets to transcript.final.

Allowed values:
thresholdstringOptionalDefaults to 0.3

Speech-detection sensitivity (0.0–1.0; higher = needs louder/clearer speech). Only applies when endpointing=vad.

prefix_padding_msstringOptionalDefaults to 300

Audio (ms, ≥0) included before the detected speech onset. Only applies when endpointing=vad.

silence_duration_msstringOptionalDefaults to 500

Silence (ms, ≥0) after speech that marks end-of-turn. Only applies when endpointing=vad.

min_speech_duration_msstringOptionalDefaults to 250

Minimum speech duration (ms, ≥0) for a valid utterance. Only applies when endpointing=vad.

Send

Realtime Audio InputobjectRequired
Send an audio chunk for real-time transcription (saaras:v3-realtime, saaras:v4)
OR
Realtime Speech StartobjectRequired
Signal the start of an utterance (endpointing=manual only)
OR
Realtime Speech EndobjectRequired
Signal the end of an utterance (endpointing=manual only)
OR
Realtime FlushobjectRequired
OR
Realtime Config UpdateobjectRequired
Update session config mid-stream
OR
Realtime End SessionobjectRequired
Gracefully end the realtime session
OR
Realtime PingobjectRequired
Send a keepalive ping to the realtime WebSocket

Receive

Realtime Session BeginobjectRequired
Receive the session.begin event with the resolved session config
OR
Realtime VAD EventsobjectRequired
Receive VAD speech start/end events (endpointing=vad only)
OR
Realtime VAD EventsobjectRequired
Receive VAD speech start/end events (endpointing=vad only)
OR
Realtime Partial TranscriptobjectRequired
Receive streaming partial transcripts during an utterance
OR
Realtime Final TranscriptobjectRequired
Receive the final transcript for a completed utterance
OR
Realtime Config UpdatedobjectRequired
Receive the config.updated acknowledgement after a config.update
OR
Realtime PongobjectRequired
Receive the pong response to a client ping
OR
Realtime Session EndobjectRequired
Receive the session.end summary when the session closes
OR
Realtime ErrorobjectRequired
Receive non-fatal or fatal error notifications