Python
For real-time voice, the Python SDK (sarvam-conv-ai-sdk) wraps the WebSocket interface in a typed client, AsyncSamvaadAgent, so you don’t have to manage frames, reconnects, or message parsing yourself.
Install
The [all] extra pulls in PyAudio for microphone capture and speaker playback, which needs PortAudio on the host:
- macOS:
brew install portaudio - Ubuntu/Debian:
sudo apt-get install portaudio19-dev - Windows: install from portaudio.com
Skip the extra if you’re bringing your own audio I/O (see Headless mode); the base package works without PyAudio.
Start a voice session
With AsyncDefaultAudioInterface attached, the SDK captures the microphone and plays agent audio back for you (16-bit PCM mono at 16kHz). agent.start() fetches a signed WebSocket URL, sends the session start message, and begins streaming in the background.
InteractionConfig
Callbacks
Pass any of these to AsyncSamvaadAgent to react to what happens during the call:
Headless mode (bring your own audio)
Omit audio_interface and push raw 16-bit PCM mono audio yourself, for example from a browser mic stream proxied through your backend, or a telephony leg:
This is also the pattern for a backend proxy: terminate the caller’s audio stream on your server, forward frames to send_audio, and relay transcript_callback / audio_callback output back to the client.
Session methods
Call await agent.stop() in a finally block so the WebSocket and audio interface are cleaned up even if the session errors out.
Never embed your API key in client-side code. Fetch the signed WebSocket URL from a backend you control, or run AsyncSamvaadAgent itself on your server and proxy audio to the frontend, as in the headless pattern above.