Deploy with Code (API & SDK)
Deploy with Code (API & SDK)
Integrate Voice Agents programmatically with APIs and SDKs. For platform speech APIs, see the API Reference.
What will be covered
Authenticate all requests with your API key.
Getting started (preview)
Use the dashboard API recipes as a starting point:
- Create or update an agent configuration: set up agents programmatically instead of through the UI
- Start a conversation session: initiate a voice or text session via API
- Trigger outbound calls: kick off calls to specific numbers with variable payloads
- Fetch transcripts and outcomes: pull conversation data into your systems for analysis or CRM updates
WebSocket integration
For real-time voice sessions, connect over the Voice Agents WebSocket interface. Handle audio frames, turn events, and tool-call signals according to the API contract.
Never expose production API keys in browser code. Keep secrets server-side and proxy WebSocket connections through your backend when building client-facing applications.
Embed the agent (widget)
To put a voice or chat agent on your website, embed the widget:
- Open your agent and go to the Embed option in Deploy with Code.
- Copy the embed snippet from the dashboard.
- Paste it into your site, typically before
</body>. - Load the page and confirm the widget initializes.
Customize the launcher position, theme, and which agent the widget uses from the dashboard, and re-test in test agent after changes.
SDK usage
Python
Web (TypeScript)
React Native
Flutter
For real-time voice, the Python SDK (sarvam-conv-ai-sdk) wraps the WebSocket interface above in a typed client, AsyncSamvaadAgent, so you don’t have to manage frames, reconnects, or message parsing yourself.
Install
The [all] extra pulls in PyAudio for microphone capture and speaker playback, which needs PortAudio on the host:
- macOS:
brew install portaudio - Ubuntu/Debian:
sudo apt-get install portaudio19-dev - Windows: install from portaudio.com
Skip the extra if you’re bringing your own audio I/O (see Headless mode); the base package works without PyAudio.
Start a voice session
With AsyncDefaultAudioInterface attached, the SDK captures the microphone and plays agent audio back for you (16-bit PCM mono at 16kHz). agent.start() fetches a signed WebSocket URL, sends the session start message, and begins streaming in the background.
InteractionConfig
Callbacks
Pass any of these to AsyncSamvaadAgent to react to what happens during the call:
Headless mode (bring your own audio)
Omit audio_interface and push raw 16-bit PCM mono audio yourself, for example from a browser mic stream proxied through your backend, or a telephony leg:
This is also the pattern for a backend proxy: terminate the caller’s audio stream on your server, forward frames to send_audio, and relay transcript_callback / audio_callback output back to the client.
Session methods
Call await agent.stop() in a finally block so the WebSocket and audio interface are cleaned up even if the session errors out.
Never embed your API key in client-side code. Fetch the signed WebSocket URL from a backend you control, or run AsyncSamvaadAgent itself on your server and proxy audio to the frontend, as in the headless pattern above.