> For clean Markdown of any page, append `.md` to the page URL. > For a complete documentation index, see https://docs.sarvam.ai/llms.txt. > For full documentation content in one file, see https://docs.sarvam.ai/llms-full.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.sarvam.ai/_mcp/server. # Python For real-time voice, the Python SDK (`sarvam-conv-ai-sdk`) wraps the [WebSocket interface](/conversations/deploy/deploy-with-code#websocket-integration) in a typed client, `AsyncSamvaadAgent`, so you don't have to manage frames, reconnects, or message parsing yourself. ## Install ```bash pip install "sarvam-conv-ai-sdk[all]" ``` The `[all]` extra pulls in PyAudio for microphone capture and speaker playback, which needs PortAudio on the host: * **macOS**: `brew install portaudio` * **Ubuntu/Debian**: `sudo apt-get install portaudio19-dev` * **Windows**: install from [portaudio.com](http://www.portaudio.com/download.html) Skip the extra if you're bringing your own audio I/O (see [Headless mode](#headless-mode-bring-your-own-audio)); the base package works without PyAudio. ## Start a voice session ```python import asyncio from pydantic import SecretStr from sarvam_conv_ai_sdk import ( AsyncSamvaadAgent, AsyncDefaultAudioInterface, InteractionConfig, InteractionType, ServerTranscriptMsg, Role, ) from sarvam_conv_ai_sdk.messages.types import UserIdentifierType async def handle_transcript(msg: ServerTranscriptMsg): speaker = "User" if msg.role == Role.USER else "Agent" print(f"{speaker}: {msg.content}") async def main(app_id: str, api_key: str): config = InteractionConfig( user_identifier_type=UserIdentifierType.CUSTOM, user_identifier="demo_user", org_id="your_org_id", workspace_id="your_workspace_id", app_id=app_id, interaction_type=InteractionType.CALL, sample_rate=16000, ) agent = AsyncSamvaadAgent( api_key=SecretStr(api_key), config=config, audio_interface=AsyncDefaultAudioInterface(input_sample_rate=16000), transcript_callback=handle_transcript, ) await agent.start() await agent.wait_for_connect(timeout=5.0) await agent.wait_for_disconnect() await agent.stop() if __name__ == "__main__": asyncio.run(main(app_id="your_app_id", api_key="your_api_key")) ``` With `AsyncDefaultAudioInterface` attached, the SDK captures the microphone and plays agent audio back for you (16-bit PCM mono at 16kHz). `agent.start()` fetches a signed WebSocket URL, sends the session start message, and begins streaming in the background. ## InteractionConfig | Field | Required | Description | | ----------------------- | -------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `user_identifier_type` | Yes | `CUSTOM`, `EMAIL`, `PHONE_NUMBER`, or `UNKNOWN` | | `user_identifier` | Yes | The identifier value; also what you search by in [Log Analyser](/conversations/monitor/agent-analytics/log-analyser) | | `org_id` | Yes | Your organization ID | | `workspace_id` | Yes | Your workspace ID | | `app_id` | Yes | The agent to connect to | | `interaction_type` | Yes | `InteractionType.CALL` for a voice session | | `sample_rate` | Yes | `16000` (16-bit PCM mono) | | `version` | No | Pins a specific committed agent version. If omitted, the SDK uses the latest committed version, and the connection fails if the agent has no committed version | | `agent_variables` | No | Seed [agent variables](/conversations/build/variables-personalization) at session start | | `initial_language_name` | No | Starting language; must be one of the agent's allowed languages | | `initial_state_name` | No | Starting state, if the agent uses [states](/conversations/build/states-conversation-flow) | ## Callbacks Pass any of these to `AsyncSamvaadAgent` to react to what happens during the call: | Callback | Fires with | Use it for | | --------------------- | -------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------- | | `transcript_callback` | `ServerTranscriptMsg` (`role`: `Role.USER` or `Role.BOT`, `content`) | A live transcript of both sides of the call | | `audio_callback` | `ServerAudioChunkMsg` | Raw agent audio, if you're handling playback yourself instead of using `audio_interface` | | `event_callback` | `ServerEventBase` | Session events, such as `interaction_connected`, `user_interrupt` (barge-in), and `interaction_end` | ## Headless mode (bring your own audio) Omit `audio_interface` and push raw 16-bit PCM mono audio yourself, for example from a browser mic stream proxied through your backend, or a telephony leg: ```python agent = AsyncSamvaadAgent( api_key=SecretStr(api_key), config=config, transcript_callback=handle_transcript, ) await agent.start() await agent.wait_for_connect() await agent.send_audio(raw_pcm_bytes) # 16-bit PCM mono at config.sample_rate await agent.stop() ``` This is also the pattern for a backend proxy: terminate the caller's audio stream on your server, forward frames to `send_audio`, and relay `transcript_callback` / `audio_callback` output back to the client. ## Session methods | Method | Description | | ------------------------------------------- | ------------------------------------------------- | | `await agent.start()` | Fetch a signed WebSocket URL and connect | | `await agent.wait_for_connect(timeout=5.0)` | Block until the session is connected | | `await agent.send_audio(bytes)` | Send a raw PCM audio chunk | | `agent.is_connected()` | Current connection status | | `agent.get_interaction_id()` | The current interaction (call) ID, once connected | | `await agent.wait_for_disconnect()` | Block until the session ends | | `await agent.stop()` | Close the connection and clean up | > **Note** > > Call `await agent.stop()` in a `finally` block so the WebSocket and audio interface are cleaned up even if the session errors out. > **Warning** > > Never embed your API key in client-side code. Fetch the signed WebSocket URL from a backend you control, or run `AsyncSamvaadAgent` itself on your server and proxy audio to the frontend, as in the headless pattern above.