> For clean Markdown of any page, append `.md` to the page URL.
> For a complete documentation index, see https://docs.sarvam.ai/llms.txt.
> For full documentation content in one file, see https://docs.sarvam.ai/llms-full.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.sarvam.ai/_mcp/server.

# Python

For real-time voice, the Python SDK (`sarvam-conv-ai-sdk`) wraps the [WebSocket interface](/conversations/deploy/deploy-with-code#websocket-integration) in a typed client, `AsyncSamvaadAgent`, so you don't have to manage frames, reconnects, or message parsing yourself.

## Install

```bash
pip install "sarvam-conv-ai-sdk[all]"
```

The `[all]` extra pulls in PyAudio for microphone capture and speaker playback, which needs PortAudio on the host:

* **macOS**: `brew install portaudio`
* **Ubuntu/Debian**: `sudo apt-get install portaudio19-dev`
* **Windows**: install from [portaudio.com](http://www.portaudio.com/download.html)

Skip the extra if you're bringing your own audio I/O (see [Headless mode](#headless-mode-bring-your-own-audio)); the base package works without PyAudio.

## Start a voice session

```python
import asyncio
from pydantic import SecretStr
from sarvam_conv_ai_sdk import (
    AsyncSamvaadAgent,
    AsyncDefaultAudioInterface,
    InteractionConfig,
    InteractionType,
    ServerTranscriptMsg,
    Role,
)
from sarvam_conv_ai_sdk.messages.types import UserIdentifierType

async def handle_transcript(msg: ServerTranscriptMsg):
    speaker = "User" if msg.role == Role.USER else "Agent"
    print(f"{speaker}: {msg.content}")

async def main(app_id: str, api_key: str):
    config = InteractionConfig(
        user_identifier_type=UserIdentifierType.CUSTOM,
        user_identifier="demo_user",
        org_id="your_org_id",
        workspace_id="your_workspace_id",
        app_id=app_id,
        interaction_type=InteractionType.CALL,
        sample_rate=16000,
    )

    agent = AsyncSamvaadAgent(
        api_key=SecretStr(api_key),
        config=config,
        audio_interface=AsyncDefaultAudioInterface(input_sample_rate=16000),
        transcript_callback=handle_transcript,
    )

    await agent.start()
    await agent.wait_for_connect(timeout=5.0)
    await agent.wait_for_disconnect()
    await agent.stop()

if __name__ == "__main__":
    asyncio.run(main(app_id="your_app_id", api_key="your_api_key"))
```

With `AsyncDefaultAudioInterface` attached, the SDK captures the microphone and plays agent audio back for you (16-bit PCM mono at 16kHz). `agent.start()` fetches a signed WebSocket URL, sends the session start message, and begins streaming in the background.

## InteractionConfig

| Field                   | Required | Description                                                                                                                                                    |
| ----------------------- | -------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `user_identifier_type`  | Yes      | `CUSTOM`, `EMAIL`, `PHONE_NUMBER`, or `UNKNOWN`                                                                                                                |
| `user_identifier`       | Yes      | The identifier value; also what you search by in [Log Analyser](/conversations/monitor/agent-analytics/log-analyser)                                           |
| `org_id`                | Yes      | Your organization ID                                                                                                                                           |
| `workspace_id`          | Yes      | Your workspace ID                                                                                                                                              |
| `app_id`                | Yes      | The agent to connect to                                                                                                                                        |
| `interaction_type`      | Yes      | `InteractionType.CALL` for a voice session                                                                                                                     |
| `sample_rate`           | Yes      | `16000` (16-bit PCM mono)                                                                                                                                      |
| `version`               | No       | Pins a specific committed agent version. If omitted, the SDK uses the latest committed version, and the connection fails if the agent has no committed version |
| `agent_variables`       | No       | Seed [agent variables](/conversations/build/variables-personalization) at session start                                                                        |
| `initial_language_name` | No       | Starting language; must be one of the agent's allowed languages                                                                                                |
| `initial_state_name`    | No       | Starting state, if the agent uses [states](/conversations/build/states-conversation-flow)                                                                      |

## Callbacks

Pass any of these to `AsyncSamvaadAgent` to react to what happens during the call:

| Callback              | Fires with                                                           | Use it for                                                                                          |
| --------------------- | -------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------- |
| `transcript_callback` | `ServerTranscriptMsg` (`role`: `Role.USER` or `Role.BOT`, `content`) | A live transcript of both sides of the call                                                         |
| `audio_callback`      | `ServerAudioChunkMsg`                                                | Raw agent audio, if you're handling playback yourself instead of using `audio_interface`            |
| `event_callback`      | `ServerEventBase`                                                    | Session events, such as `interaction_connected`, `user_interrupt` (barge-in), and `interaction_end` |

## Headless mode (bring your own audio)

Omit `audio_interface` and push raw 16-bit PCM mono audio yourself, for example from a browser mic stream proxied through your backend, or a telephony leg:

```python
agent = AsyncSamvaadAgent(
    api_key=SecretStr(api_key),
    config=config,
    transcript_callback=handle_transcript,
)

await agent.start()
await agent.wait_for_connect()

await agent.send_audio(raw_pcm_bytes)  # 16-bit PCM mono at config.sample_rate

await agent.stop()
```

This is also the pattern for a backend proxy: terminate the caller's audio stream on your server, forward frames to `send_audio`, and relay `transcript_callback` / `audio_callback` output back to the client.

## Session methods

| Method                                      | Description                                       |
| ------------------------------------------- | ------------------------------------------------- |
| `await agent.start()`                       | Fetch a signed WebSocket URL and connect          |
| `await agent.wait_for_connect(timeout=5.0)` | Block until the session is connected              |
| `await agent.send_audio(bytes)`             | Send a raw PCM audio chunk                        |
| `agent.is_connected()`                      | Current connection status                         |
| `agent.get_interaction_id()`                | The current interaction (call) ID, once connected |
| `await agent.wait_for_disconnect()`         | Block until the session ends                      |
| `await agent.stop()`                        | Close the connection and clean up                 |

> **Note**
>
> Call `await agent.stop()` in a `finally` block so the WebSocket and audio interface are cleaned up even if the session errors out.

> **Warning**
>
> Never embed your API key in client-side code. Fetch the signed WebSocket URL from a backend you control, or run `AsyncSamvaadAgent` itself on your server and proxy audio to the frontend, as in the headless pattern above.