Skip to navigation

Python

View as Markdown

For real-time voice, the Python SDK (sarvam-conv-ai-sdk) wraps the WebSocket interface in a typed client, AsyncSamvaadAgent, so you don’t have to manage frames, reconnects, or message parsing yourself.

Install

pip install "sarvam-conv-ai-sdk[all]"

The [all] extra pulls in PyAudio for microphone capture and speaker playback, which needs PortAudio on the host:

  • macOS: brew install portaudio
  • Ubuntu/Debian: sudo apt-get install portaudio19-dev
  • Windows: install from portaudio.com

Skip the extra if you’re bringing your own audio I/O (see Headless mode); the base package works without PyAudio.

Start a voice session

import asyncio
from pydantic import SecretStr
from sarvam_conv_ai_sdk import (
AsyncSamvaadAgent,
AsyncDefaultAudioInterface,
InteractionConfig,
InteractionType,
ServerTranscriptMsg,
Role,
)
from sarvam_conv_ai_sdk.messages.types import UserIdentifierType
async def handle_transcript(msg: ServerTranscriptMsg):
speaker = "User" if msg.role == Role.USER else "Agent"
print(f"{speaker}: {msg.content}")
async def main(app_id: str, api_key: str):
config = InteractionConfig(
user_identifier_type=UserIdentifierType.CUSTOM,
user_identifier="demo_user",
org_id="your_org_id",
workspace_id="your_workspace_id",
app_id=app_id,
interaction_type=InteractionType.CALL,
sample_rate=16000,
)
agent = AsyncSamvaadAgent(
api_key=SecretStr(api_key),
config=config,
audio_interface=AsyncDefaultAudioInterface(input_sample_rate=16000),
transcript_callback=handle_transcript,
)
await agent.start()
await agent.wait_for_connect(timeout=5.0)
await agent.wait_for_disconnect()
await agent.stop()
if __name__ == "__main__":
asyncio.run(main(app_id="your_app_id", api_key="your_api_key"))

With AsyncDefaultAudioInterface attached, the SDK captures the microphone and plays agent audio back for you (16-bit PCM mono at 16kHz). agent.start() fetches a signed WebSocket URL, sends the session start message, and begins streaming in the background.

InteractionConfig

FieldRequiredDescription
user_identifier_typeYesCUSTOM, EMAIL, PHONE_NUMBER, or UNKNOWN
user_identifierYesThe identifier value; also what you search by in Log Analyser
org_idYesYour organization ID
workspace_idYesYour workspace ID
app_idYesThe agent to connect to
interaction_typeYesInteractionType.CALL for a voice session
sample_rateYes16000 (16-bit PCM mono)
versionNoPins a specific committed agent version. If omitted, the SDK uses the latest committed version, and the connection fails if the agent has no committed version
agent_variablesNoSeed agent variables at session start
initial_language_nameNoStarting language; must be one of the agent’s allowed languages
initial_state_nameNoStarting state, if the agent uses states

Callbacks

Pass any of these to AsyncSamvaadAgent to react to what happens during the call:

CallbackFires withUse it for
transcript_callbackServerTranscriptMsg (role: Role.USER or Role.BOT, content)A live transcript of both sides of the call
audio_callbackServerAudioChunkMsgRaw agent audio, if you’re handling playback yourself instead of using audio_interface
event_callbackServerEventBaseSession events, such as interaction_connected, user_interrupt (barge-in), and interaction_end

Headless mode (bring your own audio)

Omit audio_interface and push raw 16-bit PCM mono audio yourself, for example from a browser mic stream proxied through your backend, or a telephony leg:

agent = AsyncSamvaadAgent(
api_key=SecretStr(api_key),
config=config,
transcript_callback=handle_transcript,
)
await agent.start()
await agent.wait_for_connect()
await agent.send_audio(raw_pcm_bytes) # 16-bit PCM mono at config.sample_rate
await agent.stop()

This is also the pattern for a backend proxy: terminate the caller’s audio stream on your server, forward frames to send_audio, and relay transcript_callback / audio_callback output back to the client.

Session methods

MethodDescription
await agent.start()Fetch a signed WebSocket URL and connect
await agent.wait_for_connect(timeout=5.0)Block until the session is connected
await agent.send_audio(bytes)Send a raw PCM audio chunk
agent.is_connected()Current connection status
agent.get_interaction_id()The current interaction (call) ID, once connected
await agent.wait_for_disconnect()Block until the session ends
await agent.stop()Close the connection and clean up

Call await agent.stop() in a finally block so the WebSocket and audio interface are cleaned up even if the session errors out.

Never embed your API key in client-side code. Fetch the signed WebSocket URL from a backend you control, or run AsyncSamvaadAgent itself on your server and proxy audio to the frontend, as in the headless pattern above.