Deploy with Code (API & SDK)

View as Markdown

Integrate Voice Agents programmatically with APIs and SDKs. For platform speech APIs, see the API Reference.

What will be covered

TopicDescription
API recipesReady-to-use code for common operations, creating agents, starting sessions, triggering outbound calls, fetching transcripts
WebSocket integrationReal-time voice sessions with audio frames, turn events, and tool-call signals
SDK usageTyped clients for real-time voice conversations: audio streaming, transcripts, and session lifecycle, without managing the WebSocket protocol yourself. Python, Web (TypeScript), React Native, and Flutter are available
AuthenticationAPI key usage and best practices for keeping secrets server-side

Authenticate all requests with your API key.

Getting started (preview)

Use the dashboard API recipes as a starting point:

  1. Create or update an agent configuration: set up agents programmatically instead of through the UI
  2. Start a conversation session: initiate a voice or text session via API
  3. Trigger outbound calls: kick off calls to specific numbers with variable payloads
  4. Fetch transcripts and outcomes: pull conversation data into your systems for analysis or CRM updates

WebSocket integration

For real-time voice sessions, connect over the Voice Agents WebSocket interface. Handle audio frames, turn events, and tool-call signals according to the API contract.

Never expose production API keys in browser code. Keep secrets server-side and proxy WebSocket connections through your backend when building client-facing applications.

Embed the agent (widget)

To put a voice or chat agent on your website, embed the widget:

  1. Open your agent and go to the Embed option in Deploy with Code.
  2. Copy the embed snippet from the dashboard.
  3. Paste it into your site, typically before </body>.
  4. Load the page and confirm the widget initializes.

Customize the launcher position, theme, and which agent the widget uses from the dashboard, and re-test in test agent after changes.

SDK usage

For real-time voice, the Python SDK (sarvam-conv-ai-sdk) wraps the WebSocket interface above in a typed client, AsyncSamvaadAgent, so you don’t have to manage frames, reconnects, or message parsing yourself.

Install

pip install "sarvam-conv-ai-sdk[all]"

The [all] extra pulls in PyAudio for microphone capture and speaker playback, which needs PortAudio on the host:

  • macOS: brew install portaudio
  • Ubuntu/Debian: sudo apt-get install portaudio19-dev
  • Windows: install from portaudio.com

Skip the extra if you’re bringing your own audio I/O (see Headless mode); the base package works without PyAudio.

Start a voice session

import asyncio
from pydantic import SecretStr
from sarvam_conv_ai_sdk import (
AsyncSamvaadAgent,
AsyncDefaultAudioInterface,
InteractionConfig,
InteractionType,
ServerTranscriptMsg,
Role,
)
from sarvam_conv_ai_sdk.messages.types import UserIdentifierType
async def handle_transcript(msg: ServerTranscriptMsg):
speaker = "User" if msg.role == Role.USER else "Agent"
print(f"{speaker}: {msg.content}")
async def main(app_id: str, api_key: str):
config = InteractionConfig(
user_identifier_type=UserIdentifierType.CUSTOM,
user_identifier="demo_user",
org_id="your_org_id",
workspace_id="your_workspace_id",
app_id=app_id,
interaction_type=InteractionType.CALL,
sample_rate=16000,
)
agent = AsyncSamvaadAgent(
api_key=SecretStr(api_key),
config=config,
audio_interface=AsyncDefaultAudioInterface(input_sample_rate=16000),
transcript_callback=handle_transcript,
)
await agent.start()
await agent.wait_for_connect(timeout=5.0)
await agent.wait_for_disconnect()
await agent.stop()
if __name__ == "__main__":
asyncio.run(main(app_id="your_app_id", api_key="your_api_key"))

With AsyncDefaultAudioInterface attached, the SDK captures the microphone and plays agent audio back for you (16-bit PCM mono at 16kHz). agent.start() fetches a signed WebSocket URL, sends the session start message, and begins streaming in the background.

InteractionConfig

FieldRequiredDescription
user_identifier_typeYesCUSTOM, EMAIL, PHONE_NUMBER, or UNKNOWN
user_identifierYesThe identifier value; also what you search by in Log Analyser
org_idYesYour organization ID
workspace_idYesYour workspace ID
app_idYesThe agent to connect to
interaction_typeYesInteractionType.CALL for a voice session
sample_rateYes16000 (16-bit PCM mono)
versionNoPins a specific committed agent version. If omitted, the SDK uses the latest committed version, and the connection fails if the agent has no committed version
agent_variablesNoSeed agent variables at session start
initial_language_nameNoStarting language; must be one of the agent’s allowed languages
initial_state_nameNoStarting state, if the agent uses states

Callbacks

Pass any of these to AsyncSamvaadAgent to react to what happens during the call:

CallbackFires withUse it for
transcript_callbackServerTranscriptMsg (role: Role.USER or Role.BOT, content)A live transcript of both sides of the call
audio_callbackServerAudioChunkMsgRaw agent audio, if you’re handling playback yourself instead of using audio_interface
event_callbackServerEventBaseSession events, such as interaction_connected, user_interrupt (barge-in), and interaction_end

Headless mode (bring your own audio)

Omit audio_interface and push raw 16-bit PCM mono audio yourself, for example from a browser mic stream proxied through your backend, or a telephony leg:

agent = AsyncSamvaadAgent(
api_key=SecretStr(api_key),
config=config,
transcript_callback=handle_transcript,
)
await agent.start()
await agent.wait_for_connect()
await agent.send_audio(raw_pcm_bytes) # 16-bit PCM mono at config.sample_rate
await agent.stop()

This is also the pattern for a backend proxy: terminate the caller’s audio stream on your server, forward frames to send_audio, and relay transcript_callback / audio_callback output back to the client.

Session methods

MethodDescription
await agent.start()Fetch a signed WebSocket URL and connect
await agent.wait_for_connect(timeout=5.0)Block until the session is connected
await agent.send_audio(bytes)Send a raw PCM audio chunk
agent.is_connected()Current connection status
agent.get_interaction_id()The current interaction (call) ID, once connected
await agent.wait_for_disconnect()Block until the session ends
await agent.stop()Close the connection and clean up

Call await agent.stop() in a finally block so the WebSocket and audio interface are cleaned up even if the session errors out.

Never embed your API key in client-side code. Fetch the signed WebSocket URL from a backend you control, or run AsyncSamvaadAgent itself on your server and proxy audio to the frontend, as in the headless pattern above.