> For clean Markdown of any page, append `.md` to the page URL. > For a complete documentation index, see https://docs.sarvam.ai/llms.txt. > For full documentation content in one file, see https://docs.sarvam.ai/llms-full.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.sarvam.ai/_mcp/server. # Web (TypeScript) For real-time voice in the browser, the TypeScript SDK (`sarvam-conv-ai-sdk/browser`) wraps the [WebSocket interface](/conversations/deploy/deploy-with-code#websocket-integration) in a typed client, `ConversationAgent`, with `BrowserAudioInterface` for microphone capture and playback. ## Install ```bash npm install sarvam-conv-ai-sdk ``` Use the `/browser` entry point for a smaller bundle: ```typescript import { ConversationAgent, BrowserAudioInterface, } from "sarvam-conv-ai-sdk/browser"; ``` ## Start a voice session ```typescript import { ConversationAgent, BrowserAudioInterface, InteractionType, } from "sarvam-conv-ai-sdk/browser"; const audioInterface = new BrowserAudioInterface(); const agent = new ConversationAgent({ apiKey: "your_api_key", config: { org_id: "your_org_id", workspace_id: "your_workspace_id", app_id: "your_app_id", user_identifier: "user123", user_identifier_type: "custom", interaction_type: InteractionType.CALL, input_sample_rate: 16000, output_sample_rate: 16000, }, audioInterface, transcriptCallback: async (msg) => { console.log(`${msg.role}: ${msg.content}`); }, }); await agent.start(); await agent.waitForConnect(10); ``` `BrowserAudioInterface` captures the microphone and plays agent audio (16-bit PCM mono). `agent.start()` fetches a signed WebSocket URL, sends the session start message, and begins streaming. Call `start()` from a user gesture (for example a button click) so the browser allows audio playback. ## InteractionConfig | Field | Required | Description | | ----------------------- | -------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `user_identifier_type` | Yes | `'custom'`, `'email'`, `'phone_number'`, or `'unknown'` | | `user_identifier` | Yes | The identifier value; also what you search by in [Log Analyser](/conversations/monitor/agent-analytics/log-analyser) | | `org_id` | Yes | Your organization ID | | `workspace_id` | Yes | Your workspace ID | | `app_id` | Yes | The agent to connect to | | `interaction_type` | Yes | `InteractionType.CALL` for a voice session | | `input_sample_rate` | Yes | `8000` or `16000` Hz | | `output_sample_rate` | Yes | `16000` or `22050` Hz | | `version` | No | Pins a specific committed agent version. If omitted, the SDK uses the latest committed version, and the connection fails if the agent has no committed version | | `agent_variables` | No | Seed [agent variables](/conversations/build/variables-personalization) at session start | | `initial_language_name` | No | Starting language; must be one of the agent's allowed languages | | `initial_state_name` | No | Starting state, if the agent uses [states](/conversations/build/states-conversation-flow) | | `initial_bot_message` | No | First message from the agent | ## Callbacks Pass any of these to `ConversationAgent` to react to what happens during the call: | Callback | Fires with | Use it for | | -------------------- | ----------------------------------------------------------------- | ------------------------------------------------------------------------------------ | | `transcriptCallback` | transcript message (`role`: `Role.USER` or `Role.BOT`, `content`) | A live transcript of both sides of the call | | `audioCallback` | audio chunk | Raw agent audio, if you're handling playback yourself | | `audioLevelCallback` | output level | Real-time level for a speaking indicator | | `eventCallback` | session event | Session events such as connect, barge-in, and end | | `stateCallback` | `AgentState` | Drive UI from `IDLE`, `CONNECTING`, `CONNECTED`, `LISTENING`, `SPEAKING`, or `ERROR` | | `startCallback` | — | Session start | | `endCallback` | — | Session end and cleanup | ## BrowserAudioInterface ```typescript new BrowserAudioInterface( sampleRate?: number, // Default: 16000 options?: { prebufferMs?: number; // Default: 700 bufferSizeMs?: number; // Default: 8000 outputLevelCallback?: (level: { rms: number; peak: number; db: number }) => void; } ) ``` `BrowserAudioInterface` uses adaptive buffering to handle network jitter. Adjust `prebufferMs` for your network: | Network condition | `prebufferMs` | Notes | | ----------------- | --------------- | --------------------------------- | | Good/stable | `300`–`500` | Lower latency; may gap on jitter | | Normal | `700` (default) | Balanced latency and stability | | Unstable | `1000`–`1500` | Higher latency, smoother playback | Use `audioLevelCallback` on the agent (or `outputLevelCallback` on the interface) to drive a speaking indicator: ```typescript const agent = new ConversationAgent({ // ...config audioLevelCallback: (level) => { updateAgentSpeakingIndicator(level); }, }); ``` ## Session methods | Method | Description | | -------------------------------------- | ----------------------------------------------------------------------------------------------------------- | | `await agent.start()` | Fetch a signed WebSocket URL and connect | | `await agent.waitForConnect(timeout?)` | Wait until connected; `timeout` is in seconds (returns `false` on timeout) | | `await agent.sendAudio(data)` | Send a raw PCM audio chunk | | `agent.isConnected()` | Current connection status | | `agent.getInteractionId()` | The current interaction (call) ID, once connected | | `agent.getState()` | Current `AgentState` | | `agent.mute()` / `agent.unmute()` | Mute or unmute the microphone without disconnecting. While muted, the SDK sends silence so VAD stays stable | | `agent.isMuted()` | Current mute status | | `await agent.waitForDisconnect()` | Wait until the session ends | | `await agent.stop()` | Close the connection and clean up | > **Note** > > Call `await agent.stop()` on unmount (or in a `finally` block) so the WebSocket and audio interface are cleaned up even if the session errors out. Reconnection is not supported: each WebSocket URL is single-use. If the connection drops, call `stop()` and create a new `ConversationAgent`. ## React example ```typescript import React, { useRef, useState, useEffect } from "react"; import { ConversationAgent, BrowserAudioInterface, InteractionType, AgentState, Role, } from "sarvam-conv-ai-sdk/browser"; interface Message { role: "user" | "bot"; content: string; } function VoiceChat() { const [state, setState] = useState(AgentState.IDLE); const [messages, setMessages] = useState([]); const [isMuted, setIsMuted] = useState(false); const agentRef = useRef(null); useEffect(() => { return () => { agentRef.current?.stop().catch(console.error); }; }, []); const startConversation = async () => { const audioInterface = new BrowserAudioInterface(16000, { outputLevelCallback: (level) => { // Update volume visualization }, }); const agent = new ConversationAgent({ apiKey: "your_api_key", config: { org_id: "your_org_id", workspace_id: "your_workspace_id", app_id: "your_app_id", user_identifier: "user123", user_identifier_type: "custom", interaction_type: InteractionType.CALL, input_sample_rate: 16000, output_sample_rate: 16000, }, audioInterface, stateCallback: (newState) => setState(newState), transcriptCallback: async (msg) => { setMessages((prev) => [ ...prev, { role: msg.role === Role.USER ? "user" : "bot", content: msg.content }, ]); }, endCallback: async () => { agentRef.current = null; setState(AgentState.IDLE); }, }); agentRef.current = agent; await agent.start(); }; const toggleMute = () => { if (!agentRef.current) return; if (agentRef.current.isMuted()) { agentRef.current.unmute(); setIsMuted(false); } else { agentRef.current.mute(); setIsMuted(true); } }; return (

State: {state}

{messages.map((msg, i) => (

{msg.role === "user" ? "You" : "Agent"}: {msg.content}

))}
); } export default VoiceChat; ``` ## Troubleshooting **`audioInterface` is required for CALL interactions.** Pass a `BrowserAudioInterface`: ```typescript const agent = new ConversationAgent({ audioInterface: new BrowserAudioInterface(), // ... }); ``` **Audio doesn't play in Firefox.** Firefox requires a user gesture before playing audio. Call `start()` from a click handler. **HTTPS is required for the microphone** in production. `localhost` is enough for development. **Microphone permission errors:** ```typescript try { const audioInterface = new BrowserAudioInterface(); const agent = new ConversationAgent({ audioInterface, // ... }); await agent.start(); } catch (err) { if (err.name === "NotAllowedError") { console.error("Microphone permission denied"); } else if (err.name === "NotFoundError") { console.error("No microphone found"); } else if (err.name === "NotReadableError") { console.error("Microphone is in use"); } } ``` > **Warning** > > Never embed your API key in client-side code. Point `baseUrl` at a backend you control and leave `apiKey` empty so the proxy adds the Sarvam key server-side. ```typescript const agent = new ConversationAgent({ apiKey: "", baseUrl: "/api/sarvam/", config: { // ... }, audioInterface: new BrowserAudioInterface(), }); ``` For a cross-origin proxy, pass your app's auth in `customHeaders` (for example `Authorization: Bearer` plus your session token).