Web (TypeScript)
For real-time voice in the browser, the TypeScript SDK (sarvam-conv-ai-sdk/browser) wraps the WebSocket interface in a typed client, ConversationAgent, with BrowserAudioInterface for microphone capture and playback.
Install
Use the /browser entry point for a smaller bundle:
Start a voice session
BrowserAudioInterface captures the microphone and plays agent audio (16-bit PCM mono). agent.start() fetches a signed WebSocket URL, sends the session start message, and begins streaming. Call start() from a user gesture (for example a button click) so the browser allows audio playback.
InteractionConfig
Callbacks
Pass any of these to ConversationAgent to react to what happens during the call:
BrowserAudioInterface
BrowserAudioInterface uses adaptive buffering to handle network jitter. Adjust prebufferMs for your network:
Use audioLevelCallback on the agent (or outputLevelCallback on the interface) to drive a speaking indicator:
Session methods
Call await agent.stop() on unmount (or in a finally block) so the WebSocket and audio interface are cleaned up even if the session errors out. Reconnection is not supported: each WebSocket URL is single-use. If the connection drops, call stop() and create a new ConversationAgent.
React example
Troubleshooting
audioInterface is required for CALL interactions. Pass a BrowserAudioInterface:
Audio doesn’t play in Firefox. Firefox requires a user gesture before playing audio. Call start() from a click handler.
HTTPS is required for the microphone in production. localhost is enough for development.
Microphone permission errors:
Never embed your API key in client-side code. Point baseUrl at a backend you control and leave apiKey empty so the proxy adds the Sarvam key server-side.
For a cross-origin proxy, pass your app’s auth in customHeaders (for example Authorization: Bearer plus your session token).