Skip to navigation

Web (TypeScript)

View as Markdown

For real-time voice in the browser, the TypeScript SDK (sarvam-conv-ai-sdk/browser) wraps the WebSocket interface in a typed client, ConversationAgent, with BrowserAudioInterface for microphone capture and playback.

Install

npm install sarvam-conv-ai-sdk

Use the /browser entry point for a smaller bundle:

import {
ConversationAgent,
BrowserAudioInterface,
} from "sarvam-conv-ai-sdk/browser";

Start a voice session

import {
ConversationAgent,
BrowserAudioInterface,
InteractionType,
} from "sarvam-conv-ai-sdk/browser";
const audioInterface = new BrowserAudioInterface();
const agent = new ConversationAgent({
apiKey: "your_api_key",
config: {
org_id: "your_org_id",
workspace_id: "your_workspace_id",
app_id: "your_app_id",
user_identifier: "user123",
user_identifier_type: "custom",
interaction_type: InteractionType.CALL,
input_sample_rate: 16000,
output_sample_rate: 16000,
},
audioInterface,
transcriptCallback: async (msg) => {
console.log(`${msg.role}: ${msg.content}`);
},
});
await agent.start();
await agent.waitForConnect(10);

BrowserAudioInterface captures the microphone and plays agent audio (16-bit PCM mono). agent.start() fetches a signed WebSocket URL, sends the session start message, and begins streaming. Call start() from a user gesture (for example a button click) so the browser allows audio playback.

InteractionConfig

FieldRequiredDescription
user_identifier_typeYes'custom', 'email', 'phone_number', or 'unknown'
user_identifierYesThe identifier value; also what you search by in Log Analyser
org_idYesYour organization ID
workspace_idYesYour workspace ID
app_idYesThe agent to connect to
interaction_typeYesInteractionType.CALL for a voice session
input_sample_rateYes8000 or 16000 Hz
output_sample_rateYes16000 or 22050 Hz
versionNoPins a specific committed agent version. If omitted, the SDK uses the latest committed version, and the connection fails if the agent has no committed version
agent_variablesNoSeed agent variables at session start
initial_language_nameNoStarting language; must be one of the agent’s allowed languages
initial_state_nameNoStarting state, if the agent uses states
initial_bot_messageNoFirst message from the agent

Callbacks

Pass any of these to ConversationAgent to react to what happens during the call:

CallbackFires withUse it for
transcriptCallbacktranscript message (role: Role.USER or Role.BOT, content)A live transcript of both sides of the call
audioCallbackaudio chunkRaw agent audio, if you’re handling playback yourself
audioLevelCallbackoutput levelReal-time level for a speaking indicator
eventCallbacksession eventSession events such as connect, barge-in, and end
stateCallbackAgentStateDrive UI from IDLE, CONNECTING, CONNECTED, LISTENING, SPEAKING, or ERROR
startCallback—Session start
endCallback—Session end and cleanup

BrowserAudioInterface

new BrowserAudioInterface(
sampleRate?: number, // Default: 16000
options?: {
prebufferMs?: number; // Default: 700
bufferSizeMs?: number; // Default: 8000
outputLevelCallback?: (level: { rms: number; peak: number; db: number }) => void;
}
)

BrowserAudioInterface uses adaptive buffering to handle network jitter. Adjust prebufferMs for your network:

Network conditionprebufferMsNotes
Good/stable300–500Lower latency; may gap on jitter
Normal700 (default)Balanced latency and stability
Unstable1000–1500Higher latency, smoother playback

Use audioLevelCallback on the agent (or outputLevelCallback on the interface) to drive a speaking indicator:

const agent = new ConversationAgent({
// ...config
audioLevelCallback: (level) => {
updateAgentSpeakingIndicator(level);
},
});

Session methods

MethodDescription
await agent.start()Fetch a signed WebSocket URL and connect
await agent.waitForConnect(timeout?)Wait until connected; timeout is in seconds (returns false on timeout)
await agent.sendAudio(data)Send a raw PCM audio chunk
agent.isConnected()Current connection status
agent.getInteractionId()The current interaction (call) ID, once connected
agent.getState()Current AgentState
agent.mute() / agent.unmute()Mute or unmute the microphone without disconnecting. While muted, the SDK sends silence so VAD stays stable
agent.isMuted()Current mute status
await agent.waitForDisconnect()Wait until the session ends
await agent.stop()Close the connection and clean up

Call await agent.stop() on unmount (or in a finally block) so the WebSocket and audio interface are cleaned up even if the session errors out. Reconnection is not supported: each WebSocket URL is single-use. If the connection drops, call stop() and create a new ConversationAgent.

React example

import React, { useRef, useState, useEffect } from "react";
import {
ConversationAgent,
BrowserAudioInterface,
InteractionType,
AgentState,
Role,
} from "sarvam-conv-ai-sdk/browser";
interface Message {
role: "user" | "bot";
content: string;
}
function VoiceChat() {
const [state, setState] = useState<AgentState>(AgentState.IDLE);
const [messages, setMessages] = useState<Message[]>([]);
const [isMuted, setIsMuted] = useState(false);
const agentRef = useRef<ConversationAgent | null>(null);
useEffect(() => {
return () => {
agentRef.current?.stop().catch(console.error);
};
}, []);
const startConversation = async () => {
const audioInterface = new BrowserAudioInterface(16000, {
outputLevelCallback: (level) => {
// Update volume visualization
},
});
const agent = new ConversationAgent({
apiKey: "your_api_key",
config: {
org_id: "your_org_id",
workspace_id: "your_workspace_id",
app_id: "your_app_id",
user_identifier: "user123",
user_identifier_type: "custom",
interaction_type: InteractionType.CALL,
input_sample_rate: 16000,
output_sample_rate: 16000,
},
audioInterface,
stateCallback: (newState) => setState(newState),
transcriptCallback: async (msg) => {
setMessages((prev) => [
...prev,
{ role: msg.role === Role.USER ? "user" : "bot", content: msg.content },
]);
},
endCallback: async () => {
agentRef.current = null;
setState(AgentState.IDLE);
},
});
agentRef.current = agent;
await agent.start();
};
const toggleMute = () => {
if (!agentRef.current) return;
if (agentRef.current.isMuted()) {
agentRef.current.unmute();
setIsMuted(false);
} else {
agentRef.current.mute();
setIsMuted(true);
}
};
return (
<div>
<p>State: {state}</p>
<button onClick={startConversation} disabled={state !== AgentState.IDLE}>
Start
</button>
<button onClick={() => agentRef.current?.stop()} disabled={state === AgentState.IDLE}>
Stop
</button>
<button onClick={toggleMute} disabled={state === AgentState.IDLE}>
{isMuted ? "Unmute" : "Mute"}
</button>
<div>
{messages.map((msg, i) => (
<p key={i}>
<strong>{msg.role === "user" ? "You" : "Agent"}:</strong> {msg.content}
</p>
))}
</div>
</div>
);
}
export default VoiceChat;

Troubleshooting

audioInterface is required for CALL interactions. Pass a BrowserAudioInterface:

const agent = new ConversationAgent({
audioInterface: new BrowserAudioInterface(),
// ...
});

Audio doesn’t play in Firefox. Firefox requires a user gesture before playing audio. Call start() from a click handler.

HTTPS is required for the microphone in production. localhost is enough for development.

Microphone permission errors:

try {
const audioInterface = new BrowserAudioInterface();
const agent = new ConversationAgent({
audioInterface,
// ...
});
await agent.start();
} catch (err) {
if (err.name === "NotAllowedError") {
console.error("Microphone permission denied");
} else if (err.name === "NotFoundError") {
console.error("No microphone found");
} else if (err.name === "NotReadableError") {
console.error("Microphone is in use");
}
}

Never embed your API key in client-side code. Point baseUrl at a backend you control and leave apiKey empty so the proxy adds the Sarvam key server-side.

const agent = new ConversationAgent({
apiKey: "",
baseUrl: "/api/sarvam/",
config: {
// ...
},
audioInterface: new BrowserAudioInterface(),
});

For a cross-origin proxy, pass your app’s auth in customHeaders (for example Authorization: Bearer plus your session token).