> For clean Markdown of any page, append `.md` to the page URL.
> For a complete documentation index, see https://docs.sarvam.ai/llms.txt.
> For full documentation content in one file, see https://docs.sarvam.ai/llms-full.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.sarvam.ai/_mcp/server.

# Web (TypeScript)

For real-time voice in the browser, the TypeScript SDK (`sarvam-conv-ai-sdk/browser`) wraps the [WebSocket interface](/conversations/deploy/deploy-with-code#websocket-integration) in a typed client, `ConversationAgent`, with `BrowserAudioInterface` for microphone capture and playback.

## Install

```bash
npm install sarvam-conv-ai-sdk
```

Use the `/browser` entry point for a smaller bundle:

```typescript
import {
  ConversationAgent,
  BrowserAudioInterface,
} from "sarvam-conv-ai-sdk/browser";
```

## Start a voice session

```typescript
import {
  ConversationAgent,
  BrowserAudioInterface,
  InteractionType,
} from "sarvam-conv-ai-sdk/browser";

const audioInterface = new BrowserAudioInterface();

const agent = new ConversationAgent({
  apiKey: "your_api_key",
  config: {
    org_id: "your_org_id",
    workspace_id: "your_workspace_id",
    app_id: "your_app_id",
    user_identifier: "user123",
    user_identifier_type: "custom",
    interaction_type: InteractionType.CALL,
    input_sample_rate: 16000,
    output_sample_rate: 16000,
  },
  audioInterface,
  transcriptCallback: async (msg) => {
    console.log(`${msg.role}: ${msg.content}`);
  },
});

await agent.start();
await agent.waitForConnect(10);
```

`BrowserAudioInterface` captures the microphone and plays agent audio (16-bit PCM mono). `agent.start()` fetches a signed WebSocket URL, sends the session start message, and begins streaming. Call `start()` from a user gesture (for example a button click) so the browser allows audio playback.

## InteractionConfig

| Field                   | Required | Description                                                                                                                                                    |
| ----------------------- | -------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `user_identifier_type`  | Yes      | `'custom'`, `'email'`, `'phone_number'`, or `'unknown'`                                                                                                        |
| `user_identifier`       | Yes      | The identifier value; also what you search by in [Log Analyser](/conversations/monitor/agent-analytics/log-analyser)                                           |
| `org_id`                | Yes      | Your organization ID                                                                                                                                           |
| `workspace_id`          | Yes      | Your workspace ID                                                                                                                                              |
| `app_id`                | Yes      | The agent to connect to                                                                                                                                        |
| `interaction_type`      | Yes      | `InteractionType.CALL` for a voice session                                                                                                                     |
| `input_sample_rate`     | Yes      | `8000` or `16000` Hz                                                                                                                                           |
| `output_sample_rate`    | Yes      | `16000` or `22050` Hz                                                                                                                                          |
| `version`               | No       | Pins a specific committed agent version. If omitted, the SDK uses the latest committed version, and the connection fails if the agent has no committed version |
| `agent_variables`       | No       | Seed [agent variables](/conversations/build/variables-personalization) at session start                                                                        |
| `initial_language_name` | No       | Starting language; must be one of the agent's allowed languages                                                                                                |
| `initial_state_name`    | No       | Starting state, if the agent uses [states](/conversations/build/states-conversation-flow)                                                                      |
| `initial_bot_message`   | No       | First message from the agent                                                                                                                                   |

## Callbacks

Pass any of these to `ConversationAgent` to react to what happens during the call:

| Callback             | Fires with                                                        | Use it for                                                                           |
| -------------------- | ----------------------------------------------------------------- | ------------------------------------------------------------------------------------ |
| `transcriptCallback` | transcript message (`role`: `Role.USER` or `Role.BOT`, `content`) | A live transcript of both sides of the call                                          |
| `audioCallback`      | audio chunk                                                       | Raw agent audio, if you're handling playback yourself                                |
| `audioLevelCallback` | output level                                                      | Real-time level for a speaking indicator                                             |
| `eventCallback`      | session event                                                     | Session events such as connect, barge-in, and end                                    |
| `stateCallback`      | `AgentState`                                                      | Drive UI from `IDLE`, `CONNECTING`, `CONNECTED`, `LISTENING`, `SPEAKING`, or `ERROR` |
| `startCallback`      | —                                                                 | Session start                                                                        |
| `endCallback`        | —                                                                 | Session end and cleanup                                                              |

## BrowserAudioInterface

```typescript
new BrowserAudioInterface(
  sampleRate?: number,  // Default: 16000
  options?: {
    prebufferMs?: number;      // Default: 700
    bufferSizeMs?: number;     // Default: 8000
    outputLevelCallback?: (level: { rms: number; peak: number; db: number }) => void;
  }
)
```

`BrowserAudioInterface` uses adaptive buffering to handle network jitter. Adjust `prebufferMs` for your network:

| Network condition | `prebufferMs`   | Notes                             |
| ----------------- | --------------- | --------------------------------- |
| Good/stable       | `300`–`500`     | Lower latency; may gap on jitter  |
| Normal            | `700` (default) | Balanced latency and stability    |
| Unstable          | `1000`–`1500`   | Higher latency, smoother playback |

Use `audioLevelCallback` on the agent (or `outputLevelCallback` on the interface) to drive a speaking indicator:

```typescript
const agent = new ConversationAgent({
  // ...config
  audioLevelCallback: (level) => {
    updateAgentSpeakingIndicator(level);
  },
});
```

## Session methods

| Method                                 | Description                                                                                                 |
| -------------------------------------- | ----------------------------------------------------------------------------------------------------------- |
| `await agent.start()`                  | Fetch a signed WebSocket URL and connect                                                                    |
| `await agent.waitForConnect(timeout?)` | Wait until connected; `timeout` is in seconds (returns `false` on timeout)                                  |
| `await agent.sendAudio(data)`          | Send a raw PCM audio chunk                                                                                  |
| `agent.isConnected()`                  | Current connection status                                                                                   |
| `agent.getInteractionId()`             | The current interaction (call) ID, once connected                                                           |
| `agent.getState()`                     | Current `AgentState`                                                                                        |
| `agent.mute()` / `agent.unmute()`      | Mute or unmute the microphone without disconnecting. While muted, the SDK sends silence so VAD stays stable |
| `agent.isMuted()`                      | Current mute status                                                                                         |
| `await agent.waitForDisconnect()`      | Wait until the session ends                                                                                 |
| `await agent.stop()`                   | Close the connection and clean up                                                                           |

> **Note**
>
> Call `await agent.stop()` on unmount (or in a `finally` block) so the WebSocket and audio interface are cleaned up even if the session errors out. Reconnection is not supported: each WebSocket URL is single-use. If the connection drops, call `stop()` and create a new `ConversationAgent`.

## React example

```typescript
import React, { useRef, useState, useEffect } from "react";
import {
  ConversationAgent,
  BrowserAudioInterface,
  InteractionType,
  AgentState,
  Role,
} from "sarvam-conv-ai-sdk/browser";

interface Message {
  role: "user" | "bot";
  content: string;
}

function VoiceChat() {
  const [state, setState] = useState<AgentState>(AgentState.IDLE);
  const [messages, setMessages] = useState<Message[]>([]);
  const [isMuted, setIsMuted] = useState(false);
  const agentRef = useRef<ConversationAgent | null>(null);

  useEffect(() => {
    return () => {
      agentRef.current?.stop().catch(console.error);
    };
  }, []);

  const startConversation = async () => {
    const audioInterface = new BrowserAudioInterface(16000, {
      outputLevelCallback: (level) => {
        // Update volume visualization
      },
    });

    const agent = new ConversationAgent({
      apiKey: "your_api_key",
      config: {
        org_id: "your_org_id",
        workspace_id: "your_workspace_id",
        app_id: "your_app_id",
        user_identifier: "user123",
        user_identifier_type: "custom",
        interaction_type: InteractionType.CALL,
        input_sample_rate: 16000,
        output_sample_rate: 16000,
      },
      audioInterface,
      stateCallback: (newState) => setState(newState),
      transcriptCallback: async (msg) => {
        setMessages((prev) => [
          ...prev,
          { role: msg.role === Role.USER ? "user" : "bot", content: msg.content },
        ]);
      },
      endCallback: async () => {
        agentRef.current = null;
        setState(AgentState.IDLE);
      },
    });

    agentRef.current = agent;
    await agent.start();
  };

  const toggleMute = () => {
    if (!agentRef.current) return;
    if (agentRef.current.isMuted()) {
      agentRef.current.unmute();
      setIsMuted(false);
    } else {
      agentRef.current.mute();
      setIsMuted(true);
    }
  };

  return (
    <div>
      <p>State: {state}</p>
      <button onClick={startConversation} disabled={state !== AgentState.IDLE}>
        Start
      </button>
      <button onClick={() => agentRef.current?.stop()} disabled={state === AgentState.IDLE}>
        Stop
      </button>
      <button onClick={toggleMute} disabled={state === AgentState.IDLE}>
        {isMuted ? "Unmute" : "Mute"}
      </button>
      <div>
        {messages.map((msg, i) => (
          <p key={i}>
            <strong>{msg.role === "user" ? "You" : "Agent"}:</strong> {msg.content}
          </p>
        ))}
      </div>
    </div>
  );
}

export default VoiceChat;
```

## Troubleshooting

**`audioInterface` is required for CALL interactions.** Pass a `BrowserAudioInterface`:

```typescript
const agent = new ConversationAgent({
  audioInterface: new BrowserAudioInterface(),
  // ...
});
```

**Audio doesn't play in Firefox.** Firefox requires a user gesture before playing audio. Call `start()` from a click handler.

**HTTPS is required for the microphone** in production. `localhost` is enough for development.

**Microphone permission errors:**

```typescript
try {
  const audioInterface = new BrowserAudioInterface();
  const agent = new ConversationAgent({
    audioInterface,
    // ...
  });
  await agent.start();
} catch (err) {
  if (err.name === "NotAllowedError") {
    console.error("Microphone permission denied");
  } else if (err.name === "NotFoundError") {
    console.error("No microphone found");
  } else if (err.name === "NotReadableError") {
    console.error("Microphone is in use");
  }
}
```

> **Warning**
>
> Never embed your API key in client-side code. Point `baseUrl` at a backend you control and leave `apiKey` empty so the proxy adds the Sarvam key server-side.

```typescript
const agent = new ConversationAgent({
  apiKey: "",
  baseUrl: "/api/sarvam/",
  config: {
    // ...
  },
  audioInterface: new BrowserAudioInterface(),
});
```

For a cross-origin proxy, pass your app's auth in `customHeaders` (for example `Authorization: Bearer` plus your session token).