> For clean Markdown of any page, append `.md` to the page URL. > For a complete documentation index, see https://docs.sarvam.ai/llms.txt. > For full documentation content in one file, see https://docs.sarvam.ai/llms-full.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.sarvam.ai/_mcp/server. # Live Video Transcription > Build a real-time transcription and translation demo with Sarvam AI's Streaming Speech-to-Text API, Flask-SocketIO, and the browser's Web Audio API. ## **Overview** This cookbook builds a real-time transcription and translation demo. Play or upload any video with speech in the browser, and watch a live transcript and English translation appear as it plays, powered by Sarvam's **Streaming Speech-to-Text API** over a persistent WebSocket connection. **What you'll build:** * A browser page that captures a playing video's audio and streams it to a server, no file upload round-trip required * A Flask-SocketIO server that keeps one persistent Sarvam streaming connection open per client per mode * Live transcription in the source language, running concurrently with live translation to English * Results pushed back to the browser over WebSocket as each speech segment finalizes #### [View on GitHub](https://github.com/sarvamai/sarvam-ai-cookbook/tree/main/examples/Live_Video_Transcription) Browse the full source for this app in the sarvam-ai-cookbook repo. #### [Run Code Locally](#2-clone-and-install) Jump to Clone and Install to set this demo up on your machine. ### Business Value * Live captions for lectures, webinars, and internal recordings * Real-time subtitling for video content in a viewer's own language * Accessibility support for hearing-impaired viewers watching pre-recorded or live video #### EdTech Caption recorded lectures live as students watch, with an English translation running alongside the original-language transcript. #### Media & Broadcast Add live subtitles to a video feed without waiting for the whole file to be transcribed first. #### Enterprise Transcribe town halls, training videos, or internal recordings for search and accessibility as they're played back. ## **How It Works** #### Capture audio in the browser The browser reads the video element's audio track via the Web Audio API, resamples it to 16kHz mono, and encodes it as WAV, entirely client-side, no server upload of the raw video file. #### Stream \~1 second frames over Socket.IO Every second, the buffered audio is packaged into one WAV chunk, base64-encoded, and emitted to the Flask server over a Socket.IO event. #### Forward into a persistent Sarvam connection For each client, the server opens one long-lived Sarvam streaming connection per active mode (transcribe and/or translate) and forwards incoming frames into it as they arrive. #### Relay results back live Sarvam emits a finalized transcript or translation as each speech segment completes. The server relays it back to the same browser tab over Socket.IO, and it's appended to the transcript panel immediately. Transcription and translation run as two independent streaming connections, so you can start one, both, or switch between them without interrupting the other. ## **1. Prerequisites** * Python 3.8 or higher * A Sarvam API key, sign up on the [Sarvam AI Dashboard](https://dashboard.sarvam.ai/) to get one * A modern browser (Web Audio API and `HTMLMediaElement.captureStream()` support) * A video file with speech to try it on ## **2. Clone and Install** #### macOS/Linux ```bash git clone https://github.com/sarvamai/sarvam-ai-cookbook.git cd sarvam-ai-cookbook/examples/Live_Video_Transcription python -m venv venv source venv/bin/activate pip install -r requirements.txt ``` #### Windows ```bash git clone https://github.com/sarvamai/sarvam-ai-cookbook.git cd sarvam-ai-cookbook/examples/Live_Video_Transcription python -m venv venv venv\Scripts\activate pip install -r requirements.txt ``` ## **3. Configure Environment Variables** Create a `.env` file in the project root: ```env SARVAM_API_KEY=sk_xxxxxxxxxxxxxxxxxxxxxxxx ``` `config.py` reads this key and also holds a few settings worth knowing about before you run the app: | Setting | Default | Meaning | | ---------------------- | -------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `HOST` / `PORT` | `127.0.0.1` / `5001` | Where the Flask server binds | | `API_LANGUAGE` | `"unknown"` | Source language passed to the transcription stream; `"unknown"` auto-detects. Set to a specific code (e.g. `"hi-IN"`) for more consistent accuracy on a known language | | `CORS_ALLOWED_ORIGINS` | `"*"` | Fine for local testing; restrict this before deploying anywhere public | ## **4. Run It** ```bash python app.py ``` Open [http://localhost:5001](http://localhost:5001), upload a video with speech, then click **Start Transcription** and/or **Start Translation**. Text appears in the corresponding panel as each segment finalizes. ## **5. Pipeline Walkthrough** ### Step 1: Capture and encode audio in the browser `templates/index.html` pulls the audio track off the `