> For clean Markdown of any page, append `.md` to the page URL. > For a complete documentation index, see https://docs.sarvam.ai/llms.txt. > For full documentation content in one file, see https://docs.sarvam.ai/llms-full.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.sarvam.ai/_mcp/server. # Developer Quickstart > Learn how to make your first API request with Sarvam AI in under 5 minutes. Complete guide with code examples for chat completion, speech-to-text, and translation APIs. **Get started with Sarvam APIs in 30 seconds:** create a key, install the SDK, and make your first call with just a few lines of code. **Prerequisites:** * A Sarvam AI account and API key. [Create one on the dashboard](https://dashboard.sarvam.ai) if you don't have one yet. * Python 3.9+ or Node.js 18+ (or skip both and call the REST API directly with `curl`). **By the end of this quickstart**, you'll have made a live call to a Sarvam AI API and gotten back a real result: a transcript, a playable audio file, a translation, or a chat reply, depending on which API you try below. ## Setup #### Create an API Key Visit the [Sarvam AI dashboard](https://dashboard.sarvam.ai) and create a new API key. Keep this key secure; you'll need it to authenticate your requests. #### Set Up Your Environment Export your API key as an environment variable: #### macOS / Linux ```bash export SARVAM_API_KEY="YOUR_SARVAM_API_KEY" ``` #### Windows ```powershell $env:SARVAM_API_KEY="YOUR_SARVAM_API_KEY" ``` #### Install the SDK Choose your preferred language and install our SDK. Need a different language or an older version? See [Libraries & SDKs](/api/getting-started/sdks). #### Python ```bash pip install -U sarvamai ``` #### JavaScript ```bash npm install sarvamai@latest ``` #### Make Your First API Call Here's a Chat Completion example using the key and SDK you just set up. Paste it in and run it: #### Python ```python from sarvamai import SarvamAI client = SarvamAI( api_subscription_key="YOUR_SARVAM_API_KEY", ) response = client.chat.completions( model="sarvam-105b", messages=[ {"role": "user", "content": "What is the capital of India?"} ], ) print(response.choices[0].message.content) # → "The capital of India is New Delhi." ``` #### JavaScript ```javascript import { SarvamAIClient } from "sarvamai"; const client = new SarvamAIClient({ apiSubscriptionKey: "YOUR_SARVAM_API_KEY" }); const response = await client.chat.completions({ model: "sarvam-105b", messages: [ { role: "user", content: "What is the capital of India?" } ] }); console.log(response.choices[0].message.content); // → "The capital of India is New Delhi." ``` #### cURL ```bash curl -X POST https://api.sarvam.ai/v1/chat/completions \ -H "api-subscription-key: " \ -H "Content-Type: application/json" \ -d '{ "model": "sarvam-105b", "messages": [ { "role": "user", "content": "What is the capital of India?" } ] }' ``` Prefer Speech to Text, Text to Speech, or Translation instead? Jump to [Try an API](#try-an-api) below, each comes with its own Python, JavaScript, and cURL example. > **Note** > > **Keep your key safe.** Store it in an environment variable or a secrets manager, and never commit it to version control. Getting a `403`? See [Authentication](/api-reference/authentication#status-codes-for-authentication-failures). Getting a `429`? See [Credits & Rate Limits](/api/getting-started/ratelimits). Any other error code is listed in [Errors & Troubleshooting](/api/getting-started/errors-troubleshooting). ## Try an API Sarvam AI covers seven core capabilities below. Pick a tab, and copy-paste your way to a working call. #### Speech to Text Saaras v3 is our latest speech recognition model. It supports multiple output modes: `transcribe` (default, original language), `translate` (to English), `verbatim` (word-for-word), `translit` (romanization), and `codemix` (mixed script). #### Python ```python from sarvamai import SarvamAI client = SarvamAI( api_subscription_key="YOUR_SARVAM_API_KEY", ) try: response = client.speech_to_text.transcribe( file=open("audio.wav", "rb"), # path to your own audio file model="saaras:v4", mode="transcribe", # or "translate", "verbatim", "translit", "codemix" ) print(response.transcript) except Exception as e: print(f"Error: {e}") ``` #### JavaScript ```javascript import {SarvamAIClient} from "sarvamai"; import fs from 'fs'; const client = new SarvamAIClient({ apiSubscriptionKey: process.env.SARVAM_API_KEY }); const audioFile = fs.createReadStream("audio.wav"); // path to your own audio file try { const response = await client.speechToText.transcribe({ file: audioFile, model: "saaras:v4", mode: "transcribe", // or "translate", "verbatim", "translit", "codemix" }); console.log(response.transcript); } catch (error) { console.error("Error:", error); } ``` #### cURL ```bash # Replace audio.wav with the path to your own audio file curl -X POST https://api.sarvam.ai/speech-to-text \ -H "api-subscription-key: " \ -H "Content-Type: multipart/form-data" \ -F model="saaras:v4" \ -F mode="transcribe" \ -F file=@audio.wav ``` > **Note** > > Use `transcribe` for same-language output, `translate` to convert audio to English text, or `verbatim`/`translit`/`codemix` for specialized formatting. `mode` defaults to `transcribe` when omitted. #### Text to Speech Bulbul v3 converts text into natural-sounding speech across Indian languages. The API returns **base64-encoded audio**, so decode it before saving or playing the file. #### Python ```python from sarvamai import SarvamAI from sarvamai.play import save client = SarvamAI( api_subscription_key="YOUR_SARVAM_API_KEY", ) try: audio = client.text_to_speech.convert( text="Welcome to Sarvam AI!", language_code="hi-IN", model="bulbul:v3", speaker="shubh", ) save(audio, "output.wav") except Exception as e: print(f"Error: {e}") ``` #### JavaScript ```javascript import { SarvamAIClient } from "sarvamai"; import fs from "fs"; const client = new SarvamAIClient({ apiSubscriptionKey: "YOUR_SARVAM_API_KEY" }); try { const response = await client.textToSpeech.convert({ text: "Welcome to Sarvam AI!", language_code: "hi-IN", model: "bulbul:v3", speaker: "shubh", }); // Decode the base64 audio before saving, since writing the raw string produces a corrupted file const audio = Buffer.from(response.audios.join(""), "base64"); fs.writeFileSync("output.wav", audio); } catch (error) { console.error("Error:", error); } ``` #### cURL ```bash curl -X POST https://api.sarvam.ai/text-to-speech \ -H "api-subscription-key: " \ -H "Content-Type: application/json" \ -d '{ "text": "Welcome to Sarvam AI!", "language_code": "hi-IN", "model": "bulbul:v3", "speaker": "shubh" }' -o response.json # Decode the base64 audio into a playable file python3 -c "import json, base64; d = json.load(open('response.json')); open('output.wav', 'wb').write(base64.b64decode(''.join(d['audios'])))" ``` > **Note** > > The `audios` field is base64-encoded, so always decode it before writing to a file, or the audio will be corrupted. See the [TTS best practices guide](/api/api-guides-tutorials/text-to-speech/best-practices) for streaming and more voice options. #### Voice Cloning First create a saved voice from a reference clip. Then generate speech with its `voice_id`. #### cURL ```bash # Step 1: create a voice curl -s -X POST "https://api.sarvam.ai/voices/create" \ -H "api-subscription-key: " \ -F "name=my-voice" \ -F "language=en-IN" \ -F "file=@reference.wav" # Step 2: generate speech with data.voice_id from step 1 curl -s -X POST "https://api.sarvam.ai/voices/clone" \ -H "api-subscription-key: " \ -F "voice_id=svc-your-voice-id" \ -F "text=नमस्ते, मैं आपकी कैसे मदद कर सकता हूँ?" \ -F "language_code=hi-IN" \ -o response.json # Decode the base64 audio into a playable file python3 -c "import base64, json; open('output.wav', 'wb').write(base64.b64decode(json.load(open('response.json'))['audio']))" ``` > **Note** > > Voice cloning requires the `text_to_speech_voice_cloning` capability on your subscription. See the [Voice Cloning overview](/api/api-guides-tutorials/voice-cloning/overview) for the full request schema, quality control, and cross-lingual cloning. #### Text Translation Sarvam Translate converts text across English and 22 Indian languages, with automatic source-language detection when `source_language_code` is `"auto"`. #### Python ```python from sarvamai import SarvamAI client = SarvamAI( api_subscription_key="YOUR_SARVAM_API_KEY", ) try: response = client.text.translate( input="Hello, how are you?", source_language_code="auto", target_language_code="hi-IN", speaker_gender="Male" ) print(response.translated_text) # → e.g. "नमस्ते, आप कैसे हैं?" except Exception as e: print(f"Error: {e}") ``` #### JavaScript ```javascript import { SarvamAIClient } from "sarvamai"; const client = new SarvamAIClient({ apiSubscriptionKey: "YOUR_SARVAM_API_KEY" }); try { const response = await client.text.translate({ input: "Hello, how are you?", source_language_code: "auto", target_language_code: "hi-IN", speaker_gender: "Male" }); console.log(response.translatedText); // → e.g. "नमस्ते, आप कैसे हैं?" } catch (error) { console.error("Error:", error); } ``` #### cURL ```bash curl -X POST https://api.sarvam.ai/translate \ -H "api-subscription-key: " \ -H "Content-Type: application/json" \ -d '{ "input": "Hello, how are you?", "source_language_code": "auto", "target_language_code": "hi-IN", "speaker_gender": "Male" }' ``` > **Note** > > `speaker_gender` adjusts grammatical gender agreement in the translation for languages where verb forms change by gender, such as Hindi. See the [Sarvam Translate model page](/api/getting-started/models/sarvam-translate) for the full list of supported languages. #### Chat Completion Sarvam-105B is a chat model with hybrid reasoning and native support for Indian languages. #### Python ```python from sarvamai import SarvamAI client = SarvamAI( api_subscription_key="YOUR_SARVAM_API_KEY" ) try: response = client.chat.completions( model="sarvam-105b", messages=[ {"role": "user", "content": "What is the capital of India?"} ], ) print(response.choices[0].message.content) # → "The capital of India is New Delhi." except Exception as e: print(f"Error: {e}") ``` #### JavaScript ```javascript import { SarvamAIClient } from "sarvamai"; // Initialize the SarvamAI client with your API key const client = new SarvamAIClient({ apiSubscriptionKey: "YOUR_SARVAM_API_KEY" }); async function main() { try { const response = await client.chat.completions({ model: "sarvam-105b", messages: [ { role: "user", content: "What is the capital of India?" } ] }); console.log(response.choices[0].message.content); // → "The capital of India is New Delhi." } catch (error) { console.error("Error:", error); } } main(); ``` #### cURL ```bash curl -X POST https://api.sarvam.ai/v1/chat/completions \ -H "api-subscription-key: " \ -H "Content-Type: application/json" \ -d '{ "model": "sarvam-105b", "messages": [ { "role": "user", "content": "What is the capital of India?" } ] }' ``` > **Note** > > `sarvam-105b` is the flagship model for the highest-quality reasoning. See [Models](/api/getting-started/models) for the full lineup and thinking modes. #### Document Translation Document Translation translates a whole PDF, Word, PowerPoint, or web page file into up to 12 Indian languages per job, preserving the original layout. It's an asynchronous job: create, upload, start, then poll and export. #### Python ```python from pathlib import Path import httpx from sarvamai import SarvamAI client = SarvamAI( api_subscription_key="YOUR_SARVAM_API_KEY", ) document = Path("chapter1.pdf") # path to your own document # 1. Create the job (job_name is optional; defaults to original_filename) job = client.document_translation.create( job_name="chapter1", source_language_code="en-IN", target_language_codes=["hi-IN", "ta-IN"], original_filename=document.name, ) # 2. Upload the file to the signed URL (required; the SDK has no upload helper yet) with document.open("rb") as f: httpx.put( job.upload_url, content=f.read(), headers={"Content-Type": "application/pdf", "x-ms-blob-type": "BlockBlob"}, timeout=120.0, ) # 3. Start the pipeline client.document_translation.start(job_id=job.job_id) print(job.job_id) ``` #### JavaScript ```javascript import fs from "fs"; import { SarvamAIClient } from "sarvamai"; const client = new SarvamAIClient({ apiSubscriptionKey: "YOUR_SARVAM_API_KEY" }); const documentPath = "chapter1.pdf"; // path to your own document // 1. Create the job (job_name is optional, defaults to original_filename) const job = await client.documentTranslation.create({ job_name: "chapter1", source_language_code: "en-IN", target_language_codes: ["hi-IN", "ta-IN"], original_filename: documentPath, }); // 2. Upload the file to the signed URL (required; the SDK has no upload helper yet) await fetch(job.upload_url, { method: "PUT", headers: { "Content-Type": "application/pdf", "x-ms-blob-type": "BlockBlob" }, body: fs.readFileSync(documentPath), }); // 3. Start the pipeline await client.documentTranslation.start(job.job_id); console.log(job.job_id); ``` > **Note** > > This kicks off the job; translation happens asynchronously. Poll `GET .../live-status` and export each target language once its state is `Completed`. See [Poll and Export](/api/api-guides-tutorials/doc-translation/how-to/poll-and-export) for the full loop. No cURL SDK helper exists for the upload step, so both examples call the signed URL directly. #### Document Intelligence Document Intelligence (Sarvam Vision) turns a PDF or scanned document into structured Markdown, HTML, or JSON. `digitise()` creates and submits the job in one call. #### Python ```python from sarvamai import SarvamAI client = SarvamAI( api_subscription_key="YOUR_SARVAM_API_KEY", ) with open("./sample-document.pdf", "rb") as f: job = client.doc_ai.digitise( file=[("sample-document.pdf", f, "application/pdf")], language="en-IN", output_format="md", ) print("job created:", job.job_id, "| status:", job.status) ``` #### JavaScript ```javascript import fs from "fs"; import { SarvamAIClient } from "sarvamai"; const client = new SarvamAIClient({ apiSubscriptionKey: "YOUR_SARVAM_API_KEY" }); const job = await client.docAi.digitise({ file: [fs.createReadStream("./sample-document.pdf")], language: "en-IN", output_format: "md", }); console.log("job created:", job.job_id, "| status:", job.status); ``` #### cURL ```bash curl -X POST https://api.sarvam.ai/doc-ai/v1/job/digitise \ -H "api-subscription-key: " \ -F file=@./sample-document.pdf \ -F language="en-IN" \ -F output_format="md" ``` > **Note** > > Poll `get_status()` until the job reaches a terminal state (`completed`, `partially_completed`, `failed`, or `rejected`), then fetch the file with `get_download_url()`. See the [Document Intelligence guide](/api/api-guides-tutorials/document-intelligence/overview) for the full polling loop and the schema-based `extract()` alternative. Use `output_format="md"`, not `"markdown"`, it returns a `400`. #### Dubbing Dubbing localizes video or audio into another Indian language, with optional voice cloning. It's an asynchronous job: create, upload, start, then poll for downloadable exports. #### Python ```python from pathlib import Path from sarvamai import SarvamAI client = SarvamAI( api_subscription_key="YOUR_SARVAM_API_KEY", ) media = Path("sample.mp4") # path to your own video or audio file created = client.dubbing.create( source_language_code="en-IN", target_language_codes=["hi-IN", "ta-IN"], export_options=["video", "srt"], voice_cloning=True, num_speakers=1, job_name=media.name, ) client.dubbing.upload(created.data.upload_url, media) # requires sarvamai>=0.1.31a1 client.dubbing.start(job_id=created.data.job_id) print(created.data.job_id) ``` > **Note** > > That starts the dub job running asynchronously. Poll `get_live_status()` for progress, then `get_export_status()` for a signed download URL per (language, format). See [Job Lifecycle](/api/api-guides-tutorials/dubbing/job-lifecycle) for the full polling script. JavaScript and cURL examples for this flow aren't published yet; see the [Create Dub API reference](/api-reference/creative-agents-dubbing/create-dub) for other languages. ## Next Steps Pick up where you left off, based on the API you just tried: #### [Speech to Text](/api/api-guides-tutorials/speech-to-text/overview) REST, Batch, and Streaming APIs, output modes, and language support. #### [Text to Speech](/api/api-guides-tutorials/text-to-speech/best-practices) Voices, pacing, audio formats, and streaming output. #### [Voice Cloning](/api/api-guides-tutorials/voice-cloning/overview) Generate speech in a cloned voice from a short reference clip. #### [Chat Completion](/api/api-guides-tutorials/chat-completion/overview) Available models, reasoning modes, and streaming responses. #### [Text Translation](/api/getting-started/models/sarvam-translate) Supported languages and translation model options. #### [Document Translation](/api/api-guides-tutorials/doc-translation/overview) Job lifecycle, style guidelines, and native numeral formatting. #### [Document Intelligence](/api/api-guides-tutorials/document-intelligence/overview) Digitise vs. Extract, config IDs, and structured output formats. #### [Dubbing](/api/api-guides-tutorials/dubbing/overview) Voice cloning, job lifecycle, and export formats. #### [Browse the Cookbook](/api/cookbook) Step-by-step tutorials and end-to-end example projects. #### [Join our Discord](https://discord.com/invite/5rAsykttcs) Early access to new releases, biweekly challenges, office hours with the team, and India's AI builder community. > Learn how to make your first API request with Sarvam AI in under 5 minutes. Complete guide with code examples for chat completion, speech-to-text, and translation APIs.