Building for Indian Languages
Indian-language products have to handle realities that English-first stacks ignore: people mix English into every sentence, the same language is written in multiple scripts, a lot of audio arrives over 8kHz phone lines, and documents come as scans in regional scripts. This page maps those realities to the exact Sarvam speech, voice, and vision APIs, with code that’s been run against the live API.
Language coverage at a glance
Coverage differs by capability, so pick the model that matches your language set.
Speech-to-Text: transcribe real Indian audio
saaras:v3 exposes a mode parameter that controls how speech is written down. This is where most Indic-specific decisions happen:
The difference is easiest to see on one clip. For audio of “मुझे flight book करनी है” (a typical Hinglish sentence), the same audio produces:
Use codemix for chat/agent transcripts that feel natural, transcribe for clean native-script records, and translit when a downstream system only handles Latin script. The REST endpoint handles clips up to 30 seconds, for longer audio use the Batch API.
Telephony and 8kHz audio
A large share of Indian voice traffic is phone audio: 8kHz, mono, often µ-law/A-law encoded. Two rules keep transcription quality high:
- Match the sample rate everywhere. For 8kHz audio, set
sample_rate=8000both when opening the streaming connection and when sending each audio chunk. Mismatched rates cause garbled output. - Use a supported streaming codec. Streaming STT accepts WAV and raw PCM (
pcm_s16le,pcm_l16,pcm_raw) only. - Signal end-of-audio. Call
ws.flush()after the last chunk so the server knows to emit a transcript instead of waiting indefinitely for more audio.
saaras:v3 is tuned for telephony, so prefer it for call audio.
See the Streaming STT guide for the full WebSocket lifecycle, VAD, reconnection, and voice-agent barge-in.
Text-to-Speech: natural Indian voices
bulbul:v3 speaks 11 languages (10 Indian + English) with 30+ voices, and it’s built for the way Indians actually write, so you usually don’t pre-process anything.
- Pass code-mixed text directly. Mixed Hindi-English (“आपका OTP 4321 है। Please use it…”) is spoken naturally; no need to romanize or split it.
- Normalize English words and numbers with
enable_preprocessing=truewhen your text has lots of abbreviations or digits. - Pick a voice with
speaker(e.g.shubh,priya,kavya). See the voice list. - Output to a phone line by requesting a telephony codec,
output_audio_codecsupportsmulawandalawalongsidemp3,wav,opus, etc.
Pronunciation control
Brand names, abbreviations, and regional terms (“NEFT”, “KYC”, a company name) don’t always come out right by default. Create a Pronunciation Dictionary with Bulbul v3 to pin how specific words are spoken, then pass its dict_id to any TTS call (REST, HTTP stream, or WebSocket):
Document AI: Indic OCR with Sarvam Vision
Indian documents arrive as scans, photos, and PDFs in regional scripts, often with tables and mixed languages. Sarvam Vision extracts text and structure across 23 languages (22 Indian + English), preserving the original script. Digitise returns the whole document as html or md; Extract pulls just the fields you define.
It’s an asynchronous job API: create the job in one call, poll status, then fetch the output.
See the Document AI guide for schema-based Extract and the full job lifecycle. The older document_intelligence job API (create → upload → start → poll → download) still works for existing integrations, but new work should use doc_ai.
Document translation
When the deliverable is a translated file (not plain text), use the Document Translation API. Upload once, pick up to 12 target languages per job, and export each language in the same format as the source (PDF stays PDF, DOCX stays DOCX). The pipeline runs OCR or structural extraction, then translation — layout and formatting are preserved.
The flow matches other async document APIs: create job → PUT file to signed upload_url → start → poll live-status → export per language.
See Poll and export for per-language export and download URLs.
Use Document AI when you only need digitised text or structured fields. Use Document Translation when stakeholders need an editable or printable file in another language.
Text translation and transliteration
For short text (not whole files), mayura:v1 (11 languages, 10 Indian + English, colloquial modes, output-script and native-numeral control) and sarvam-translate:v1 (23 languages, 22 Indian + English, formal) handle translation, and the Transliteration API converts between scripts. See the Text translation guide for output_script and numerals_format options.