Skip to navigation

Building for Indian Languages

The concepts that make Indian-language speech and document apps work in production: code-mixing, scripts, telephony audio, voices, and Indic OCR.
View as Markdown

Indian-language products have to handle realities that English-first stacks ignore: people mix English into every sentence, the same language is written in multiple scripts, a lot of audio arrives over 8kHz phone lines, and documents come as scans in regional scripts. This page maps those realities to the exact Sarvam speech, voice, and vision APIs, with code that’s been run against the live API.

Language coverage at a glance

Coverage differs by capability, so pick the model that matches your language set.

CapabilityModelLanguages
Speech-to-TextSaaras v323 (22 Indian + English)
Speech-to-Text-TranslateSaaras v323 (22 Indian + English) → English
Text-to-SpeechBulbul v311 (10 Indian + English)
Document DigitizationSarvam Vision23 (22 Indian + English)
Document TranslationDocument Translation API23 (22 Indian + English); PDF, Office, and web formats
Text translationMayura v1 / Sarvam-Translate v111 (10 Indian + English) / 23 (22 Indian + English)
Chat / reasoningSarvam-105B11 (10 Indian + English)

Speech-to-Text: transcribe real Indian audio

saaras:v3 exposes a mode parameter that controls how speech is written down. This is where most Indic-specific decisions happen:

modeWhat you get
transcribeTranscription in the native script
translateEnglish translation of the speech
verbatimWord-for-word, including fillers and repetitions
translitRomanized (Latin-script) output
codemixNatural code-mixed text, English words stay in English

The difference is easiest to see on one clip. For audio of “मुझे flight book करनी है” (a typical Hinglish sentence), the same audio produces:

modeTranscript (live output)
transcribeमुझे फ्लाइट बुक करनी है।
codemixमुझे flight बुक करनी है।
translitMujhe flight book karni hai.
from sarvamai import SarvamAI
client = SarvamAI(api_subscription_key="YOUR_SARVAM_API_KEY")
with open("audio.wav", "rb") as f:
response = client.speech_to_text.transcribe(
file=f,
model="saaras:v4",
mode="codemix", # keep English words in English
language_code="hi-IN",
)
print(response.transcript) # मुझे flight बुक करनी है।

Use codemix for chat/agent transcripts that feel natural, transcribe for clean native-script records, and translit when a downstream system only handles Latin script. The REST endpoint handles clips up to 30 seconds, for longer audio use the Batch API.

Telephony and 8kHz audio

A large share of Indian voice traffic is phone audio: 8kHz, mono, often µ-law/A-law encoded. Two rules keep transcription quality high:

  1. Match the sample rate everywhere. For 8kHz audio, set sample_rate=8000 both when opening the streaming connection and when sending each audio chunk. Mismatched rates cause garbled output.
  2. Use a supported streaming codec. Streaming STT accepts WAV and raw PCM (pcm_s16le, pcm_l16, pcm_raw) only.
  3. Signal end-of-audio. Call ws.flush() after the last chunk so the server knows to emit a transcript instead of waiting indefinitely for more audio.

saaras:v3 is tuned for telephony, so prefer it for call audio.

import asyncio
import base64
from sarvamai import AsyncSarvamAI
# Load 8kHz call audio and base64-encode it
with open("call_recording_8khz.wav", "rb") as f:
audio_chunk = base64.b64encode(f.read()).decode("utf-8")
async def transcribe_call_audio():
client = AsyncSarvamAI(api_subscription_key="YOUR_SARVAM_API_KEY")
async with client.speech_to_text_streaming.connect(
model="saaras:v4",
mode="transcribe",
language_code="hi-IN",
sample_rate=8000, # match the phone line
input_audio_codec="pcm_s16le",
flush_signal=True, # required to enable ws.flush()
) as ws:
await ws.transcribe(audio=audio_chunk, encoding="audio/wav", sample_rate=8000)
await ws.flush() # signal end-of-audio so the server emits a transcript
print(await ws.recv())
asyncio.run(transcribe_call_audio())

See the Streaming STT guide for the full WebSocket lifecycle, VAD, reconnection, and voice-agent barge-in.

Text-to-Speech: natural Indian voices

bulbul:v3 speaks 11 languages (10 Indian + English) with 30+ voices, and it’s built for the way Indians actually write, so you usually don’t pre-process anything.

  • Pass code-mixed text directly. Mixed Hindi-English (“आपका OTP 4321 है। Please use it…”) is spoken naturally; no need to romanize or split it.
  • Normalize English words and numbers with enable_preprocessing=true when your text has lots of abbreviations or digits.
  • Pick a voice with speaker (e.g. shubh, priya, kavya). See the voice list.
  • Output to a phone line by requesting a telephony codec, output_audio_codec supports mulaw and alaw alongside mp3, wav, opus, etc.
from sarvamai import SarvamAI
client = SarvamAI(api_subscription_key="YOUR_SARVAM_API_KEY")
with open("otp.mp3", "wb") as f:
for chunk in client.text_to_speech.convert_stream(
text="आपका OTP 4321 है। Please use it within 5 minutes.",
language_code="hi-IN",
speaker="shubh",
model="bulbul:v3",
output_audio_codec="mp3",
):
f.write(chunk)

Pronunciation control

Brand names, abbreviations, and regional terms (“NEFT”, “KYC”, a company name) don’t always come out right by default. Create a Pronunciation Dictionary with Bulbul v3 to pin how specific words are spoken, then pass its dict_id to any TTS call (REST, HTTP stream, or WebSocket):

for chunk in client.text_to_speech.convert_stream(
text="NEFT transfer karein aur KYC complete karein",
language_code="hi-IN",
speaker="shubh",
model="bulbul:v3",
dict_id="p_5cb7faa6", # your pronunciation dictionary
output_audio_codec="mp3",
):
... # write/play chunk

Document AI: Indic OCR with Sarvam Vision

Indian documents arrive as scans, photos, and PDFs in regional scripts, often with tables and mixed languages. Sarvam Vision extracts text and structure across 23 languages (22 Indian + English), preserving the original script. Digitise returns the whole document as html or md; Extract pulls just the fields you define.

It’s an asynchronous job API: create the job in one call, poll status, then fetch the output.

import os, time
from sarvamai import SarvamAI
client = SarvamAI(api_subscription_key=os.environ["SARVAM_API_KEY"])
with open("document.pdf", "rb") as f:
job = client.doc_ai.digitise(
file=[("document.pdf", f, "application/pdf")],
language="hi-IN", # primary document language (BCP-47)
output_format="md", # "html" (default) or "md"
)
TERMINAL = {"completed", "partially_completed", "failed", "rejected"}
while True:
st = client.doc_ai.get_status(job_id=job.job_id)
if st.status.lower() in TERMINAL:
break
time.sleep(5)
# ZIP containing the .md output plus page-level JSON
dl = client.doc_ai.get_download_url(job_id=job.job_id)
print("download:", dl.method, dl.url)

See the Document AI guide for schema-based Extract and the full job lifecycle. The older document_intelligence job API (create → upload → start → poll → download) still works for existing integrations, but new work should use doc_ai.

Document translation

When the deliverable is a translated file (not plain text), use the Document Translation API. Upload once, pick up to 12 target languages per job, and export each language in the same format as the source (PDF stays PDF, DOCX stays DOCX). The pipeline runs OCR or structural extraction, then translation — layout and formatting are preserved.

The flow matches other async document APIs: create job → PUT file to signed upload_url → start → poll live-status → export per language.

import os, time
from pathlib import Path
import httpx
from sarvamai import SarvamAI
client = SarvamAI(api_subscription_key=os.environ["SARVAM_API_KEY"])
document = Path("handbook.pdf")
job = client.document_translation.create(
job_name="handbook",
source_language_code="en-IN",
target_language_codes=["hi-IN", "ta-IN"],
original_filename=document.name,
)
with document.open("rb") as f:
httpx.put(
job.upload_url,
content=f.read(),
headers={
"Content-Type": "application/pdf",
"x-ms-blob-type": "BlockBlob",
},
timeout=120.0,
).raise_for_status()
client.document_translation.start(job_id=job.job_id)
TERMINAL = {"Completed", "PartiallyCompleted", "Failed"}
while True:
status = client.document_translation.get_live_status(job_id=job.job_id)
if status.job_state in TERMINAL:
break
time.sleep(12)
# Export each completed language — see Poll and export for the download loop
print("job_id:", job.job_id, "state:", status.job_state)

See Poll and export for per-language export and download URLs.

Use Document AI when you only need digitised text or structured fields. Use Document Translation when stakeholders need an editable or printable file in another language.

Text translation and transliteration

For short text (not whole files), mayura:v1 (11 languages, 10 Indian + English, colloquial modes, output-script and native-numeral control) and sarvam-translate:v1 (23 languages, 22 Indian + English, formal) handle translation, and the Transliteration API converts between scripts. See the Text translation guide for output_script and numerals_format options.

Where to go next