Skip to navigation

Models

View as Markdown

Sarvam AI provides a purpose-built AI stack for building applications in Indian languages. Our models span speech-to-text, speech translation, text translation, and high-quality text-to-speech, designed specifically for India’s linguistic diversity, accents, and real-world usage patterns.

Each model is trained and evaluated on Indian languages and culturally grounded data, enabling higher accuracy in production scenarios. With simple, well-documented APIs and predictable performance, developers can build, deploy, and scale India-first AI experiences without managing model complexity.

New to building for Indian languages? Start with Building for Indian Languages, a practical guide to language coverage, code-mixing, scripts, native numerals, 8kHz telephony audio, and pronunciation control.

Model Selection Guide

Sarvam Models: trained for Indian languages:

ModelAPIDescription
Saaras v4Speech to TextState-of-the-art ASR with 23 language support (22 Indian + Global/Indian English) and multiple output modes: transcribe, translate, verbatim, translit, codemix. Default, recommended model; v3 remains available.
Bulbul v3Text to SpeechNatural-sounding voices for 11 languages (10 Indian + English) with customizable pitch, pace, and speaker options.
Sarvam Voice CloningVoice CloningCreate a voice from a short reference clip and synthesize speech with it across 12 Indian languages - clone inline per request, or save the voice once and reuse its ID.
Sarvam-105BChat Completion105B parameter flagship model, sarvam-105b for complex reasoning and agentic tasks, plus sarvam-105b-conversations for real-time dialogue and voice agents.
MayuraText TranslationHigh-quality translation between 11 languages (10 Indian + English) with context preservation.
Sarvam-TranslateText TranslationExtended translation support for all 23 languages (22 Indian + English) with superior accuracy.
Sarvam VisionDocument IntelligenceExtract and digitize content from documents in 23 languages with accurate OCR and structured output.

Open-Weight Models

Open-weight models are available on /v2 using the same Sarvam API key and credits.

GLM-5.3, Gemma 4 31B, and DeepSeek V4 Flash are available in beta and rolling out gradually. See Access to Beta APIs.

ModelContext windowModalityReasoning
GLM-5.31,048,576 tokensText → textYes
Gemma 4 31B131,072 tokensText + image → textNo
DeepSeek V4 Flash1,048,576 tokensText → textYes

Read more about the open-weight models →

Language Support Overview

Language coverage varies by model. Check the table below before picking one. Full per-model tables are linked from each model’s own page.

ModelLanguagesStatus
Saaras v4 (Speech to Text)23 (22 Indian + Global/Indian English), full listRecommended
Sarvam Translate (Text Translation)23 (22 Indian + English), full listActive
Sarvam Vision (Document Intelligence)23 (22 Indian + English), full listActive
Bulbul v3 (Text to Speech)11 (10 Indian + English), full listActive
Sarvam Voice Cloning12 Indian languages, full listActive
Mayura (Text Translation)11 (10 Indian + English), full listActive
Sarvam-105B (Chat LLM)11 (10 Indian + English), sarvam-105b, sarvam-105b-conversationsActive
GLM-5.3 (Chat LLM, open-weight)See the model page for language coverageBeta
Gemma 4 31B (Chat LLM, open-weight)See the model page for language coverageBeta
DeepSeek V4 Flash (Chat LLM, open-weight)See the model page for language coverageBeta

23-language set (Saaras v4, Sarvam Translate, Sarvam Vision)

LanguageCodeLanguageCode
Hindihi-INAssameseas-IN
Bengalibn-INUrduur-IN
Kannadakn-INNepaline-IN
Malayalamml-INKonkanikok-IN
Marathimr-INKashmiriks-IN
Odiaod-INSindhisd-IN
Punjabipa-INSanskritsa-IN
Tamilta-INSantalisat-IN
Telugute-INManipurimni-IN
Englishen-INBodobrx-IN
Gujaratigu-INMaithilimai-IN
Dogridoi-IN

11-language set (Bulbul v3, Mayura, Sarvam-105B)

LanguageCodeLanguageCode
Hindihi-INKannadakn-IN
Bengalibn-INMalayalamml-IN
Tamilta-INMarathimr-IN
Telugute-INPunjabipa-IN
Gujaratigu-INOdiaod-IN
Englishen-IN

Use Cases

Build a multilingual voice assistant

  1. Speech Input: Use Saaras v4 with mode="transcribe" to convert user speech to text
  2. Understanding: Process with Sarvam-105B for intelligent responses
  3. Speech Output: Convert responses to speech with Bulbul

Perfect for customer service, smart home devices, and accessibility applications.

Learn how to build a voice agent with LiveKit →