Sarvam Vision
Sarvam Vision is a 3B parameter state-space Vision Language Model (VLM) purpose-built for high-accuracy Document Intelligence. It powers our Document Intelligence pipeline.
At a Glance
Why Sarvam Vision?
One of the most challenging problems in vision AI today is accurate document intelligence for Indian languages. Much of India’s knowledge—historical texts, government records, academic papers, and cultural archives—remains locked in libraries, scanned collections, and legacy documents. Unlocking this vast repository is essential for preserving cultural heritage and making knowledge accessible.
While frontier Vision Language Models have set a high bar for processing modern English documents, a significant gap remains: most global models treat Indian languages as secondary, often resulting in lower accuracy for regional scripts. Sarvam Vision bridges this gap with native support for 22 Indian languages, delivering world-class accuracy where others fall short.
Want to learn more about how we built Sarvam Vision? Check out our blog post.
What You Can Do
- Text Extraction: Extract text from PDFs and scanned documents in 23 languages (22 Indian + English)
- Tables: Convert complex tables to HTML or Markdown
- Structure Preservation: Maintain document layout, reading order, and hierarchies
Supported Languages
All 22 official Indian languages plus English:
Capabilities
Text Extraction
Tables
Multilingual
High-Fidelity Document Intelligence
Sarvam Vision extracts text from documents with exceptional accuracy, preserving the original structure and reading order across 23 languages (22 Indian + English).
Features:
- High-accuracy text extraction from PDFs and scanned documents
- Preserves document layout and reading order
- Native support for all Indian scripts
- Outputs clean HTML or Markdown
Quick Start
Get started with Document AI for high-fidelity text extraction across all supported languages. Create the job in one call, poll until it reaches a terminal state, then fetch the output.
See the Document AI overview for schema-based Extract, output formats, and the full job lifecycle.
Legacy: the document_intelligence job API
Superseded by doc_ai above. The older Document Digitization API used a multi-step
job flow (create → upload → start → poll → download) and is still supported for
existing integrations. New integrations should use doc_ai.
Model Specifications
- Model Size: 3B parameters
- Supported Input Formats: PDF, PNG, JPG, ZIP (flat archive with JPG/PNG document pages)
- Output Formats: HTML, Markdown (md) (delivered as ZIP file). JSON with structured page-level data is always included by default, regardless of the chosen output format.
- Languages: 23 languages (22 Indian + English)