Sarvam-30B (Deprecated)
Sarvam-30B (Deprecated)
Deprecated Model: Sarvam-30B has been deprecated. Please migrate to Sarvam-105B for improved performance across all tasks. The information below is retained for reference only.
Sarvam-30B (Chat LLM)
A 30B parameter Mixture-of-Experts reasoning model trained from scratch, optimized for Indian languages with only 2.4B active parameters per token. Delivers strong reasoning, coding, and conversational capabilities while remaining efficient to deploy.
Highlights:
- 30B total parameters, 2.4B active: efficient MoE architecture with Grouped Query Attention
- Pre-trained on 16 trillion tokens across code, math, multilingual, and web data
- State-of-the-art Indian language performance across native and romanized scripts
- Optimized inference for H100, L40S, and Apple Silicon (MXFP4)
- OpenAI-compatible chat completions API
At a Glance
Key Features
Learn More
For detailed information on architecture, training methodology, performance benchmarks, and inference optimizations, visit our blog.
Model Specifications
Key Capabilities
Basic Chat Completion
Multi-turn Conversation
Streaming
Simple, one-turn interaction where the user asks a question and the model replies with a single, direct response.