Skip to navigation

Text-to-Speech Rest API

View as Markdown

Provides a synchronous REST endpoint where a POST request with text returns base64-encoded audio as response.

The JSON response contains an audios array of base64-encoded WAV strings, not raw binary. Decode before saving or playing:

import base64
combined = "".join(response.audios)
wav_bytes = base64.b64decode(combined)
with open("output.wav", "wb") as f:
f.write(wav_bytes)

See TTS best practices for JavaScript and streaming examples.

Common use cases:

  • Story narration: Generate expressive audio for audiobooks and narratives
  • Podcast generation: Create natural-sounding voiceovers for episodes at scale
  • Content creation: Add voice to blogs, articles, and social media posts
  • E-learning: Build multilingual course material with clear pronunciation

What You Can Do

30+ Voices

Pick from male and female speakers, each with distinct tone and style.
Pass the speaker param to switch instantly.

11 Languages (10 Indian + English)

Hindi, Bengali, Tamil, Telugu, Kannada, Malayalam, Marathi, Gujarati, Punjabi, Odia, and English (Indian accent).
Set via language_code.

Up to 2500 Characters

Send long-form text in a single request (v3). No need to chunk or paginate your input.

Pace Control

Speed up or slow down speech with the pace parameter, range 0.5 to 2.0 for v3.

Flexible Sample Rates

8kHz to 48kHz output. Higher rates (32k, 44.1k, 48k) available in v3 REST API only. Default: 24kHz.

Multiple Audio Formats

Response is base64-encoded. Supports WAV, MP3, Linear16, Mulaw, Alaw, Opus, FLAC, and AAC.

Model: Bulbul v3

Bulbul v3 is purpose-built for Indian languages and accents. It handles code-mixed text (e.g., Hinglish), number normalization, and natural prosody out of the box, with minimal preprocessing needed.

Text to Speech Features

Basic Text to Speech Synthesis

Convert text to natural-sounding speech with high quality. Features include:

  • Multiple voice options
  • Support for Indian languages
  • Natural prosody and intonation
  • High-quality audio output
from sarvamai import SarvamAI
from sarvamai.play import save
client = SarvamAI(api_subscription_key="YOUR_SARVAM_API_KEY")
# Convert text to speech
audio = client.text_to_speech.convert(
language_code="en-IN",
text="Welcome to Sarvam AI!",
model="bulbul:v3",
speaker="shubh"
)
save(audio, "output1.wav")

API Response Format

FieldTypeDescription
request_idstringUnique identifier for the request
audiosarrayBase64-encoded audio files. Each element corresponds to an input text

Supported audio formats: WAV (default), MP3, Linear16, Mulaw, Alaw, Opus, FLAC, AAC

{
"request_id": "20241115_12345678-1234-5678-1234-567812345678",
"audios": [
"UklGRiQAAABXQVZFZm10IBAAAAABAAEAQB8AAAB9AAACABAAZGF0YQAAAAA..."
]
}

Python:

import base64
audio_base64 = response.audios[0]
audio_bytes = base64.b64decode(audio_base64)
with open("output.wav", "wb") as f:
f.write(audio_bytes)

JavaScript:

import fs from "fs";
const audioBase64 = response.audios[0];
const audioBuffer = Buffer.from(audioBase64, 'base64');
fs.writeFileSync('output.wav', audioBuffer);

Error Responses

All errors return a JSON object with an error field (message, code, request_id). The full error-code table, retry guidance, and SDK exception reference live on the central Errors & Troubleshooting page.

Errors specific to this endpoint:

HTTP StatusError CodeWhen This HappensWhat To Do
400invalid_request_errorMissing required parameters or malformed requestCheck text and language_code fields
422unprocessable_entity_errorText too long or invalid speaker/modelKeep text under 1500 chars (v2) or 2500 chars (v3)
from sarvamai import SarvamAI
from sarvamai.core.api_error import ApiError
client = SarvamAI(api_subscription_key="YOUR_SARVAM_API_KEY")
try:
response = client.text_to_speech.convert(
text="Welcome to Sarvam AI!",
language_code="en-IN",
speaker="shubh",
model="bulbul:v3"
)
# Process audio...
except ApiError as e:
if e.status_code == 400:
print(f"Bad request: {e.body}")
elif e.status_code == 403:
print("Invalid API key. Check your credentials.")
elif e.status_code == 422:
print(f"Invalid parameters: {e.body}")
elif e.status_code == 429:
print("Rate limit exceeded. Wait and retry.")
else:
print(f"Error {e.status_code}: {e.body}")

Check out our detailed API Reference to explore Text to Speech and all available options.

Need help? Contact us on discord for guidance.