Skip to navigation

Updated Pricing for Language Models

Language model prices have changed. sarvam-105b and sarvam-105b-conversations now cost ₹15 input, ₹5 cached input, and ₹60 output per 1M tokens. gemma4 is now ₹35 / ₹10 / ₹90. glm5.3 cached input and output drop to ₹20 and ₹390, and input stays at ₹126. deepseekv4.1-flash is now ₹25 / ₹0.5 / ₹105. See Pricing for the full table.

Session Affinity for Chat Completion V2

POST /v2/chat/completions now accepts an optional x-session-affinity header. Send the same session ID on every turn of a conversation, and later turns can reuse cached prompt context for lower latency. Session IDs are scoped to your API subscription, and requests without the header work as before. See Session affinity.