> For clean Markdown of any page, append `.md` to the page URL. > For a complete documentation index, see https://docs.sarvam.ai/llms.txt. > For full documentation content in one file, see https://docs.sarvam.ai/llms-full.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.sarvam.ai/_mcp/server. # Large Language Models ## Docs - [Chat Completion API](https://docs.sarvam.ai/api/api-guides-tutorials/chat-completion/overview.md): Get started with Sarvam AI LLM models for conversational AI. Build intelligent chat applications with native Indian language support and deep contextual reasoning capabilities. - [Responses API](https://docs.sarvam.ai/api/api-guides-tutorials/chat-completion/responses-api.md): Create stateless text, reasoning, structured-output, and tool-calling responses with Sarvam's OpenAI-compatible Responses API. - [How to list your chat messages](https://docs.sarvam.ai/api/api-guides-tutorials/chat-completion/how-to/list-your-chat-messages.md): Defines your entire conversation. - [How to control response randomness with `temperature`](https://docs.sarvam.ai/api/api-guides-tutorials/chat-completion/how-to/control-response-randomness.md): Control how focused or varied model responses are. - [How to control response diversity with `top_p`](https://docs.sarvam.ai/api/api-guides-tutorials/chat-completion/how-to/control-response-diversity.md): Method used to generate text by limiting the possibilities of the next word - [How to adjust the model's thinking level with `reasoning_effort`](https://docs.sarvam.ai/api/api-guides-tutorials/chat-completion/how-to/adjust-the-models-thinking-level.md): controls **how much effort the model puts into reasoning. - [How to encourage new topics with `presence_penalty`](https://docs.sarvam.ai/api/api-guides-tutorials/chat-completion/how-to/encourage-new-topics-in-response.md): Helps you steer the model toward introducing new concepts or topics. - [How to reduce repetition with `frequency_penalty`](https://docs.sarvam.ai/api/api-guides-tutorials/chat-completion/how-to/reduce-repetition-words-or-phrases-in-response.md): Helps you control how often the model repeats words or phrases - [How to get repeatable results using `seed`](https://docs.sarvam.ai/api/api-guides-tutorials/chat-completion/how-to/get-repeatable-results.md): Reduce output variation with best-effort seeded sampling - [How to control the response length with `max_tokens`](https://docs.sarvam.ai/api/api-guides-tutorials/chat-completion/how-to/control-the-response-length.md): control how long the model's response can be - [How to control where the model stops using `stop`](https://docs.sarvam.ai/api/api-guides-tutorials/chat-completion/how-to/control-where-the-model-stops.md): Tell the model to **stop generating further tokens.