--- library_name: transformers license: apache-2.0 language: - en - bn - hi - kn - gu - mr - ml - or - pa - ta - te base_model: - mistralai/Mistral-Small-3.1-24B-Instruct-2503 --- ## Model Information Sarvam-M multilingual hybrid reasoning llm is an instruction tuned generative model in 24B (text in/text out) post trained over Mistral 3.1 24B. It significantly improves on the base Mistral model: +20% average improvement on Indian language benchmarks, +21.6% on math benchmarks, and +17.6% on programming benchmarks. The gains in tasks in the intersectionality of Indian languages and math are even higher, e.g., +86% improvement in a romanized Indian language GSM-8K benchmark. Learn in detail about sarvam-M in our [blog post](link) ## Key Features - **Hybrid thinking mode** A single model supports both "think" and "non-think" modes. Use the think mode for tasks requiring complex logical reasoning, math, and coding, and switch to the non-think mode for efficient, general-purpose conversation. - **Indic Skills** Specifically post-trained on Indian languages alongside English, the model also embodies a character that reflects and emphasizes Indian cultural values. - **Reasoning capabilities** Sarvam-M outperforms most models of similar size on coding and math benchmarks, demonstrating strong reasoning capabilities. - **Chatting Experience** With support for both Indic scripts and romanized versions of Indian languages, Sarvam-M offers a smooth and accessible multilingual chat experience. ## Quickstart The following contains a code snippet illustrating how to use the model generate content based on given inputs. ```python from transformers import AutoModelForCausalLM, AutoTokenizer model_name = "sarvamai/sarvam-M" # load the tokenizer and the model tokenizer = AutoTokenizer.from_pretrained(model_name) model = AutoModelForCausalLM.from_pretrained( model_name, torch_dtype="auto", device_map="auto" ) # prepare the model input prompt = "Who are you and what is your purpose on this planet?" messages = [{"role": "user", "content": prompt}] text = tokenizer.apply_chat_template( messages, tokenize=False, enable_thinking=True, # Switches between thinking and non-thinking modes. Default is True. ) model_inputs = tokenizer([text], return_tensors="pt").to(model.device) # conduct text completion generated_ids = model.generate(**model_inputs, max_new_tokens=8192) output_ids = generated_ids[0][len(model_inputs.input_ids[0]) :].tolist() output_text = tokenizer.decode(output_ids) if "" in output_text: reasoning_content = output_text.split("")[0].rstrip("\n") content = output_text.split("")[-1].lstrip("\n").rstrip("") else: reasoning_content = "" content = output_text.rstrip("") print("reasoning content:", reasoning_content) print("content:", content) ``` ## VLLM Deployment For deployment, you can use `vllm>=0.8.5` to create an OpenAI-compatible API endpoint: ```shell vllm serve sarvamai/sarvam-M ``` For inference and switching between thinking and non-thinking mode, refer to the below python code: ```python from openai import OpenAI # Modify OpenAI's API key and API base to use vLLM's API server. openai_api_key = "EMPTY" openai_api_base = "http://localhost:8000/v1" client = OpenAI( api_key=openai_api_key, base_url=openai_api_base, ) models = client.models.list() model = models.data[0].id messages = [{"role": "user", "content": "How many letter r in word strawberry?"}] # By default, the model is in thinking mode. # If you want to disable thinking, add: # extra_body={"chat_template_kwargs": {"enable_thinking": False}} response = client.chat.completions.create(model=model, messages=messages) output_text = response.choices[0].message.content if "" in output_text: reasoning_content = output_text.split("")[0].rstrip("\n") content = output_text.split("")[-1].lstrip("\n").rstrip("") else: reasoning_content = "" content = output_text.rstrip("") print("reasoning content:", reasoning_content) print("content:", content) # For the next round, add the assistant's response and reasoning to the messages. messages.append( {"role": "assistant", "content": content, "reasoning_content": reasoning_content} ) ``` The above example also shows how to add assistant turns in the messages for multiturn conversation.