rahular commited on
Commit
8f43073
·
verified ·
1 Parent(s): 6919098

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +8 -12
README.md CHANGED
@@ -98,7 +98,7 @@ print("content:", content)
98
  > For thinking mode, we recommend `temperature=0.4`; for no-think mode, `temperature=0.2`.
99
 
100
 
101
- # How to use with Sarvam APIs
102
 
103
  ```python
104
  from openai import OpenAI
@@ -120,7 +120,7 @@ messages = [
120
  response1 = client.chat.completions.create(
121
  model=model_name,
122
  messages=messages,
123
- reasoning_effort="medium", # Enable thinking mode
124
  max_completion_tokens=4096,
125
  )
126
  print("First response:", response1.choices[0].message.content)
@@ -151,9 +151,9 @@ Refer to API docs here: [sarvam API docs](https://docs.sarvam.ai/api-reference-d
151
 
152
  # VLLM Deployment
153
 
154
- For easy deployment, we can use `vllm>=0.8.5` and create an OpenAI-compatible API endpoint with `vllm serve sarvamai/sarvam-m`
155
 
156
- For more control, we can use vllm in Python. That way, we can explicitly enable or disable thinking mode.
157
 
158
  ```python
159
  from openai import OpenAI
@@ -172,7 +172,7 @@ model = models.data[0].id
172
 
173
  messages = [{"role": "user", "content": "Why is 42 the best number?"}]
174
 
175
- # By default, the model is in thinking mode.
176
  # If you want to disable thinking, add:
177
  # extra_body={"chat_template_kwargs": {"enable_thinking": False}}
178
  response = client.chat.completions.create(model=model, messages=messages)
@@ -196,13 +196,9 @@ messages.append(
196
 
197
  # Running the model on a CPU
198
 
199
- The repo contains bf16 and q8 gguf files built using https://github.com/ggml-org/llama.cpp/blob/master/docs/build.md#cpu-build
200
-
201
- You can use the model using cli as explained in docs https://github.com/ggml-org/llama.cpp/tree/master/tools/main
202
 
203
  Example Command:
204
  ```
205
- ./build/bin/llama-cli -i -m /projects/data/romit_sarvam_ai/models/gguf/sarvam-m-q8_0.gguf -c 8192 -t 16
206
- ```
207
-
208
- We got about 4 tokens per second on 16 cores with q8 model.
 
98
  > For thinking mode, we recommend `temperature=0.4`; for no-think mode, `temperature=0.2`.
99
 
100
 
101
+ # With Sarvam APIs
102
 
103
  ```python
104
  from openai import OpenAI
 
120
  response1 = client.chat.completions.create(
121
  model=model_name,
122
  messages=messages,
123
+ reasoning_effort="medium", # Enable thinking mode. `None` for disable.
124
  max_completion_tokens=4096,
125
  )
126
  print("First response:", response1.choices[0].message.content)
 
151
 
152
  # VLLM Deployment
153
 
154
+ For easy deployment, we can use `vllm>=0.8.5` and create an OpenAI-compatible API endpoint with `vllm serve sarvamai/sarvam-m`.
155
 
156
+ If you want to use vLLM with python, you can do the following.
157
 
158
  ```python
159
  from openai import OpenAI
 
172
 
173
  messages = [{"role": "user", "content": "Why is 42 the best number?"}]
174
 
175
+ # By default, thinking mode is enabled.
176
  # If you want to disable thinking, add:
177
  # extra_body={"chat_template_kwargs": {"enable_thinking": False}}
178
  response = client.chat.completions.create(model=model, messages=messages)
 
196
 
197
  # Running the model on a CPU
198
 
199
+ This repo contains quantized (q8) version of the model as well. You can use the model on your local machine (without gpu) as explained [here](docs https://github.com/ggml-org/llama.cpp/tree/master/tools/main).
 
 
200
 
201
  Example Command:
202
  ```
203
+ ./build/bin/llama-cli -i -m /your/folder/path/sarvam-m-q8_0.gguf -c 8192 -t 16
204
+ ```