mohit-sarvam commited on
Commit
2bf88b5
·
verified ·
1 Parent(s): a198288

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +8 -5
README.md CHANGED
@@ -116,7 +116,7 @@ messages = [
116
  response1 = client.chat.completions.create(
117
  model=model_name,
118
  messages=messages,
119
- reasoning_effort="medium", # Optional reasoning mode
120
  max_completion_tokens=4096,
121
  )
122
  print("First response:", response1.choices[0].message.content)
@@ -135,7 +135,7 @@ messages.extend(
135
  response2 = client.chat.completions.create(
136
  model=model_name,
137
  messages=messages,
138
- reasoning_effort="high",
139
  max_completion_tokens=8192,
140
  )
141
  print("Follow-up response:", response2.choices[0].message.content)
@@ -143,6 +143,8 @@ print("Follow-up response:", response2.choices[0].message.content)
143
 
144
  Refer to API docs here: [sarvam API docs](https://docs.sarvam.ai/api-reference-docs/introduction)
145
 
 
 
146
  # VLLM Deployment
147
 
148
  For easy deployment, we can use `vllm>=0.8.5` and create an OpenAI-compatible API endpoint with `vllm serve sarvamai/sarvam-m`
@@ -195,7 +197,8 @@ The repo contains bf16 and q8 gguf files built using https://github.com/ggml-org
195
  You can use the model using cli as explained in docs https://github.com/ggml-org/llama.cpp/tree/master/tools/main
196
 
197
  Example Command:
 
 
 
198
 
199
- `./build/bin/llama-cli -i -m /projects/data/romit_sarvam_ai/models/gguf/sarvam-m-q8_0.gguf -c 8192 -t 16`
200
-
201
- We got about 4 tokens per second on 16 cores.
 
116
  response1 = client.chat.completions.create(
117
  model=model_name,
118
  messages=messages,
119
+ reasoning_effort="medium", # Enable thinking mode
120
  max_completion_tokens=4096,
121
  )
122
  print("First response:", response1.choices[0].message.content)
 
135
  response2 = client.chat.completions.create(
136
  model=model_name,
137
  messages=messages,
138
+ reasoning_effort="medium",
139
  max_completion_tokens=8192,
140
  )
141
  print("Follow-up response:", response2.choices[0].message.content)
 
143
 
144
  Refer to API docs here: [sarvam API docs](https://docs.sarvam.ai/api-reference-docs/introduction)
145
 
146
+ The model has only two modes: `think` and `no-think`. If `reasoning_effort` is set then thinking mode is on otherwise it is off.
147
+
148
  # VLLM Deployment
149
 
150
  For easy deployment, we can use `vllm>=0.8.5` and create an OpenAI-compatible API endpoint with `vllm serve sarvamai/sarvam-m`
 
197
  You can use the model using cli as explained in docs https://github.com/ggml-org/llama.cpp/tree/master/tools/main
198
 
199
  Example Command:
200
+ ```
201
+ ./build/bin/llama-cli -i -m /projects/data/romit_sarvam_ai/models/gguf/sarvam-m-q8_0.gguf -c 8192 -t 16
202
+ ```
203
 
204
+ We got about 4 tokens per second on 16 cores with q8 model.