mohit-sarvam kurianbenoy commited on
Commit
729f7b8
·
verified ·
1 Parent(s): 75102ae

Update README.md (#2)

Browse files

- Update README.md (f7012ed104a7f950fea4b3ec74e222e40a337c5c)
- Update README.md (33143c2dc11d0670921789cdf9f5d36620337f29)


Co-authored-by: Kurian Benoy <kurianbenoy@users.noreply.huggingface.co>

Files changed (1) hide show
  1. README.md +53 -0
README.md CHANGED
@@ -94,6 +94,59 @@ print("reasoning content:", reasoning_content)
94
  print("content:", content)
95
  ```
96
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
97
  # VLLM Deployment
98
 
99
  For easy deployment, we can use `vllm>=0.8.5` and create an OpenAI-compatible API endpoint with `vllm serve sarvamai/sarvam-m`
 
94
  print("content:", content)
95
  ```
96
 
97
+ # How to use with Sarvam APIs
98
+
99
+ ```python
100
+ from openai import OpenAI
101
+
102
+ base_url = "https://api.sarvam.ai/v1"
103
+ model_name = "sarvam-m"
104
+ api_key = "Your-API-Key" # get it from https://dashboard.sarvam.ai/
105
+
106
+
107
+ client = OpenAI(
108
+ base_url=base_url,
109
+ api_key=api_key,
110
+ ).with_options(max_retries=1)
111
+
112
+ response = client.chat.completions.create(
113
+ model=model_name,
114
+ messages=[
115
+ {"role": "system", "content": "say hi"},
116
+ {"role": "user", "content": "say hi"},
117
+ ],
118
+ stream=False,
119
+ max_completion_tokens=2048,
120
+ # reasoning_effort="low", # set either of 3 values to enable reasoning
121
+ )
122
+ print(response.choices[0].message.content)
123
+
124
+ response1 = client.chat.completions.create(
125
+ model=model_name,
126
+ messages=[
127
+ {"role": "system", "content": "You're a helpful AI assistant"},
128
+ {"role": "user", "content": "Explain quantum computing in simple terms"}
129
+ ],
130
+ max_completion_tokens=4096,
131
+ reasoning_effort="medium" # Optional reasoning mode
132
+ )
133
+ print("First response:", response1.choices[0].message.content)
134
+
135
+ # Second turn (using previous response as context)
136
+ response2 = client.chat.completions.create(
137
+ model=model_name,
138
+ messages=[
139
+ {"role": "system", "content": "You're a helpful AI assistant"},
140
+ {"role": "user", "content": "Explain quantum computing in simple terms"},
141
+ {"role": "assistant", "content": response1.choices[0].message.content}, # Previous response
142
+ {"role": "user", "content": "Can you give an analogy for superposition?"}
143
+ ],
144
+ reasoning_effort="high",
145
+ max_completion_tokens=8192,
146
+ )
147
+ print("Follow-up response:", response2.choices[0].message.content)
148
+ ```
149
+
150
  # VLLM Deployment
151
 
152
  For easy deployment, we can use `vllm>=0.8.5` and create an OpenAI-compatible API endpoint with `vllm serve sarvamai/sarvam-m`