Instructions to use ai-sage/GigaChat3-10B-A1.8B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use ai-sage/GigaChat3-10B-A1.8B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="ai-sage/GigaChat3-10B-A1.8B") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("ai-sage/GigaChat3-10B-A1.8B") model = AutoModelForCausalLM.from_pretrained("ai-sage/GigaChat3-10B-A1.8B", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use ai-sage/GigaChat3-10B-A1.8B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "ai-sage/GigaChat3-10B-A1.8B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ai-sage/GigaChat3-10B-A1.8B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/ai-sage/GigaChat3-10B-A1.8B
- SGLang
How to use ai-sage/GigaChat3-10B-A1.8B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "ai-sage/GigaChat3-10B-A1.8B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ai-sage/GigaChat3-10B-A1.8B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "ai-sage/GigaChat3-10B-A1.8B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ai-sage/GigaChat3-10B-A1.8B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use ai-sage/GigaChat3-10B-A1.8B with Docker Model Runner:
docker model run hf.co/ai-sage/GigaChat3-10B-A1.8B
Parsing Non‑Standard Function Calls from GigaChat3‑10B‑A1.8B Model Responses
I am trying to use the OpenAI‑compatible API (the v1/chat/completions endpoint) with the model ai-sage/GigaChat3-10B-A1.8B hosted on Hugging Face. In the request I include a tools section (function‑calling) according to the OpenAI specification:
{
"model": "ai-sage/GigaChat3-10B-A1.8B",
"messages": [
{"role": "user", "content": "Save to my personal memory: prefers short answers"}
],
"tools": [
{
"type": "function",
"function": {
"name": "manage_user_memory",
"description": "...",
"parameters": {
"type": "object",
"properties": {
"content": {"anyOf":[{"type":"string"},{"type":"null"}],"default":null},
"action": {"type":"string","enum":["create","update","delete"],"default":"create"},
"id": {"anyOf":[{"type":"string","format":"uuid"},{"type":"null"}],"default":null}
}
}
}
}
]
}
The server’s response looks like this:
{
"id": "...",
"model": "ai-sage/GigaChat3-10B-A1.8B",
"choices": [
{
"message": {
"role": "assistant",
"content": "<|message_sep|>\n\nfunction call<|role_sep|>\n{\"name\": \"manage_user_memory\", \"arguments\": {\"action\": \"create\", \"content\": \"Prefers short answers\"}}",
"function_call": null,
"tool_calls": []
},
"finish_reason": "stop"
}
]
}
The model does produce a function call, but it does not populate the official function_call or tool_calls fields defined by OpenAI. Instead, it embeds a serialized function call as a plain string inside the content field. Because of this, client libraries such as openai, LangChain, or vLLM/sglang cannot automatically detect and handle the tool call; it has to be parsed manually
Could anyone share a reliable code snippet (Python, JavaScript, or another language) that extracts the function name and arguments from the content field returned by the model ai-sage/GigaChat3-10B-A1.8B? The response embeds a serialized function call inside a plain string, e.g.:
{
"role": "assistant",
"content": "<|message_sep|>\n\nfunction call<|role_sep|>\n{\"name\": \"manage_user_memory\", \"arguments\": {\"action\": \"create\", \"content\": \"Prefers short answers\"}}"
}
Hi! Please use this function to extract the function name and arguments, make sure there is no EOS at the end:
import json
import re
REGEX_FUNCTION_CALL_V3 = re.compile(r"function call<\|role_sep\|>\n(.*)$", re.DOTALL)
REGEX_CONTENT_PATTERN = re.compile(r"^(.*?)<\|message_sep\|>", re.DOTALL)
def parse_function_and_content(completion_str: str):
"""
Using the regexes the user provided, attempt to extract function call and content.
Returns (function_call_str_or_None, content_str_or_None)
"""
function_call = None
content = None
m_func = REGEX_FUNCTION_CALL_V3.search(completion_str)
if m_func:
try:
function_call = json.loads(m_func.group(1))
if isinstance(function_call, dict) and "name" in function_call and "arguments" in function_call:
if not isinstance(function_call["arguments"], dict):
function_call = None
else:
function_call = None
except json.JSONDecodeError:
function_call = None
# will return raw string in failed attempt of function calling
return function_call, completion_str
m_content = REGEX_CONTENT_PATTERN.search(completion_str)
if m_content:
content = m_content.group(1)
else:
# as a fallback, everything before the first message_sep marker if present
if "<|message_sep|>" in completion_str:
content = completion_str.split("<|message_sep|>")[0]
else:
content = completion_str
return function_call, content
Which --tool-call-parser should I use to launch vllm so that it parses this specific format?
UPD: It appears there is no existing parser in vllm for this format. I have vibe-coded a custom plugin for the required format (use it at your own risk).
https://gist.github.com/bulatovv/b9b5116a0af14fe09164146ed8eabafb
To use it, pass the full path to the plugin with these parameters:
--enable-auto-tool-choice \
--tool-parser-plugin /.../.../gigachat3_tool_parser.py \
--tool-call-parser gigachat3 \
Full support in vLLM, SGLang and llama.cpp It's on the way.
Our temporary solution is available for vLLM in a separate branch - https://github.com/vllm-project/vllm/pull/29905 .
In llama.cpp function calls are also available, but with a number of technical limitations — we have added detailed instructions to the description of the GGUF model.
(APIServer pid=1) ERROR 12-16 21:17:24 [serving_chat.py:1292] Error in chat completion stream generator.
(APIServer pid=1) ERROR 12-16 21:17:24 [serving_chat.py:1292] Traceback (most recent call last):
(APIServer pid=1) ERROR 12-16 21:17:24 [serving_chat.py:1292] File "/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/serving_chat.py", line 1168, in chat_completion_stream_generator
(APIServer pid=1) ERROR 12-16 21:17:24 [serving_chat.py:1292] actual_call = tool_parser.streamed_args_for_tool[index]
(APIServer pid=1) ERROR 12-16 21:17:24 [serving_chat.py:1292] ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~^^^^^^^
(APIServer pid=1) ERROR 12-16 21:17:24 [serving_chat.py:1292] IndexError: list index out of range
@gdagil could you share an example? Also you can try new, more robust version - https://github.com/vllm-project/vllm/pull/30338