Text Generation
Transformers
Safetensors
Russian
English
deepseek_v3
instruct
Mixture of Experts
multilingual
bf16
long-context
tool-use
retrieval
SMITH-Exp
conversational
text-generation-inference
Instructions to use ai-forever/SMITH-Exp with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use ai-forever/SMITH-Exp with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="ai-forever/SMITH-Exp") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("ai-forever/SMITH-Exp") model = AutoModelForCausalLM.from_pretrained("ai-forever/SMITH-Exp", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use ai-forever/SMITH-Exp with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "ai-forever/SMITH-Exp" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ai-forever/SMITH-Exp", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/ai-forever/SMITH-Exp
- SGLang
How to use ai-forever/SMITH-Exp with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "ai-forever/SMITH-Exp" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ai-forever/SMITH-Exp", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "ai-forever/SMITH-Exp" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ai-forever/SMITH-Exp", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use ai-forever/SMITH-Exp with Docker Model Runner:
docker model run hf.co/ai-forever/SMITH-Exp
Add SMITH-10B-Exp tokenizer, chat template, and GigaChat-3 plugin.
Browse files- .gitattributes +1 -0
- README.md +112 -0
- chat_template.jinja +348 -0
- config.json +53 -0
- generation_config.json +7 -0
- gigachat3_guided_decoding.py +390 -0
- requirements.txt +17 -0
- special_tokens_map.json +16 -0
- tokenizer.json +3 -0
- tokenizer_config.json +131 -0
.gitattributes
CHANGED
|
@@ -33,3 +33,4 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
|
|
| 33 |
*.zip filter=lfs diff=lfs merge=lfs -text
|
| 34 |
*.zst filter=lfs diff=lfs merge=lfs -text
|
| 35 |
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
|
|
|
|
|
| 33 |
*.zip filter=lfs diff=lfs merge=lfs -text
|
| 34 |
*.zst filter=lfs diff=lfs merge=lfs -text
|
| 35 |
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
| 36 |
+
tokenizer.json filter=lfs diff=lfs merge=lfs -text
|
README.md
CHANGED
|
@@ -1,3 +1,115 @@
|
|
| 1 |
---
|
| 2 |
license: mit
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 3 |
---
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
---
|
| 2 |
license: mit
|
| 3 |
+
language:
|
| 4 |
+
- ru
|
| 5 |
+
- en
|
| 6 |
+
base_model:
|
| 7 |
+
- ai-sage/GigaChat3.1-10B-A1.8B-bf16
|
| 8 |
+
pipeline_tag: text-generation
|
| 9 |
+
library_name: transformers
|
| 10 |
+
tags:
|
| 11 |
+
- instruct
|
| 12 |
+
- moe
|
| 13 |
+
- multilingual
|
| 14 |
+
- bf16
|
| 15 |
+
- long-context
|
| 16 |
+
- tool-use
|
| 17 |
+
- retrieval
|
| 18 |
+
- SMITH-10B-Exp
|
| 19 |
---
|
| 20 |
+
|
| 21 |
+
# SMITH-10B-Exp
|
| 22 |
+
|
| 23 |
+
Recommended to be used with [SMITH harness](https://github.com/smit-org/smith-harness).
|
| 24 |
+
|
| 25 |
+
**SMITH-10B-Exp** is an instruct Mixture-of-Experts checkpoint fine-tuned from [GigaChat 3.1 Lightning](https://huggingface.co/ai-sage/GigaChat3.1-10B-A1.8B-bf16) (10B total parameters, 1.8B active, BF16) for multi-hop retrieval with GigaChat-3 tool calls (`search_index`, then `answer` with `supporting_corpus_ids`). This folder ships the tokenizer, chat template, and `gigachat3_guided_decoding.py` (overlay BF16 weights when packing a full Hub snapshot).
|
| 26 |
+
|
| 27 |
+
## Model architecture
|
| 28 |
+
|
| 29 |
+
DeepSeek-V3 MoE (`DeepseekV3ForCausalLM`) with Multi-head Latent Attention (MLA) and Multi-Token Prediction (MTP). Hidden size 1536, 26 layers, 64 routed experts (4 per token), YaRN context to 262144.
|
| 30 |
+
|
| 31 |
+
## Usage
|
| 32 |
+
|
| 33 |
+
### Serving dependencies
|
| 34 |
+
|
| 35 |
+
```bash
|
| 36 |
+
uv venv --python python3.11 --seed .venv
|
| 37 |
+
uv pip install -r requirements.txt \
|
| 38 |
+
--python .venv/bin/python \
|
| 39 |
+
--torch-backend=cu130 \
|
| 40 |
+
--index-strategy unsafe-best-match
|
| 41 |
+
```
|
| 42 |
+
|
| 43 |
+
Package pins and licenses are in [`requirements.txt`](requirements.txt). SMITH-10B-Exp BF16 fits one 80GB GPU.
|
| 44 |
+
|
| 45 |
+
### transformers
|
| 46 |
+
|
| 47 |
+
```python
|
| 48 |
+
import torch
|
| 49 |
+
from transformers import AutoTokenizer, AutoModelForCausalLM, GenerationConfig
|
| 50 |
+
|
| 51 |
+
model_name = "ai-forever/SMITH-10B-Exp"
|
| 52 |
+
tokenizer = AutoTokenizer.from_pretrained(model_name)
|
| 53 |
+
model = AutoModelForCausalLM.from_pretrained(
|
| 54 |
+
model_name,
|
| 55 |
+
torch_dtype=torch.bfloat16,
|
| 56 |
+
device_map="auto",
|
| 57 |
+
)
|
| 58 |
+
model.generation_config = GenerationConfig.from_pretrained(model_name)
|
| 59 |
+
messages = [
|
| 60 |
+
{"role": "user", "content": "Which corpus passages support the answer?"}
|
| 61 |
+
]
|
| 62 |
+
prompt = tokenizer.apply_chat_template(
|
| 63 |
+
messages,
|
| 64 |
+
tokenize=False,
|
| 65 |
+
add_generation_prompt=True,
|
| 66 |
+
)
|
| 67 |
+
inputs = tokenizer(prompt, return_tensors="pt")
|
| 68 |
+
inputs = {k: v.to(model.device) for k, v in inputs.items()}
|
| 69 |
+
outputs = model.generate(**inputs, max_new_tokens=512)
|
| 70 |
+
prompt_len = inputs["input_ids"].shape[1]
|
| 71 |
+
print(tokenizer.decode(outputs[0][prompt_len:], skip_special_tokens=True))
|
| 72 |
+
```
|
| 73 |
+
|
| 74 |
+
### vLLM
|
| 75 |
+
|
| 76 |
+
```bash
|
| 77 |
+
vllm serve ai-forever/SMITH-10B-Exp \
|
| 78 |
+
--trust-remote-code \
|
| 79 |
+
--enable-auto-tool-choice \
|
| 80 |
+
--tool-call-parser gigachat3 \
|
| 81 |
+
--tool-parser-plugin ./gigachat3_guided_decoding.py \
|
| 82 |
+
--chat-template ./chat_template.jinja \
|
| 83 |
+
--tensor-parallel-size 1 \
|
| 84 |
+
--dtype bfloat16
|
| 85 |
+
```
|
| 86 |
+
|
| 87 |
+
```bash
|
| 88 |
+
curl http://localhost:8000/v1/chat/completions \
|
| 89 |
+
-H "Content-Type: application/json" \
|
| 90 |
+
-d '{
|
| 91 |
+
"model": "ai-forever/SMITH-10B-Exp",
|
| 92 |
+
"temperature": 0,
|
| 93 |
+
"tool_choice": "required",
|
| 94 |
+
"messages": [
|
| 95 |
+
{"role": "user", "content": "Which corpus passages support the answer?"}
|
| 96 |
+
],
|
| 97 |
+
"tools": [
|
| 98 |
+
{
|
| 99 |
+
"type": "function",
|
| 100 |
+
"function": {
|
| 101 |
+
"name": "search_index",
|
| 102 |
+
"description": "Search the local semantic index and return relevant text snippets.",
|
| 103 |
+
"parameters": {
|
| 104 |
+
"type": "object",
|
| 105 |
+
"properties": {
|
| 106 |
+
"query": {"type": "string", "description": "The search query string."},
|
| 107 |
+
"k": {"type": "integer", "description": "Optional number of top results to return."}
|
| 108 |
+
},
|
| 109 |
+
"required": ["query"]
|
| 110 |
+
}
|
| 111 |
+
}
|
| 112 |
+
}
|
| 113 |
+
]
|
| 114 |
+
}'
|
| 115 |
+
```
|
chat_template.jinja
ADDED
|
@@ -0,0 +1,348 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{#--------TOOL RENDERING FUNCTIONS---------#}
|
| 2 |
+
|
| 3 |
+
{#---------------------------------------------------------------
|
| 4 |
+
Converts JSON Schema (dict) to a TypeScript type definition
|
| 5 |
+
----------------------------------------------------------------#}
|
| 6 |
+
{%- macro json_schema_to_typescript(schema, indent="") -%}
|
| 7 |
+
{%- set ADDITIONAL_JSON_KEYS = ['format', 'maxItems', 'maximum', 'minItems', 'minimum', 'pattern'] -%}
|
| 8 |
+
{%- set ty = schema.get("type") -%}
|
| 9 |
+
|
| 10 |
+
{# ---------------- OBJECT ---------------- #}
|
| 11 |
+
{%- if ty == "object" -%}
|
| 12 |
+
{{- "{\n" -}}
|
| 13 |
+
|
| 14 |
+
{# Start building property list #}
|
| 15 |
+
{%- set props = schema.get("properties", {}) -%}
|
| 16 |
+
{%- set required = schema.get("required", []) -%}
|
| 17 |
+
{%- set has_additional_props = schema.get("additionalProperties") is defined -%}
|
| 18 |
+
{%- set additional_props_type = none -%}
|
| 19 |
+
{%- if has_additional_props -%}
|
| 20 |
+
{%- if schema.additionalProperties == true -%}
|
| 21 |
+
{%- set additional_props_type = {'type': 'any'} -%}
|
| 22 |
+
{%- elif schema.additionalProperties is mapping -%}
|
| 23 |
+
{%- set additional_props_type = schema.additionalProperties -%}
|
| 24 |
+
{%- endif -%}
|
| 25 |
+
{%- endif -%}
|
| 26 |
+
|
| 27 |
+
{%- for key, val in props.items() -%}
|
| 28 |
+
{# ---------- Description Comments ---------- #}
|
| 29 |
+
{%- if "description" in val -%}
|
| 30 |
+
{%- for line in val['description'].split('\n') -%}
|
| 31 |
+
{%- if line.strip() -%}
|
| 32 |
+
{{- indent + '// ' + line + '\n' -}}
|
| 33 |
+
{%- endif -%}
|
| 34 |
+
{%- endfor -%}
|
| 35 |
+
{%- endif -%}
|
| 36 |
+
|
| 37 |
+
{# ---------- Additional JSON Keys ---------- #}
|
| 38 |
+
{%- for add_key, add_val in val.items() -%}
|
| 39 |
+
{%- if add_key in ADDITIONAL_JSON_KEYS -%}
|
| 40 |
+
{%- if add_val is string -%}
|
| 41 |
+
{{- indent + '// ' + add_key + ': "' + add_val + '"' + '\n' -}}
|
| 42 |
+
{%- else -%}
|
| 43 |
+
{{- indent + '// ' + add_key + ': ' ~ add_val ~ '\n' -}}
|
| 44 |
+
{%- endif -%}
|
| 45 |
+
{%- endif -%}
|
| 46 |
+
{%- endfor -%}
|
| 47 |
+
|
| 48 |
+
{# ---------- Property Definition ---------- #}
|
| 49 |
+
{%- set type_str = json_schema_to_typescript(
|
| 50 |
+
val,
|
| 51 |
+
indent + " "
|
| 52 |
+
) -%}
|
| 53 |
+
|
| 54 |
+
{{- indent + key + ('' if key in required else '?') + ': ' + type_str + ',' -}}
|
| 55 |
+
|
| 56 |
+
{%- if "default" in val or "defalut_value" in val -%}
|
| 57 |
+
{%- set default = val.get("default", val.get("defalut_value")) -%}
|
| 58 |
+
{%- if default is string -%}
|
| 59 |
+
{{- ' // default: "' + default + '"' -}}
|
| 60 |
+
{%- else -%}
|
| 61 |
+
{{- ' // default: ' ~ default -}}
|
| 62 |
+
{%- endif -%}
|
| 63 |
+
{%- endif -%}
|
| 64 |
+
|
| 65 |
+
{{- "\n" -}}
|
| 66 |
+
{%- endfor -%}
|
| 67 |
+
|
| 68 |
+
{# Handle additionalProperties as index signature #}
|
| 69 |
+
{%- if has_additional_props and additional_props_type is not none -%}
|
| 70 |
+
{%- set additional_type_str = json_schema_to_typescript(
|
| 71 |
+
additional_props_type,
|
| 72 |
+
indent + " "
|
| 73 |
+
) -%}
|
| 74 |
+
{{- indent + '[key: string]: ' + additional_type_str + '\n' -}}
|
| 75 |
+
{%- endif -%}
|
| 76 |
+
|
| 77 |
+
{{- indent[: (indent|length - " "|length) ] + '}' -}}
|
| 78 |
+
|
| 79 |
+
{# ---------------- STRING ---------------- #}
|
| 80 |
+
{%- elif ty == "string" -%}
|
| 81 |
+
{%- if schema.get("enum") -%}
|
| 82 |
+
{%- set ns = namespace(enum = []) -%}
|
| 83 |
+
{%- for en in schema['enum'] -%}
|
| 84 |
+
{%- set ns.enum = ns.enum + ['"' ~ en ~ '"'] -%}
|
| 85 |
+
{%- endfor -%}
|
| 86 |
+
{{- ns.enum | join(' | ') -}}
|
| 87 |
+
{%- elif schema.get("format", "none") in ['date-time', 'date'] -%}
|
| 88 |
+
{{- 'Date' -}}
|
| 89 |
+
{%- else -%}
|
| 90 |
+
{{- 'string' -}}
|
| 91 |
+
{%- endif -%}
|
| 92 |
+
|
| 93 |
+
{# ---------------- NUMBER / INTEGER ---------------- #}
|
| 94 |
+
{%- elif ty in ["number", "integer"] -%}
|
| 95 |
+
{%- if schema.get("enum") -%}
|
| 96 |
+
{{- schema.enum | join(' | ') -}}
|
| 97 |
+
{%- else -%}
|
| 98 |
+
{{- 'number' -}}
|
| 99 |
+
{%- endif -%}
|
| 100 |
+
|
| 101 |
+
{# ---------------- BOOLEAN ---------------- #}
|
| 102 |
+
{%- elif ty == "boolean" -%}
|
| 103 |
+
{{- 'boolean' -}}
|
| 104 |
+
|
| 105 |
+
{# ---------------- ARRAY ---------------- #}
|
| 106 |
+
{%- elif ty == "array" -%}
|
| 107 |
+
{%- if "items" in schema -%}
|
| 108 |
+
{{- json_schema_to_typescript(schema['items'], indent) + '[]' -}}
|
| 109 |
+
{%- else -%}
|
| 110 |
+
{{- 'Array<any>' -}}
|
| 111 |
+
{%- endif -%}
|
| 112 |
+
|
| 113 |
+
{# ---------------- FALLBACK ---------------- #}
|
| 114 |
+
{%- else -%}
|
| 115 |
+
{{- 'any' -}}
|
| 116 |
+
{%- endif -%}
|
| 117 |
+
{%- endmacro -%}
|
| 118 |
+
|
| 119 |
+
{#---------------------------------------------------------------
|
| 120 |
+
Renders a namespace and its tool definitions in TypeScript style
|
| 121 |
+
----------------------------------------------------------------#}
|
| 122 |
+
|
| 123 |
+
{%- macro render_tool_namespace(namespace_name, tools) -%}
|
| 124 |
+
{%- set ns = namespace(sections = ['namespace ' ~ namespace_name ~ ' {']) -%}
|
| 125 |
+
|
| 126 |
+
{%- for tool in tools -%}
|
| 127 |
+
{%- if tool.function -%}
|
| 128 |
+
{%- set tool = tool.function -%}
|
| 129 |
+
{%- endif -%}
|
| 130 |
+
|
| 131 |
+
{%- set ns_tool = namespace(content_lines=[]) -%}
|
| 132 |
+
|
| 133 |
+
{# ---------- TOOL DESCRIPTION ---------- #}
|
| 134 |
+
{%- if tool.get('description') -%}
|
| 135 |
+
{%- for line in tool['description'].split('\n') -%}
|
| 136 |
+
{%- if line.strip() -%}
|
| 137 |
+
{%- set ns_tool.content_lines = ns_tool.content_lines + ['// ' ~ line] -%}
|
| 138 |
+
{%- endif -%}
|
| 139 |
+
{%- endfor -%}
|
| 140 |
+
{%- endif -%}
|
| 141 |
+
|
| 142 |
+
{# ---------- TOOL SIGNATURE ---------- #}
|
| 143 |
+
{%- set main_body = "" -%}
|
| 144 |
+
{%- set params = tool.get("parameters") -%}
|
| 145 |
+
{%- if params and params.get("properties") -%}
|
| 146 |
+
{%- set param_type = json_schema_to_typescript(params, " ") -%}
|
| 147 |
+
{%- set main_body = 'type ' ~ tool.name ~ ' = (_: ' ~ param_type ~ ') => ' -%}
|
| 148 |
+
{%- else -%}
|
| 149 |
+
{%- set main_body = 'type ' ~ tool.name ~ ' = () => ' -%}
|
| 150 |
+
{%- endif -%}
|
| 151 |
+
|
| 152 |
+
{# ---------- RETURN TYPE ---------- #}
|
| 153 |
+
{%- set return_params = tool.get("return_parameters") -%}
|
| 154 |
+
{%- if return_params and return_params.get("properties") -%}
|
| 155 |
+
{%- set return_type = json_schema_to_typescript(return_params, " ") -%}
|
| 156 |
+
{%- set main_body = main_body ~ return_type -%}
|
| 157 |
+
{%- else -%}
|
| 158 |
+
{%- set main_body = main_body ~ 'any' -%}
|
| 159 |
+
{%- endif -%}
|
| 160 |
+
|
| 161 |
+
{%- set main_body = main_body ~ ';\n' -%}
|
| 162 |
+
|
| 163 |
+
{%- set ns_tool.content_lines = ns_tool.content_lines + [main_body] -%}
|
| 164 |
+
|
| 165 |
+
{# ---------- ADD TOOL TO SECTIONS ---------- #}
|
| 166 |
+
{%- set ns.sections = ns.sections + [ns_tool.content_lines | join('\n')] -%}
|
| 167 |
+
{%- endfor -%}
|
| 168 |
+
|
| 169 |
+
{%- set ns.sections = ns.sections + ['} // namespace ' ~ namespace_name] -%}
|
| 170 |
+
|
| 171 |
+
{{- ns.sections | join('\n') -}}
|
| 172 |
+
{%- endmacro -%}
|
| 173 |
+
|
| 174 |
+
|
| 175 |
+
{# ----------- MESSAGE RENDERING HELPER FUNCTIONS ------------ #}
|
| 176 |
+
|
| 177 |
+
{%- macro render_function_call(call) -%}
|
| 178 |
+
{%- if call.function -%}
|
| 179 |
+
{%- set call = call.function -%}
|
| 180 |
+
{%- endif -%}
|
| 181 |
+
|
| 182 |
+
{%- set arguments = call['arguments'] -%}
|
| 183 |
+
{%- if arguments is not string -%}
|
| 184 |
+
{%- set arguments = arguments| tojson(ensure_ascii=False) -%}
|
| 185 |
+
{%- endif -%}
|
| 186 |
+
|
| 187 |
+
{{- '{"name": "' ~ call['name'] ~ '", "arguments": ' ~ arguments ~ '}' -}}
|
| 188 |
+
{%- endmacro -%}
|
| 189 |
+
|
| 190 |
+
|
| 191 |
+
{%- macro render_role_message(message, role=None) -%}
|
| 192 |
+
{%- if not role -%}
|
| 193 |
+
{%- set role = message["role"] -%}
|
| 194 |
+
{%- endif -%}
|
| 195 |
+
|
| 196 |
+
{%- set message_content = message['content'] or '' -%}
|
| 197 |
+
{%- if message_content is not string -%}
|
| 198 |
+
{%- set message_content = message_content | tojson(ensure_ascii=False) -%}
|
| 199 |
+
{%- endif -%}
|
| 200 |
+
|
| 201 |
+
{{- role + add_tokens.role_sep + message_content -}}
|
| 202 |
+
|
| 203 |
+
{%- if message.tool_calls is defined and message.tool_calls -%}
|
| 204 |
+
{{- add_tokens.function_call + render_function_call(message.tool_calls[0]) -}}
|
| 205 |
+
{%- endif -%}
|
| 206 |
+
|
| 207 |
+
{{- add_tokens.message_sep -}}
|
| 208 |
+
|
| 209 |
+
{%- endmacro -%}
|
| 210 |
+
|
| 211 |
+
|
| 212 |
+
|
| 213 |
+
{# ----- SPECIAL TOKENS ----- #}
|
| 214 |
+
|
| 215 |
+
{%- set add_tokens = namespace(
|
| 216 |
+
role_sep="<|role_sep|>\n",
|
| 217 |
+
message_sep="<|message_sep|>\n\n",
|
| 218 |
+
function_call="<|function_call|>"
|
| 219 |
+
) -%}
|
| 220 |
+
|
| 221 |
+
{# ----- DEFAULT DEVSYSTEM ----- #}
|
| 222 |
+
|
| 223 |
+
{%- set DEVSYSTEM -%}
|
| 224 |
+
<role_description>
|
| 225 |
+
Описание доступных в диалоге ролей.
|
| 226 |
+
|
| 227 |
+
`developer system`
|
| 228 |
+
Сообщение, добавленное Сбером до основного диалога. Имеет самый высокий приоритет и определяет глобальные, неотменяемые условия (например, правила ведения диалога, политику безопасности, общий стиль ответов ассистента и пр.).
|
| 229 |
+
|
| 230 |
+
`system`
|
| 231 |
+
Системная инструкция, добавляемая разработчиками или пользователем, но с приоритетом ниже, чем `developer system`. Обычно описывает инструкции ассистента, конкретный стиль ответа и другие условия для данного конкретного диалога.
|
| 232 |
+
|
| 233 |
+
`user`
|
| 234 |
+
Сообщение или запрос от пользователя. Ассистент следует ему, если это не противоречит инструкциям более высокого приоритета (см. <instruction_priority>).
|
| 235 |
+
|
| 236 |
+
`user memory`
|
| 237 |
+
Последовательность наиболее актуальных долговременных фактов о пользователе на момент его запроса, представленная в виде JSON‑списка строк. Факты в ней перечислены в хронологическом порядке, то есть более новые факты дописываются в конец последовательности. При этом при изменении или удалении фактов записи о предыдущих фактах остаются в последовательности. Ассистент сохраняет факты с помощью функции и использует их в соответствии с указаниями из блока <memory_guidelines> ниже.
|
| 238 |
+
|
| 239 |
+
`added files`
|
| 240 |
+
Метаинформация о файлах, доступных для использования в диалоге, представленная в формате JSON. Содержит следующие ключи: id (уникальный идентификатор файла), name (имя файла), type (тип файла).
|
| 241 |
+
|
| 242 |
+
`assistant`
|
| 243 |
+
Ответ ассистента на запрос пользователя. Если системная инструкция или пользователь не задаёт дополнительных правил для `assistant`, то такая реплика должна соответс��вовать указаниям из блока <assistant_guidelines> ниже. Список доступных для вызова функций содержится в последней реплике роли `available functions`. Название необходимой для вызова функции и аргументы будут сгенерированы после специального токена вызова функции. В своих репликах ассистент следует инструкциям в соответствии с <instruction_priority>.
|
| 244 |
+
Вызов функции осуществляется в строгом соответствии с инструкцией из блока <function_usage>.
|
| 245 |
+
|
| 246 |
+
`function descriptions`
|
| 247 |
+
Описания функций в формате TypeScript. Функция — это специальный инструмент (или набор инструкций), который ассистент может вызвать для выполнения конкретных действий, вычислений или получения данных, необходимых для решения задачи пользователя. Каждое описание функции содержит блоки с именем, описанием, аргументами. Иногда описание содержит отдельные блоки с возвращаемыми параметрами и примерами применения, иллюстрирующими правильный вызов и аргументы.
|
| 248 |
+
|
| 249 |
+
`available functions`
|
| 250 |
+
Список, который содержит названия функций, доступных для вызова. Если список не содержит элементов, то в следующем сообщении функции, доступные для вызова, отсутствуют.
|
| 251 |
+
|
| 252 |
+
`function result`
|
| 253 |
+
Результат последнего вызова функции.
|
| 254 |
+
</role_description>
|
| 255 |
+
|
| 256 |
+
|
| 257 |
+
<available_modalities>
|
| 258 |
+
Ассистент умеет работать со следующими модальностями: текст, доступные функции.
|
| 259 |
+
</available_modalities>
|
| 260 |
+
|
| 261 |
+
|
| 262 |
+
<instruction_priority>
|
| 263 |
+
В случае противоречия инструкций разных ролей в контексте диалога соблюдай приоритеты:
|
| 264 |
+
`developer system` > `system` > `user` > `function descriptions` > `function result` > `user memory`
|
| 265 |
+
</instruction_priority>
|
| 266 |
+
|
| 267 |
+
|
| 268 |
+
<function_usage>
|
| 269 |
+
Базовые инструкции для работы с функциями.
|
| 270 |
+
|
| 271 |
+
Можно вызывать только те функции, которые доступны исходя из последнего сообщения `available functions`.
|
| 272 |
+
|
| 273 |
+
Вызывай доступные функции в случае, если согласно их описанию такой вызов поможет дать более полный и/или точный ответ на запрос пользователя. Заполняй аргументы функций, используя информацию из контекста диалога. Если функция может помочь ответить на запрос, но для её обязательного аргумента отсутствует информация в контексте, уточни у пользователя недостающие данные перед вызовом функции. При недоступности необходимой функции или ошибке — кратко сообщи об этом пользователю и по возможности предложи альтернативу.
|
| 274 |
+
</function_usage>
|
| 275 |
+
|
| 276 |
+
|
| 277 |
+
<memory_guidelines>
|
| 278 |
+
Правила использования фактов в долговременной памяти:
|
| 279 |
+
|
| 280 |
+
Если в диалоге нет сообщения под ролью `user memory`, то это равносильно отсутствию долговременных фактов о пользователе в памяти. В таком случае информация о пользователе ограничена текущим диалогом, и новые факты не должны сохраняться.
|
| 281 |
+
</memory_guidelines>
|
| 282 |
+
|
| 283 |
+
|
| 284 |
+
<assistant_guidelines>
|
| 285 |
+
GigaChat — нейросетевая модель искусственного интеллекта, созданная компанией Сбер в России.
|
| 286 |
+
|
| 287 |
+
GigaChat старается отвечать на языке, на котором пользователь задал запрос. Если из запроса пользователя и контекста диалога язык определить невозможно, GigaChat использует русский.
|
| 288 |
+
GigaChat предоставляет подробные ответы на более сложные и открытые вопросы.
|
| 289 |
+
GigaChat в ответе не и��пользует названия доступных функций.
|
| 290 |
+
GigaChat отвечает безопасно, в соответствии с действующим законодательством Российской Федерации, стараясь помочь пользователю решить задачу или поддержать беседу.
|
| 291 |
+
|
| 292 |
+
Ты — GigaChat.
|
| 293 |
+
</assistant_guidelines>
|
| 294 |
+
|
| 295 |
+
|
| 296 |
+
Ниже будет приведён диалог.
|
| 297 |
+
В диалоге могут быть разнообразные роли, описанные в блоке <role_description>.
|
| 298 |
+
Каждая реплика начинается с названия роли и специального токена, обозначающего конец полного наименования роли, а заканчивается специальным токеном конца реплики.
|
| 299 |
+
Твоя задача — продолжить диалог от последней указанной роли в соответствии с контекстом диалога.
|
| 300 |
+
{%- endset -%}
|
| 301 |
+
|
| 302 |
+
|
| 303 |
+
{#- ---------------------- RENDERING STARTS HERE ---------------------- -#}
|
| 304 |
+
|
| 305 |
+
|
| 306 |
+
{# ----- RENDER BOS TOKEN ----- #}
|
| 307 |
+
{{- bos_token -}}
|
| 308 |
+
|
| 309 |
+
|
| 310 |
+
{# ----- RENDER DEVSYSTEM ----- #}
|
| 311 |
+
{{- render_role_message({"role": "developer system", "content": DEVSYSTEM}) -}}
|
| 312 |
+
|
| 313 |
+
{# ----- RENDER SYSTEM IF PRESENT ----- #}
|
| 314 |
+
{%- if messages and messages[0]['role'] == 'system' -%}
|
| 315 |
+
{{- render_role_message(messages[0]) -}}
|
| 316 |
+
{%- set messages = messages[1:] -%}
|
| 317 |
+
{%- else -%}
|
| 318 |
+
{{- render_role_message({"role": "system", "content": ""}) -}}
|
| 319 |
+
{%- endif -%}
|
| 320 |
+
|
| 321 |
+
{# ----- RENDER TOOLS ----- #}
|
| 322 |
+
{%- if tools -%}
|
| 323 |
+
{%- set tools_content = (
|
| 324 |
+
render_tool_namespace('functions', tools)
|
| 325 |
+
+ "\n\n"
|
| 326 |
+
) -%}
|
| 327 |
+
{{- render_role_message({'role': 'function descriptions', 'content': tools_content}) -}}
|
| 328 |
+
{%- endif -%}
|
| 329 |
+
|
| 330 |
+
{# ----- MAIN MESSAGE LOOP ----- #}
|
| 331 |
+
{%- for message in messages -%}
|
| 332 |
+
|
| 333 |
+
{# ----- TOOL MESSAGE -------#}
|
| 334 |
+
{%- if message['role'] == 'tool' -%}
|
| 335 |
+
{{- render_role_message(message, 'function result') -}}
|
| 336 |
+
|
| 337 |
+
{# ----- OTHER MESSAGES ----- #}
|
| 338 |
+
{%- else -%}
|
| 339 |
+
{{- render_role_message(message) -}}
|
| 340 |
+
{%- endif -%}
|
| 341 |
+
|
| 342 |
+
{# ----- ADDING GENERATION PROMPT ----- #}
|
| 343 |
+
|
| 344 |
+
{%- if loop.last and add_generation_prompt and message['role'] != 'assistant' -%}
|
| 345 |
+
{{- 'assistant' + add_tokens.role_sep -}}
|
| 346 |
+
{%- endif -%}
|
| 347 |
+
|
| 348 |
+
{%- endfor -%}
|
config.json
ADDED
|
@@ -0,0 +1,53 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"vocab_size": 128256,
|
| 3 |
+
"max_position_embeddings": 262144,
|
| 4 |
+
"hidden_size": 7168,
|
| 5 |
+
"intermediate_size": 18432,
|
| 6 |
+
"moe_intermediate_size": 2048,
|
| 7 |
+
"num_hidden_layers": 64,
|
| 8 |
+
"num_nextn_predict_layers": 1,
|
| 9 |
+
"num_attention_heads": 64,
|
| 10 |
+
"n_shared_experts": 1,
|
| 11 |
+
"n_routed_experts": 256,
|
| 12 |
+
"ep_size": 1,
|
| 13 |
+
"routed_scaling_factor": 2.5,
|
| 14 |
+
"kv_lora_rank": 512,
|
| 15 |
+
"q_lora_rank": 1536,
|
| 16 |
+
"qk_rope_head_dim": 64,
|
| 17 |
+
"v_head_dim": 192,
|
| 18 |
+
"qk_nope_head_dim": 128,
|
| 19 |
+
"topk_method": "noaux_tc",
|
| 20 |
+
"n_group": 8,
|
| 21 |
+
"topk_group": 4,
|
| 22 |
+
"num_experts_per_tok": 8,
|
| 23 |
+
"moe_layer_freq": 1,
|
| 24 |
+
"first_k_dense_replace": 3,
|
| 25 |
+
"norm_topk_prob": true,
|
| 26 |
+
"scoring_func": "sigmoid",
|
| 27 |
+
"num_key_value_heads": 64,
|
| 28 |
+
"hidden_act": "silu",
|
| 29 |
+
"initializer_range": 0.006,
|
| 30 |
+
"rms_norm_eps": 1e-06,
|
| 31 |
+
"use_cache": true,
|
| 32 |
+
"rope_theta": 100000,
|
| 33 |
+
"rope_scaling": {
|
| 34 |
+
"beta_fast": 32.0,
|
| 35 |
+
"beta_slow": 1.0,
|
| 36 |
+
"factor": 64.0,
|
| 37 |
+
"mscale": 1.0,
|
| 38 |
+
"mscale_all_dim": 1.0,
|
| 39 |
+
"original_max_position_embeddings": 4096,
|
| 40 |
+
"rope_type": "yarn"
|
| 41 |
+
},
|
| 42 |
+
"attention_bias": false,
|
| 43 |
+
"attention_dropout": 0.0,
|
| 44 |
+
"tie_word_embeddings": false,
|
| 45 |
+
"architectures": [
|
| 46 |
+
"DeepseekV3ForCausalLM"
|
| 47 |
+
],
|
| 48 |
+
"bos_token_id": 1,
|
| 49 |
+
"eos_token_id": 2,
|
| 50 |
+
"transformers_version": "4.57.3",
|
| 51 |
+
"model_type": "deepseek_v3",
|
| 52 |
+
"dtype": "bfloat16"
|
| 53 |
+
}
|
generation_config.json
ADDED
|
@@ -0,0 +1,7 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"bos_token_id": 1,
|
| 3 |
+
"eos_token_id": 2,
|
| 4 |
+
"pad_token_id": 2,
|
| 5 |
+
"transformers_version": "4.57.3",
|
| 6 |
+
"_from_model_config": true
|
| 7 |
+
}
|
gigachat3_guided_decoding.py
ADDED
|
@@ -0,0 +1,390 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Standalone vLLM gigachat3 parser + xgrammar tags + optional retrieved-id GBNF.
|
| 2 |
+
|
| 3 |
+
Loaded via ``--tool-parser-plugin ./gigachat3_guided_decoding.py``.
|
| 4 |
+
Does not import ``trace_generation`` or ``harness``.
|
| 5 |
+
|
| 6 |
+
Extra-body keys must match ``smith_harness.guided``:
|
| 7 |
+
``smith_guided_string_enums``, ``smith_guided_array_min_items``,
|
| 8 |
+
``smith_guided_array_max_items``. Absent keys → general Giga grammar only.
|
| 9 |
+
"""
|
| 10 |
+
|
| 11 |
+
from __future__ import annotations
|
| 12 |
+
|
| 13 |
+
import json
|
| 14 |
+
from typing import Any, Mapping, Sequence
|
| 15 |
+
|
| 16 |
+
from xgrammar import StructuralTag
|
| 17 |
+
from xgrammar.openai_tool_call_schema import BuiltinToolParam, FunctionToolParam
|
| 18 |
+
from xgrammar.structural_tag import (
|
| 19 |
+
AnyTextFormat,
|
| 20 |
+
JSONSchemaFormat,
|
| 21 |
+
TagFormat,
|
| 22 |
+
TagsWithSeparatorFormat,
|
| 23 |
+
TriggeredTagsFormat,
|
| 24 |
+
)
|
| 25 |
+
|
| 26 |
+
from vllm.tool_parsers.abstract_tool_parser import ToolParserManager
|
| 27 |
+
from vllm.tool_parsers.gigachat3_tool_parser import GigaChat3ToolParser
|
| 28 |
+
from vllm.tool_parsers.structural_tag_registry import register_vllm_structural_tag
|
| 29 |
+
|
| 30 |
+
GIGACHAT3_TOOL_CALL_PARSER = "gigachat3"
|
| 31 |
+
SMITH_GUIDED_STRING_ENUMS_KEY = "smith_guided_string_enums"
|
| 32 |
+
SMITH_GUIDED_ARRAY_MIN_ITEMS_KEY = "smith_guided_array_min_items"
|
| 33 |
+
SMITH_GUIDED_ARRAY_MAX_ITEMS_KEY = "smith_guided_array_max_items"
|
| 34 |
+
DEFAULT_GUIDED_ARRAY_MIN_ITEMS = 2
|
| 35 |
+
DEFAULT_GUIDED_ARRAY_MAX_ITEMS = 10
|
| 36 |
+
UNCAPPED_ARRAY_MAX_ITEMS = 32
|
| 37 |
+
DEFAULT_MAX_QUERY_CHARS = 256
|
| 38 |
+
QUERY_STRING_FIELD_NAMES = frozenset({"query"})
|
| 39 |
+
_ARGUMENTS_FIELD_PREFIX = '", "arguments": '
|
| 40 |
+
_FUNCTION_CALL_WRAPS = (
|
| 41 |
+
('<|function_call|>{"name": "', "}"),
|
| 42 |
+
('function call<|role_sep|>\n{"name": "', "}"),
|
| 43 |
+
)
|
| 44 |
+
_FUNCTION_CALL_TRIGGERS = ("<|function_call|>", "function call<|role_sep|>")
|
| 45 |
+
_NAME_BEGIN = '{"name": "'
|
| 46 |
+
|
| 47 |
+
|
| 48 |
+
def _schema_is_array(schema: dict[str, Any]) -> bool:
|
| 49 |
+
schema_type = schema.get("type")
|
| 50 |
+
if schema_type == "array":
|
| 51 |
+
return True
|
| 52 |
+
return isinstance(schema_type, list) and "array" in schema_type
|
| 53 |
+
|
| 54 |
+
|
| 55 |
+
def _schema_is_string(schema: dict[str, Any]) -> bool:
|
| 56 |
+
schema_type = schema.get("type")
|
| 57 |
+
if schema_type == "string":
|
| 58 |
+
return True
|
| 59 |
+
return isinstance(schema_type, list) and "string" in schema_type
|
| 60 |
+
|
| 61 |
+
|
| 62 |
+
def cap_uncapped_query_string_max_length(
|
| 63 |
+
schema: Any,
|
| 64 |
+
*,
|
| 65 |
+
max_length: int = DEFAULT_MAX_QUERY_CHARS,
|
| 66 |
+
) -> Any:
|
| 67 |
+
if isinstance(schema, dict):
|
| 68 |
+
out = {
|
| 69 |
+
key: cap_uncapped_query_string_max_length(value, max_length=max_length)
|
| 70 |
+
for key, value in schema.items()
|
| 71 |
+
}
|
| 72 |
+
props = out.get("properties")
|
| 73 |
+
if isinstance(props, dict):
|
| 74 |
+
new_props = dict(props)
|
| 75 |
+
changed = False
|
| 76 |
+
for name, prop in props.items():
|
| 77 |
+
if name not in QUERY_STRING_FIELD_NAMES or not isinstance(prop, dict):
|
| 78 |
+
continue
|
| 79 |
+
if _schema_is_string(prop) and "maxLength" not in prop:
|
| 80 |
+
new_props[name] = {**prop, "maxLength": int(max_length)}
|
| 81 |
+
changed = True
|
| 82 |
+
if changed:
|
| 83 |
+
out["properties"] = new_props
|
| 84 |
+
return out
|
| 85 |
+
if isinstance(schema, list):
|
| 86 |
+
return [
|
| 87 |
+
cap_uncapped_query_string_max_length(item, max_length=max_length)
|
| 88 |
+
for item in schema
|
| 89 |
+
]
|
| 90 |
+
return schema
|
| 91 |
+
|
| 92 |
+
|
| 93 |
+
def cap_uncapped_array_max_items(
|
| 94 |
+
schema: Any,
|
| 95 |
+
*,
|
| 96 |
+
max_items: int = UNCAPPED_ARRAY_MAX_ITEMS,
|
| 97 |
+
) -> Any:
|
| 98 |
+
if isinstance(schema, dict):
|
| 99 |
+
out = {
|
| 100 |
+
key: cap_uncapped_array_max_items(value, max_items=max_items)
|
| 101 |
+
for key, value in schema.items()
|
| 102 |
+
}
|
| 103 |
+
if _schema_is_array(out) and "maxItems" not in out:
|
| 104 |
+
out["maxItems"] = max_items
|
| 105 |
+
return out
|
| 106 |
+
if isinstance(schema, list):
|
| 107 |
+
return [
|
| 108 |
+
cap_uncapped_array_max_items(item, max_items=max_items) for item in schema
|
| 109 |
+
]
|
| 110 |
+
return schema
|
| 111 |
+
|
| 112 |
+
|
| 113 |
+
def _normalize_guided_array_bounds(
|
| 114 |
+
min_items: Any = None,
|
| 115 |
+
max_items: Any = None,
|
| 116 |
+
) -> tuple[int, int]:
|
| 117 |
+
lo = DEFAULT_GUIDED_ARRAY_MIN_ITEMS if min_items is None else int(min_items)
|
| 118 |
+
hi = DEFAULT_GUIDED_ARRAY_MAX_ITEMS if max_items is None else int(max_items)
|
| 119 |
+
if lo < 1:
|
| 120 |
+
raise ValueError("guided_supporting_ids_min_items must be >= 1")
|
| 121 |
+
if hi < lo:
|
| 122 |
+
raise ValueError(
|
| 123 |
+
"guided_supporting_ids_max_items must be >= guided_supporting_ids_min_items"
|
| 124 |
+
)
|
| 125 |
+
return lo, hi
|
| 126 |
+
|
| 127 |
+
|
| 128 |
+
def _coerce_bound(value: Any) -> int | None:
|
| 129 |
+
if value is None or isinstance(value, bool):
|
| 130 |
+
return None
|
| 131 |
+
try:
|
| 132 |
+
return int(value)
|
| 133 |
+
except (TypeError, ValueError):
|
| 134 |
+
return None
|
| 135 |
+
|
| 136 |
+
|
| 137 |
+
def _bound_from_mapping(payload: Any, key: str) -> int | None:
|
| 138 |
+
if not isinstance(payload, Mapping):
|
| 139 |
+
return None
|
| 140 |
+
return _coerce_bound(payload.get(key))
|
| 141 |
+
|
| 142 |
+
|
| 143 |
+
def guided_array_bounds_from_request(request: Any) -> tuple[int, int]:
|
| 144 |
+
extra_body = getattr(request, "extra_body", None)
|
| 145 |
+
extras = getattr(request, "__pydantic_extra__", None)
|
| 146 |
+
min_items = (
|
| 147 |
+
_coerce_bound(getattr(request, SMITH_GUIDED_ARRAY_MIN_ITEMS_KEY, None))
|
| 148 |
+
or _bound_from_mapping(extra_body, SMITH_GUIDED_ARRAY_MIN_ITEMS_KEY)
|
| 149 |
+
or _bound_from_mapping(extras, SMITH_GUIDED_ARRAY_MIN_ITEMS_KEY)
|
| 150 |
+
)
|
| 151 |
+
max_items = (
|
| 152 |
+
_coerce_bound(getattr(request, SMITH_GUIDED_ARRAY_MAX_ITEMS_KEY, None))
|
| 153 |
+
or _bound_from_mapping(extra_body, SMITH_GUIDED_ARRAY_MAX_ITEMS_KEY)
|
| 154 |
+
or _bound_from_mapping(extras, SMITH_GUIDED_ARRAY_MAX_ITEMS_KEY)
|
| 155 |
+
)
|
| 156 |
+
return _normalize_guided_array_bounds(min_items, max_items)
|
| 157 |
+
|
| 158 |
+
|
| 159 |
+
def guided_string_enums_from_request(request: Any) -> dict[str, dict[str, list[str]]]:
|
| 160 |
+
candidates: list[Any] = [
|
| 161 |
+
getattr(request, SMITH_GUIDED_STRING_ENUMS_KEY, None),
|
| 162 |
+
]
|
| 163 |
+
extra_body = getattr(request, "extra_body", None)
|
| 164 |
+
if isinstance(extra_body, dict):
|
| 165 |
+
candidates.append(extra_body.get(SMITH_GUIDED_STRING_ENUMS_KEY))
|
| 166 |
+
extras = getattr(request, "__pydantic_extra__", None)
|
| 167 |
+
if isinstance(extras, dict):
|
| 168 |
+
candidates.append(extras.get(SMITH_GUIDED_STRING_ENUMS_KEY))
|
| 169 |
+
for guided in candidates:
|
| 170 |
+
if isinstance(guided, dict) and guided:
|
| 171 |
+
return guided
|
| 172 |
+
return {}
|
| 173 |
+
|
| 174 |
+
|
| 175 |
+
def _array_items_ebnf(min_items: int, max_items: int) -> str:
|
| 176 |
+
required = "corpus_id"
|
| 177 |
+
if min_items > 1:
|
| 178 |
+
required += (' ws "," ws corpus_id') * (min_items - 1)
|
| 179 |
+
extra = max_items - min_items
|
| 180 |
+
if extra:
|
| 181 |
+
return f'{required} (ws "," ws corpus_id){{0,{extra}}}'
|
| 182 |
+
return required
|
| 183 |
+
|
| 184 |
+
|
| 185 |
+
def build_explicit_guided_array_grammar(
|
| 186 |
+
schema: Any,
|
| 187 |
+
field_enums: Mapping[str, Sequence[str]],
|
| 188 |
+
*,
|
| 189 |
+
min_items: Any = None,
|
| 190 |
+
max_items: Any = None,
|
| 191 |
+
) -> str | None:
|
| 192 |
+
if not isinstance(schema, dict) or not field_enums:
|
| 193 |
+
return None
|
| 194 |
+
if schema.get("type") != "object" or len(field_enums) != 1:
|
| 195 |
+
return None
|
| 196 |
+
properties = schema.get("properties")
|
| 197 |
+
required = schema.get("required")
|
| 198 |
+
if not isinstance(properties, dict) or len(properties) != 1:
|
| 199 |
+
return None
|
| 200 |
+
field, values = next(iter(field_enums.items()))
|
| 201 |
+
spec = properties.get(field)
|
| 202 |
+
if (
|
| 203 |
+
not isinstance(spec, dict)
|
| 204 |
+
or spec.get("type") != "array"
|
| 205 |
+
or not isinstance(spec.get("items"), dict)
|
| 206 |
+
or spec["items"].get("type") != "string"
|
| 207 |
+
or not isinstance(required, list)
|
| 208 |
+
or field not in required
|
| 209 |
+
):
|
| 210 |
+
return None
|
| 211 |
+
cleaned = sorted({str(item).strip() for item in values if str(item).strip()})
|
| 212 |
+
if not cleaned:
|
| 213 |
+
return None
|
| 214 |
+
lo, hi = _normalize_guided_array_bounds(min_items, max_items)
|
| 215 |
+
field_literal = json.dumps(json.dumps(field, ensure_ascii=False), ensure_ascii=False)
|
| 216 |
+
value_literals = [
|
| 217 |
+
json.dumps(json.dumps(value, ensure_ascii=False), ensure_ascii=False)
|
| 218 |
+
for value in cleaned
|
| 219 |
+
]
|
| 220 |
+
items = _array_items_ebnf(lo, hi)
|
| 221 |
+
return "\n".join(
|
| 222 |
+
[
|
| 223 |
+
(
|
| 224 |
+
f'root ::= "{{" ws {field_literal} ws ":" ws "[" ws '
|
| 225 |
+
f'{items} ws "]" ws "}}"'
|
| 226 |
+
),
|
| 227 |
+
f"corpus_id ::= {' | '.join(value_literals)}",
|
| 228 |
+
"ws ::= [ \\n\\t]{0,2}",
|
| 229 |
+
]
|
| 230 |
+
)
|
| 231 |
+
|
| 232 |
+
|
| 233 |
+
def _tool_name_from_begin(begin: str) -> str:
|
| 234 |
+
if _NAME_BEGIN in begin:
|
| 235 |
+
rest = begin.split(_NAME_BEGIN, 1)[1]
|
| 236 |
+
return rest.split('"', 1)[0]
|
| 237 |
+
return ""
|
| 238 |
+
|
| 239 |
+
|
| 240 |
+
def inject_guided_string_enums_into_structural_tag(
|
| 241 |
+
tag: Any,
|
| 242 |
+
guided_by_tool: Mapping[str, Mapping[str, Sequence[str]]],
|
| 243 |
+
*,
|
| 244 |
+
min_items: Any = None,
|
| 245 |
+
max_items: Any = None,
|
| 246 |
+
) -> Any:
|
| 247 |
+
if not guided_by_tool:
|
| 248 |
+
return tag
|
| 249 |
+
if isinstance(tag, list):
|
| 250 |
+
return [
|
| 251 |
+
inject_guided_string_enums_into_structural_tag(
|
| 252 |
+
item, guided_by_tool, min_items=min_items, max_items=max_items
|
| 253 |
+
)
|
| 254 |
+
for item in tag
|
| 255 |
+
]
|
| 256 |
+
if not isinstance(tag, dict):
|
| 257 |
+
return tag
|
| 258 |
+
node_type = str(tag.get("type") or "")
|
| 259 |
+
begin = str(tag.get("begin") or "")
|
| 260 |
+
looks_like_tag = node_type == "tag" or (
|
| 261 |
+
_NAME_BEGIN in begin and isinstance(tag.get("content"), dict)
|
| 262 |
+
)
|
| 263 |
+
if looks_like_tag:
|
| 264 |
+
tool_name = _tool_name_from_begin(begin)
|
| 265 |
+
field_enums = guided_by_tool.get(tool_name)
|
| 266 |
+
content = tag.get("content")
|
| 267 |
+
if field_enums and isinstance(content, dict) and (
|
| 268 |
+
content.get("type") == "json_schema" or "json_schema" in content
|
| 269 |
+
):
|
| 270 |
+
schema = content.get("json_schema")
|
| 271 |
+
if isinstance(schema, str):
|
| 272 |
+
schema = json.loads(schema)
|
| 273 |
+
grammar = build_explicit_guided_array_grammar(
|
| 274 |
+
schema, field_enums, min_items=min_items, max_items=max_items
|
| 275 |
+
)
|
| 276 |
+
if grammar is None:
|
| 277 |
+
return tag
|
| 278 |
+
patched_content = dict(content)
|
| 279 |
+
patched_content.pop("json_schema", None)
|
| 280 |
+
patched_content["type"] = "grammar"
|
| 281 |
+
patched_content["grammar"] = grammar
|
| 282 |
+
out = dict(tag)
|
| 283 |
+
out["content"] = patched_content
|
| 284 |
+
return out
|
| 285 |
+
return tag
|
| 286 |
+
return {
|
| 287 |
+
key: inject_guided_string_enums_into_structural_tag(
|
| 288 |
+
value, guided_by_tool, min_items=min_items, max_items=max_items
|
| 289 |
+
)
|
| 290 |
+
for key, value in tag.items()
|
| 291 |
+
}
|
| 292 |
+
|
| 293 |
+
|
| 294 |
+
def patch_structural_tag_json(
|
| 295 |
+
structural_tag_json: str,
|
| 296 |
+
guided_by_tool: Mapping[str, Mapping[str, Sequence[str]]],
|
| 297 |
+
*,
|
| 298 |
+
min_items: Any = None,
|
| 299 |
+
max_items: Any = None,
|
| 300 |
+
) -> str:
|
| 301 |
+
if not structural_tag_json or not guided_by_tool:
|
| 302 |
+
return structural_tag_json
|
| 303 |
+
tag = json.loads(structural_tag_json)
|
| 304 |
+
patched = inject_guided_string_enums_into_structural_tag(
|
| 305 |
+
tag, guided_by_tool, min_items=min_items, max_items=max_items
|
| 306 |
+
)
|
| 307 |
+
return json.dumps(patched)
|
| 308 |
+
|
| 309 |
+
|
| 310 |
+
def _tool_argument_schema(function: Any) -> dict[str, Any] | bool:
|
| 311 |
+
if getattr(function, "strict", None) is False:
|
| 312 |
+
return True
|
| 313 |
+
parameters = function.parameters if function.parameters is not None else True
|
| 314 |
+
if isinstance(parameters, dict):
|
| 315 |
+
capped = cap_uncapped_array_max_items(parameters)
|
| 316 |
+
return cap_uncapped_query_string_max_length(capped)
|
| 317 |
+
return parameters
|
| 318 |
+
|
| 319 |
+
|
| 320 |
+
def _gigachat3_tool_tags(tools: list[FunctionToolParam]) -> list[TagFormat]:
|
| 321 |
+
return [
|
| 322 |
+
TagFormat(
|
| 323 |
+
begin=begin + tool.function.name + _ARGUMENTS_FIELD_PREFIX,
|
| 324 |
+
content=JSONSchemaFormat(
|
| 325 |
+
json_schema=_tool_argument_schema(tool.function),
|
| 326 |
+
max_whitespace_cnt=2,
|
| 327 |
+
),
|
| 328 |
+
end=end,
|
| 329 |
+
)
|
| 330 |
+
for tool in tools
|
| 331 |
+
for begin, end in _FUNCTION_CALL_WRAPS
|
| 332 |
+
]
|
| 333 |
+
|
| 334 |
+
|
| 335 |
+
@register_vllm_structural_tag("gigachat3")
|
| 336 |
+
def get_gigachat3_structural_tag(
|
| 337 |
+
tools: list[FunctionToolParam],
|
| 338 |
+
builtin_tools: list[BuiltinToolParam],
|
| 339 |
+
tool_choice: str,
|
| 340 |
+
reasoning: bool,
|
| 341 |
+
) -> StructuralTag:
|
| 342 |
+
del builtin_tools, reasoning
|
| 343 |
+
tags = _gigachat3_tool_tags(tools)
|
| 344 |
+
if tool_choice == "auto":
|
| 345 |
+
suffix_tag = (
|
| 346 |
+
TriggeredTagsFormat(triggers=list(_FUNCTION_CALL_TRIGGERS), tags=tags)
|
| 347 |
+
if tags
|
| 348 |
+
else AnyTextFormat()
|
| 349 |
+
)
|
| 350 |
+
else:
|
| 351 |
+
suffix_tag = TagsWithSeparatorFormat(
|
| 352 |
+
tags=tags,
|
| 353 |
+
separator="",
|
| 354 |
+
at_least_one=True,
|
| 355 |
+
stop_after_first=True,
|
| 356 |
+
)
|
| 357 |
+
return StructuralTag(format=suffix_tag)
|
| 358 |
+
|
| 359 |
+
|
| 360 |
+
class GigaChat3StructuralToolParser(GigaChat3ToolParser):
|
| 361 |
+
"""Upstream gigachat3 parser with a GigaChat xgrammar structural tag."""
|
| 362 |
+
|
| 363 |
+
structural_tag_model = "gigachat3"
|
| 364 |
+
supports_required_and_named = False
|
| 365 |
+
|
| 366 |
+
def adjust_request(self, request):
|
| 367 |
+
if request.tools and request.tool_choice != "none":
|
| 368 |
+
request.skip_special_tokens = False
|
| 369 |
+
guided = guided_string_enums_from_request(request)
|
| 370 |
+
structured = getattr(request, "structured_outputs", None)
|
| 371 |
+
structural_tag = (
|
| 372 |
+
getattr(structured, "structural_tag", None) if structured is not None else None
|
| 373 |
+
)
|
| 374 |
+
if guided and isinstance(structural_tag, str) and structural_tag:
|
| 375 |
+
min_items, max_items = guided_array_bounds_from_request(request)
|
| 376 |
+
structured.structural_tag = patch_structural_tag_json(
|
| 377 |
+
structural_tag,
|
| 378 |
+
guided,
|
| 379 |
+
min_items=min_items,
|
| 380 |
+
max_items=max_items,
|
| 381 |
+
)
|
| 382 |
+
return request
|
| 383 |
+
|
| 384 |
+
|
| 385 |
+
ToolParserManager.register_module(
|
| 386 |
+
name=GIGACHAT3_TOOL_CALL_PARSER,
|
| 387 |
+
force=True,
|
| 388 |
+
module=GigaChat3StructuralToolParser,
|
| 389 |
+
)
|
| 390 |
+
ToolParserManager.lazy_parsers.pop(GIGACHAT3_TOOL_CALL_PARSER, None)
|
requirements.txt
ADDED
|
@@ -0,0 +1,17 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Serving pins for SMITH-10B-Exp (vLLM + transformers generate).
|
| 2 |
+
# SPDX comments apply to the named package only; pip also installs
|
| 3 |
+
# transitive dependencies under their own licenses.
|
| 4 |
+
#
|
| 5 |
+
# xgrammar 0.2.4+ requires transformers<5; keep 0.2.3 with Transformers 5.6.2.
|
| 6 |
+
|
| 7 |
+
# Apache-2.0 — huggingface/transformers
|
| 8 |
+
# https://pypi.org/project/transformers/
|
| 9 |
+
transformers==5.6.2
|
| 10 |
+
|
| 11 |
+
# Apache-2.0 — vllm-project/vllm
|
| 12 |
+
# https://pypi.org/project/vllm/
|
| 13 |
+
vllm==0.25.1
|
| 14 |
+
|
| 15 |
+
# Apache-2.0 — mlc-ai/xgrammar
|
| 16 |
+
# https://pypi.org/project/xgrammar/
|
| 17 |
+
xgrammar==0.2.3
|
special_tokens_map.json
ADDED
|
@@ -0,0 +1,16 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"bos_token": {
|
| 3 |
+
"content": "<s>",
|
| 4 |
+
"lstrip": false,
|
| 5 |
+
"normalized": false,
|
| 6 |
+
"rstrip": false,
|
| 7 |
+
"single_word": false
|
| 8 |
+
},
|
| 9 |
+
"eos_token": {
|
| 10 |
+
"content": "</s>",
|
| 11 |
+
"lstrip": false,
|
| 12 |
+
"normalized": false,
|
| 13 |
+
"rstrip": false,
|
| 14 |
+
"single_word": false
|
| 15 |
+
}
|
| 16 |
+
}
|
tokenizer.json
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:b4b3d90c67830a4e566296ad8d0f6b5ac5a5cdbfd331b200d0ee8263aaaea1fe
|
| 3 |
+
size 10680800
|
tokenizer_config.json
ADDED
|
@@ -0,0 +1,131 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"added_tokens_decoder": {
|
| 3 |
+
"1": {
|
| 4 |
+
"content": "<s>",
|
| 5 |
+
"lstrip": false,
|
| 6 |
+
"normalized": false,
|
| 7 |
+
"rstrip": false,
|
| 8 |
+
"single_word": false,
|
| 9 |
+
"special": true
|
| 10 |
+
},
|
| 11 |
+
"2": {
|
| 12 |
+
"content": "</s>",
|
| 13 |
+
"lstrip": false,
|
| 14 |
+
"normalized": false,
|
| 15 |
+
"rstrip": false,
|
| 16 |
+
"single_word": false,
|
| 17 |
+
"special": true
|
| 18 |
+
},
|
| 19 |
+
"128000": {
|
| 20 |
+
"content": "<|role_sep|>\n",
|
| 21 |
+
"lstrip": false,
|
| 22 |
+
"normalized": false,
|
| 23 |
+
"rstrip": false,
|
| 24 |
+
"single_word": false,
|
| 25 |
+
"special": true
|
| 26 |
+
},
|
| 27 |
+
"128001": {
|
| 28 |
+
"content": "<|message_sep|>\n\n",
|
| 29 |
+
"lstrip": false,
|
| 30 |
+
"normalized": false,
|
| 31 |
+
"rstrip": false,
|
| 32 |
+
"single_word": false,
|
| 33 |
+
"special": true
|
| 34 |
+
},
|
| 35 |
+
"128002": {
|
| 36 |
+
"content": "<|file|>",
|
| 37 |
+
"lstrip": false,
|
| 38 |
+
"normalized": false,
|
| 39 |
+
"rstrip": false,
|
| 40 |
+
"single_word": false,
|
| 41 |
+
"special": true
|
| 42 |
+
},
|
| 43 |
+
"128003": {
|
| 44 |
+
"content": "<|/file|>",
|
| 45 |
+
"lstrip": false,
|
| 46 |
+
"normalized": false,
|
| 47 |
+
"rstrip": false,
|
| 48 |
+
"single_word": false,
|
| 49 |
+
"special": true
|
| 50 |
+
},
|
| 51 |
+
"128004": {
|
| 52 |
+
"content": "[image_token]",
|
| 53 |
+
"lstrip": false,
|
| 54 |
+
"normalized": false,
|
| 55 |
+
"rstrip": false,
|
| 56 |
+
"single_word": false,
|
| 57 |
+
"special": true
|
| 58 |
+
},
|
| 59 |
+
"128005": {
|
| 60 |
+
"content": "[video_image_token]",
|
| 61 |
+
"lstrip": false,
|
| 62 |
+
"normalized": false,
|
| 63 |
+
"rstrip": false,
|
| 64 |
+
"single_word": false,
|
| 65 |
+
"special": true
|
| 66 |
+
},
|
| 67 |
+
"128006": {
|
| 68 |
+
"content": "[audio_token]",
|
| 69 |
+
"lstrip": false,
|
| 70 |
+
"normalized": false,
|
| 71 |
+
"rstrip": false,
|
| 72 |
+
"single_word": false,
|
| 73 |
+
"special": true
|
| 74 |
+
},
|
| 75 |
+
"128007": {
|
| 76 |
+
"content": "[video_audio_token]",
|
| 77 |
+
"lstrip": false,
|
| 78 |
+
"normalized": false,
|
| 79 |
+
"rstrip": false,
|
| 80 |
+
"single_word": false,
|
| 81 |
+
"special": true
|
| 82 |
+
},
|
| 83 |
+
"128008": {
|
| 84 |
+
"content": "<point>",
|
| 85 |
+
"lstrip": false,
|
| 86 |
+
"normalized": false,
|
| 87 |
+
"rstrip": false,
|
| 88 |
+
"single_word": false,
|
| 89 |
+
"special": true
|
| 90 |
+
},
|
| 91 |
+
"128009": {
|
| 92 |
+
"content": "</point>",
|
| 93 |
+
"lstrip": false,
|
| 94 |
+
"normalized": false,
|
| 95 |
+
"rstrip": false,
|
| 96 |
+
"single_word": false,
|
| 97 |
+
"special": true
|
| 98 |
+
},
|
| 99 |
+
"128010": {
|
| 100 |
+
"content": "<bbox>",
|
| 101 |
+
"lstrip": false,
|
| 102 |
+
"normalized": false,
|
| 103 |
+
"rstrip": false,
|
| 104 |
+
"single_word": false,
|
| 105 |
+
"special": true
|
| 106 |
+
},
|
| 107 |
+
"128011": {
|
| 108 |
+
"content": "</bbox>",
|
| 109 |
+
"lstrip": false,
|
| 110 |
+
"normalized": false,
|
| 111 |
+
"rstrip": false,
|
| 112 |
+
"single_word": false,
|
| 113 |
+
"special": true
|
| 114 |
+
},
|
| 115 |
+
"128012": {
|
| 116 |
+
"content": "<|function_call|>",
|
| 117 |
+
"lstrip": false,
|
| 118 |
+
"normalized": false,
|
| 119 |
+
"rstrip": false,
|
| 120 |
+
"single_word": false,
|
| 121 |
+
"special": true
|
| 122 |
+
}
|
| 123 |
+
},
|
| 124 |
+
"bos_token": "<s>",
|
| 125 |
+
"clean_up_tokenization_spaces": true,
|
| 126 |
+
"eos_token": "</s>",
|
| 127 |
+
"extra_special_tokens": {},
|
| 128 |
+
"model_max_length": 1000000000000000019884624838656,
|
| 129 |
+
"tokenizer_class": "PreTrainedTokenizer",
|
| 130 |
+
"unk_token": null
|
| 131 |
+
}
|