nmmursit commited on
Commit
7860bd8
·
verified ·
1 Parent(s): af20758

Initial model upload - clean repository

Browse files
.gitattributes CHANGED
@@ -33,3 +33,7 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
 
 
 
 
 
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
+ tokenizer.json filter=lfs diff=lfs merge=lfs -text
37
+ 4b_qwen_armo.png filter=lfs diff=lfs merge=lfs -text
38
+ comparison_rewards_by_token_length-filtered.png filter=lfs diff=lfs merge=lfs -text
39
+ qwen4b_loss.png filter=lfs diff=lfs merge=lfs -text
4b_qwen_armo.png ADDED

Git LFS Details

  • SHA256: e3d0225aae1e69ada1e46b3cacb4af5298239b80cf3e0ffd73fd3ede6402a651
  • Pointer size: 131 Bytes
  • Size of remote file: 109 kB
README.md ADDED
@@ -0,0 +1,197 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ language:
3
+ - tr
4
+ - en
5
+ license: apache-2.0
6
+ tags:
7
+ - text-generation
8
+ - turkish
9
+ - legal
10
+ - turkish-legal
11
+ - mecellem
12
+ - qwen
13
+ - decoder-only
14
+ - continual-pretraining
15
+ - TRUBA
16
+ - MN5
17
+ base_model: Qwen/Qwen3-4B
18
+ pipeline_tag: text-generation
19
+ ---
20
+
21
+ # Mecellem-Qwen3-4B-TR
22
+
23
+ [![License](https://img.shields.io/badge/License-Apache%202.0-blue.svg)](https://opensource.org/licenses/Apache-2.0)
24
+
25
+ ## Model Description
26
+
27
+ Mecellem-Qwen3-4B-TR is a Turkish legal language model adapted through Continual Pre-training (CPT) on Turkish legal and official texts. The model is based on Qwen3-4B decoder architecture (4B parameters) and trained using a single-phase, large-scale CPT process. Unlike the 1.7B model's four-phase curriculum learning, this model employs a single-phase training strategy on a comprehensive dataset, demonstrating that larger model capacity can benefit from direct large-scale domain adaptation.
28
+
29
+ **Key Features:**
30
+ - Continual pre-training on approximately 270.8 billion tokens in a single phase
31
+ - Single-phase large-scale CPT process (270,791,712,595 tokens)
32
+ - Dataset includes Turkish legal sources (Yargıtay, Danıştay, YÖKTEZ) and general Turkish web data (FineWeb2, CulturaX)
33
+ - Preserves general language capabilities while injecting domain-specific legal knowledge
34
+
35
+ **Model Type:** Decoder-only Language Model
36
+ **Parameters:** 4B
37
+ **Base Model:** Qwen/Qwen3-4B
38
+ **Architecture:** Qwen3 decoder with grouped query attention (GQA)
39
+
40
+ ### Architecture Details
41
+
42
+ - **Max Position Embeddings:** 40,960 tokens
43
+ - **Number of Layers:** 36 transformer layers
44
+ - **Hidden Size:** 2,560
45
+ - **FFN Hidden Size:** 9,728
46
+ - **Number of Heads:** 32
47
+ - **Number of KV Heads (GQA):** 8
48
+ - **Activation Function:** SwiGLU
49
+ - **Position Encodings:** RoPE (Rotary Position Embeddings)
50
+ - **Layer Norm:** RMSNorm
51
+
52
+ ### Training Details
53
+
54
+ **Continual Pre-training (CPT):**
55
+ - **Total Training Tokens:** ~270.8 billion tokens (270,791,712,595 tokens)
56
+ - **Training Method:** Single-phase large-scale CPT
57
+ - **Framework:** NVIDIA NeMo with Megatron-Core
58
+ - **Precision:** BF16 mixed precision
59
+ - **Hardware Infrastructure:**
60
+ - **System:** MareNostrum 5 ACC partition at Barcelona Supercomputing Center (BSC)
61
+ - **Compute Nodes:** 100 nodes
62
+ - **GPUs:** 400× NVIDIA Hopper H100 64GB GPUs (SXM) (4 GPUs per node)
63
+ - **Node Configuration:** Each node equipped with 4× H100 GPUs, 80 CPU cores, 512GB DDR5 memory
64
+ - **Interconnect:** 800 Gb/s InfiniBand for distributed training
65
+ - **GPU Interconnect:** NVLink for intra-node GPU communication (4 GPUs per node connected via NVLink)
66
+ - **Distributed Training:** Data-parallel multi-node and multi-GPU distributed architecture with 4 GPUs per node
67
+ - **InfiniBand Network:** Enabled efficient processing of large-scale token flow and ensured high scalability and training stability in long-term CPT training
68
+ - **Hardware Utilization:** 18.7% median MFU, 2.57M tokens/sec throughput
69
+
70
+ **Dataset Composition:**
71
+ - **Legal Sources:**
72
+ - Court of Cassation (Yargıtay): 10.3M sequences, ~3.43B tokens
73
+ - Council of State (Danıştay): 151K sequences, ~0.11B tokens
74
+ - Academic theses (YÖKTEZ): 21.1M sequences, ~9.61B tokens (after DocsOCR processing)
75
+ - **General Turkish Sources:**
76
+ - FineWeb2: General Turkish web data
77
+ - CulturaX: Multilingual corpus (Turkish subset)
78
+ - Total general Turkish: 212M sequences, ~96.17B tokens
79
+ - **Additional Categories:** English, Mathematics, Python code, multilingual content (Spanish, Arabic, Russian, Chinese)
80
+
81
+ **Training Hyperparameters:**
82
+ - Sequence Length: 4,096 tokens
83
+ - Optimizer: Adam with cosine learning rate schedule
84
+ - Max Learning Rate: 5×10⁻⁵
85
+ - Min Learning Rate: 5×10⁻⁶
86
+ - Weight Decay: 0.01
87
+ - Warmup Steps: 7,675 steps
88
+ - Max Steps: 153,508 steps
89
+ - Global Batch Size: 400
90
+ - Per-GPU Batch Size: 1
91
+ - Gradient Accumulation: 16
92
+
93
+ ### Training Visualization
94
+
95
+ The following visualizations show the model's training progress and dataset distribution:
96
+
97
+ ![Dataset Distribution](qwen4b_dataset.png)
98
+
99
+ *Qwen3-4B CPT Dataset Distribution Single Phase. The model was trained using a single-phase, large-scale CPT process.*
100
+
101
+ ![Training Loss](qwen4b_loss.png)
102
+
103
+ *Qwen3-4B CPT Training and Validation Loss Curves. The model shows consistent improvement throughout training.*
104
+
105
+ ### Benchmark Performance
106
+
107
+ The model was evaluated using the Muhakim reward model on Turkish legal tasks:
108
+
109
+ ![Benchmark Performance](4b_qwen_armo.png)
110
+
111
+ *Benchmark Performance of 4B Decoder-Only Models Across Context Lengths Using the Muhakim Reward Model. Mecellem-Qwen3-4B-TR consistently outperforms the base Qwen3-4B model across all five legal quality objectives.*
112
+
113
+ ### Rewards Comparison Analysis
114
+
115
+ The following visualization compares rewards across different token lengths for base vs CPT models:
116
+
117
+ ![Rewards Comparison](comparison_rewards_by_token_length-filtered.png)
118
+
119
+ *Rewards Comparison: Base vs CPT Models Across Token Lengths. Mecellem-Qwen3-4B-TR shows consistent improvements over the base model across all context length settings, demonstrating the effectiveness of Turkish legal domain adaptation.*
120
+
121
+
122
+ ## Usage
123
+
124
+ ### Installation
125
+
126
+ ```bash
127
+ pip install transformers torch
128
+ ```
129
+
130
+ ### Text Generation
131
+
132
+ ```python
133
+ from transformers import AutoTokenizer, AutoModelForCausalLM
134
+ import torch
135
+
136
+ # Load model and tokenizer
137
+ tokenizer = AutoTokenizer.from_pretrained("newmindai/Mecellem-Qwen3-4B-TR")
138
+ model = AutoModelForCausalLM.from_pretrained("newmindai/Mecellem-Qwen3-4B-TR")
139
+
140
+ # Example prompt
141
+ prompt = "Türk hukuk sisteminde sözleşme feshi"
142
+ inputs = tokenizer(prompt, return_tensors="pt")
143
+
144
+ # Generate
145
+ with torch.no_grad():
146
+ outputs = model.generate(
147
+ **inputs,
148
+ max_new_tokens=256,
149
+ temperature=0.7,
150
+ do_sample=True,
151
+ top_p=0.9
152
+ )
153
+
154
+ generated_text = tokenizer.decode(outputs[0], skip_special_tokens=True)
155
+ print(generated_text)
156
+ ```
157
+
158
+ ## Use Cases
159
+
160
+ - Turkish legal text generation
161
+ - Legal document summarization
162
+ - Legal question answering
163
+ - Legal text completion
164
+ - Domain-specific language modeling for Turkish legal domain
165
+ - Retrieval-Augmented Generation (RAG) applications
166
+
167
+ ## Acknowledgments
168
+
169
+ This work was supported by the EuroHPC Joint Undertaking through project etur46 with access to the MareNostrum 5 supercomputer, hosted by Barcelona Supercomputing Center (BSC), Spain. MareNostrum 5 is owned by EuroHPC JU and operated by BSC. We are grateful to the BSC support team for their assistance with job scheduling, environment configuration, and technical guidance throughout the project.
170
+
171
+ The numerical calculations reported in this work were fully/partially performed at TÜBİTAK ULAKBİM, High Performance and Grid Computing Center (TRUBA resources). The authors gratefully acknowledge the know-how provided by the MINERVA Support for expert guidance and collaboration opportunities in HPC-AI integration.
172
+
173
+ ## References
174
+
175
+ If you use this model, please cite our paper:
176
+
177
+ ```bibtex
178
+ @article{mecellem2026,
179
+ title={Mecellem Models: Turkish Models Trained from Scratch and Continually Pre-trained for the Legal Domain},
180
+ author={Uğur, Özgür and Göksu, Mahmut and Şavirdi, Esra and Çimen, Mahmut and Yılmaz, Musa and Demir, Alp Talha and Güllüce, Rumeysa and Çetin, İclal and Sağbaş, Ömer Can},
181
+ journal={Procedia Computer Science},
182
+ year={2026},
183
+ publisher={Elsevier}
184
+ }
185
+ ```
186
+ ### Base Model References
187
+
188
+ ```bibtex
189
+ @article{qwen2024,
190
+ title={Qwen3: A Large Language Model Series},
191
+ author={Qwen Team},
192
+ journal={arXiv preprint arXiv:2409.00000},
193
+ year={2024}
194
+ }
195
+ ```
196
+
197
+ <!-- Updated: 2026-01-15 09:38:43 -->
added_tokens.json ADDED
@@ -0,0 +1,28 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "</think>": 151668,
3
+ "</tool_call>": 151658,
4
+ "</tool_response>": 151666,
5
+ "<think>": 151667,
6
+ "<tool_call>": 151657,
7
+ "<tool_response>": 151665,
8
+ "<|box_end|>": 151649,
9
+ "<|box_start|>": 151648,
10
+ "<|endoftext|>": 151643,
11
+ "<|file_sep|>": 151664,
12
+ "<|fim_middle|>": 151660,
13
+ "<|fim_pad|>": 151662,
14
+ "<|fim_prefix|>": 151659,
15
+ "<|fim_suffix|>": 151661,
16
+ "<|im_end|>": 151645,
17
+ "<|im_start|>": 151644,
18
+ "<|image_pad|>": 151655,
19
+ "<|object_ref_end|>": 151647,
20
+ "<|object_ref_start|>": 151646,
21
+ "<|quad_end|>": 151651,
22
+ "<|quad_start|>": 151650,
23
+ "<|repo_name|>": 151663,
24
+ "<|video_pad|>": 151656,
25
+ "<|vision_end|>": 151653,
26
+ "<|vision_pad|>": 151654,
27
+ "<|vision_start|>": 151652
28
+ }
chat_template.jinja ADDED
@@ -0,0 +1,85 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {%- if tools %}
2
+ {{- '<|im_start|>system\n' }}
3
+ {%- if messages[0].role == 'system' %}
4
+ {{- messages[0].content + '\n\n' }}
5
+ {%- endif %}
6
+ {{- "# Tools\n\nYou may call one or more functions to assist with the user query.\n\nYou are provided with function signatures within <tools></tools> XML tags:\n<tools>" }}
7
+ {%- for tool in tools %}
8
+ {{- "\n" }}
9
+ {{- tool | tojson }}
10
+ {%- endfor %}
11
+ {{- "\n</tools>\n\nFor each function call, return a json object with function name and arguments within <tool_call></tool_call> XML tags:\n<tool_call>\n{\"name\": <function-name>, \"arguments\": <args-json-object>}\n</tool_call><|im_end|>\n" }}
12
+ {%- else %}
13
+ {%- if messages[0].role == 'system' %}
14
+ {{- '<|im_start|>system\n' + messages[0].content + '<|im_end|>\n' }}
15
+ {%- endif %}
16
+ {%- endif %}
17
+ {%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
18
+ {%- for message in messages[::-1] %}
19
+ {%- set index = (messages|length - 1) - loop.index0 %}
20
+ {%- if ns.multi_step_tool and message.role == "user" and not(message.content.startswith('<tool_response>') and message.content.endswith('</tool_response>')) %}
21
+ {%- set ns.multi_step_tool = false %}
22
+ {%- set ns.last_query_index = index %}
23
+ {%- endif %}
24
+ {%- endfor %}
25
+ {%- for message in messages %}
26
+ {%- if (message.role == "user") or (message.role == "system" and not loop.first) %}
27
+ {{- '<|im_start|>' + message.role + '\n' + message.content + '<|im_end|>' + '\n' }}
28
+ {%- elif message.role == "assistant" %}
29
+ {%- set content = message.content %}
30
+ {%- set reasoning_content = '' %}
31
+ {%- if message.reasoning_content is defined and message.reasoning_content is not none %}
32
+ {%- set reasoning_content = message.reasoning_content %}
33
+ {%- else %}
34
+ {%- if '</think>' in message.content %}
35
+ {%- set content = message.content.split('</think>')[-1].lstrip('\n') %}
36
+ {%- set reasoning_content = message.content.split('</think>')[0].rstrip('\n').split('<think>')[-1].lstrip('\n') %}
37
+ {%- endif %}
38
+ {%- endif %}
39
+ {%- if loop.index0 > ns.last_query_index %}
40
+ {%- if loop.last or (not loop.last and reasoning_content) %}
41
+ {{- '<|im_start|>' + message.role + '\n<think>\n' + reasoning_content.strip('\n') + '\n</think>\n\n' + content.lstrip('\n') }}
42
+ {%- else %}
43
+ {{- '<|im_start|>' + message.role + '\n' + content }}
44
+ {%- endif %}
45
+ {%- else %}
46
+ {{- '<|im_start|>' + message.role + '\n' + content }}
47
+ {%- endif %}
48
+ {%- if message.tool_calls %}
49
+ {%- for tool_call in message.tool_calls %}
50
+ {%- if (loop.first and content) or (not loop.first) %}
51
+ {{- '\n' }}
52
+ {%- endif %}
53
+ {%- if tool_call.function %}
54
+ {%- set tool_call = tool_call.function %}
55
+ {%- endif %}
56
+ {{- '<tool_call>\n{"name": "' }}
57
+ {{- tool_call.name }}
58
+ {{- '", "arguments": ' }}
59
+ {%- if tool_call.arguments is string %}
60
+ {{- tool_call.arguments }}
61
+ {%- else %}
62
+ {{- tool_call.arguments | tojson }}
63
+ {%- endif %}
64
+ {{- '}\n</tool_call>' }}
65
+ {%- endfor %}
66
+ {%- endif %}
67
+ {{- '<|im_end|>\n' }}
68
+ {%- elif message.role == "tool" %}
69
+ {%- if loop.first or (messages[loop.index0 - 1].role != "tool") %}
70
+ {{- '<|im_start|>user' }}
71
+ {%- endif %}
72
+ {{- '\n<tool_response>\n' }}
73
+ {{- message.content }}
74
+ {{- '\n</tool_response>' }}
75
+ {%- if loop.last or (messages[loop.index0 + 1].role != "tool") %}
76
+ {{- '<|im_end|>\n' }}
77
+ {%- endif %}
78
+ {%- endif %}
79
+ {%- endfor %}
80
+ {%- if add_generation_prompt %}
81
+ {{- '<|im_start|>assistant\n' }}
82
+ {%- if enable_thinking is defined and enable_thinking is false %}
83
+ {{- '<think>\n\n</think>\n\n' }}
84
+ {%- endif %}
85
+ {%- endif %}
comparison_rewards_by_token_length-filtered.png ADDED

Git LFS Details

  • SHA256: 9f877585b9786dc0db2a661c5b7711eb3db83b0d222b13b39cc21f8eca0944da
  • Pointer size: 131 Bytes
  • Size of remote file: 266 kB
config.json ADDED
@@ -0,0 +1,68 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "architectures": [
3
+ "Qwen3ForCausalLM"
4
+ ],
5
+ "attention_bias": false,
6
+ "attention_dropout": 0.0,
7
+ "bos_token_id": 151643,
8
+ "eos_token_id": 151645,
9
+ "head_dim": 128,
10
+ "hidden_act": "silu",
11
+ "hidden_size": 2560,
12
+ "initializer_range": 0.02,
13
+ "intermediate_size": 9728,
14
+ "layer_types": [
15
+ "full_attention",
16
+ "full_attention",
17
+ "full_attention",
18
+ "full_attention",
19
+ "full_attention",
20
+ "full_attention",
21
+ "full_attention",
22
+ "full_attention",
23
+ "full_attention",
24
+ "full_attention",
25
+ "full_attention",
26
+ "full_attention",
27
+ "full_attention",
28
+ "full_attention",
29
+ "full_attention",
30
+ "full_attention",
31
+ "full_attention",
32
+ "full_attention",
33
+ "full_attention",
34
+ "full_attention",
35
+ "full_attention",
36
+ "full_attention",
37
+ "full_attention",
38
+ "full_attention",
39
+ "full_attention",
40
+ "full_attention",
41
+ "full_attention",
42
+ "full_attention",
43
+ "full_attention",
44
+ "full_attention",
45
+ "full_attention",
46
+ "full_attention",
47
+ "full_attention",
48
+ "full_attention",
49
+ "full_attention",
50
+ "full_attention"
51
+ ],
52
+ "max_position_embeddings": 40960,
53
+ "max_window_layers": 36,
54
+ "model_type": "qwen3",
55
+ "num_attention_heads": 32,
56
+ "num_hidden_layers": 36,
57
+ "num_key_value_heads": 8,
58
+ "rms_norm_eps": 1e-06,
59
+ "rope_scaling": null,
60
+ "rope_theta": 1000000.0,
61
+ "sliding_window": null,
62
+ "tie_word_embeddings": true,
63
+ "torch_dtype": "bfloat16",
64
+ "transformers_version": "4.53.0",
65
+ "use_cache": true,
66
+ "use_sliding_window": false,
67
+ "vocab_size": 151936
68
+ }
generation_config.json ADDED
@@ -0,0 +1,6 @@
 
 
 
 
 
 
 
1
+ {
2
+ "_from_model_config": true,
3
+ "bos_token_id": 151643,
4
+ "eos_token_id": 151645,
5
+ "transformers_version": "4.53.0"
6
+ }
merges.txt ADDED
The diff for this file is too large to render. See raw diff
 
model-00001-of-00002.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:db0a3241bb2798ec48e2a8e792c0a282991aafb1f6d73a49ac74a60a77caaaf2
3
+ size 4967215360
model-00002-of-00002.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:d9d33e063e8613a50deee514e53ebeb3f0fca11c2a3029abb22d7d3bbbbd022c
3
+ size 3855679144
model.safetensors.index.json ADDED
@@ -0,0 +1,407 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "metadata": {
3
+ "total_parameters": 4411424256,
4
+ "total_size": 8822848512
5
+ },
6
+ "weight_map": {
7
+ "lm_head.weight": "model-00002-of-00002.safetensors",
8
+ "model.embed_tokens.weight": "model-00001-of-00002.safetensors",
9
+ "model.layers.0.input_layernorm.weight": "model-00001-of-00002.safetensors",
10
+ "model.layers.0.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
11
+ "model.layers.0.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
12
+ "model.layers.0.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
13
+ "model.layers.0.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
14
+ "model.layers.0.self_attn.k_norm.weight": "model-00001-of-00002.safetensors",
15
+ "model.layers.0.self_attn.k_proj.weight": "model-00001-of-00002.safetensors",
16
+ "model.layers.0.self_attn.o_proj.weight": "model-00001-of-00002.safetensors",
17
+ "model.layers.0.self_attn.q_norm.weight": "model-00001-of-00002.safetensors",
18
+ "model.layers.0.self_attn.q_proj.weight": "model-00001-of-00002.safetensors",
19
+ "model.layers.0.self_attn.v_proj.weight": "model-00001-of-00002.safetensors",
20
+ "model.layers.1.input_layernorm.weight": "model-00001-of-00002.safetensors",
21
+ "model.layers.1.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
22
+ "model.layers.1.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
23
+ "model.layers.1.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
24
+ "model.layers.1.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
25
+ "model.layers.1.self_attn.k_norm.weight": "model-00001-of-00002.safetensors",
26
+ "model.layers.1.self_attn.k_proj.weight": "model-00001-of-00002.safetensors",
27
+ "model.layers.1.self_attn.o_proj.weight": "model-00001-of-00002.safetensors",
28
+ "model.layers.1.self_attn.q_norm.weight": "model-00001-of-00002.safetensors",
29
+ "model.layers.1.self_attn.q_proj.weight": "model-00001-of-00002.safetensors",
30
+ "model.layers.1.self_attn.v_proj.weight": "model-00001-of-00002.safetensors",
31
+ "model.layers.10.input_layernorm.weight": "model-00001-of-00002.safetensors",
32
+ "model.layers.10.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
33
+ "model.layers.10.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
34
+ "model.layers.10.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
35
+ "model.layers.10.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
36
+ "model.layers.10.self_attn.k_norm.weight": "model-00001-of-00002.safetensors",
37
+ "model.layers.10.self_attn.k_proj.weight": "model-00001-of-00002.safetensors",
38
+ "model.layers.10.self_attn.o_proj.weight": "model-00001-of-00002.safetensors",
39
+ "model.layers.10.self_attn.q_norm.weight": "model-00001-of-00002.safetensors",
40
+ "model.layers.10.self_attn.q_proj.weight": "model-00001-of-00002.safetensors",
41
+ "model.layers.10.self_attn.v_proj.weight": "model-00001-of-00002.safetensors",
42
+ "model.layers.11.input_layernorm.weight": "model-00001-of-00002.safetensors",
43
+ "model.layers.11.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
44
+ "model.layers.11.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
45
+ "model.layers.11.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
46
+ "model.layers.11.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
47
+ "model.layers.11.self_attn.k_norm.weight": "model-00001-of-00002.safetensors",
48
+ "model.layers.11.self_attn.k_proj.weight": "model-00001-of-00002.safetensors",
49
+ "model.layers.11.self_attn.o_proj.weight": "model-00001-of-00002.safetensors",
50
+ "model.layers.11.self_attn.q_norm.weight": "model-00001-of-00002.safetensors",
51
+ "model.layers.11.self_attn.q_proj.weight": "model-00001-of-00002.safetensors",
52
+ "model.layers.11.self_attn.v_proj.weight": "model-00001-of-00002.safetensors",
53
+ "model.layers.12.input_layernorm.weight": "model-00001-of-00002.safetensors",
54
+ "model.layers.12.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
55
+ "model.layers.12.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
56
+ "model.layers.12.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
57
+ "model.layers.12.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
58
+ "model.layers.12.self_attn.k_norm.weight": "model-00001-of-00002.safetensors",
59
+ "model.layers.12.self_attn.k_proj.weight": "model-00001-of-00002.safetensors",
60
+ "model.layers.12.self_attn.o_proj.weight": "model-00001-of-00002.safetensors",
61
+ "model.layers.12.self_attn.q_norm.weight": "model-00001-of-00002.safetensors",
62
+ "model.layers.12.self_attn.q_proj.weight": "model-00001-of-00002.safetensors",
63
+ "model.layers.12.self_attn.v_proj.weight": "model-00001-of-00002.safetensors",
64
+ "model.layers.13.input_layernorm.weight": "model-00001-of-00002.safetensors",
65
+ "model.layers.13.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
66
+ "model.layers.13.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
67
+ "model.layers.13.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
68
+ "model.layers.13.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
69
+ "model.layers.13.self_attn.k_norm.weight": "model-00001-of-00002.safetensors",
70
+ "model.layers.13.self_attn.k_proj.weight": "model-00001-of-00002.safetensors",
71
+ "model.layers.13.self_attn.o_proj.weight": "model-00001-of-00002.safetensors",
72
+ "model.layers.13.self_attn.q_norm.weight": "model-00001-of-00002.safetensors",
73
+ "model.layers.13.self_attn.q_proj.weight": "model-00001-of-00002.safetensors",
74
+ "model.layers.13.self_attn.v_proj.weight": "model-00001-of-00002.safetensors",
75
+ "model.layers.14.input_layernorm.weight": "model-00001-of-00002.safetensors",
76
+ "model.layers.14.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
77
+ "model.layers.14.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
78
+ "model.layers.14.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
79
+ "model.layers.14.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
80
+ "model.layers.14.self_attn.k_norm.weight": "model-00001-of-00002.safetensors",
81
+ "model.layers.14.self_attn.k_proj.weight": "model-00001-of-00002.safetensors",
82
+ "model.layers.14.self_attn.o_proj.weight": "model-00001-of-00002.safetensors",
83
+ "model.layers.14.self_attn.q_norm.weight": "model-00001-of-00002.safetensors",
84
+ "model.layers.14.self_attn.q_proj.weight": "model-00001-of-00002.safetensors",
85
+ "model.layers.14.self_attn.v_proj.weight": "model-00001-of-00002.safetensors",
86
+ "model.layers.15.input_layernorm.weight": "model-00001-of-00002.safetensors",
87
+ "model.layers.15.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
88
+ "model.layers.15.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
89
+ "model.layers.15.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
90
+ "model.layers.15.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
91
+ "model.layers.15.self_attn.k_norm.weight": "model-00001-of-00002.safetensors",
92
+ "model.layers.15.self_attn.k_proj.weight": "model-00001-of-00002.safetensors",
93
+ "model.layers.15.self_attn.o_proj.weight": "model-00001-of-00002.safetensors",
94
+ "model.layers.15.self_attn.q_norm.weight": "model-00001-of-00002.safetensors",
95
+ "model.layers.15.self_attn.q_proj.weight": "model-00001-of-00002.safetensors",
96
+ "model.layers.15.self_attn.v_proj.weight": "model-00001-of-00002.safetensors",
97
+ "model.layers.16.input_layernorm.weight": "model-00001-of-00002.safetensors",
98
+ "model.layers.16.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
99
+ "model.layers.16.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
100
+ "model.layers.16.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
101
+ "model.layers.16.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
102
+ "model.layers.16.self_attn.k_norm.weight": "model-00001-of-00002.safetensors",
103
+ "model.layers.16.self_attn.k_proj.weight": "model-00001-of-00002.safetensors",
104
+ "model.layers.16.self_attn.o_proj.weight": "model-00001-of-00002.safetensors",
105
+ "model.layers.16.self_attn.q_norm.weight": "model-00001-of-00002.safetensors",
106
+ "model.layers.16.self_attn.q_proj.weight": "model-00001-of-00002.safetensors",
107
+ "model.layers.16.self_attn.v_proj.weight": "model-00001-of-00002.safetensors",
108
+ "model.layers.17.input_layernorm.weight": "model-00001-of-00002.safetensors",
109
+ "model.layers.17.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
110
+ "model.layers.17.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
111
+ "model.layers.17.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
112
+ "model.layers.17.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
113
+ "model.layers.17.self_attn.k_norm.weight": "model-00001-of-00002.safetensors",
114
+ "model.layers.17.self_attn.k_proj.weight": "model-00001-of-00002.safetensors",
115
+ "model.layers.17.self_attn.o_proj.weight": "model-00001-of-00002.safetensors",
116
+ "model.layers.17.self_attn.q_norm.weight": "model-00001-of-00002.safetensors",
117
+ "model.layers.17.self_attn.q_proj.weight": "model-00001-of-00002.safetensors",
118
+ "model.layers.17.self_attn.v_proj.weight": "model-00001-of-00002.safetensors",
119
+ "model.layers.18.input_layernorm.weight": "model-00001-of-00002.safetensors",
120
+ "model.layers.18.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
121
+ "model.layers.18.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
122
+ "model.layers.18.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
123
+ "model.layers.18.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
124
+ "model.layers.18.self_attn.k_norm.weight": "model-00001-of-00002.safetensors",
125
+ "model.layers.18.self_attn.k_proj.weight": "model-00001-of-00002.safetensors",
126
+ "model.layers.18.self_attn.o_proj.weight": "model-00001-of-00002.safetensors",
127
+ "model.layers.18.self_attn.q_norm.weight": "model-00001-of-00002.safetensors",
128
+ "model.layers.18.self_attn.q_proj.weight": "model-00001-of-00002.safetensors",
129
+ "model.layers.18.self_attn.v_proj.weight": "model-00001-of-00002.safetensors",
130
+ "model.layers.19.input_layernorm.weight": "model-00001-of-00002.safetensors",
131
+ "model.layers.19.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
132
+ "model.layers.19.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
133
+ "model.layers.19.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
134
+ "model.layers.19.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
135
+ "model.layers.19.self_attn.k_norm.weight": "model-00001-of-00002.safetensors",
136
+ "model.layers.19.self_attn.k_proj.weight": "model-00001-of-00002.safetensors",
137
+ "model.layers.19.self_attn.o_proj.weight": "model-00001-of-00002.safetensors",
138
+ "model.layers.19.self_attn.q_norm.weight": "model-00001-of-00002.safetensors",
139
+ "model.layers.19.self_attn.q_proj.weight": "model-00001-of-00002.safetensors",
140
+ "model.layers.19.self_attn.v_proj.weight": "model-00001-of-00002.safetensors",
141
+ "model.layers.2.input_layernorm.weight": "model-00001-of-00002.safetensors",
142
+ "model.layers.2.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
143
+ "model.layers.2.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
144
+ "model.layers.2.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
145
+ "model.layers.2.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
146
+ "model.layers.2.self_attn.k_norm.weight": "model-00001-of-00002.safetensors",
147
+ "model.layers.2.self_attn.k_proj.weight": "model-00001-of-00002.safetensors",
148
+ "model.layers.2.self_attn.o_proj.weight": "model-00001-of-00002.safetensors",
149
+ "model.layers.2.self_attn.q_norm.weight": "model-00001-of-00002.safetensors",
150
+ "model.layers.2.self_attn.q_proj.weight": "model-00001-of-00002.safetensors",
151
+ "model.layers.2.self_attn.v_proj.weight": "model-00001-of-00002.safetensors",
152
+ "model.layers.20.input_layernorm.weight": "model-00002-of-00002.safetensors",
153
+ "model.layers.20.mlp.down_proj.weight": "model-00002-of-00002.safetensors",
154
+ "model.layers.20.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
155
+ "model.layers.20.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
156
+ "model.layers.20.post_attention_layernorm.weight": "model-00002-of-00002.safetensors",
157
+ "model.layers.20.self_attn.k_norm.weight": "model-00001-of-00002.safetensors",
158
+ "model.layers.20.self_attn.k_proj.weight": "model-00001-of-00002.safetensors",
159
+ "model.layers.20.self_attn.o_proj.weight": "model-00001-of-00002.safetensors",
160
+ "model.layers.20.self_attn.q_norm.weight": "model-00001-of-00002.safetensors",
161
+ "model.layers.20.self_attn.q_proj.weight": "model-00001-of-00002.safetensors",
162
+ "model.layers.20.self_attn.v_proj.weight": "model-00001-of-00002.safetensors",
163
+ "model.layers.21.input_layernorm.weight": "model-00002-of-00002.safetensors",
164
+ "model.layers.21.mlp.down_proj.weight": "model-00002-of-00002.safetensors",
165
+ "model.layers.21.mlp.gate_proj.weight": "model-00002-of-00002.safetensors",
166
+ "model.layers.21.mlp.up_proj.weight": "model-00002-of-00002.safetensors",
167
+ "model.layers.21.post_attention_layernorm.weight": "model-00002-of-00002.safetensors",
168
+ "model.layers.21.self_attn.k_norm.weight": "model-00002-of-00002.safetensors",
169
+ "model.layers.21.self_attn.k_proj.weight": "model-00002-of-00002.safetensors",
170
+ "model.layers.21.self_attn.o_proj.weight": "model-00002-of-00002.safetensors",
171
+ "model.layers.21.self_attn.q_norm.weight": "model-00002-of-00002.safetensors",
172
+ "model.layers.21.self_attn.q_proj.weight": "model-00002-of-00002.safetensors",
173
+ "model.layers.21.self_attn.v_proj.weight": "model-00002-of-00002.safetensors",
174
+ "model.layers.22.input_layernorm.weight": "model-00002-of-00002.safetensors",
175
+ "model.layers.22.mlp.down_proj.weight": "model-00002-of-00002.safetensors",
176
+ "model.layers.22.mlp.gate_proj.weight": "model-00002-of-00002.safetensors",
177
+ "model.layers.22.mlp.up_proj.weight": "model-00002-of-00002.safetensors",
178
+ "model.layers.22.post_attention_layernorm.weight": "model-00002-of-00002.safetensors",
179
+ "model.layers.22.self_attn.k_norm.weight": "model-00002-of-00002.safetensors",
180
+ "model.layers.22.self_attn.k_proj.weight": "model-00002-of-00002.safetensors",
181
+ "model.layers.22.self_attn.o_proj.weight": "model-00002-of-00002.safetensors",
182
+ "model.layers.22.self_attn.q_norm.weight": "model-00002-of-00002.safetensors",
183
+ "model.layers.22.self_attn.q_proj.weight": "model-00002-of-00002.safetensors",
184
+ "model.layers.22.self_attn.v_proj.weight": "model-00002-of-00002.safetensors",
185
+ "model.layers.23.input_layernorm.weight": "model-00002-of-00002.safetensors",
186
+ "model.layers.23.mlp.down_proj.weight": "model-00002-of-00002.safetensors",
187
+ "model.layers.23.mlp.gate_proj.weight": "model-00002-of-00002.safetensors",
188
+ "model.layers.23.mlp.up_proj.weight": "model-00002-of-00002.safetensors",
189
+ "model.layers.23.post_attention_layernorm.weight": "model-00002-of-00002.safetensors",
190
+ "model.layers.23.self_attn.k_norm.weight": "model-00002-of-00002.safetensors",
191
+ "model.layers.23.self_attn.k_proj.weight": "model-00002-of-00002.safetensors",
192
+ "model.layers.23.self_attn.o_proj.weight": "model-00002-of-00002.safetensors",
193
+ "model.layers.23.self_attn.q_norm.weight": "model-00002-of-00002.safetensors",
194
+ "model.layers.23.self_attn.q_proj.weight": "model-00002-of-00002.safetensors",
195
+ "model.layers.23.self_attn.v_proj.weight": "model-00002-of-00002.safetensors",
196
+ "model.layers.24.input_layernorm.weight": "model-00002-of-00002.safetensors",
197
+ "model.layers.24.mlp.down_proj.weight": "model-00002-of-00002.safetensors",
198
+ "model.layers.24.mlp.gate_proj.weight": "model-00002-of-00002.safetensors",
199
+ "model.layers.24.mlp.up_proj.weight": "model-00002-of-00002.safetensors",
200
+ "model.layers.24.post_attention_layernorm.weight": "model-00002-of-00002.safetensors",
201
+ "model.layers.24.self_attn.k_norm.weight": "model-00002-of-00002.safetensors",
202
+ "model.layers.24.self_attn.k_proj.weight": "model-00002-of-00002.safetensors",
203
+ "model.layers.24.self_attn.o_proj.weight": "model-00002-of-00002.safetensors",
204
+ "model.layers.24.self_attn.q_norm.weight": "model-00002-of-00002.safetensors",
205
+ "model.layers.24.self_attn.q_proj.weight": "model-00002-of-00002.safetensors",
206
+ "model.layers.24.self_attn.v_proj.weight": "model-00002-of-00002.safetensors",
207
+ "model.layers.25.input_layernorm.weight": "model-00002-of-00002.safetensors",
208
+ "model.layers.25.mlp.down_proj.weight": "model-00002-of-00002.safetensors",
209
+ "model.layers.25.mlp.gate_proj.weight": "model-00002-of-00002.safetensors",
210
+ "model.layers.25.mlp.up_proj.weight": "model-00002-of-00002.safetensors",
211
+ "model.layers.25.post_attention_layernorm.weight": "model-00002-of-00002.safetensors",
212
+ "model.layers.25.self_attn.k_norm.weight": "model-00002-of-00002.safetensors",
213
+ "model.layers.25.self_attn.k_proj.weight": "model-00002-of-00002.safetensors",
214
+ "model.layers.25.self_attn.o_proj.weight": "model-00002-of-00002.safetensors",
215
+ "model.layers.25.self_attn.q_norm.weight": "model-00002-of-00002.safetensors",
216
+ "model.layers.25.self_attn.q_proj.weight": "model-00002-of-00002.safetensors",
217
+ "model.layers.25.self_attn.v_proj.weight": "model-00002-of-00002.safetensors",
218
+ "model.layers.26.input_layernorm.weight": "model-00002-of-00002.safetensors",
219
+ "model.layers.26.mlp.down_proj.weight": "model-00002-of-00002.safetensors",
220
+ "model.layers.26.mlp.gate_proj.weight": "model-00002-of-00002.safetensors",
221
+ "model.layers.26.mlp.up_proj.weight": "model-00002-of-00002.safetensors",
222
+ "model.layers.26.post_attention_layernorm.weight": "model-00002-of-00002.safetensors",
223
+ "model.layers.26.self_attn.k_norm.weight": "model-00002-of-00002.safetensors",
224
+ "model.layers.26.self_attn.k_proj.weight": "model-00002-of-00002.safetensors",
225
+ "model.layers.26.self_attn.o_proj.weight": "model-00002-of-00002.safetensors",
226
+ "model.layers.26.self_attn.q_norm.weight": "model-00002-of-00002.safetensors",
227
+ "model.layers.26.self_attn.q_proj.weight": "model-00002-of-00002.safetensors",
228
+ "model.layers.26.self_attn.v_proj.weight": "model-00002-of-00002.safetensors",
229
+ "model.layers.27.input_layernorm.weight": "model-00002-of-00002.safetensors",
230
+ "model.layers.27.mlp.down_proj.weight": "model-00002-of-00002.safetensors",
231
+ "model.layers.27.mlp.gate_proj.weight": "model-00002-of-00002.safetensors",
232
+ "model.layers.27.mlp.up_proj.weight": "model-00002-of-00002.safetensors",
233
+ "model.layers.27.post_attention_layernorm.weight": "model-00002-of-00002.safetensors",
234
+ "model.layers.27.self_attn.k_norm.weight": "model-00002-of-00002.safetensors",
235
+ "model.layers.27.self_attn.k_proj.weight": "model-00002-of-00002.safetensors",
236
+ "model.layers.27.self_attn.o_proj.weight": "model-00002-of-00002.safetensors",
237
+ "model.layers.27.self_attn.q_norm.weight": "model-00002-of-00002.safetensors",
238
+ "model.layers.27.self_attn.q_proj.weight": "model-00002-of-00002.safetensors",
239
+ "model.layers.27.self_attn.v_proj.weight": "model-00002-of-00002.safetensors",
240
+ "model.layers.28.input_layernorm.weight": "model-00002-of-00002.safetensors",
241
+ "model.layers.28.mlp.down_proj.weight": "model-00002-of-00002.safetensors",
242
+ "model.layers.28.mlp.gate_proj.weight": "model-00002-of-00002.safetensors",
243
+ "model.layers.28.mlp.up_proj.weight": "model-00002-of-00002.safetensors",
244
+ "model.layers.28.post_attention_layernorm.weight": "model-00002-of-00002.safetensors",
245
+ "model.layers.28.self_attn.k_norm.weight": "model-00002-of-00002.safetensors",
246
+ "model.layers.28.self_attn.k_proj.weight": "model-00002-of-00002.safetensors",
247
+ "model.layers.28.self_attn.o_proj.weight": "model-00002-of-00002.safetensors",
248
+ "model.layers.28.self_attn.q_norm.weight": "model-00002-of-00002.safetensors",
249
+ "model.layers.28.self_attn.q_proj.weight": "model-00002-of-00002.safetensors",
250
+ "model.layers.28.self_attn.v_proj.weight": "model-00002-of-00002.safetensors",
251
+ "model.layers.29.input_layernorm.weight": "model-00002-of-00002.safetensors",
252
+ "model.layers.29.mlp.down_proj.weight": "model-00002-of-00002.safetensors",
253
+ "model.layers.29.mlp.gate_proj.weight": "model-00002-of-00002.safetensors",
254
+ "model.layers.29.mlp.up_proj.weight": "model-00002-of-00002.safetensors",
255
+ "model.layers.29.post_attention_layernorm.weight": "model-00002-of-00002.safetensors",
256
+ "model.layers.29.self_attn.k_norm.weight": "model-00002-of-00002.safetensors",
257
+ "model.layers.29.self_attn.k_proj.weight": "model-00002-of-00002.safetensors",
258
+ "model.layers.29.self_attn.o_proj.weight": "model-00002-of-00002.safetensors",
259
+ "model.layers.29.self_attn.q_norm.weight": "model-00002-of-00002.safetensors",
260
+ "model.layers.29.self_attn.q_proj.weight": "model-00002-of-00002.safetensors",
261
+ "model.layers.29.self_attn.v_proj.weight": "model-00002-of-00002.safetensors",
262
+ "model.layers.3.input_layernorm.weight": "model-00001-of-00002.safetensors",
263
+ "model.layers.3.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
264
+ "model.layers.3.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
265
+ "model.layers.3.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
266
+ "model.layers.3.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
267
+ "model.layers.3.self_attn.k_norm.weight": "model-00001-of-00002.safetensors",
268
+ "model.layers.3.self_attn.k_proj.weight": "model-00001-of-00002.safetensors",
269
+ "model.layers.3.self_attn.o_proj.weight": "model-00001-of-00002.safetensors",
270
+ "model.layers.3.self_attn.q_norm.weight": "model-00001-of-00002.safetensors",
271
+ "model.layers.3.self_attn.q_proj.weight": "model-00001-of-00002.safetensors",
272
+ "model.layers.3.self_attn.v_proj.weight": "model-00001-of-00002.safetensors",
273
+ "model.layers.30.input_layernorm.weight": "model-00002-of-00002.safetensors",
274
+ "model.layers.30.mlp.down_proj.weight": "model-00002-of-00002.safetensors",
275
+ "model.layers.30.mlp.gate_proj.weight": "model-00002-of-00002.safetensors",
276
+ "model.layers.30.mlp.up_proj.weight": "model-00002-of-00002.safetensors",
277
+ "model.layers.30.post_attention_layernorm.weight": "model-00002-of-00002.safetensors",
278
+ "model.layers.30.self_attn.k_norm.weight": "model-00002-of-00002.safetensors",
279
+ "model.layers.30.self_attn.k_proj.weight": "model-00002-of-00002.safetensors",
280
+ "model.layers.30.self_attn.o_proj.weight": "model-00002-of-00002.safetensors",
281
+ "model.layers.30.self_attn.q_norm.weight": "model-00002-of-00002.safetensors",
282
+ "model.layers.30.self_attn.q_proj.weight": "model-00002-of-00002.safetensors",
283
+ "model.layers.30.self_attn.v_proj.weight": "model-00002-of-00002.safetensors",
284
+ "model.layers.31.input_layernorm.weight": "model-00002-of-00002.safetensors",
285
+ "model.layers.31.mlp.down_proj.weight": "model-00002-of-00002.safetensors",
286
+ "model.layers.31.mlp.gate_proj.weight": "model-00002-of-00002.safetensors",
287
+ "model.layers.31.mlp.up_proj.weight": "model-00002-of-00002.safetensors",
288
+ "model.layers.31.post_attention_layernorm.weight": "model-00002-of-00002.safetensors",
289
+ "model.layers.31.self_attn.k_norm.weight": "model-00002-of-00002.safetensors",
290
+ "model.layers.31.self_attn.k_proj.weight": "model-00002-of-00002.safetensors",
291
+ "model.layers.31.self_attn.o_proj.weight": "model-00002-of-00002.safetensors",
292
+ "model.layers.31.self_attn.q_norm.weight": "model-00002-of-00002.safetensors",
293
+ "model.layers.31.self_attn.q_proj.weight": "model-00002-of-00002.safetensors",
294
+ "model.layers.31.self_attn.v_proj.weight": "model-00002-of-00002.safetensors",
295
+ "model.layers.32.input_layernorm.weight": "model-00002-of-00002.safetensors",
296
+ "model.layers.32.mlp.down_proj.weight": "model-00002-of-00002.safetensors",
297
+ "model.layers.32.mlp.gate_proj.weight": "model-00002-of-00002.safetensors",
298
+ "model.layers.32.mlp.up_proj.weight": "model-00002-of-00002.safetensors",
299
+ "model.layers.32.post_attention_layernorm.weight": "model-00002-of-00002.safetensors",
300
+ "model.layers.32.self_attn.k_norm.weight": "model-00002-of-00002.safetensors",
301
+ "model.layers.32.self_attn.k_proj.weight": "model-00002-of-00002.safetensors",
302
+ "model.layers.32.self_attn.o_proj.weight": "model-00002-of-00002.safetensors",
303
+ "model.layers.32.self_attn.q_norm.weight": "model-00002-of-00002.safetensors",
304
+ "model.layers.32.self_attn.q_proj.weight": "model-00002-of-00002.safetensors",
305
+ "model.layers.32.self_attn.v_proj.weight": "model-00002-of-00002.safetensors",
306
+ "model.layers.33.input_layernorm.weight": "model-00002-of-00002.safetensors",
307
+ "model.layers.33.mlp.down_proj.weight": "model-00002-of-00002.safetensors",
308
+ "model.layers.33.mlp.gate_proj.weight": "model-00002-of-00002.safetensors",
309
+ "model.layers.33.mlp.up_proj.weight": "model-00002-of-00002.safetensors",
310
+ "model.layers.33.post_attention_layernorm.weight": "model-00002-of-00002.safetensors",
311
+ "model.layers.33.self_attn.k_norm.weight": "model-00002-of-00002.safetensors",
312
+ "model.layers.33.self_attn.k_proj.weight": "model-00002-of-00002.safetensors",
313
+ "model.layers.33.self_attn.o_proj.weight": "model-00002-of-00002.safetensors",
314
+ "model.layers.33.self_attn.q_norm.weight": "model-00002-of-00002.safetensors",
315
+ "model.layers.33.self_attn.q_proj.weight": "model-00002-of-00002.safetensors",
316
+ "model.layers.33.self_attn.v_proj.weight": "model-00002-of-00002.safetensors",
317
+ "model.layers.34.input_layernorm.weight": "model-00002-of-00002.safetensors",
318
+ "model.layers.34.mlp.down_proj.weight": "model-00002-of-00002.safetensors",
319
+ "model.layers.34.mlp.gate_proj.weight": "model-00002-of-00002.safetensors",
320
+ "model.layers.34.mlp.up_proj.weight": "model-00002-of-00002.safetensors",
321
+ "model.layers.34.post_attention_layernorm.weight": "model-00002-of-00002.safetensors",
322
+ "model.layers.34.self_attn.k_norm.weight": "model-00002-of-00002.safetensors",
323
+ "model.layers.34.self_attn.k_proj.weight": "model-00002-of-00002.safetensors",
324
+ "model.layers.34.self_attn.o_proj.weight": "model-00002-of-00002.safetensors",
325
+ "model.layers.34.self_attn.q_norm.weight": "model-00002-of-00002.safetensors",
326
+ "model.layers.34.self_attn.q_proj.weight": "model-00002-of-00002.safetensors",
327
+ "model.layers.34.self_attn.v_proj.weight": "model-00002-of-00002.safetensors",
328
+ "model.layers.35.input_layernorm.weight": "model-00002-of-00002.safetensors",
329
+ "model.layers.35.mlp.down_proj.weight": "model-00002-of-00002.safetensors",
330
+ "model.layers.35.mlp.gate_proj.weight": "model-00002-of-00002.safetensors",
331
+ "model.layers.35.mlp.up_proj.weight": "model-00002-of-00002.safetensors",
332
+ "model.layers.35.post_attention_layernorm.weight": "model-00002-of-00002.safetensors",
333
+ "model.layers.35.self_attn.k_norm.weight": "model-00002-of-00002.safetensors",
334
+ "model.layers.35.self_attn.k_proj.weight": "model-00002-of-00002.safetensors",
335
+ "model.layers.35.self_attn.o_proj.weight": "model-00002-of-00002.safetensors",
336
+ "model.layers.35.self_attn.q_norm.weight": "model-00002-of-00002.safetensors",
337
+ "model.layers.35.self_attn.q_proj.weight": "model-00002-of-00002.safetensors",
338
+ "model.layers.35.self_attn.v_proj.weight": "model-00002-of-00002.safetensors",
339
+ "model.layers.4.input_layernorm.weight": "model-00001-of-00002.safetensors",
340
+ "model.layers.4.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
341
+ "model.layers.4.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
342
+ "model.layers.4.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
343
+ "model.layers.4.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
344
+ "model.layers.4.self_attn.k_norm.weight": "model-00001-of-00002.safetensors",
345
+ "model.layers.4.self_attn.k_proj.weight": "model-00001-of-00002.safetensors",
346
+ "model.layers.4.self_attn.o_proj.weight": "model-00001-of-00002.safetensors",
347
+ "model.layers.4.self_attn.q_norm.weight": "model-00001-of-00002.safetensors",
348
+ "model.layers.4.self_attn.q_proj.weight": "model-00001-of-00002.safetensors",
349
+ "model.layers.4.self_attn.v_proj.weight": "model-00001-of-00002.safetensors",
350
+ "model.layers.5.input_layernorm.weight": "model-00001-of-00002.safetensors",
351
+ "model.layers.5.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
352
+ "model.layers.5.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
353
+ "model.layers.5.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
354
+ "model.layers.5.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
355
+ "model.layers.5.self_attn.k_norm.weight": "model-00001-of-00002.safetensors",
356
+ "model.layers.5.self_attn.k_proj.weight": "model-00001-of-00002.safetensors",
357
+ "model.layers.5.self_attn.o_proj.weight": "model-00001-of-00002.safetensors",
358
+ "model.layers.5.self_attn.q_norm.weight": "model-00001-of-00002.safetensors",
359
+ "model.layers.5.self_attn.q_proj.weight": "model-00001-of-00002.safetensors",
360
+ "model.layers.5.self_attn.v_proj.weight": "model-00001-of-00002.safetensors",
361
+ "model.layers.6.input_layernorm.weight": "model-00001-of-00002.safetensors",
362
+ "model.layers.6.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
363
+ "model.layers.6.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
364
+ "model.layers.6.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
365
+ "model.layers.6.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
366
+ "model.layers.6.self_attn.k_norm.weight": "model-00001-of-00002.safetensors",
367
+ "model.layers.6.self_attn.k_proj.weight": "model-00001-of-00002.safetensors",
368
+ "model.layers.6.self_attn.o_proj.weight": "model-00001-of-00002.safetensors",
369
+ "model.layers.6.self_attn.q_norm.weight": "model-00001-of-00002.safetensors",
370
+ "model.layers.6.self_attn.q_proj.weight": "model-00001-of-00002.safetensors",
371
+ "model.layers.6.self_attn.v_proj.weight": "model-00001-of-00002.safetensors",
372
+ "model.layers.7.input_layernorm.weight": "model-00001-of-00002.safetensors",
373
+ "model.layers.7.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
374
+ "model.layers.7.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
375
+ "model.layers.7.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
376
+ "model.layers.7.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
377
+ "model.layers.7.self_attn.k_norm.weight": "model-00001-of-00002.safetensors",
378
+ "model.layers.7.self_attn.k_proj.weight": "model-00001-of-00002.safetensors",
379
+ "model.layers.7.self_attn.o_proj.weight": "model-00001-of-00002.safetensors",
380
+ "model.layers.7.self_attn.q_norm.weight": "model-00001-of-00002.safetensors",
381
+ "model.layers.7.self_attn.q_proj.weight": "model-00001-of-00002.safetensors",
382
+ "model.layers.7.self_attn.v_proj.weight": "model-00001-of-00002.safetensors",
383
+ "model.layers.8.input_layernorm.weight": "model-00001-of-00002.safetensors",
384
+ "model.layers.8.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
385
+ "model.layers.8.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
386
+ "model.layers.8.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
387
+ "model.layers.8.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
388
+ "model.layers.8.self_attn.k_norm.weight": "model-00001-of-00002.safetensors",
389
+ "model.layers.8.self_attn.k_proj.weight": "model-00001-of-00002.safetensors",
390
+ "model.layers.8.self_attn.o_proj.weight": "model-00001-of-00002.safetensors",
391
+ "model.layers.8.self_attn.q_norm.weight": "model-00001-of-00002.safetensors",
392
+ "model.layers.8.self_attn.q_proj.weight": "model-00001-of-00002.safetensors",
393
+ "model.layers.8.self_attn.v_proj.weight": "model-00001-of-00002.safetensors",
394
+ "model.layers.9.input_layernorm.weight": "model-00001-of-00002.safetensors",
395
+ "model.layers.9.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
396
+ "model.layers.9.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
397
+ "model.layers.9.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
398
+ "model.layers.9.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
399
+ "model.layers.9.self_attn.k_norm.weight": "model-00001-of-00002.safetensors",
400
+ "model.layers.9.self_attn.k_proj.weight": "model-00001-of-00002.safetensors",
401
+ "model.layers.9.self_attn.o_proj.weight": "model-00001-of-00002.safetensors",
402
+ "model.layers.9.self_attn.q_norm.weight": "model-00001-of-00002.safetensors",
403
+ "model.layers.9.self_attn.q_proj.weight": "model-00001-of-00002.safetensors",
404
+ "model.layers.9.self_attn.v_proj.weight": "model-00001-of-00002.safetensors",
405
+ "model.norm.weight": "model-00002-of-00002.safetensors"
406
+ }
407
+ }
qwen4b_dataset.png ADDED
qwen4b_loss.png ADDED
special_tokens_map.json ADDED
@@ -0,0 +1,38 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "additional_special_tokens": [
3
+ "<|im_start|>",
4
+ "<|im_end|>",
5
+ "<|object_ref_start|>",
6
+ "<|object_ref_end|>",
7
+ "<|box_start|>",
8
+ "<|box_end|>",
9
+ "<|quad_start|>",
10
+ "<|quad_end|>",
11
+ "<|vision_start|>",
12
+ "<|vision_end|>",
13
+ "<|vision_pad|>",
14
+ "<|image_pad|>",
15
+ "<|video_pad|>"
16
+ ],
17
+ "eos_token": {
18
+ "content": "<|endoftext|>",
19
+ "lstrip": false,
20
+ "normalized": false,
21
+ "rstrip": false,
22
+ "single_word": false
23
+ },
24
+ "pad_token": {
25
+ "content": "<|endoftext|>",
26
+ "lstrip": false,
27
+ "normalized": false,
28
+ "rstrip": false,
29
+ "single_word": false
30
+ },
31
+ "sep_token": {
32
+ "content": "<|endoftext|>",
33
+ "lstrip": false,
34
+ "normalized": false,
35
+ "rstrip": false,
36
+ "single_word": false
37
+ }
38
+ }
tokenizer.json ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:aeb13307a71acd8fe81861d94ad54ab689df773318809eed3cbe794b4492dae4
3
+ size 11422654
tokenizer_config.json ADDED
@@ -0,0 +1,240 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "add_bos_token": false,
3
+ "add_prefix_space": false,
4
+ "added_tokens_decoder": {
5
+ "151643": {
6
+ "content": "<|endoftext|>",
7
+ "lstrip": false,
8
+ "normalized": false,
9
+ "rstrip": false,
10
+ "single_word": false,
11
+ "special": true
12
+ },
13
+ "151644": {
14
+ "content": "<|im_start|>",
15
+ "lstrip": false,
16
+ "normalized": false,
17
+ "rstrip": false,
18
+ "single_word": false,
19
+ "special": true
20
+ },
21
+ "151645": {
22
+ "content": "<|im_end|>",
23
+ "lstrip": false,
24
+ "normalized": false,
25
+ "rstrip": false,
26
+ "single_word": false,
27
+ "special": true
28
+ },
29
+ "151646": {
30
+ "content": "<|object_ref_start|>",
31
+ "lstrip": false,
32
+ "normalized": false,
33
+ "rstrip": false,
34
+ "single_word": false,
35
+ "special": true
36
+ },
37
+ "151647": {
38
+ "content": "<|object_ref_end|>",
39
+ "lstrip": false,
40
+ "normalized": false,
41
+ "rstrip": false,
42
+ "single_word": false,
43
+ "special": true
44
+ },
45
+ "151648": {
46
+ "content": "<|box_start|>",
47
+ "lstrip": false,
48
+ "normalized": false,
49
+ "rstrip": false,
50
+ "single_word": false,
51
+ "special": true
52
+ },
53
+ "151649": {
54
+ "content": "<|box_end|>",
55
+ "lstrip": false,
56
+ "normalized": false,
57
+ "rstrip": false,
58
+ "single_word": false,
59
+ "special": true
60
+ },
61
+ "151650": {
62
+ "content": "<|quad_start|>",
63
+ "lstrip": false,
64
+ "normalized": false,
65
+ "rstrip": false,
66
+ "single_word": false,
67
+ "special": true
68
+ },
69
+ "151651": {
70
+ "content": "<|quad_end|>",
71
+ "lstrip": false,
72
+ "normalized": false,
73
+ "rstrip": false,
74
+ "single_word": false,
75
+ "special": true
76
+ },
77
+ "151652": {
78
+ "content": "<|vision_start|>",
79
+ "lstrip": false,
80
+ "normalized": false,
81
+ "rstrip": false,
82
+ "single_word": false,
83
+ "special": true
84
+ },
85
+ "151653": {
86
+ "content": "<|vision_end|>",
87
+ "lstrip": false,
88
+ "normalized": false,
89
+ "rstrip": false,
90
+ "single_word": false,
91
+ "special": true
92
+ },
93
+ "151654": {
94
+ "content": "<|vision_pad|>",
95
+ "lstrip": false,
96
+ "normalized": false,
97
+ "rstrip": false,
98
+ "single_word": false,
99
+ "special": true
100
+ },
101
+ "151655": {
102
+ "content": "<|image_pad|>",
103
+ "lstrip": false,
104
+ "normalized": false,
105
+ "rstrip": false,
106
+ "single_word": false,
107
+ "special": true
108
+ },
109
+ "151656": {
110
+ "content": "<|video_pad|>",
111
+ "lstrip": false,
112
+ "normalized": false,
113
+ "rstrip": false,
114
+ "single_word": false,
115
+ "special": true
116
+ },
117
+ "151657": {
118
+ "content": "<tool_call>",
119
+ "lstrip": false,
120
+ "normalized": false,
121
+ "rstrip": false,
122
+ "single_word": false,
123
+ "special": false
124
+ },
125
+ "151658": {
126
+ "content": "</tool_call>",
127
+ "lstrip": false,
128
+ "normalized": false,
129
+ "rstrip": false,
130
+ "single_word": false,
131
+ "special": false
132
+ },
133
+ "151659": {
134
+ "content": "<|fim_prefix|>",
135
+ "lstrip": false,
136
+ "normalized": false,
137
+ "rstrip": false,
138
+ "single_word": false,
139
+ "special": false
140
+ },
141
+ "151660": {
142
+ "content": "<|fim_middle|>",
143
+ "lstrip": false,
144
+ "normalized": false,
145
+ "rstrip": false,
146
+ "single_word": false,
147
+ "special": false
148
+ },
149
+ "151661": {
150
+ "content": "<|fim_suffix|>",
151
+ "lstrip": false,
152
+ "normalized": false,
153
+ "rstrip": false,
154
+ "single_word": false,
155
+ "special": false
156
+ },
157
+ "151662": {
158
+ "content": "<|fim_pad|>",
159
+ "lstrip": false,
160
+ "normalized": false,
161
+ "rstrip": false,
162
+ "single_word": false,
163
+ "special": false
164
+ },
165
+ "151663": {
166
+ "content": "<|repo_name|>",
167
+ "lstrip": false,
168
+ "normalized": false,
169
+ "rstrip": false,
170
+ "single_word": false,
171
+ "special": false
172
+ },
173
+ "151664": {
174
+ "content": "<|file_sep|>",
175
+ "lstrip": false,
176
+ "normalized": false,
177
+ "rstrip": false,
178
+ "single_word": false,
179
+ "special": false
180
+ },
181
+ "151665": {
182
+ "content": "<tool_response>",
183
+ "lstrip": false,
184
+ "normalized": false,
185
+ "rstrip": false,
186
+ "single_word": false,
187
+ "special": false
188
+ },
189
+ "151666": {
190
+ "content": "</tool_response>",
191
+ "lstrip": false,
192
+ "normalized": false,
193
+ "rstrip": false,
194
+ "single_word": false,
195
+ "special": false
196
+ },
197
+ "151667": {
198
+ "content": "<think>",
199
+ "lstrip": false,
200
+ "normalized": false,
201
+ "rstrip": false,
202
+ "single_word": false,
203
+ "special": false
204
+ },
205
+ "151668": {
206
+ "content": "</think>",
207
+ "lstrip": false,
208
+ "normalized": false,
209
+ "rstrip": false,
210
+ "single_word": false,
211
+ "special": false
212
+ }
213
+ },
214
+ "additional_special_tokens": [
215
+ "<|im_start|>",
216
+ "<|im_end|>",
217
+ "<|object_ref_start|>",
218
+ "<|object_ref_end|>",
219
+ "<|box_start|>",
220
+ "<|box_end|>",
221
+ "<|quad_start|>",
222
+ "<|quad_end|>",
223
+ "<|vision_start|>",
224
+ "<|vision_end|>",
225
+ "<|vision_pad|>",
226
+ "<|image_pad|>",
227
+ "<|video_pad|>"
228
+ ],
229
+ "bos_token": null,
230
+ "clean_up_tokenization_spaces": false,
231
+ "eos_token": "<|endoftext|>",
232
+ "errors": "replace",
233
+ "extra_special_tokens": {},
234
+ "model_max_length": 131072,
235
+ "pad_token": "<|endoftext|>",
236
+ "sep_token": "<|endoftext|>",
237
+ "split_special_tokens": false,
238
+ "tokenizer_class": "Qwen2Tokenizer",
239
+ "unk_token": null
240
+ }
vocab.json ADDED
The diff for this file is too large to render. See raw diff