kimsan0622 commited on
Commit
50b17ad
·
verified ·
1 Parent(s): 9f97a34

Upload KETI Llama 7B v0.1

Browse files
.gitattributes CHANGED
@@ -33,3 +33,4 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
 
 
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
+ tokenizer.json filter=lfs diff=lfs merge=lfs -text
README.md ADDED
@@ -0,0 +1,127 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: other
3
+ language:
4
+ - en
5
+ - ko
6
+ library_name: transformers
7
+ pipeline_tag: text-generation
8
+ tags:
9
+ - llama
10
+ - causal-lm
11
+ - sft
12
+ - dpo
13
+ - grpo
14
+ - long-context
15
+ - keti
16
+ ---
17
+
18
+ # KETI Llama 7B v0.1
19
+
20
+ `KETI Llama 7B v0.1` is a long-context causal language model released under the
21
+ Hugging Face repository `KETI-AIR/keti-llama-7b-v0.1`.
22
+
23
+ This checkpoint was produced from the local base
24
+ model with the following training pipeline:
25
+
26
+ 1. SFT on instruction data with 32K sequence packing.
27
+ 2. DPO on preference data.
28
+ 3. RL/GRPO on RL data.
29
+
30
+ The uploaded artifact is the merged Hugging Face model from:
31
+
32
+ ```text
33
+ outputs/llama-8b-keti-dpo-rl-merged
34
+ ```
35
+
36
+ ## Model Details
37
+
38
+ - Architecture: `LlamaForCausalLM`
39
+ - Parameters: 8B-class
40
+ - Context length in config: 131,072 tokens
41
+ - Hidden size: 4096
42
+ - Layers: 32
43
+ - Attention heads: 32
44
+ - KV heads: 8
45
+ - Vocabulary size: 128,256
46
+ - Recommended dtype: `bfloat16`
47
+
48
+ ## Evaluation
49
+
50
+ Evaluation timestamp: `20260604_202553`
51
+
52
+ | Category | Dataset | Version | Metric | Mode | Score |
53
+ | --- | --- | --- | --- | --- | ---: |
54
+ | Core | core_average | - | naive_average | gen | 27.77 |
55
+ | Instruction Following | IFEval | 353ae7 | Prompt-level-strict-accuracy | gen | 50.65 |
56
+ | Math Calculation | aime2024 | bc6078 | accuracy | gen | 16.67 |
57
+ | Math Calculation | aime2025 | 5e9f4f | accuracy | gen | 3.33 |
58
+ | Math Calculation | math_prm800k_500 | 11c4b5 | accuracy | gen | 60.20 |
59
+ | General Reasoning | bbh | - | naive_average | gen | 11.87 |
60
+ | General Reasoning | GPQA_diamond | 5aeece | accuracy | gen | 20.71 |
61
+ | Knowledge | mmlu_pro | - | naive_average | gen | 28.26 |
62
+ | Code | openai_humaneval | dcae0e | humaneval_pass@1 | gen | 60.98 |
63
+ | Code | lcb_code_generation | b5b6c5 | pass@1 | gen | 6.00 |
64
+ | Long Context Reasoning | leval | - | naive_average | gen | 39.37 |
65
+ | Long Context Reasoning | longbench | - | naive_average | gen | 20.57 |
66
+ | Long Context Reasoning | LongBenchv2 | 75fbba | accuracy | gen | 24.85 |
67
+ | Long Context Reasoning | keti_long_ctx_gutenberg | - | naive_average | gen | 17.62 |
68
+
69
+
70
+ ## Quick Start
71
+
72
+ ```python
73
+ import torch
74
+ from transformers import AutoModelForCausalLM, AutoTokenizer
75
+
76
+ model_id = "KETI-AIR/keti-llama-7b-v0.1"
77
+
78
+ tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
79
+ model = AutoModelForCausalLM.from_pretrained(
80
+ model_id,
81
+ torch_dtype=torch.bfloat16,
82
+ device_map="auto",
83
+ trust_remote_code=True,
84
+ )
85
+
86
+ messages = [
87
+ {"role": "user", "content": "Explain why long-context reasoning is useful."}
88
+ ]
89
+ inputs = tokenizer.apply_chat_template(
90
+ messages,
91
+ add_generation_prompt=True,
92
+ return_tensors="pt",
93
+ ).to(model.device)
94
+
95
+ outputs = model.generate(
96
+ inputs,
97
+ max_new_tokens=512,
98
+ do_sample=True,
99
+ temperature=0.7,
100
+ top_p=0.9,
101
+ )
102
+ print(tokenizer.decode(outputs[0][inputs.shape[-1]:], skip_special_tokens=True))
103
+ ```
104
+
105
+ ## Intended Use
106
+
107
+ This model is intended for research and development on instruction following,
108
+ code generation, mathematical reasoning, and long-context generation tasks.
109
+
110
+ ## Limitations
111
+
112
+ The model can generate incorrect, unsafe, or biased content. Users should
113
+ evaluate the model for their own deployment setting and apply appropriate safety
114
+ filters and human review where needed.
115
+
116
+ ## Training Framework
117
+
118
+ - Transformers: 5.8.1
119
+ - PyTorch: 2.11.0+cu130
120
+ - Datasets: 4.8.5
121
+ - Tokenizers: 0.22.2
122
+ - TRL: 1.4.0
123
+
124
+ ## Citation
125
+
126
+ If you use this model, please cite the corresponding KETI-AIR release and the
127
+ training/evaluation resources used in your work.
chat_template.jinja ADDED
@@ -0,0 +1,250 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {{- bos_token if bos_token is defined else '' }}
2
+ {%- if custom_tools is defined and custom_tools is not none %}
3
+ {%- set tools = custom_tools %}
4
+ {%- endif %}
5
+ {%- if tools is not defined %}
6
+ {%- set tools = none %}
7
+ {%- endif %}
8
+ {%- if tools is not none and tools is iterable and tools is not mapping and tools|length == 0 %}
9
+ {%- set tools = none %}
10
+ {%- endif %}
11
+ {%- if tools_in_user_message is not defined %}
12
+ {#- Llama 3.1-style tool prompting is most stable when tools are injected into the first user turn. #}
13
+ {%- set tools_in_user_message = true %}
14
+ {%- endif %}
15
+ {%- if date_string is not defined %}
16
+ {%- if strftime_now is defined %}
17
+ {%- set date_string = strftime_now("%d %b %Y") %}
18
+ {%- else %}
19
+ {#- Avoid baking a stale Today Date into newer data. Pass date_string explicitly if needed. #}
20
+ {%- set date_string = none %}
21
+ {%- endif %}
22
+ {%- endif %}
23
+ {%- if cutoff_date_string is not defined %}
24
+ {%- set cutoff_date_string = "December 2023" %}
25
+ {%- endif %}
26
+ {%- if image_token is not defined %}
27
+ {%- set image_token = "<|image|>" %}
28
+ {%- endif %}
29
+ {%- if video_token is not defined %}
30
+ {%- set video_token = "<|video|>" %}
31
+ {%- endif %}
32
+
33
+ {#- Modern content renderer inspired by Qwen-style templates.
34
+ Supports:
35
+ - string content
36
+ - None / missing content
37
+ - list content blocks: {type: text, text: ...}, {type: image/image_url/input_image, ...}, {type: video/input_video, ...}
38
+ - direct {text: ...}, {image: ...}, {image_url: ...}, {video: ...}
39
+ #}
40
+ {%- set image_count = namespace(value=0) %}
41
+ {%- set video_count = namespace(value=0) %}
42
+
43
+ {%- macro render_content_item(item, do_vision_count, is_system_content=false) %}
44
+ {%- if item is string %}
45
+ {{- item }}
46
+ {%- elif item is mapping %}
47
+ {%- set item_type = item.get('type', '') %}
48
+ {%- if item_type in ['image', 'image_url', 'input_image'] or 'image' in item or 'image_url' in item %}
49
+ {%- if is_system_content %}
50
+ {{- raise_exception('System message cannot contain images.') }}
51
+ {%- endif %}
52
+ {%- if do_vision_count %}
53
+ {%- set image_count.value = image_count.value + 1 %}
54
+ {%- endif %}
55
+ {%- if add_vision_id is defined and add_vision_id %}
56
+ {{- 'Picture ' ~ image_count.value ~ ': ' }}
57
+ {%- endif %}
58
+ {{- image_token }}
59
+ {%- elif item_type in ['video', 'input_video'] or 'video' in item %}
60
+ {%- if is_system_content %}
61
+ {{- raise_exception('System message cannot contain videos.') }}
62
+ {%- endif %}
63
+ {%- if do_vision_count %}
64
+ {%- set video_count.value = video_count.value + 1 %}
65
+ {%- endif %}
66
+ {%- if add_vision_id is defined and add_vision_id %}
67
+ {{- 'Video ' ~ video_count.value ~ ': ' }}
68
+ {%- endif %}
69
+ {{- video_token }}
70
+ {%- elif 'text' in item %}
71
+ {{- item.text }}
72
+ {%- elif item_type in ['text', 'input_text'] and 'content' in item %}
73
+ {{- item.content }}
74
+ {%- elif 'content' in item %}
75
+ {{- render_content(item.content, do_vision_count, is_system_content) }}
76
+ {%- else %}
77
+ {{- raise_exception('Unexpected item type in content.') }}
78
+ {%- endif %}
79
+ {%- elif item is none or item is undefined %}
80
+ {{- '' }}
81
+ {%- else %}
82
+ {{- raise_exception('Unexpected item type in content.') }}
83
+ {%- endif %}
84
+ {%- endmacro %}
85
+
86
+ {%- macro render_content(content, do_vision_count, is_system_content=false) %}
87
+ {%- if content is string %}
88
+ {{- content }}
89
+ {%- elif content is none or content is undefined %}
90
+ {{- '' }}
91
+ {%- elif content is mapping %}
92
+ {{- render_content_item(content, do_vision_count, is_system_content) }}
93
+ {%- elif content is iterable %}
94
+ {%- for item in content %}
95
+ {{- render_content_item(item, do_vision_count, is_system_content) }}
96
+ {%- endfor %}
97
+ {%- else %}
98
+ {{- raise_exception('Unexpected content type.') }}
99
+ {%- endif %}
100
+ {%- endmacro %}
101
+
102
+ {%- macro render_tool_call_json(tool_call) %}
103
+ {%- if tool_call.function is defined %}
104
+ {%- set fn = tool_call.function %}
105
+ {%- else %}
106
+ {%- set fn = tool_call %}
107
+ {%- endif %}
108
+ {%- if fn.name is not defined %}
109
+ {{- raise_exception('Tool call is missing function name.') }}
110
+ {%- endif %}
111
+ {%- set args = fn.arguments if fn.arguments is defined and fn.arguments is not none else {} %}
112
+ {{- '{"name": ' }}{{- fn.name | tojson }}{{- ', "parameters": ' }}
113
+ {%- if args is string and args|trim %}
114
+ {{- args }}
115
+ {%- else %}
116
+ {{- args | tojson }}
117
+ {%- endif %}
118
+ {{- '}' }}
119
+ {%- endmacro %}
120
+
121
+ {%- macro render_tool_calls_json(tool_calls) %}
122
+ {%- if tool_calls|length == 1 %}
123
+ {{- render_tool_call_json(tool_calls[0]) }}
124
+ {%- else %}
125
+ {{- '[' }}
126
+ {%- for tool_call in tool_calls %}
127
+ {{- render_tool_call_json(tool_call) }}
128
+ {%- if not loop.last %}
129
+ {{- ', ' }}
130
+ {%- endif %}
131
+ {%- endfor %}
132
+ {{- ']' }}
133
+ {%- endif %}
134
+ {%- endmacro %}
135
+
136
+ {%- if not messages %}
137
+ {{- raise_exception('No messages provided.') }}
138
+ {%- endif %}
139
+
140
+ {#- Extract an initial system/developer message, so we can slot it into the Llama system header. #}
141
+ {%- if messages[0]['role'] == 'system' or messages[0]['role'] == 'developer' %}
142
+ {%- set system_message = render_content(messages[0]['content'], false, true)|trim %}
143
+ {%- set messages = messages[1:] %}
144
+ {%- else %}
145
+ {%- if tools is not none %}
146
+ {%- set system_message = "You are a helpful assistant with tool calling capabilities. Only reply with a tool call if the function exists in the library provided by the user. If it doesn't exist, just reply directly in natural language. When you receive a tool call response, use the output to format an answer to the original user question." %}
147
+ {%- else %}
148
+ {%- set system_message = "" %}
149
+ {%- endif %}
150
+ {%- endif %}
151
+
152
+ {#- System message #}
153
+ {{- "<|start_header_id|>system<|end_header_id|>\n\n" }}
154
+ {%- if tools is not none %}
155
+ {{- "Environment: ipython\n" }}
156
+ {%- endif %}
157
+ {%- if cutoff_date_string %}
158
+ {{- "Cutting Knowledge Date: " + cutoff_date_string + "\n" }}
159
+ {%- endif %}
160
+ {%- if date_string %}
161
+ {{- "Today Date: " + date_string + "\n" }}
162
+ {%- endif %}
163
+ {%- if tools is not none or cutoff_date_string or date_string %}
164
+ {{- "\n" }}
165
+ {%- endif %}
166
+ {%- if tools is not none and not tools_in_user_message %}
167
+ {{- "You have access to the following functions. To call a function, please respond with JSON for a function call. " }}
168
+ {{- 'For a single function call, respond in the format {"name": function name, "parameters": dictionary of argument name and its value}. ' }}
169
+ {{- 'If multiple function calls are required, respond with a JSON array of objects in that same format. ' }}
170
+ {{- "Do not use variables.\n\n" }}
171
+ {%- for t in tools %}
172
+ {{- t | tojson(indent=4) }}
173
+ {{- "\n\n" }}
174
+ {%- endfor %}
175
+ {%- endif %}
176
+ {{- system_message }}
177
+ {{- "<|eot_id|>" }}
178
+
179
+ {#- Custom tools can be passed in the first user message with extra guidance. #}
180
+ {%- if tools_in_user_message and tools is not none %}
181
+ {%- if messages | length != 0 %}
182
+ {%- set first_user_message = render_content(messages[0]['content'], true)|trim %}
183
+ {%- set messages = messages[1:] %}
184
+ {%- else %}
185
+ {{- raise_exception("Cannot put tools in the first user message when there's no first user message!") }}
186
+ {%- endif %}
187
+ {{- '<|start_header_id|>user<|end_header_id|>\n\n' -}}
188
+ {{- "Given the following functions, please respond with JSON for a function call " }}
189
+ {{- "with its proper arguments that best answers the given prompt.\n\n" }}
190
+ {{- 'For a single function call, respond in the format {"name": function name, "parameters": dictionary of argument name and its value}. ' }}
191
+ {{- 'If multiple function calls are required, respond with a JSON array of objects in that same format. ' }}
192
+ {{- "Do not use variables.\n\n" }}
193
+ {%- for t in tools %}
194
+ {{- t | tojson(indent=4) }}
195
+ {{- "\n\n" }}
196
+ {%- endfor %}
197
+ {{- first_user_message + "<|eot_id|>" }}
198
+ {%- endif %}
199
+
200
+ {%- for message in messages %}
201
+ {%- set role = message.role %}
202
+ {%- if role == 'system' or role == 'developer' %}
203
+ {{- '<|start_header_id|>system<|end_header_id|>\n\n' }}
204
+ {{- render_content(message.content, true, true)|trim }}
205
+ {{- '<|eot_id|>' }}
206
+ {%- elif role == 'user' %}
207
+ {{- '<|start_header_id|>user<|end_header_id|>\n\n' }}
208
+ {{- render_content(message.content, true)|trim }}
209
+ {{- '<|eot_id|>' }}
210
+ {%- elif role == 'assistant' %}
211
+ {%- set content = render_content(message.content, true)|trim %}
212
+ {%- set reasoning_content = '' %}
213
+ {%- if message.reasoning_content is defined and message.reasoning_content is string %}
214
+ {%- set reasoning_content = message.reasoning_content|trim %}
215
+ {%- elif '</think>' in content %}
216
+ {%- set reasoning_content = content.split('</think>')[0].rstrip('\n').split('<think>')[-1].lstrip('\n')|trim %}
217
+ {%- set content = content.split('</think>')[-1].lstrip('\n')|trim %}
218
+ {%- endif %}
219
+ {{- '<|start_header_id|>assistant<|end_header_id|>\n\n' }}
220
+ {%- if reasoning_content %}
221
+ {{- '<think>\n' + reasoning_content + '\n</think>\n\n' }}
222
+ {%- endif %}
223
+ {{- content }}
224
+ {%- if message.tool_calls is defined and message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
225
+ {%- if content or reasoning_content %}
226
+ {{- '\n\n' }}
227
+ {%- endif %}
228
+ {{- render_tool_calls_json(message.tool_calls) }}
229
+ {%- endif %}
230
+ {{- '<|eot_id|>' }}
231
+ {%- elif role == 'tool' or role == 'ipython' or role == 'function' %}
232
+ {{- '<|start_header_id|>ipython<|end_header_id|>\n\n' }}
233
+ {%- if message.content is mapping %}
234
+ {{- {"output": message.content} | tojson }}
235
+ {%- else %}
236
+ {%- set tool_content = render_content(message.content, true)|trim %}
237
+ {{- {"output": tool_content} | tojson }}
238
+ {%- endif %}
239
+ {{- '<|eot_id|>' }}
240
+ {%- else %}
241
+ {{- raise_exception('Unexpected message role: ' + role) }}
242
+ {%- endif %}
243
+ {%- endfor %}
244
+
245
+ {%- if add_generation_prompt is defined and add_generation_prompt %}
246
+ {{- '<|start_header_id|>assistant<|end_header_id|>\n\n' }}
247
+ {%- if enable_thinking is defined and enable_thinking %}
248
+ {{- '<think>\n' }}
249
+ {%- endif %}
250
+ {%- endif %}
config.json ADDED
@@ -0,0 +1,36 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "architectures": [
3
+ "LlamaForCausalLM"
4
+ ],
5
+ "attention_bias": false,
6
+ "attention_dropout": 0.0,
7
+ "bos_token_id": 128000,
8
+ "dtype": "float32",
9
+ "eos_token_id": 128001,
10
+ "head_dim": 128,
11
+ "hidden_act": "silu",
12
+ "hidden_size": 4096,
13
+ "initializer_range": 0.02,
14
+ "intermediate_size": 14336,
15
+ "max_position_embeddings": 131072,
16
+ "mlp_bias": false,
17
+ "model_type": "llama",
18
+ "num_attention_heads": 32,
19
+ "num_hidden_layers": 32,
20
+ "num_key_value_heads": 8,
21
+ "pad_token_id": 128001,
22
+ "pretraining_tp": 1,
23
+ "rms_norm_eps": 1e-05,
24
+ "rope_parameters": {
25
+ "factor": 8.0,
26
+ "high_freq_factor": 4.0,
27
+ "low_freq_factor": 1.0,
28
+ "original_max_position_embeddings": 8192,
29
+ "rope_theta": 500000.0,
30
+ "rope_type": "llama3"
31
+ },
32
+ "tie_word_embeddings": false,
33
+ "transformers_version": "5.8.1",
34
+ "use_cache": false,
35
+ "vocab_size": 128256
36
+ }
generation_config.json ADDED
@@ -0,0 +1,12 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "_from_model_config": true,
3
+ "bos_token_id": 128000,
4
+ "eos_token_id": [
5
+ 128001
6
+ ],
7
+ "output_attentions": false,
8
+ "output_hidden_states": false,
9
+ "pad_token_id": 128001,
10
+ "transformers_version": "5.8.1",
11
+ "use_cache": true
12
+ }
model.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:ca40dc6a43271c21057f48b1a49d0c2adaf74022584a49d9683bde34b0460d2c
3
+ size 32121079032
tokenizer.json ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:6b9e4e7fb171f92fd137b777cc2714bf87d11576700a1dcd7a399e7bbe39537b
3
+ size 17209920
tokenizer_config.json ADDED
@@ -0,0 +1,15 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "backend": "tokenizers",
3
+ "bos_token": "<|begin_of_text|>",
4
+ "clean_up_tokenization_spaces": true,
5
+ "eos_token": "<|end_of_text|>",
6
+ "is_local": true,
7
+ "local_files_only": false,
8
+ "model_input_names": [
9
+ "input_ids",
10
+ "attention_mask"
11
+ ],
12
+ "model_max_length": 131072,
13
+ "pad_token": "<|end_of_text|>",
14
+ "tokenizer_class": "TokenizersBackend"
15
+ }