bhxdianzhang commited on
Commit
ea420d1
·
verified ·
1 Parent(s): 9576927

Upload ParaDesigner-SFT model card and final weights

Browse files
.gitattributes CHANGED
@@ -33,3 +33,4 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
 
 
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
+ tokenizer.json filter=lfs diff=lfs merge=lfs -text
CITATION.cff ADDED
@@ -0,0 +1,16 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ cff-version: 1.2.0
2
+ title: "ParaDesigner-SFT: A ParadoxGPT Designer Specialist Model"
3
+ message: "If you use or reference this private model, please cite it as below."
4
+ type: software
5
+ authors:
6
+ - family-names: Zhang
7
+ given-names: Heng
8
+ year: 2026
9
+ repository-code: "https://huggingface.co/bhxdianzhang/ParaDesigner-SFT"
10
+ abstract: "A private ParadoxGPT supervised fine-tuned Designer specialist model for scientific-paper workflows."
11
+ keywords:
12
+ - ParadoxGPT
13
+ - Designer
14
+ - supervised fine-tuning
15
+ - scientific writing
16
+ license: other
README.md ADDED
@@ -0,0 +1,220 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: other
3
+ library_name: transformers
4
+ base_model: Qwen/Qwen3.5-4B
5
+ language:
6
+ - zh
7
+ - en
8
+ tags:
9
+ - paradoxgpt
10
+ - designer
11
+ - sft
12
+ - supervised-fine-tuning
13
+ - scientific-writing
14
+ - research-agents
15
+ - qwen
16
+ pipeline_tag: text-generation
17
+ model-index:
18
+ - name: ParaDesigner-SFT
19
+ results:
20
+ - task:
21
+ type: text-generation
22
+ name: Designer specialist modeling
23
+ dataset:
24
+ name: ParaSFT-designer
25
+ type: bhxdianzhang/ParaSFT-designer
26
+ metrics:
27
+ - type: loss
28
+ value: 0.9128507972
29
+ name: eval_loss
30
+ ---
31
+
32
+ # ParaDesigner-SFT
33
+
34
+ English | [中文](#中文)
35
+
36
+ ## Overview
37
+
38
+ ParaDesigner-SFT is a private ParadoxGPT Designer specialist model fine-tuned from a Qwen 4B base model. It is designed for research-paper workflows rather than general chat.
39
+
40
+ ## Intended Use
41
+
42
+ Use ParaDesigner-SFT for:
43
+
44
+ - paper-level scientific writing and review-support workflows;
45
+ - converting paper context into structured reasoning traces;
46
+ - internal ParadoxGPT research-agent training and evaluation;
47
+ - assisting human researchers with evidence, argumentation, and risk analysis.
48
+
49
+ Outputs should be checked against the original paper context. This model is not a substitute for human scientific judgment.
50
+
51
+ ## Training Data
52
+
53
+ The model was supervised-fine-tuned on private dataset `bhxdianzhang/ParaSFT-designer`.
54
+
55
+ | Split | Samples |
56
+ |---|---:|
57
+ | Train | 35,389 |
58
+ | Dev | 2,062 |
59
+ | Test | 2,296 |
60
+ | Total | 39,747 |
61
+
62
+ Task distribution:
63
+
64
+ | Task type | Samples |
65
+ |---|---:|
66
+ | `D1_claim_to_evidence` | 7,954 |
67
+ | `D2_experiment_argument_plan` | 7,951 |
68
+ | `D3_ablation_analysis_design` | 7,942 |
69
+ | `D4_sufficiency_critique` | 7,944 |
70
+ | `D5_interpretation_boundary` | 7,956 |
71
+
72
+ Quality filter: each target answer contains exactly one balanced `<think>...</think>` reasoning block followed by the final answer.
73
+
74
+ ## Training Summary
75
+
76
+ Base model: Qwen 4B local base model.
77
+
78
+ | Field | Value |
79
+ |---|---:|
80
+ | Epochs | 3.0 |
81
+ | Learning rate | 3e-6 |
82
+ | Scheduler | cosine |
83
+ | Distributed devices | 2 |
84
+ | Gradient accumulation | 8 |
85
+ | Total train batch size | 16 |
86
+ | Final train loss | 0.8653310693 |
87
+ | Final eval loss | 0.9128507972 |
88
+
89
+ ## Qualitative Behavior
90
+
91
+ Local side-by-side checks show strong task alignment across D1-D5. Compared with the same 4B backbone before SFT, ParaDesigner-SFT is much more concise, more Chinese-final-answer focused, and better at turning claims into evidence plans, experiment argumentation, ablation/mechanism needs, sufficiency critiques, and interpretation boundaries.
92
+
93
+ ParaDesigner-SFT is best used with rich paper context matching the training shape, such as title, abstract, introduction, selected body sections, figure/table captions, claims, reviews, or task-specific paper context depending on the skill.
94
+
95
+ ### Example Prompt Shape
96
+
97
+ ```text
98
+ 你是顶会论文分析专家。给定论文上下文,请按指定任务做中文分析。
99
+
100
+ 论文标题: ...
101
+
102
+ 摘要: ...
103
+
104
+ 引言: ...
105
+
106
+ 正文关键片段: ...
107
+
108
+ 图表/实验/claim 信息: ...
109
+
110
+ 任务: ...
111
+ ```
112
+
113
+ ## Limitations
114
+
115
+ - The model is optimized for a narrow ParadoxGPT specialist workflow, not general chat.
116
+ - It may inherit teacher-model annotation errors from the SFT data.
117
+ - It should not be used as an authority on paper correctness without source verification.
118
+
119
+
120
+ ## License and Use Restrictions
121
+
122
+ This repository is marked with `license: other`.
123
+
124
+ The model was trained on private ParadoxGPT SFT data derived from parsed academic papers, reviews, and teacher annotations. Rights to original papers and reviews remain with their respective authors, reviewers, venues, and publishers. This private model is provided for internal research and engineering use only. Redistribution or public release should be reviewed separately against source venue policies and applicable copyright rules.
125
+
126
+ ## Citation
127
+
128
+ ```bibtex
129
+ @model{zhang2026paradesignersft,
130
+ title = {ParaDesigner-SFT: A ParadoxGPT Designer Specialist Model},
131
+ author = {Heng Zhang},
132
+ year = {2026},
133
+ publisher = {Hugging Face},
134
+ howpublished = {https://huggingface.co/bhxdianzhang/ParaDesigner-SFT},
135
+ note = {Private ParadoxGPT supervised fine-tuned Designer model}
136
+ }
137
+ ```
138
+
139
+ ---
140
+
141
+ # 中文
142
+
143
+ ## 概述
144
+
145
+ ParaDesigner-SFT 是 ParadoxGPT 的 Designer 专家模型,由 Qwen 4B 基座监督微调而来。它面向科研论文工作流,不是通用聊天模型。
146
+
147
+ ## 适用场景
148
+
149
+ ParaDesigner-SFT 适合用于:
150
+
151
+ - 论文级科研写作与 review-support 工作流;
152
+ - 将论文上下文转成结构化 reasoning trace;
153
+ - ParadoxGPT 内部 research-agent 训练与评测;
154
+ - 辅助研究者做证据、论证和风险分析。
155
+
156
+ 输出应结合原论文上下文复核,不能替代人类科研判断。
157
+
158
+ ## 训练数据
159
+
160
+ 模型使用私有数据集 `bhxdianzhang/ParaSFT-designer` 进行监督微调。
161
+
162
+ | Split | 样本数 |
163
+ |---|---:|
164
+ | Train | 35,389 |
165
+ | Dev | 2,062 |
166
+ | Test | 2,296 |
167
+ | Total | 39,747 |
168
+
169
+ 任务分布:
170
+
171
+ | Task type | Samples |
172
+ |---|---:|
173
+ | `D1_claim_to_evidence` | 7,954 |
174
+ | `D2_experiment_argument_plan` | 7,951 |
175
+ | `D3_ablation_analysis_design` | 7,942 |
176
+ | `D4_sufficiency_critique` | 7,944 |
177
+ | `D5_interpretation_boundary` | 7,956 |
178
+
179
+ 质量过滤:每条目标答案都包含且只包含一个配平的 `<think>...</think>` 思考块,后接最终答案。
180
+
181
+ ## 训练摘要
182
+
183
+ 基座模型:本地 Qwen 4B base model。
184
+
185
+ | 字段 | 数值 |
186
+ |---|---:|
187
+ | Epochs | 3.0 |
188
+ | Learning rate | 3e-6 |
189
+ | Scheduler | cosine |
190
+ | 分布式设备数 | 2 |
191
+ | Gradient accumulation | 8 |
192
+ | Total train batch size | 16 |
193
+ | Final train loss | 0.8653310693 |
194
+ | Final eval loss | 0.9128507972 |
195
+
196
+ ## 定性行为
197
+
198
+ 本地 side-by-side 检查显示,ParaDesigner-SFT 在 D1-D5 上任务对齐稳定。相比同一 4B backbone 的 SFT 前模型,它更收敛、更聚焦中文最终答案,也更擅长把 claim 转成证据规划、实验论证、消融/机制分析需求、充分性批判和解释边界。
199
+
200
+ 为了获得最佳效果,请使用与训练数据一致的丰富论文上下文,例如 title、abstract、introduction、正文关键片段、figure/table captions、claims、reviews 或具体任务所需的 paper context。
201
+
202
+ ## 局限
203
+
204
+ - 模型针对 ParadoxGPT 窄域专家工作流优化,不是通用聊天模型。
205
+ - 模型可能继承 SFT 数据中的 teacher-model 标注误差。
206
+ - 不应在未检查原论文上下文的情况下,把输出当成科研事实。
207
+
208
+
209
+ ## 引用
210
+
211
+ ```bibtex
212
+ @model{zhang2026paradesignersft,
213
+ title = {ParaDesigner-SFT: A ParadoxGPT Designer Specialist Model},
214
+ author = {Heng Zhang},
215
+ year = {2026},
216
+ publisher = {Hugging Face},
217
+ howpublished = {https://huggingface.co/bhxdianzhang/ParaDesigner-SFT},
218
+ note = {Private ParadoxGPT supervised fine-tuned Designer model}
219
+ }
220
+ ```
all_results.json ADDED
@@ -0,0 +1,12 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "epoch": 3.0,
3
+ "eval_loss": 0.9128507971763611,
4
+ "eval_runtime": 448.7493,
5
+ "eval_samples_per_second": 4.595,
6
+ "eval_steps_per_second": 2.297,
7
+ "total_flos": 1.7279135623491355e+19,
8
+ "train_loss": 0.8653310692547458,
9
+ "train_runtime": 153956.3557,
10
+ "train_samples_per_second": 0.69,
11
+ "train_steps_per_second": 0.043
12
+ }
chat_template.jinja ADDED
@@ -0,0 +1,154 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {%- set image_count = namespace(value=0) %}
2
+ {%- set video_count = namespace(value=0) %}
3
+ {%- macro render_content(content, do_vision_count, is_system_content=false) %}
4
+ {%- if content is string %}
5
+ {{- content }}
6
+ {%- elif content is iterable and content is not mapping %}
7
+ {%- for item in content %}
8
+ {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
9
+ {%- if is_system_content %}
10
+ {{- raise_exception('System message cannot contain images.') }}
11
+ {%- endif %}
12
+ {%- if do_vision_count %}
13
+ {%- set image_count.value = image_count.value + 1 %}
14
+ {%- endif %}
15
+ {%- if add_vision_id %}
16
+ {{- 'Picture ' ~ image_count.value ~ ': ' }}
17
+ {%- endif %}
18
+ {{- '<|vision_start|><|image_pad|><|vision_end|>' }}
19
+ {%- elif 'video' in item or item.type == 'video' %}
20
+ {%- if is_system_content %}
21
+ {{- raise_exception('System message cannot contain videos.') }}
22
+ {%- endif %}
23
+ {%- if do_vision_count %}
24
+ {%- set video_count.value = video_count.value + 1 %}
25
+ {%- endif %}
26
+ {%- if add_vision_id %}
27
+ {{- 'Video ' ~ video_count.value ~ ': ' }}
28
+ {%- endif %}
29
+ {{- '<|vision_start|><|video_pad|><|vision_end|>' }}
30
+ {%- elif 'text' in item %}
31
+ {{- item.text }}
32
+ {%- else %}
33
+ {{- raise_exception('Unexpected item type in content.') }}
34
+ {%- endif %}
35
+ {%- endfor %}
36
+ {%- elif content is none or content is undefined %}
37
+ {{- '' }}
38
+ {%- else %}
39
+ {{- raise_exception('Unexpected content type.') }}
40
+ {%- endif %}
41
+ {%- endmacro %}
42
+ {%- if not messages %}
43
+ {{- raise_exception('No messages provided.') }}
44
+ {%- endif %}
45
+ {%- if tools and tools is iterable and tools is not mapping %}
46
+ {{- '<|im_start|>system\n' }}
47
+ {{- "# Tools\n\nYou have access to the following functions:\n\n<tools>" }}
48
+ {%- for tool in tools %}
49
+ {{- "\n" }}
50
+ {{- tool | tojson }}
51
+ {%- endfor %}
52
+ {{- "\n</tools>" }}
53
+ {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n<tool_call>\n<function=example_function_name>\n<parameter=example_parameter_1>\nvalue_1\n</parameter>\n<parameter=example_parameter_2>\nThis is the value for the second parameter\nthat can span\nmultiple lines\n</parameter>\n</function>\n</tool_call>\n\n<IMPORTANT>\nReminder:\n- Function calls MUST follow the specified format: an inner <function=...></function> block must be nested within <tool_call></tool_call> XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n</IMPORTANT>' }}
54
+ {%- if messages[0].role == 'system' %}
55
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
56
+ {%- if content %}
57
+ {{- '\n\n' + content }}
58
+ {%- endif %}
59
+ {%- endif %}
60
+ {{- '<|im_end|>\n' }}
61
+ {%- else %}
62
+ {%- if messages[0].role == 'system' %}
63
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
64
+ {{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
65
+ {%- endif %}
66
+ {%- endif %}
67
+ {%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
68
+ {%- for message in messages[::-1] %}
69
+ {%- set index = (messages|length - 1) - loop.index0 %}
70
+ {%- if ns.multi_step_tool and message.role == "user" %}
71
+ {%- set content = render_content(message.content, false)|trim %}
72
+ {%- if not(content.startswith('<tool_response>') and content.endswith('</tool_response>')) %}
73
+ {%- set ns.multi_step_tool = false %}
74
+ {%- set ns.last_query_index = index %}
75
+ {%- endif %}
76
+ {%- endif %}
77
+ {%- endfor %}
78
+ {%- if ns.multi_step_tool %}
79
+ {{- raise_exception('No user query found in messages.') }}
80
+ {%- endif %}
81
+ {%- for message in messages %}
82
+ {%- set content = render_content(message.content, true)|trim %}
83
+ {%- if message.role == "system" %}
84
+ {%- if not loop.first %}
85
+ {{- raise_exception('System message must be at the beginning.') }}
86
+ {%- endif %}
87
+ {%- elif message.role == "user" %}
88
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
89
+ {%- elif message.role == "assistant" %}
90
+ {%- set reasoning_content = '' %}
91
+ {%- if message.reasoning_content is string %}
92
+ {%- set reasoning_content = message.reasoning_content %}
93
+ {%- else %}
94
+ {%- if '</think>' in content %}
95
+ {%- set reasoning_content = content.split('</think>')[0].rstrip('\n').split('<think>')[-1].lstrip('\n') %}
96
+ {%- set content = content.split('</think>')[-1].lstrip('\n') %}
97
+ {%- endif %}
98
+ {%- endif %}
99
+ {%- set reasoning_content = reasoning_content|trim %}
100
+ {%- if loop.index0 > ns.last_query_index %}
101
+ {{- '<|im_start|>' + message.role + '\n<think>\n' + reasoning_content + '\n</think>\n\n' + content }}
102
+ {%- else %}
103
+ {{- '<|im_start|>' + message.role + '\n' + content }}
104
+ {%- endif %}
105
+ {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
106
+ {%- for tool_call in message.tool_calls %}
107
+ {%- if tool_call.function is defined %}
108
+ {%- set tool_call = tool_call.function %}
109
+ {%- endif %}
110
+ {%- if loop.first %}
111
+ {%- if content|trim %}
112
+ {{- '\n\n<tool_call>\n<function=' + tool_call.name + '>\n' }}
113
+ {%- else %}
114
+ {{- '<tool_call>\n<function=' + tool_call.name + '>\n' }}
115
+ {%- endif %}
116
+ {%- else %}
117
+ {{- '\n<tool_call>\n<function=' + tool_call.name + '>\n' }}
118
+ {%- endif %}
119
+ {%- if tool_call.arguments is defined %}
120
+ {%- for args_name, args_value in tool_call.arguments|items %}
121
+ {{- '<parameter=' + args_name + '>\n' }}
122
+ {%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}
123
+ {{- args_value }}
124
+ {{- '\n</parameter>\n' }}
125
+ {%- endfor %}
126
+ {%- endif %}
127
+ {{- '</function>\n</tool_call>' }}
128
+ {%- endfor %}
129
+ {%- endif %}
130
+ {{- '<|im_end|>\n' }}
131
+ {%- elif message.role == "tool" %}
132
+ {%- if loop.previtem and loop.previtem.role != "tool" %}
133
+ {{- '<|im_start|>user' }}
134
+ {%- endif %}
135
+ {{- '\n<tool_response>\n' }}
136
+ {{- content }}
137
+ {{- '\n</tool_response>' }}
138
+ {%- if not loop.last and loop.nextitem.role != "tool" %}
139
+ {{- '<|im_end|>\n' }}
140
+ {%- elif loop.last %}
141
+ {{- '<|im_end|>\n' }}
142
+ {%- endif %}
143
+ {%- else %}
144
+ {{- raise_exception('Unexpected message role.') }}
145
+ {%- endif %}
146
+ {%- endfor %}
147
+ {%- if add_generation_prompt %}
148
+ {{- '<|im_start|>assistant\n' }}
149
+ {%- if enable_thinking is defined and enable_thinking is false %}
150
+ {{- '<think>\n\n</think>\n\n' }}
151
+ {%- else %}
152
+ {{- '<think>\n' }}
153
+ {%- endif %}
154
+ {%- endif %}
config.json ADDED
@@ -0,0 +1,113 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "architectures": [
3
+ "Qwen3_5ForConditionalGeneration"
4
+ ],
5
+ "dtype": "bfloat16",
6
+ "eos_token_id": 248046,
7
+ "hidden_size": 2560,
8
+ "image_token_id": 248056,
9
+ "model_type": "qwen3_5",
10
+ "pad_token_id": 248044,
11
+ "text_config": {
12
+ "attention_bias": false,
13
+ "attention_dropout": 0.0,
14
+ "attn_output_gate": true,
15
+ "bos_token_id": null,
16
+ "dtype": "bfloat16",
17
+ "eos_token_id": 248044,
18
+ "full_attention_interval": 4,
19
+ "head_dim": 256,
20
+ "hidden_act": "silu",
21
+ "hidden_size": 2560,
22
+ "initializer_range": 0.02,
23
+ "intermediate_size": 9216,
24
+ "layer_types": [
25
+ "linear_attention",
26
+ "linear_attention",
27
+ "linear_attention",
28
+ "full_attention",
29
+ "linear_attention",
30
+ "linear_attention",
31
+ "linear_attention",
32
+ "full_attention",
33
+ "linear_attention",
34
+ "linear_attention",
35
+ "linear_attention",
36
+ "full_attention",
37
+ "linear_attention",
38
+ "linear_attention",
39
+ "linear_attention",
40
+ "full_attention",
41
+ "linear_attention",
42
+ "linear_attention",
43
+ "linear_attention",
44
+ "full_attention",
45
+ "linear_attention",
46
+ "linear_attention",
47
+ "linear_attention",
48
+ "full_attention",
49
+ "linear_attention",
50
+ "linear_attention",
51
+ "linear_attention",
52
+ "full_attention",
53
+ "linear_attention",
54
+ "linear_attention",
55
+ "linear_attention",
56
+ "full_attention"
57
+ ],
58
+ "linear_conv_kernel_dim": 4,
59
+ "linear_key_head_dim": 128,
60
+ "linear_num_key_heads": 16,
61
+ "linear_num_value_heads": 32,
62
+ "linear_value_head_dim": 128,
63
+ "mamba_ssm_dtype": "float32",
64
+ "max_position_embeddings": 262144,
65
+ "mlp_only_layers": [],
66
+ "model_type": "qwen3_5_text",
67
+ "mtp_num_hidden_layers": 1,
68
+ "mtp_use_dedicated_embeddings": false,
69
+ "num_attention_heads": 16,
70
+ "num_hidden_layers": 32,
71
+ "num_key_value_heads": 4,
72
+ "pad_token_id": null,
73
+ "partial_rotary_factor": 0.25,
74
+ "rms_norm_eps": 1e-06,
75
+ "rope_parameters": {
76
+ "mrope_interleaved": true,
77
+ "mrope_section": [
78
+ 11,
79
+ 11,
80
+ 10
81
+ ],
82
+ "partial_rotary_factor": 0.25,
83
+ "rope_theta": 10000000,
84
+ "rope_type": "default"
85
+ },
86
+ "tie_word_embeddings": true,
87
+ "use_cache": false,
88
+ "vocab_size": 248320
89
+ },
90
+ "tie_word_embeddings": true,
91
+ "transformers_version": "5.6.0",
92
+ "use_cache": false,
93
+ "video_token_id": 248057,
94
+ "vision_config": {
95
+ "deepstack_visual_indexes": [],
96
+ "depth": 24,
97
+ "dtype": "bfloat16",
98
+ "hidden_act": "gelu_pytorch_tanh",
99
+ "hidden_size": 1024,
100
+ "in_channels": 3,
101
+ "initializer_range": 0.02,
102
+ "intermediate_size": 4096,
103
+ "model_type": "qwen3_5_vision",
104
+ "num_heads": 16,
105
+ "num_position_embeddings": 2304,
106
+ "out_hidden_size": 2560,
107
+ "patch_size": 16,
108
+ "spatial_merge_size": 2,
109
+ "temporal_patch_size": 2
110
+ },
111
+ "vision_end_token_id": 248054,
112
+ "vision_start_token_id": 248053
113
+ }
eval_results.json ADDED
@@ -0,0 +1,7 @@
 
 
 
 
 
 
 
 
1
+ {
2
+ "epoch": 3.0,
3
+ "eval_loss": 0.9128507971763611,
4
+ "eval_runtime": 448.7493,
5
+ "eval_samples_per_second": 4.595,
6
+ "eval_steps_per_second": 2.297
7
+ }
generation_config.json ADDED
@@ -0,0 +1,10 @@
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "_from_model_config": true,
3
+ "eos_token_id": [
4
+ 248046,
5
+ 248044
6
+ ],
7
+ "pad_token_id": 248044,
8
+ "transformers_version": "5.6.0",
9
+ "use_cache": true
10
+ }
model.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:133cd343799b31c6d8ad97e710ecdfa43ac4fdc863ffda3da024626e857920c1
3
+ size 10350019328
processor_config.json ADDED
@@ -0,0 +1,60 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "image_processor": {
3
+ "do_convert_rgb": true,
4
+ "do_normalize": true,
5
+ "do_rescale": true,
6
+ "do_resize": true,
7
+ "image_mean": [
8
+ 0.5,
9
+ 0.5,
10
+ 0.5
11
+ ],
12
+ "image_processor_type": "Qwen2VLImageProcessor",
13
+ "image_std": [
14
+ 0.5,
15
+ 0.5,
16
+ 0.5
17
+ ],
18
+ "merge_size": 2,
19
+ "patch_size": 16,
20
+ "resample": 3,
21
+ "rescale_factor": 0.00392156862745098,
22
+ "size": {
23
+ "longest_edge": 16777216,
24
+ "shortest_edge": 65536
25
+ },
26
+ "temporal_patch_size": 2
27
+ },
28
+ "processor_class": "Qwen3VLProcessor",
29
+ "video_processor": {
30
+ "do_convert_rgb": true,
31
+ "do_normalize": true,
32
+ "do_rescale": true,
33
+ "do_resize": true,
34
+ "do_sample_frames": true,
35
+ "fps": 2,
36
+ "image_mean": [
37
+ 0.5,
38
+ 0.5,
39
+ 0.5
40
+ ],
41
+ "image_std": [
42
+ 0.5,
43
+ 0.5,
44
+ 0.5
45
+ ],
46
+ "max_frames": 768,
47
+ "merge_size": 2,
48
+ "min_frames": 4,
49
+ "patch_size": 16,
50
+ "resample": 3,
51
+ "rescale_factor": 0.00392156862745098,
52
+ "return_metadata": false,
53
+ "size": {
54
+ "longest_edge": 25165824,
55
+ "shortest_edge": 4096
56
+ },
57
+ "temporal_patch_size": 2,
58
+ "video_processor_type": "Qwen3VLVideoProcessor"
59
+ }
60
+ }
tokenizer.json ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:06b9509352d2af50381ab2247e083b80d32d5c0aba91c272ca9ff729b6a0e523
3
+ size 19989325
tokenizer_config.json ADDED
@@ -0,0 +1,34 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "add_prefix_space": false,
3
+ "audio_bos_token": "<|audio_start|>",
4
+ "audio_eos_token": "<|audio_end|>",
5
+ "audio_token": "<|audio_pad|>",
6
+ "backend": "tokenizers",
7
+ "bos_token": null,
8
+ "clean_up_tokenization_spaces": false,
9
+ "eos_token": "<|im_end|>",
10
+ "errors": "replace",
11
+ "image_token": "<|image_pad|>",
12
+ "is_local": true,
13
+ "local_files_only": false,
14
+ "model_max_length": 262144,
15
+ "model_specific_special_tokens": {
16
+ "audio_bos_token": "<|audio_start|>",
17
+ "audio_eos_token": "<|audio_end|>",
18
+ "audio_token": "<|audio_pad|>",
19
+ "image_token": "<|image_pad|>",
20
+ "video_token": "<|video_pad|>",
21
+ "vision_bos_token": "<|vision_start|>",
22
+ "vision_eos_token": "<|vision_end|>"
23
+ },
24
+ "pad_token": "<|endoftext|>",
25
+ "padding_side": "right",
26
+ "pretokenize_regex": "(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\\r\\n\\p{L}\\p{N}]?[\\p{L}\\p{M}]+|\\p{N}| ?[^\\s\\p{L}\\p{M}\\p{N}]+[\\r\\n]*|\\s*[\\r\\n]+|\\s+(?!\\S)|\\s+",
27
+ "processor_class": "Qwen3VLProcessor",
28
+ "split_special_tokens": false,
29
+ "tokenizer_class": "Qwen2Tokenizer",
30
+ "unk_token": null,
31
+ "video_token": "<|video_pad|>",
32
+ "vision_bos_token": "<|vision_start|>",
33
+ "vision_eos_token": "<|vision_end|>"
34
+ }
train_results.json ADDED
@@ -0,0 +1,8 @@
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "epoch": 3.0,
3
+ "total_flos": 1.7279135623491355e+19,
4
+ "train_loss": 0.8653310692547458,
5
+ "train_runtime": 153956.3557,
6
+ "train_samples_per_second": 0.69,
7
+ "train_steps_per_second": 0.043
8
+ }
trainer_state.json ADDED
The diff for this file is too large to render. See raw diff