ramgpt commited on
Commit
ff690e8
·
verified ·
1 Parent(s): d000d60

Add files using upload-large-folder tool

Browse files
.gitattributes CHANGED
@@ -33,3 +33,4 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
 
 
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
+ tokenizer.json filter=lfs diff=lfs merge=lfs -text
LICENSE ADDED
@@ -0,0 +1,202 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+
2
+ Apache License
3
+ Version 2.0, January 2004
4
+ http://www.apache.org/licenses/
5
+
6
+ TERMS AND CONDITIONS FOR USE, REPRODUCTION, AND DISTRIBUTION
7
+
8
+ 1. Definitions.
9
+
10
+ "License" shall mean the terms and conditions for use, reproduction,
11
+ and distribution as defined by Sections 1 through 9 of this document.
12
+
13
+ "Licensor" shall mean the copyright owner or entity authorized by
14
+ the copyright owner that is granting the License.
15
+
16
+ "Legal Entity" shall mean the union of the acting entity and all
17
+ other entities that control, are controlled by, or are under common
18
+ control with that entity. For the purposes of this definition,
19
+ "control" means (i) the power, direct or indirect, to cause the
20
+ direction or management of such entity, whether by contract or
21
+ otherwise, or (ii) ownership of fifty percent (50%) or more of the
22
+ outstanding shares, or (iii) beneficial ownership of such entity.
23
+
24
+ "You" (or "Your") shall mean an individual or Legal Entity
25
+ exercising permissions granted by this License.
26
+
27
+ "Source" form shall mean the preferred form for making modifications,
28
+ including but not limited to software source code, documentation
29
+ source, and configuration files.
30
+
31
+ "Object" form shall mean any form resulting from mechanical
32
+ transformation or translation of a Source form, including but
33
+ not limited to compiled object code, generated documentation,
34
+ and conversions to other media types.
35
+
36
+ "Work" shall mean the work of authorship, whether in Source or
37
+ Object form, made available under the License, as indicated by a
38
+ copyright notice that is included in or attached to the work
39
+ (an example is provided in the Appendix below).
40
+
41
+ "Derivative Works" shall mean any work, whether in Source or Object
42
+ form, that is based on (or derived from) the Work and for which the
43
+ editorial revisions, annotations, elaborations, or other modifications
44
+ represent, as a whole, an original work of authorship. For the purposes
45
+ of this License, Derivative Works shall not include works that remain
46
+ separable from, or merely link (or bind by name) to the interfaces of,
47
+ the Work and Derivative Works thereof.
48
+
49
+ "Contribution" shall mean any work of authorship, including
50
+ the original version of the Work and any modifications or additions
51
+ to that Work or Derivative Works thereof, that is intentionally
52
+ submitted to Licensor for inclusion in the Work by the copyright owner
53
+ or by an individual or Legal Entity authorized to submit on behalf of
54
+ the copyright owner. For the purposes of this definition, "submitted"
55
+ means any form of electronic, verbal, or written communication sent
56
+ to the Licensor or its representatives, including but not limited to
57
+ communication on electronic mailing lists, source code control systems,
58
+ and issue tracking systems that are managed by, or on behalf of, the
59
+ Licensor for the purpose of discussing and improving the Work, but
60
+ excluding communication that is conspicuously marked or otherwise
61
+ designated in writing by the copyright owner as "Not a Contribution."
62
+
63
+ "Contributor" shall mean Licensor and any individual or Legal Entity
64
+ on behalf of whom a Contribution has been received by Licensor and
65
+ subsequently incorporated within the Work.
66
+
67
+ 2. Grant of Copyright License. Subject to the terms and conditions of
68
+ this License, each Contributor hereby grants to You a perpetual,
69
+ worldwide, non-exclusive, no-charge, royalty-free, irrevocable
70
+ copyright license to reproduce, prepare Derivative Works of,
71
+ publicly display, publicly perform, sublicense, and distribute the
72
+ Work and such Derivative Works in Source or Object form.
73
+
74
+ 3. Grant of Patent License. Subject to the terms and conditions of
75
+ this License, each Contributor hereby grants to You a perpetual,
76
+ worldwide, non-exclusive, no-charge, royalty-free, irrevocable
77
+ (except as stated in this section) patent license to make, have made,
78
+ use, offer to sell, sell, import, and otherwise transfer the Work,
79
+ where such license applies only to those patent claims licensable
80
+ by such Contributor that are necessarily infringed by their
81
+ Contribution(s) alone or by combination of their Contribution(s)
82
+ with the Work to which such Contribution(s) was submitted. If You
83
+ institute patent litigation against any entity (including a
84
+ cross-claim or counterclaim in a lawsuit) alleging that the Work
85
+ or a Contribution incorporated within the Work constitutes direct
86
+ or contributory patent infringement, then any patent licenses
87
+ granted to You under this License for that Work shall terminate
88
+ as of the date such litigation is filed.
89
+
90
+ 4. Redistribution. You may reproduce and distribute copies of the
91
+ Work or Derivative Works thereof in any medium, with or without
92
+ modifications, and in Source or Object form, provided that You
93
+ meet the following conditions:
94
+
95
+ (a) You must give any other recipients of the Work or
96
+ Derivative Works a copy of this License; and
97
+
98
+ (b) You must cause any modified files to carry prominent notices
99
+ stating that You changed the files; and
100
+
101
+ (c) You must retain, in the Source form of any Derivative Works
102
+ that You distribute, all copyright, patent, trademark, and
103
+ attribution notices from the Source form of the Work,
104
+ excluding those notices that do not pertain to any part of
105
+ the Derivative Works; and
106
+
107
+ (d) If the Work includes a "NOTICE" text file as part of its
108
+ distribution, then any Derivative Works that You distribute must
109
+ include a readable copy of the attribution notices contained
110
+ within such NOTICE file, excluding those notices that do not
111
+ pertain to any part of the Derivative Works, in at least one
112
+ of the following places: within a NOTICE text file distributed
113
+ as part of the Derivative Works; within the Source form or
114
+ documentation, if provided along with the Derivative Works; or,
115
+ within a display generated by the Derivative Works, if and
116
+ wherever such third-party notices normally appear. The contents
117
+ of the NOTICE file are for informational purposes only and
118
+ do not modify the License. You may add Your own attribution
119
+ notices within Derivative Works that You distribute, alongside
120
+ or as an addendum to the NOTICE text from the Work, provided
121
+ that such additional attribution notices cannot be construed
122
+ as modifying the License.
123
+
124
+ You may add Your own copyright statement to Your modifications and
125
+ may provide additional or different license terms and conditions
126
+ for use, reproduction, or distribution of Your modifications, or
127
+ for any such Derivative Works as a whole, provided Your use,
128
+ reproduction, and distribution of the Work otherwise complies with
129
+ the conditions stated in this License.
130
+
131
+ 5. Submission of Contributions. Unless You explicitly state otherwise,
132
+ any Contribution intentionally submitted for inclusion in the Work
133
+ by You to the Licensor shall be under the terms and conditions of
134
+ this License, without any additional terms or conditions.
135
+ Notwithstanding the above, nothing herein shall supersede or modify
136
+ the terms of any separate license agreement you may have executed
137
+ with Licensor regarding such Contributions.
138
+
139
+ 6. Trademarks. This License does not grant permission to use the trade
140
+ names, trademarks, service marks, or product names of the Licensor,
141
+ except as required for reasonable and customary use in describing the
142
+ origin of the Work and reproducing the content of the NOTICE file.
143
+
144
+ 7. Disclaimer of Warranty. Unless required by applicable law or
145
+ agreed to in writing, Licensor provides the Work (and each
146
+ Contributor provides its Contributions) on an "AS IS" BASIS,
147
+ WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or
148
+ implied, including, without limitation, any warranties or conditions
149
+ of TITLE, NON-INFRINGEMENT, MERCHANTABILITY, or FITNESS FOR A
150
+ PARTICULAR PURPOSE. You are solely responsible for determining the
151
+ appropriateness of using or redistributing the Work and assume any
152
+ risks associated with Your exercise of permissions under this License.
153
+
154
+ 8. Limitation of Liability. In no event and under no legal theory,
155
+ whether in tort (including negligence), contract, or otherwise,
156
+ unless required by applicable law (such as deliberate and grossly
157
+ negligent acts) or agreed to in writing, shall any Contributor be
158
+ liable to You for damages, including any direct, indirect, special,
159
+ incidental, or consequential damages of any character arising as a
160
+ result of this License or out of the use or inability to use the
161
+ Work (including but not limited to damages for loss of goodwill,
162
+ work stoppage, computer failure or malfunction, or any and all
163
+ other commercial damages or losses), even if such Contributor
164
+ has been advised of the possibility of such damages.
165
+
166
+ 9. Accepting Warranty or Additional Liability. While redistributing
167
+ the Work or Derivative Works thereof, You may choose to offer,
168
+ and charge a fee for, acceptance of support, warranty, indemnity,
169
+ or other liability obligations and/or rights consistent with this
170
+ License. However, in accepting such obligations, You may act only
171
+ on Your own behalf and on Your sole responsibility, not on behalf
172
+ of any other Contributor, and only if You agree to indemnify,
173
+ defend, and hold each Contributor harmless for any liability
174
+ incurred by, or claims asserted against, such Contributor by reason
175
+ of your accepting any such warranty or additional liability.
176
+
177
+ END OF TERMS AND CONDITIONS
178
+
179
+ APPENDIX: How to apply the Apache License to your work.
180
+
181
+ To apply the Apache License to your work, attach the following
182
+ boilerplate notice, with the fields enclosed by brackets "[]"
183
+ replaced with your own identifying information. (Don't include
184
+ the brackets!) The text should be enclosed in the appropriate
185
+ comment syntax for the file format. We also recommend that a
186
+ file or class name and description of purpose be included on the
187
+ same "printed page" as the copyright notice for easier
188
+ identification within third-party archives.
189
+
190
+ Copyright 2026 Alibaba Cloud
191
+
192
+ Licensed under the Apache License, Version 2.0 (the "License");
193
+ you may not use this file except in compliance with the License.
194
+ You may obtain a copy of the License at
195
+
196
+ http://www.apache.org/licenses/LICENSE-2.0
197
+
198
+ Unless required by applicable law or agreed to in writing, software
199
+ distributed under the License is distributed on an "AS IS" BASIS,
200
+ WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
201
+ See the License for the specific language governing permissions and
202
+ limitations under the License.
README.md ADDED
@@ -0,0 +1,81 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ base_model: Cloudflare/clef-flash
3
+ license: apache-2.0
4
+ library_name: exllamav3
5
+ tags:
6
+ - exl3
7
+ - exllamav3
8
+ - clef
9
+ - systemone
10
+ - qwen3.5
11
+ ---
12
+
13
+ # Cloudflare Clef-Flash EXL3
14
+
15
+ EXL3 conversion of `Cloudflare/clef-flash` with a functional text/JSON Clef SystemOne adapter.
16
+
17
+ Source revision: `17f0b0ad64efb65d273590632833508766b2aae6`
18
+
19
+ ## Quantization
20
+
21
+ - Qwen3.5 decoder: EXL3 4.00 bpw
22
+ - Vision tower: EXL3 6 bpw
23
+ - LM head: FP16 (`head_bits=16`)
24
+ - Cloudflare joint schema head: original BF16 weights retained in `joint_head.safetensors`; loaded as FP16 by the adapter
25
+ - Tested with ExLlamaV3 `1.5.2+cu128.torch2.10.0` on an RTX 4090
26
+
27
+ The LM head is intentionally kept at FP16 because Clef uses its output embedding vectors when scoring schema options.
28
+
29
+ ## Clef validation
30
+
31
+ The bundled `clef_exl3.py` bridges ExLlamaV3 final hidden states into Cloudflare's original `JointSchemaHead` and returns the same `noul`, `choice`, and `score` answer structures used by SystemOne.
32
+
33
+ | Check | BF16 reference | EXL3 |
34
+ |---|---:|---:|
35
+ | Invoice status: overdue | 0.9732 | 0.9753 |
36
+ | Invoice total > $1000 | 0.9732 | 0.9726 |
37
+ | Outage routing: technical | 0.9601 | 0.9572 |
38
+ | Urgency expected score | 1.7852 | 1.7966 |
39
+ | Service outage: true | 0.8339 | 0.8310 |
40
+
41
+ Across all numeric values in the two bundled validation cases, mean absolute delta was `0.004373` and maximum absolute delta was `0.0115`.
42
+
43
+ `clef-exl3-smoke.json`, `clef-bf16-reference.json`, and `VALIDATION.json` contain the validation outputs and provenance summary.
44
+
45
+ A standard generation probe on the RTX 4090 measured `75.372 tok/s`. This is a loader/generation smoke benchmark, not SystemOne decision throughput.
46
+
47
+ ## Current scope
48
+
49
+ Text and JSON state inputs are validated. The vision tower is included and quantized, but `clef_exl3.py` does not yet wire image/video inputs into the Clef decision path. Do not treat this release as validated multimodal SystemOne inference.
50
+
51
+ A `preprocessor_config.json` compatibility shim is included because this Cloudflare release stores the same image-processor metadata inside `processor_config.json`, while the tested ExLlamaV3 Qwen3.5 loader expects the standalone file.
52
+
53
+ ## Usage
54
+
55
+ ```python
56
+ import sys
57
+ from huggingface_hub import snapshot_download
58
+
59
+ path = snapshot_download("ramgpt/clef-flash-EXL3")
60
+ sys.path.insert(0, path)
61
+ from clef_exl3 import ClefEXL3
62
+
63
+ model = ClefEXL3(path)
64
+ response = model.systemone({
65
+ "model": "clef-flash-exl3",
66
+ "state": {"invoice": {"total": 1250, "status": "overdue"}},
67
+ "questions": {
68
+ "status": {
69
+ "type": "choice",
70
+ "instructions": "What is the invoice status?",
71
+ "criteria": {"paid": "Paid", "overdue": "Past due", "draft": "Not sent"}
72
+ },
73
+ "large": {"type": "noul", "instructions": "Is the total above 1000 USD?"}
74
+ }
75
+ })
76
+ print(response["answers"])
77
+ ```
78
+
79
+ ## Attribution
80
+
81
+ The base model, Clef joint schema head, and `joint_schema_model.py` originate from `Cloudflare/clef-flash` and are provided under the source model's Apache-2.0 license. This repository adds the EXL3 conversion, metadata compatibility shim, adapter, and validation artifacts.
VALIDATION.json ADDED
@@ -0,0 +1,22 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "source_model": "Cloudflare/clef-flash",
3
+ "source_revision": "17f0b0ad64efb65d273590632833508766b2aae6",
4
+ "quantization": {
5
+ "decoder_bits": 4.0,
6
+ "vision_bits": 6,
7
+ "lm_head_bits": 16,
8
+ "joint_head_source_dtype": "bfloat16",
9
+ "joint_head_runtime_dtype": "float16"
10
+ },
11
+ "runtime": {
12
+ "gpu": "NVIDIA GeForce RTX 4090",
13
+ "torch": "2.10.0+cu128",
14
+ "exllamav3": "1.5.2+cu128.torch2.10.0",
15
+ "decode_tps_probe": 75.372
16
+ },
17
+ "clef_smoke_passed": true,
18
+ "bf16_reference_passed": true,
19
+ "mean_abs_numeric_delta": 0.004373,
20
+ "max_abs_numeric_delta": 0.0115,
21
+ "notes": "Delta summary covers the bundled invoice and outage SystemOne validation cases."
22
+ }
__pycache__/joint_schema_model.cpython-312.pyc ADDED
Binary file (32.2 kB). View file
 
chat_template.jinja ADDED
@@ -0,0 +1,154 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {%- set image_count = namespace(value=0) %}
2
+ {%- set video_count = namespace(value=0) %}
3
+ {%- macro render_content(content, do_vision_count, is_system_content=false) %}
4
+ {%- if content is string %}
5
+ {{- content }}
6
+ {%- elif content is iterable and content is not mapping %}
7
+ {%- for item in content %}
8
+ {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
9
+ {%- if is_system_content %}
10
+ {{- raise_exception('System message cannot contain images.') }}
11
+ {%- endif %}
12
+ {%- if do_vision_count %}
13
+ {%- set image_count.value = image_count.value + 1 %}
14
+ {%- endif %}
15
+ {%- if add_vision_id %}
16
+ {{- 'Picture ' ~ image_count.value ~ ': ' }}
17
+ {%- endif %}
18
+ {{- '<|vision_start|><|image_pad|><|vision_end|>' }}
19
+ {%- elif 'video' in item or item.type == 'video' %}
20
+ {%- if is_system_content %}
21
+ {{- raise_exception('System message cannot contain videos.') }}
22
+ {%- endif %}
23
+ {%- if do_vision_count %}
24
+ {%- set video_count.value = video_count.value + 1 %}
25
+ {%- endif %}
26
+ {%- if add_vision_id %}
27
+ {{- 'Video ' ~ video_count.value ~ ': ' }}
28
+ {%- endif %}
29
+ {{- '<|vision_start|><|video_pad|><|vision_end|>' }}
30
+ {%- elif 'text' in item %}
31
+ {{- item.text }}
32
+ {%- else %}
33
+ {{- raise_exception('Unexpected item type in content.') }}
34
+ {%- endif %}
35
+ {%- endfor %}
36
+ {%- elif content is none or content is undefined %}
37
+ {{- '' }}
38
+ {%- else %}
39
+ {{- raise_exception('Unexpected content type.') }}
40
+ {%- endif %}
41
+ {%- endmacro %}
42
+ {%- if not messages %}
43
+ {{- raise_exception('No messages provided.') }}
44
+ {%- endif %}
45
+ {%- if tools and tools is iterable and tools is not mapping %}
46
+ {{- '<|im_start|>system\n' }}
47
+ {{- "# Tools\n\nYou have access to the following functions:\n\n<tools>" }}
48
+ {%- for tool in tools %}
49
+ {{- "\n" }}
50
+ {{- tool | tojson }}
51
+ {%- endfor %}
52
+ {{- "\n</tools>" }}
53
+ {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n<tool_call>\n<function=example_function_name>\n<parameter=example_parameter_1>\nvalue_1\n</parameter>\n<parameter=example_parameter_2>\nThis is the value for the second parameter\nthat can span\nmultiple lines\n</parameter>\n</function>\n</tool_call>\n\n<IMPORTANT>\nReminder:\n- Function calls MUST follow the specified format: an inner <function=...></function> block must be nested within <tool_call></tool_call> XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n</IMPORTANT>' }}
54
+ {%- if messages[0].role == 'system' %}
55
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
56
+ {%- if content %}
57
+ {{- '\n\n' + content }}
58
+ {%- endif %}
59
+ {%- endif %}
60
+ {{- '<|im_end|>\n' }}
61
+ {%- else %}
62
+ {%- if messages[0].role == 'system' %}
63
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
64
+ {{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
65
+ {%- endif %}
66
+ {%- endif %}
67
+ {%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
68
+ {%- for message in messages[::-1] %}
69
+ {%- set index = (messages|length - 1) - loop.index0 %}
70
+ {%- if ns.multi_step_tool and message.role == "user" %}
71
+ {%- set content = render_content(message.content, false)|trim %}
72
+ {%- if not(content.startswith('<tool_response>') and content.endswith('</tool_response>')) %}
73
+ {%- set ns.multi_step_tool = false %}
74
+ {%- set ns.last_query_index = index %}
75
+ {%- endif %}
76
+ {%- endif %}
77
+ {%- endfor %}
78
+ {%- if ns.multi_step_tool %}
79
+ {{- raise_exception('No user query found in messages.') }}
80
+ {%- endif %}
81
+ {%- for message in messages %}
82
+ {%- set content = render_content(message.content, true)|trim %}
83
+ {%- if message.role == "system" %}
84
+ {%- if not loop.first %}
85
+ {{- raise_exception('System message must be at the beginning.') }}
86
+ {%- endif %}
87
+ {%- elif message.role == "user" %}
88
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
89
+ {%- elif message.role == "assistant" %}
90
+ {%- set reasoning_content = '' %}
91
+ {%- if message.reasoning_content is string %}
92
+ {%- set reasoning_content = message.reasoning_content %}
93
+ {%- else %}
94
+ {%- if '</think>' in content %}
95
+ {%- set reasoning_content = content.split('</think>')[0].rstrip('\n').split('<think>')[-1].lstrip('\n') %}
96
+ {%- set content = content.split('</think>')[-1].lstrip('\n') %}
97
+ {%- endif %}
98
+ {%- endif %}
99
+ {%- set reasoning_content = reasoning_content|trim %}
100
+ {%- if loop.index0 > ns.last_query_index %}
101
+ {{- '<|im_start|>' + message.role + '\n<think>\n' + reasoning_content + '\n</think>\n\n' + content }}
102
+ {%- else %}
103
+ {{- '<|im_start|>' + message.role + '\n' + content }}
104
+ {%- endif %}
105
+ {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
106
+ {%- for tool_call in message.tool_calls %}
107
+ {%- if tool_call.function is defined %}
108
+ {%- set tool_call = tool_call.function %}
109
+ {%- endif %}
110
+ {%- if loop.first %}
111
+ {%- if content|trim %}
112
+ {{- '\n\n<tool_call>\n<function=' + tool_call.name + '>\n' }}
113
+ {%- else %}
114
+ {{- '<tool_call>\n<function=' + tool_call.name + '>\n' }}
115
+ {%- endif %}
116
+ {%- else %}
117
+ {{- '\n<tool_call>\n<function=' + tool_call.name + '>\n' }}
118
+ {%- endif %}
119
+ {%- if tool_call.arguments is defined %}
120
+ {%- for args_name, args_value in tool_call.arguments|items %}
121
+ {{- '<parameter=' + args_name + '>\n' }}
122
+ {%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}
123
+ {{- args_value }}
124
+ {{- '\n</parameter>\n' }}
125
+ {%- endfor %}
126
+ {%- endif %}
127
+ {{- '</function>\n</tool_call>' }}
128
+ {%- endfor %}
129
+ {%- endif %}
130
+ {{- '<|im_end|>\n' }}
131
+ {%- elif message.role == "tool" %}
132
+ {%- if loop.previtem and loop.previtem.role != "tool" %}
133
+ {{- '<|im_start|>user' }}
134
+ {%- endif %}
135
+ {{- '\n<tool_response>\n' }}
136
+ {{- content }}
137
+ {{- '\n</tool_response>' }}
138
+ {%- if not loop.last and loop.nextitem.role != "tool" %}
139
+ {{- '<|im_end|>\n' }}
140
+ {%- elif loop.last %}
141
+ {{- '<|im_end|>\n' }}
142
+ {%- endif %}
143
+ {%- else %}
144
+ {{- raise_exception('Unexpected message role.') }}
145
+ {%- endif %}
146
+ {%- endfor %}
147
+ {%- if add_generation_prompt %}
148
+ {{- '<|im_start|>assistant\n' }}
149
+ {%- if enable_thinking is defined and enable_thinking is false %}
150
+ {{- '<think>\n\n</think>\n\n' }}
151
+ {%- else %}
152
+ {{- '<think>\n' }}
153
+ {%- endif %}
154
+ {%- endif %}
clef-bf16-reference.json ADDED
@@ -0,0 +1,65 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "load_s": 4.466,
3
+ "invoice_s": 1.3387,
4
+ "outage_s": 0.0882,
5
+ "invoice": {
6
+ "model": "clef-flash",
7
+ "answers": {
8
+ "status": {
9
+ "type": "choice",
10
+ "choice": "overdue",
11
+ "confidence": 0.9732,
12
+ "probabilities": {
13
+ "paid": 0.0175,
14
+ "overdue": 0.9732,
15
+ "draft": 0.0094
16
+ }
17
+ },
18
+ "large": {
19
+ "type": "noul",
20
+ "noul": 0.9732
21
+ }
22
+ },
23
+ "usage": {
24
+ "input_tokens": 260,
25
+ "output_tokens": 0
26
+ }
27
+ },
28
+ "outage": {
29
+ "model": "clef-flash",
30
+ "answers": {
31
+ "department": {
32
+ "type": "choice",
33
+ "choice": "technical",
34
+ "confidence": 0.9601,
35
+ "probabilities": {
36
+ "billing": 0.0399,
37
+ "technical": 0.9601
38
+ }
39
+ },
40
+ "urgency": {
41
+ "type": "score",
42
+ "score": 1.7852,
43
+ "confidence": 0.8575,
44
+ "legend": {
45
+ "0": "Can wait",
46
+ "1": "This week",
47
+ "2": "Today"
48
+ },
49
+ "probabilities": {
50
+ "0": 0.0723,
51
+ "1": 0.0701,
52
+ "2": 0.8575
53
+ }
54
+ },
55
+ "outage": {
56
+ "type": "noul",
57
+ "noul": 0.8339
58
+ }
59
+ },
60
+ "usage": {
61
+ "input_tokens": 300,
62
+ "output_tokens": 0
63
+ }
64
+ }
65
+ }
clef-exl3-smoke.json ADDED
@@ -0,0 +1,67 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+
2
+ {
3
+ "passed": true,
4
+ "load_s": 2.787,
5
+ "invoice_s": 1.3508,
6
+ "outage_s": 0.0685,
7
+ "invoice": {
8
+ "model": "clef-flash-exl3",
9
+ "answers": {
10
+ "status": {
11
+ "type": "choice",
12
+ "choice": "overdue",
13
+ "confidence": 0.9753,
14
+ "probabilities": {
15
+ "paid": 0.0158,
16
+ "overdue": 0.9753,
17
+ "draft": 0.0089
18
+ }
19
+ },
20
+ "large": {
21
+ "type": "noul",
22
+ "noul": 0.9726
23
+ }
24
+ },
25
+ "usage": {
26
+ "input_tokens": 260,
27
+ "output_tokens": 0
28
+ }
29
+ },
30
+ "outage": {
31
+ "model": "clef-flash-exl3",
32
+ "answers": {
33
+ "department": {
34
+ "type": "choice",
35
+ "choice": "technical",
36
+ "confidence": 0.9572,
37
+ "probabilities": {
38
+ "billing": 0.0428,
39
+ "technical": 0.9572
40
+ }
41
+ },
42
+ "urgency": {
43
+ "type": "score",
44
+ "score": 1.7966,
45
+ "confidence": 0.869,
46
+ "legend": {
47
+ "0": "Can wait",
48
+ "1": "This week",
49
+ "2": "Today"
50
+ },
51
+ "probabilities": {
52
+ "0": 0.0724,
53
+ "1": 0.0586,
54
+ "2": 0.869
55
+ }
56
+ },
57
+ "outage": {
58
+ "type": "noul",
59
+ "noul": 0.831
60
+ }
61
+ },
62
+ "usage": {
63
+ "input_tokens": 300,
64
+ "output_tokens": 0
65
+ }
66
+ }
67
+ }
clef_exl3.py ADDED
@@ -0,0 +1,94 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ #!/usr/bin/env python3
2
+ import importlib.util
3
+ import json
4
+ import sys
5
+ from pathlib import Path
6
+ from types import SimpleNamespace
7
+
8
+ import torch
9
+ from exllamav3 import Config, Model, Tokenizer
10
+ from safetensors.torch import load_file
11
+
12
+
13
+ def load_joint_module(path):
14
+ spec = importlib.util.spec_from_file_location("clef_joint_schema_model", path / "joint_schema_model.py")
15
+ module = importlib.util.module_from_spec(spec)
16
+ sys.modules[spec.name] = module
17
+ spec.loader.exec_module(module)
18
+ return module
19
+
20
+
21
+ class TokenizerAdapter:
22
+ def __init__(self, tokenizer):
23
+ self.tokenizer = tokenizer
24
+
25
+ def __call__(self, text, add_special_tokens=False):
26
+ ids = self.tokenizer.encode(text, encode_special_tokens=True)
27
+ return SimpleNamespace(input_ids=ids[0].tolist())
28
+
29
+
30
+ class ClefEXL3:
31
+ def __init__(self, model_path, device="cuda:0"):
32
+ self.path = Path(model_path)
33
+ self.device = torch.device(device)
34
+ self.joint = load_joint_module(self.path)
35
+ config = Config.from_directory(str(self.path))
36
+ self.tokenizer = TokenizerAdapter(Tokenizer.from_config(config))
37
+ self.model = Model.from_config(config, component="text")
38
+ self.model.load(progressbar=True, device=device)
39
+ head_config = json.loads((self.path / "joint_head_config.json").read_text())
40
+ self.head = self.joint.JointSchemaHead(**head_config)
41
+ self.head.load_state_dict(load_file(self.path / "joint_head.safetensors"), strict=True)
42
+ self.head = self.head.to(device=self.device, dtype=torch.float16).eval()
43
+ lm_head = self.model.modules[self.model.logit_layer_idx]
44
+ if getattr(lm_head, "quant_type", None) != "fp16":
45
+ raise RuntimeError("Clef EXL3 requires an FP16 LM head; quantize with --head_bits 16")
46
+ raw = json.loads((self.path / "config.json").read_text())
47
+ text_config = raw.get("text_config", raw)
48
+ vocab_size = int(text_config["vocab_size"])
49
+ self.output_embedding_weight = lm_head.inner.weight.T[:vocab_size]
50
+
51
+ @torch.inference_mode()
52
+ def hidden_states(self, input_ids):
53
+ params = {"attn_mode": "flash_attn_nc"}
54
+ x = self.model.prepare_inputs(input_ids, params)
55
+ for module, instance, _ in self.model.fwd_modules:
56
+ if module.caps.get("logits_output"):
57
+ break
58
+ params["layer_instance"] = instance
59
+ x = module.prepare_for_device(x, params)
60
+ x = module.forward(x, params)
61
+ return x
62
+
63
+ @torch.inference_mode()
64
+ def systemone(self, request, max_length=16384):
65
+ if request.get("images") or request.get("videos"):
66
+ raise NotImplementedError("Clef EXL3 adapter currently supports text/JSON state only")
67
+ encoded = self.joint.encode_record(
68
+ self.tokenizer,
69
+ request,
70
+ max_length=max_length,
71
+ processor=None,
72
+ )
73
+ input_ids = torch.tensor([encoded.input_ids], dtype=torch.long, device=self.device)
74
+ attention_mask = torch.ones_like(input_ids)
75
+ hidden = self.hidden_states(input_ids).to(torch.float16)
76
+ logits = self.head(
77
+ hidden,
78
+ input_ids,
79
+ attention_mask,
80
+ [encoded],
81
+ self.output_embedding_weight,
82
+ )[0]
83
+ questions = request["questions"]
84
+ answers = {}
85
+ for question, question_logits in zip(encoded.questions, logits):
86
+ probabilities = dict(zip(question.option_ids, question_logits.float().softmax(-1).tolist()))
87
+ answers[question.question_id] = self.joint.systemone_answer(
88
+ questions[question.question_id], probabilities
89
+ )
90
+ return {
91
+ "model": request.get("model", "clef-flash-exl3"),
92
+ "answers": answers,
93
+ "usage": {"input_tokens": len(encoded.input_ids), "output_tokens": 0},
94
+ }
clef_exl3_smoke.py ADDED
@@ -0,0 +1,65 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ #!/usr/bin/env python3
2
+ import argparse
3
+ import json
4
+ import time
5
+
6
+ from clef_exl3 import ClefEXL3
7
+
8
+
9
+ def main():
10
+ ap = argparse.ArgumentParser()
11
+ ap.add_argument("--model", required=True)
12
+ args = ap.parse_args()
13
+ t0 = time.perf_counter()
14
+ model = ClefEXL3(args.model)
15
+ load_s = time.perf_counter() - t0
16
+ invoice = {
17
+ "model": "clef-flash-exl3",
18
+ "state": {"invoice": {"vendor": "Acme", "total": 1250.0, "currency": "USD", "status": "overdue"}},
19
+ "questions": {
20
+ "status": {
21
+ "type": "choice",
22
+ "instructions": "What is the invoice status?",
23
+ "criteria": {"paid": "Invoice is paid.", "overdue": "Invoice is past due.", "draft": "Not sent."}
24
+ },
25
+ "large": {"type": "noul", "instructions": "Is the total above 1000 USD?"}
26
+ }
27
+ }
28
+ outage = {
29
+ "model": "clef-flash-exl3",
30
+ "state": "Our checkout started returning errors and orders are blocked.",
31
+ "questions": {
32
+ "department": {
33
+ "type": "choice",
34
+ "instructions": "Which team should handle the message?",
35
+ "criteria": {"billing": "Payments or invoices", "technical": "Bugs or outages"}
36
+ },
37
+ "urgency": {"type": "score", "criteria": ["Can wait", "This week", "Today"]},
38
+ "outage": {"type": "noul", "instructions": "Is a service down?"}
39
+ }
40
+ }
41
+ t1 = time.perf_counter()
42
+ r1 = model.systemone(invoice)
43
+ t2 = time.perf_counter()
44
+ r2 = model.systemone(outage)
45
+ t3 = time.perf_counter()
46
+ passed = (
47
+ r1["answers"]["status"].get("choice") == "overdue"
48
+ and r1["answers"]["large"].get("noul", 0) > 0.5
49
+ and r2["answers"]["department"].get("choice") == "technical"
50
+ and r2["answers"]["outage"].get("noul", 0) > 0.5
51
+ )
52
+ result = {
53
+ "passed": passed,
54
+ "load_s": round(load_s, 3),
55
+ "invoice_s": round(t2 - t1, 4),
56
+ "outage_s": round(t3 - t2, 4),
57
+ "invoice": r1,
58
+ "outage": r2
59
+ }
60
+ print(json.dumps(result, indent=2))
61
+ raise SystemExit(0 if passed else 3)
62
+
63
+
64
+ if __name__ == "__main__":
65
+ main()
config.json ADDED
@@ -0,0 +1,122 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "architectures": [
3
+ "Qwen3_5ForConditionalGeneration"
4
+ ],
5
+ "dtype": "bfloat16",
6
+ "image_token_id": 248056,
7
+ "model_type": "qwen3_5",
8
+ "text_config": {
9
+ "attention_bias": false,
10
+ "attention_dropout": 0.0,
11
+ "attn_output_gate": true,
12
+ "bos_token_id": null,
13
+ "dtype": "bfloat16",
14
+ "eos_token_id": 248044,
15
+ "full_attention_interval": 4,
16
+ "head_dim": 256,
17
+ "hidden_act": "silu",
18
+ "hidden_size": 4096,
19
+ "initializer_range": 0.02,
20
+ "intermediate_size": 12288,
21
+ "layer_types": [
22
+ "linear_attention",
23
+ "linear_attention",
24
+ "linear_attention",
25
+ "full_attention",
26
+ "linear_attention",
27
+ "linear_attention",
28
+ "linear_attention",
29
+ "full_attention",
30
+ "linear_attention",
31
+ "linear_attention",
32
+ "linear_attention",
33
+ "full_attention",
34
+ "linear_attention",
35
+ "linear_attention",
36
+ "linear_attention",
37
+ "full_attention",
38
+ "linear_attention",
39
+ "linear_attention",
40
+ "linear_attention",
41
+ "full_attention",
42
+ "linear_attention",
43
+ "linear_attention",
44
+ "linear_attention",
45
+ "full_attention",
46
+ "linear_attention",
47
+ "linear_attention",
48
+ "linear_attention",
49
+ "full_attention",
50
+ "linear_attention",
51
+ "linear_attention",
52
+ "linear_attention",
53
+ "full_attention"
54
+ ],
55
+ "linear_conv_kernel_dim": 4,
56
+ "linear_key_head_dim": 128,
57
+ "linear_num_key_heads": 16,
58
+ "linear_num_value_heads": 32,
59
+ "linear_value_head_dim": 128,
60
+ "mamba_ssm_dtype": "float32",
61
+ "max_position_embeddings": 262144,
62
+ "mlp_only_layers": [],
63
+ "model_type": "qwen3_5_text",
64
+ "mtp_num_hidden_layers": 0,
65
+ "mtp_use_dedicated_embeddings": false,
66
+ "num_attention_heads": 16,
67
+ "num_hidden_layers": 32,
68
+ "num_key_value_heads": 4,
69
+ "pad_token_id": null,
70
+ "partial_rotary_factor": 0.25,
71
+ "rms_norm_eps": 1e-06,
72
+ "rope_parameters": {
73
+ "mrope_interleaved": true,
74
+ "mrope_section": [
75
+ 11,
76
+ 11,
77
+ 10
78
+ ],
79
+ "partial_rotary_factor": 0.25,
80
+ "rope_theta": 10000000,
81
+ "rope_type": "default"
82
+ },
83
+ "tie_word_embeddings": false,
84
+ "use_cache": true,
85
+ "vocab_size": 248320
86
+ },
87
+ "tie_word_embeddings": false,
88
+ "transformers_version": "5.10.2",
89
+ "video_token_id": 248057,
90
+ "vision_config": {
91
+ "deepstack_visual_indexes": [],
92
+ "depth": 27,
93
+ "dtype": "bfloat16",
94
+ "hidden_act": "gelu_pytorch_tanh",
95
+ "hidden_size": 1152,
96
+ "in_channels": 3,
97
+ "initializer_range": 0.02,
98
+ "intermediate_size": 4304,
99
+ "model_type": "qwen3_5_vision",
100
+ "num_heads": 16,
101
+ "num_position_embeddings": 2304,
102
+ "out_hidden_size": 4096,
103
+ "patch_size": 16,
104
+ "spatial_merge_size": 2,
105
+ "temporal_patch_size": 2
106
+ },
107
+ "vision_end_token_id": 248054,
108
+ "vision_start_token_id": 248053,
109
+ "quantization_config": {
110
+ "quant_method": "exl3",
111
+ "version": "1.5.2",
112
+ "bits": 4.0,
113
+ "head_bits": 16,
114
+ "calibration": {
115
+ "rows": 250,
116
+ "cols": 2048
117
+ },
118
+ "out_scales": "always",
119
+ "codebook": "mul1",
120
+ "vision_bits": 6
121
+ }
122
+ }
generation_config.json ADDED
@@ -0,0 +1,6 @@
 
 
 
 
 
 
 
1
+ {
2
+ "_from_model_config": true,
3
+ "eos_token_id": 248044,
4
+ "transformers_version": "5.10.2",
5
+ "use_cache": true
6
+ }
joint_head.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:19cdcec8c81dc9212be320fff47462ab342fbc1278be4368fb3da71241cf5ba0
3
+ size 243538016
joint_head_config.json ADDED
@@ -0,0 +1,8 @@
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "hidden_size": 4096,
3
+ "width": 1024,
4
+ "routing_layers": 2,
5
+ "layers": 4,
6
+ "heads": 16,
7
+ "feedforward": 4096
8
+ }
joint_schema_model.py ADDED
@@ -0,0 +1,576 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Clef: a multimodal Qwen backbone with a joint schema head for typed decisions.
2
+
3
+ A record provides a ``state`` (any JSON value), optional ``images`` and ``videos``,
4
+ and ``questions``. Each question has a ``type`` (``noul``, ``choice``, or ``score``),
5
+ ``instructions``, and, for ``choice`` and ``score``, ``criteria`` describing the
6
+ allowed options. The model returns one logit per allowed option for every question.
7
+ ``systemone`` answers a Jev/SystemOne ``/v1/systemone`` request body with the same
8
+ response body.
9
+ """
10
+
11
+ from __future__ import annotations
12
+
13
+ import json
14
+ import math
15
+ from dataclasses import dataclass
16
+ from dataclasses import field as dataclass_field
17
+ from pathlib import Path
18
+ from typing import Any
19
+
20
+ import torch
21
+ import torch.nn.functional as functional
22
+
23
+
24
+ SYSTEM_PROMPT = (
25
+ "Read the complete state and schema. Decide every field jointly. Each answer "
26
+ "must be exactly one of that field's allowed options."
27
+ )
28
+ IMAGE_PLACEHOLDER = "<|vision_start|><|image_pad|><|vision_end|>"
29
+ VIDEO_PLACEHOLDER = "<|vision_start|><|video_pad|><|vision_end|>"
30
+ MEDIA_BATCH_KEYS = ("pixel_values", "image_grid_thw", "pixel_values_videos", "video_grid_thw")
31
+ MEDIA_TOKEN_KEYS = ("mm_token_type_ids",)
32
+ QUESTION_TYPES = {"noul": 0, "choice": 1, "score": 2}
33
+
34
+
35
+ def render(value: Any) -> str:
36
+ if isinstance(value, str):
37
+ return value
38
+ return json.dumps(
39
+ value,
40
+ ensure_ascii=False,
41
+ separators=(",", ":"),
42
+ sort_keys=True,
43
+ )
44
+
45
+
46
+ def question_options(question: dict[str, Any]) -> list[tuple[str, Any]]:
47
+ question_type = str(question["type"])
48
+ if question_type == "noul":
49
+ criteria = {
50
+ "true": "The proposition is true or the answer is yes.",
51
+ "false": "The proposition is false or the answer is no.",
52
+ }
53
+ criteria.update(question.get("criteria") or {})
54
+ return [(key, criteria[key]) for key in ("true", "false")]
55
+ if question_type == "choice":
56
+ return sorted((str(key), value) for key, value in question["criteria"].items())
57
+ return [(str(index), value) for index, value in enumerate(question["criteria"])]
58
+
59
+
60
+ @dataclass(frozen=True)
61
+ class EncodedQuestion:
62
+ question_id: str
63
+ question_type: int
64
+ question_span: tuple[int, int]
65
+ option_spans: tuple[tuple[int, int], ...]
66
+ option_ids: tuple[str, ...]
67
+
68
+
69
+ @dataclass(frozen=True)
70
+ class EncodedRecord:
71
+ input_ids: tuple[int, ...]
72
+ questions: tuple[EncodedQuestion, ...]
73
+ record_id: str
74
+ media: dict[str, Any] | None = dataclass_field(default=None, compare=False, repr=False)
75
+
76
+
77
+ def _tokens(tokenizer: Any, text: str) -> list[int]:
78
+ return tokenizer(text, add_special_tokens=False).input_ids
79
+
80
+
81
+ def _encode_media(processor: Any, record: dict[str, Any]) -> tuple[list[int], dict[str, Any] | None]:
82
+ images = list(record.get("images") or [])
83
+ videos = list(record.get("videos") or [])
84
+ if not images and not videos:
85
+ return [], None
86
+ if processor is None:
87
+ raise ValueError("records with images or videos require a processor")
88
+ text = IMAGE_PLACEHOLDER * len(images) + VIDEO_PLACEHOLDER * len(videos) + "\n"
89
+ encoded = processor(
90
+ text=[text],
91
+ images=images or None,
92
+ videos=videos or None,
93
+ return_tensors="pt",
94
+ **(record.get("media_kwargs") or {}),
95
+ )
96
+ media = {key: encoded[key] for key in MEDIA_BATCH_KEYS if key in encoded}
97
+ for key in MEDIA_TOKEN_KEYS:
98
+ if key in encoded:
99
+ media[key] = encoded[key][0].tolist()
100
+ return encoded["input_ids"][0].tolist(), media
101
+
102
+
103
+ def encode_record(
104
+ tokenizer: Any,
105
+ record: dict[str, Any],
106
+ max_length: int = 16384,
107
+ max_state_tokens: int | None = None,
108
+ processor: Any | None = None,
109
+ ) -> EncodedRecord:
110
+ schema_ids = _tokens(tokenizer, "\n\nSCHEMA FIELDS:\n")
111
+ questions: list[EncodedQuestion] = []
112
+ for question_index, (question_id, question) in enumerate(record["questions"].items()):
113
+ schema_ids.extend(
114
+ _tokens(
115
+ tokenizer,
116
+ f"\nFIELD {question_index + 1}\nID: {question_id}\nTYPE: {question['type']}\nINSTRUCTION: ",
117
+ )
118
+ )
119
+ question_start = len(schema_ids)
120
+ instructions = question.get("instructions")
121
+ if instructions is None or instructions == "":
122
+ instructions = str(question_id)
123
+ schema_ids.extend(_tokens(tokenizer, render(instructions)))
124
+ question_end = len(schema_ids)
125
+ schema_ids.extend(_tokens(tokenizer, "\nALLOWED OPTIONS:\n"))
126
+
127
+ option_spans: list[tuple[int, int]] = []
128
+ option_ids: list[str] = []
129
+ for option_index, (option_id, description) in enumerate(question_options(question)):
130
+ schema_ids.extend(_tokens(tokenizer, f"OPTION {option_index + 1}: "))
131
+ option_start = len(schema_ids)
132
+ semantics = {"option_id": option_id}
133
+ if description is not None:
134
+ semantics["description"] = description
135
+ schema_ids.extend(_tokens(tokenizer, render(semantics)))
136
+ option_spans.append((option_start, len(schema_ids)))
137
+ option_ids.append(option_id)
138
+ schema_ids.extend(_tokens(tokenizer, "\n"))
139
+ schema_ids.extend(_tokens(tokenizer, "END FIELD\n"))
140
+ questions.append(
141
+ EncodedQuestion(
142
+ question_id=str(question_id),
143
+ question_type=QUESTION_TYPES[str(question["type"])],
144
+ question_span=(question_start, question_end),
145
+ option_spans=tuple(option_spans),
146
+ option_ids=tuple(option_ids),
147
+ )
148
+ )
149
+
150
+ prefix_ids = _tokens(
151
+ tokenizer,
152
+ f"<|im_start|>system\n{SYSTEM_PROMPT}<|im_end|>\n<|im_start|>user\nSTATE:\n",
153
+ )
154
+ suffix_ids = _tokens(
155
+ tokenizer,
156
+ "\n<|im_end|>\n<|im_start|>assistant\n<think>\n\n</think>\n\nJOINT SCHEMA DECISIONS:",
157
+ )
158
+ media_ids, media = _encode_media(processor, record)
159
+ if media is not None:
160
+ media["token_offset"] = len(prefix_ids)
161
+ prefix_ids = prefix_ids + media_ids
162
+ state_ids = _tokens(tokenizer, render(record["state"]))
163
+ if max_state_tokens is not None:
164
+ state_ids = state_ids[:max_state_tokens]
165
+ fixed_length = len(prefix_ids) + len(schema_ids) + len(suffix_ids)
166
+ if fixed_length > max_length:
167
+ raise ValueError(
168
+ f"schema requires {fixed_length} tokens before state; maximum is {max_length}"
169
+ )
170
+ state_ids = state_ids[: max_length - fixed_length]
171
+ schema_offset = len(prefix_ids) + len(state_ids)
172
+ shifted_questions = tuple(
173
+ EncodedQuestion(
174
+ question_id=question.question_id,
175
+ question_type=question.question_type,
176
+ question_span=(
177
+ question.question_span[0] + schema_offset,
178
+ question.question_span[1] + schema_offset,
179
+ ),
180
+ option_spans=tuple(
181
+ (start + schema_offset, end + schema_offset)
182
+ for start, end in question.option_spans
183
+ ),
184
+ option_ids=question.option_ids,
185
+ )
186
+ for question in questions
187
+ )
188
+ input_ids = tuple(prefix_ids + state_ids + schema_ids + suffix_ids)
189
+ if not input_ids or not shifted_questions:
190
+ raise ValueError("record produced no model input or questions")
191
+ return EncodedRecord(
192
+ input_ids=input_ids,
193
+ questions=shifted_questions,
194
+ record_id=str(record.get("id", "unknown")),
195
+ media=media,
196
+ )
197
+
198
+
199
+ def collate_records(
200
+ records: list[EncodedRecord],
201
+ pad_token_id: int,
202
+ device: torch.device,
203
+ ) -> dict[str, Any]:
204
+ maximum_length = max(len(record.input_ids) for record in records)
205
+ input_ids = torch.full(
206
+ (len(records), maximum_length),
207
+ pad_token_id,
208
+ dtype=torch.long,
209
+ device=device,
210
+ )
211
+ attention_mask = torch.zeros(
212
+ (len(records), maximum_length),
213
+ dtype=torch.long,
214
+ device=device,
215
+ )
216
+ for index, record in enumerate(records):
217
+ length = len(record.input_ids)
218
+ input_ids[index, :length] = torch.tensor(record.input_ids, device=device)
219
+ attention_mask[index, :length] = 1
220
+ media: dict[str, torch.Tensor] = {}
221
+ for key in MEDIA_BATCH_KEYS:
222
+ values = [record.media[key] for record in records if record.media and key in record.media]
223
+ if values:
224
+ media[key] = torch.cat(values, dim=0).to(device)
225
+ for key in MEDIA_TOKEN_KEYS:
226
+ if any(record.media and key in record.media for record in records):
227
+ token_values = torch.zeros((len(records), maximum_length), dtype=torch.long, device=device)
228
+ for index, record in enumerate(records):
229
+ if record.media and key in record.media:
230
+ offset = record.media["token_offset"]
231
+ values = torch.tensor(record.media[key], dtype=torch.long, device=device)
232
+ token_values[index, offset : offset + len(values)] = values
233
+ media[key] = token_values
234
+ return {
235
+ "input_ids": input_ids,
236
+ "attention_mask": attention_mask,
237
+ "records": records,
238
+ "media": media,
239
+ }
240
+
241
+
242
+ class EvidenceRoutingLayer(torch.nn.Module):
243
+ def __init__(
244
+ self,
245
+ width: int,
246
+ heads: int,
247
+ feedforward: int,
248
+ dropout: float = 0.0,
249
+ ) -> None:
250
+ super().__init__()
251
+ self.query_norm = torch.nn.LayerNorm(width)
252
+ self.memory_norm = torch.nn.LayerNorm(width)
253
+ self.attention = torch.nn.MultiheadAttention(
254
+ width,
255
+ heads,
256
+ dropout=dropout,
257
+ batch_first=True,
258
+ )
259
+ self.attention_dropout = torch.nn.Dropout(dropout)
260
+ self.feedforward_norm = torch.nn.LayerNorm(width)
261
+ self.feedforward = torch.nn.Sequential(
262
+ torch.nn.Linear(width, feedforward),
263
+ torch.nn.GELU(),
264
+ torch.nn.Dropout(dropout),
265
+ torch.nn.Linear(feedforward, width),
266
+ torch.nn.Dropout(dropout),
267
+ )
268
+
269
+ def forward(self, queries: torch.Tensor, memory: torch.Tensor) -> torch.Tensor:
270
+ normalized_queries = self.query_norm(queries)
271
+ routed, _ = self.attention(
272
+ normalized_queries,
273
+ self.memory_norm(memory),
274
+ self.memory_norm(memory),
275
+ need_weights=False,
276
+ )
277
+ queries = queries + self.attention_dropout(routed)
278
+ return queries + self.feedforward(self.feedforward_norm(queries))
279
+
280
+
281
+ class JointSchemaHead(torch.nn.Module):
282
+ def __init__(
283
+ self,
284
+ hidden_size: int,
285
+ width: int,
286
+ routing_layers: int,
287
+ layers: int,
288
+ heads: int,
289
+ feedforward: int,
290
+ dropout: float = 0.0,
291
+ ) -> None:
292
+ super().__init__()
293
+ self.hidden_norm = torch.nn.LayerNorm(hidden_size)
294
+ self.memory_projection = torch.nn.Linear(hidden_size, width, bias=False)
295
+ self.question_projection = torch.nn.Linear(hidden_size, width, bias=False)
296
+ self.option_question_projection = torch.nn.Linear(hidden_size, width, bias=False)
297
+ self.global_projection = torch.nn.Linear(hidden_size, width, bias=False)
298
+ self.option_context_projection = torch.nn.Linear(hidden_size, width, bias=False)
299
+ self.option_lexical_projection = torch.nn.Linear(hidden_size, width, bias=False)
300
+ self.type_embedding = torch.nn.Embedding(3, width)
301
+ self.evidence_layers = torch.nn.ModuleList(
302
+ [
303
+ EvidenceRoutingLayer(
304
+ width=width,
305
+ heads=heads,
306
+ feedforward=feedforward,
307
+ dropout=dropout,
308
+ )
309
+ for _ in range(routing_layers)
310
+ ]
311
+ )
312
+ self.option_summary_norm = torch.nn.LayerNorm(width)
313
+ self.layers = torch.nn.ModuleList(
314
+ [
315
+ torch.nn.TransformerDecoderLayer(
316
+ d_model=width,
317
+ nhead=heads,
318
+ dim_feedforward=feedforward,
319
+ dropout=dropout,
320
+ activation="gelu",
321
+ batch_first=True,
322
+ norm_first=True,
323
+ )
324
+ for _ in range(layers)
325
+ ]
326
+ )
327
+ self.field_norm = torch.nn.LayerNorm(width)
328
+ self.option_norm = torch.nn.LayerNorm(width)
329
+ self.residual_scorer = torch.nn.Sequential(
330
+ torch.nn.Linear(width * 4, width),
331
+ torch.nn.GELU(),
332
+ torch.nn.Dropout(dropout),
333
+ torch.nn.Linear(width, 1),
334
+ )
335
+ self.prior_logit_scale = torch.nn.Parameter(torch.zeros(()))
336
+ self.joint_logit_scale = torch.nn.Parameter(torch.zeros(()))
337
+ self.residual_gate = torch.nn.Parameter(torch.zeros(()))
338
+
339
+ @staticmethod
340
+ def _mean_span(values: torch.Tensor, span: tuple[int, int]) -> torch.Tensor:
341
+ start, end = span
342
+ return values[start:end].mean(dim=0)
343
+
344
+ def forward(
345
+ self,
346
+ hidden_states: torch.Tensor,
347
+ input_ids: torch.Tensor,
348
+ attention_mask: torch.Tensor,
349
+ records: list[EncodedRecord],
350
+ output_embedding_weight: torch.Tensor,
351
+ ) -> list[list[torch.Tensor]]:
352
+ results: list[list[torch.Tensor]] = []
353
+ normalized_hidden = self.hidden_norm(hidden_states)
354
+ for batch_index, record in enumerate(records):
355
+ sequence_length = int(attention_mask[batch_index].sum().item())
356
+ sequence_hidden = normalized_hidden[batch_index, :sequence_length]
357
+ memory = self.memory_projection(sequence_hidden).unsqueeze(0)
358
+ global_vector = sequence_hidden[-1]
359
+ question_vectors = torch.stack(
360
+ [
361
+ self._mean_span(sequence_hidden, question.question_span)
362
+ for question in record.questions
363
+ ]
364
+ )
365
+ type_ids = torch.tensor(
366
+ [question.question_type for question in record.questions],
367
+ device=hidden_states.device,
368
+ )
369
+ option_contexts: list[torch.Tensor] = []
370
+ lexical_options: list[torch.Tensor] = []
371
+ option_counts = []
372
+ for question in record.questions:
373
+ context_vectors = torch.stack(
374
+ [
375
+ self._mean_span(sequence_hidden, span)
376
+ for span in question.option_spans
377
+ ]
378
+ )
379
+ lexical_vectors = []
380
+ for start, end in question.option_spans:
381
+ token_ids = input_ids[batch_index, start:end]
382
+ lexical_vectors.append(output_embedding_weight[token_ids].mean(dim=0))
383
+ lexical = torch.stack(lexical_vectors)
384
+ option_contexts.append(context_vectors)
385
+ lexical_options.append(lexical)
386
+ option_counts.append(len(question.option_spans))
387
+
388
+ option_queries = []
389
+ for question_index, (context_vectors, lexical) in enumerate(
390
+ zip(option_contexts, lexical_options)
391
+ ):
392
+ option_queries.append(
393
+ self.option_context_projection(context_vectors)
394
+ + self.option_lexical_projection(lexical)
395
+ + self.option_question_projection(
396
+ question_vectors[question_index]
397
+ ).unsqueeze(0)
398
+ )
399
+ routed_options = torch.cat(option_queries, dim=0).unsqueeze(0)
400
+ for layer in self.evidence_layers:
401
+ routed_options = layer(routed_options, memory)
402
+ routed_options = routed_options[0]
403
+ split_options = list(torch.split(routed_options, option_counts, dim=0))
404
+
405
+ base_fields = self.question_projection(question_vectors)
406
+ option_summaries = []
407
+ for field, options in zip(base_fields, split_options):
408
+ routing_weights = torch.softmax(
409
+ torch.matmul(options, field) / math.sqrt(options.shape[-1]),
410
+ dim=0,
411
+ )
412
+ option_summaries.append(
413
+ torch.sum(routing_weights.unsqueeze(-1) * options, dim=0)
414
+ )
415
+ fields = (
416
+ base_fields
417
+ + self.option_summary_norm(torch.stack(option_summaries))
418
+ + self.global_projection(global_vector).unsqueeze(0)
419
+ + self.type_embedding(type_ids)
420
+ )
421
+ fields = fields.unsqueeze(0)
422
+ for layer in self.layers:
423
+ fields = layer(fields, memory)
424
+ fields = self.field_norm(fields[0])
425
+
426
+ record_logits: list[torch.Tensor] = []
427
+ for field, question, lexical, routed in zip(
428
+ fields,
429
+ record.questions,
430
+ lexical_options,
431
+ split_options,
432
+ ):
433
+ anchor = functional.normalize(
434
+ question_vectors[len(record_logits)] + global_vector,
435
+ dim=-1,
436
+ )
437
+ lexical_anchor = functional.normalize(lexical, dim=-1)
438
+ prior_scale = self.prior_logit_scale.clamp(max=math.log(100.0)).exp()
439
+ prior = prior_scale * torch.matmul(lexical_anchor, anchor)
440
+ options = self.option_norm(routed)
441
+ repeated_field = field.unsqueeze(0).expand_as(options)
442
+ cosine = functional.cosine_similarity(repeated_field, options, dim=-1)
443
+ features = torch.cat(
444
+ [
445
+ repeated_field,
446
+ options,
447
+ repeated_field * options,
448
+ torch.abs(repeated_field - options),
449
+ ],
450
+ dim=-1,
451
+ )
452
+ residual = self.residual_scorer(features).squeeze(-1)
453
+ joint_scale = self.joint_logit_scale.clamp(max=math.log(100.0)).exp()
454
+ joint = joint_scale * cosine + residual
455
+ record_logits.append(
456
+ prior + torch.sigmoid(self.residual_gate) * joint
457
+ )
458
+ results.append(record_logits)
459
+ return results
460
+
461
+
462
+ class ClefModel(torch.nn.Module):
463
+ def __init__(self, language_model: Any, head: JointSchemaHead) -> None:
464
+ super().__init__()
465
+ self.language_model = language_model
466
+ self.head = head
467
+
468
+ def forward(self, batch: dict[str, Any]) -> list[list[torch.Tensor]]:
469
+ base_model = (
470
+ self.language_model.get_base_model()
471
+ if hasattr(self.language_model, "get_base_model")
472
+ else self.language_model
473
+ )
474
+ media = batch.get("media") or {}
475
+ text_model = base_model.model
476
+ if not media and hasattr(text_model, "language_model"):
477
+ text_model = text_model.language_model
478
+ outputs = text_model(
479
+ input_ids=batch["input_ids"],
480
+ attention_mask=batch["attention_mask"],
481
+ use_cache=False,
482
+ return_dict=True,
483
+ **media,
484
+ )
485
+ return self.head(
486
+ outputs.last_hidden_state,
487
+ batch["input_ids"],
488
+ batch["attention_mask"],
489
+ batch["records"],
490
+ base_model.get_output_embeddings().weight,
491
+ )
492
+
493
+
494
+ def load_release_model(
495
+ model_path: str | Path,
496
+ device: str | torch.device = "cuda",
497
+ dtype: torch.dtype = torch.bfloat16,
498
+ **from_pretrained_kwargs: Any,
499
+ ) -> tuple[ClefModel, Any]:
500
+ """Load a Clef release (merged backbone, joint schema head, and processor)."""
501
+ from huggingface_hub import snapshot_download
502
+ from safetensors.torch import load_file
503
+ from transformers import AutoProcessor, Qwen3_5ForConditionalGeneration
504
+
505
+ path = Path(model_path)
506
+ if not path.is_dir():
507
+ path = Path(snapshot_download(str(model_path)))
508
+ backbone = Qwen3_5ForConditionalGeneration.from_pretrained(
509
+ path,
510
+ dtype=dtype,
511
+ device_map={"": str(device)},
512
+ **from_pretrained_kwargs,
513
+ )
514
+ backbone.config.use_cache = False
515
+ head_config = json.loads((path / "joint_head_config.json").read_text())
516
+ head = JointSchemaHead(**head_config)
517
+ head.load_state_dict(load_file(path / "joint_head.safetensors"), strict=True)
518
+ head = head.to(device=device, dtype=dtype)
519
+ processor = AutoProcessor.from_pretrained(path)
520
+ return ClefModel(backbone, head).eval(), processor
521
+
522
+
523
+ def systemone_answer(question: dict[str, Any], probabilities: dict[str, float]) -> dict[str, Any]:
524
+ """Convert per-option probabilities for one question into a SystemOne answer."""
525
+ if question["type"] == "noul":
526
+ return {"type": "noul", "noul": round(probabilities["true"], 4)}
527
+ if question["type"] == "choice":
528
+ options = [str(option) for option in question["criteria"]]
529
+ choice = max(options, key=probabilities.__getitem__)
530
+ return {
531
+ "type": "choice",
532
+ "choice": choice,
533
+ "confidence": round(probabilities[choice], 4),
534
+ "probabilities": {option: round(probabilities[option], 4) for option in options},
535
+ }
536
+ levels = [str(index) for index in range(len(question["criteria"]))]
537
+ return {
538
+ "type": "score",
539
+ "score": round(sum(index * probabilities[level] for index, level in enumerate(levels)), 4),
540
+ "confidence": round(max(probabilities[level] for level in levels), 4),
541
+ "legend": dict(zip(levels, question["criteria"])),
542
+ "probabilities": {level: round(probabilities[level], 4) for level in levels},
543
+ }
544
+
545
+
546
+ @torch.inference_mode()
547
+ def systemone(model: ClefModel, processor: Any, request: dict[str, Any], max_length: int = 16384) -> dict[str, Any]:
548
+ """Answer a Jev/SystemOne ``/v1/systemone`` request body with a SystemOne response body.
549
+
550
+ The request has ``model``, ``state``, and ``questions``, plus optional ``images`` and ``videos``.
551
+ """
552
+ questions = request.get("questions")
553
+ if not isinstance(request.get("model"), str) or "state" not in request:
554
+ raise ValueError("model and state are required")
555
+ if not isinstance(questions, dict) or not questions:
556
+ raise ValueError("at least one question is required")
557
+ for question_id, question in questions.items():
558
+ if question.get("type") not in QUESTION_TYPES:
559
+ raise ValueError(f"{question_id}: type must be noul, choice, or score")
560
+ if question["type"] != "noul" and not question.get("criteria"):
561
+ raise ValueError(f"{question_id}: criteria must not be empty")
562
+ encoded = encode_record(processor.tokenizer, request, max_length=max_length, processor=processor)
563
+ device = next(model.parameters()).device
564
+ logits = model(collate_records([encoded], processor.tokenizer.pad_token_id, device))[0]
565
+ answers = {
566
+ question.question_id: systemone_answer(
567
+ questions[question.question_id],
568
+ dict(zip(question.option_ids, question_logits.float().softmax(-1).tolist())),
569
+ )
570
+ for question, question_logits in zip(encoded.questions, logits)
571
+ }
572
+ return {
573
+ "model": request["model"],
574
+ "answers": answers,
575
+ "usage": {"input_tokens": len(encoded.input_ids), "output_tokens": 0},
576
+ }
model-00001-of-00002.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:0ab5bbfe488a41262ef1174298de958e78ce6461ed8d4ac6bdb8188dca8acde1
3
+ size 4206704791
model-00002-of-00002.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:a0ddf7d7711ad3487621c42cd9b89dcd494b64b622808f383e4abee647c98cb4
3
+ size 3904432105
model.safetensors.index.json ADDED
The diff for this file is too large to render. See raw diff
 
preprocessor_config.json ADDED
@@ -0,0 +1,13 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "size": {
3
+ "longest_edge": 16777216,
4
+ "shortest_edge": 65536
5
+ },
6
+ "patch_size": 16,
7
+ "temporal_patch_size": 2,
8
+ "merge_size": 2,
9
+ "image_mean": [0.5, 0.5, 0.5],
10
+ "image_std": [0.5, 0.5, 0.5],
11
+ "processor_class": "Qwen3VLProcessor",
12
+ "image_processor_type": "Qwen2VLImageProcessorFast"
13
+ }
processor_config.json ADDED
@@ -0,0 +1,60 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "image_processor": {
3
+ "do_convert_rgb": true,
4
+ "do_normalize": true,
5
+ "do_rescale": true,
6
+ "do_resize": true,
7
+ "image_mean": [
8
+ 0.5,
9
+ 0.5,
10
+ 0.5
11
+ ],
12
+ "image_processor_type": "Qwen2VLImageProcessorFast",
13
+ "image_std": [
14
+ 0.5,
15
+ 0.5,
16
+ 0.5
17
+ ],
18
+ "merge_size": 2,
19
+ "patch_size": 16,
20
+ "resample": 3,
21
+ "rescale_factor": 0.00392156862745098,
22
+ "size": {
23
+ "longest_edge": 16777216,
24
+ "shortest_edge": 65536
25
+ },
26
+ "temporal_patch_size": 2
27
+ },
28
+ "processor_class": "Qwen3VLProcessor",
29
+ "video_processor": {
30
+ "do_convert_rgb": true,
31
+ "do_normalize": true,
32
+ "do_rescale": true,
33
+ "do_resize": true,
34
+ "do_sample_frames": true,
35
+ "fps": 2,
36
+ "image_mean": [
37
+ 0.5,
38
+ 0.5,
39
+ 0.5
40
+ ],
41
+ "image_std": [
42
+ 0.5,
43
+ 0.5,
44
+ 0.5
45
+ ],
46
+ "max_frames": 768,
47
+ "merge_size": 2,
48
+ "min_frames": 4,
49
+ "patch_size": 16,
50
+ "resample": 3,
51
+ "rescale_factor": 0.00392156862745098,
52
+ "return_metadata": false,
53
+ "size": {
54
+ "longest_edge": 25165824,
55
+ "shortest_edge": 4096
56
+ },
57
+ "temporal_patch_size": 2,
58
+ "video_processor_type": "Qwen3VLVideoProcessor"
59
+ }
60
+ }
quantization_config.json ADDED
The diff for this file is too large to render. See raw diff
 
tokenizer.json ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:06b9509352d2af50381ab2247e083b80d32d5c0aba91c272ca9ff729b6a0e523
3
+ size 19989325
tokenizer_config.json ADDED
@@ -0,0 +1,30 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "add_prefix_space": false,
3
+ "audio_bos_token": "<|audio_start|>",
4
+ "audio_eos_token": "<|audio_end|>",
5
+ "audio_token": "<|audio_pad|>",
6
+ "backend": "tokenizers",
7
+ "bos_token": null,
8
+ "clean_up_tokenization_spaces": false,
9
+ "eos_token": "<|im_end|>",
10
+ "errors": "replace",
11
+ "image_token": "<|image_pad|>",
12
+ "model_max_length": 262144,
13
+ "model_specific_special_tokens": {
14
+ "audio_bos_token": "<|audio_start|>",
15
+ "audio_eos_token": "<|audio_end|>",
16
+ "audio_token": "<|audio_pad|>",
17
+ "image_token": "<|image_pad|>",
18
+ "video_token": "<|video_pad|>",
19
+ "vision_bos_token": "<|vision_start|>",
20
+ "vision_eos_token": "<|vision_end|>"
21
+ },
22
+ "pad_token": "<|endoftext|>",
23
+ "pretokenize_regex": "(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\\r\\n\\p{L}\\p{N}]?[\\p{L}\\p{M}]+|\\p{N}| ?[^\\s\\p{L}\\p{M}\\p{N}]+[\\r\\n]*|\\s*[\\r\\n]+|\\s+(?!\\S)|\\s+",
24
+ "split_special_tokens": false,
25
+ "tokenizer_class": "Qwen2Tokenizer",
26
+ "unk_token": null,
27
+ "video_token": "<|video_pad|>",
28
+ "vision_bos_token": "<|vision_start|>",
29
+ "vision_eos_token": "<|vision_end|>"
30
+ }