gump2049 commited on
Commit
4a9fded
·
verified ·
1 Parent(s): c95c554

Initial private GGUF release (BF16, Q8_0, imatrix Q4_K_M)

Browse files
.gitattributes CHANGED
@@ -33,3 +33,7 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
 
 
 
 
 
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
+ APUS-OpenJev-v1-4B-BF16.gguf filter=lfs diff=lfs merge=lfs -text
37
+ APUS-OpenJev-v1-4B-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
38
+ APUS-OpenJev-v1-4B-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
39
+ imatrix.gguf filter=lfs diff=lfs merge=lfs -text
APUS-OpenJev-v1-4B-BF16.gguf ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:843fbf2825172bd8b89b57752c6d7bcbc19c82fae6100c262aeed1f063bc06d7
3
+ size 8424393568
APUS-OpenJev-v1-4B-Q4_K_M.gguf ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:3e77f041b48a6b28081071dc4dca8b047c10acf0a4b9b9b48603cef6b06ee186
3
+ size 2708804768
APUS-OpenJev-v1-4B-Q8_0.gguf ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:5e57075a169a76de5150f5c5defd805525ce4df86ae8865f70f128e15e57416c
3
+ size 4482403168
LICENSE ADDED
@@ -0,0 +1,202 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+
2
+ Apache License
3
+ Version 2.0, January 2004
4
+ http://www.apache.org/licenses/
5
+
6
+ TERMS AND CONDITIONS FOR USE, REPRODUCTION, AND DISTRIBUTION
7
+
8
+ 1. Definitions.
9
+
10
+ "License" shall mean the terms and conditions for use, reproduction,
11
+ and distribution as defined by Sections 1 through 9 of this document.
12
+
13
+ "Licensor" shall mean the copyright owner or entity authorized by
14
+ the copyright owner that is granting the License.
15
+
16
+ "Legal Entity" shall mean the union of the acting entity and all
17
+ other entities that control, are controlled by, or are under common
18
+ control with that entity. For the purposes of this definition,
19
+ "control" means (i) the power, direct or indirect, to cause the
20
+ direction or management of such entity, whether by contract or
21
+ otherwise, or (ii) ownership of fifty percent (50%) or more of the
22
+ outstanding shares, or (iii) beneficial ownership of such entity.
23
+
24
+ "You" (or "Your") shall mean an individual or Legal Entity
25
+ exercising permissions granted by this License.
26
+
27
+ "Source" form shall mean the preferred form for making modifications,
28
+ including but not limited to software source code, documentation
29
+ source, and configuration files.
30
+
31
+ "Object" form shall mean any form resulting from mechanical
32
+ transformation or translation of a Source form, including but
33
+ not limited to compiled object code, generated documentation,
34
+ and conversions to other media types.
35
+
36
+ "Work" shall mean the work of authorship, whether in Source or
37
+ Object form, made available under the License, as indicated by a
38
+ copyright notice that is included in or attached to the work
39
+ (an example is provided in the Appendix below).
40
+
41
+ "Derivative Works" shall mean any work, whether in Source or Object
42
+ form, that is based on (or derived from) the Work and for which the
43
+ editorial revisions, annotations, elaborations, or other modifications
44
+ represent, as a whole, an original work of authorship. For the purposes
45
+ of this License, Derivative Works shall not include works that remain
46
+ separable from, or merely link (or bind by name) to the interfaces of,
47
+ the Work and Derivative Works thereof.
48
+
49
+ "Contribution" shall mean any work of authorship, including
50
+ the original version of the Work and any modifications or additions
51
+ to that Work or Derivative Works thereof, that is intentionally
52
+ submitted to Licensor for inclusion in the Work by the copyright owner
53
+ or by an individual or Legal Entity authorized to submit on behalf of
54
+ the copyright owner. For the purposes of this definition, "submitted"
55
+ means any form of electronic, verbal, or written communication sent
56
+ to the Licensor or its representatives, including but not limited to
57
+ communication on electronic mailing lists, source code control systems,
58
+ and issue tracking systems that are managed by, or on behalf of, the
59
+ Licensor for the purpose of discussing and improving the Work, but
60
+ excluding communication that is conspicuously marked or otherwise
61
+ designated in writing by the copyright owner as "Not a Contribution."
62
+
63
+ "Contributor" shall mean Licensor and any individual or Legal Entity
64
+ on behalf of whom a Contribution has been received by Licensor and
65
+ subsequently incorporated within the Work.
66
+
67
+ 2. Grant of Copyright License. Subject to the terms and conditions of
68
+ this License, each Contributor hereby grants to You a perpetual,
69
+ worldwide, non-exclusive, no-charge, royalty-free, irrevocable
70
+ copyright license to reproduce, prepare Derivative Works of,
71
+ publicly display, publicly perform, sublicense, and distribute the
72
+ Work and such Derivative Works in Source or Object form.
73
+
74
+ 3. Grant of Patent License. Subject to the terms and conditions of
75
+ this License, each Contributor hereby grants to You a perpetual,
76
+ worldwide, non-exclusive, no-charge, royalty-free, irrevocable
77
+ (except as stated in this section) patent license to make, have made,
78
+ use, offer to sell, sell, import, and otherwise transfer the Work,
79
+ where such license applies only to those patent claims licensable
80
+ by such Contributor that are necessarily infringed by their
81
+ Contribution(s) alone or by combination of their Contribution(s)
82
+ with the Work to which such Contribution(s) was submitted. If You
83
+ institute patent litigation against any entity (including a
84
+ cross-claim or counterclaim in a lawsuit) alleging that the Work
85
+ or a Contribution incorporated within the Work constitutes direct
86
+ or contributory patent infringement, then any patent licenses
87
+ granted to You under this License for that Work shall terminate
88
+ as of the date such litigation is filed.
89
+
90
+ 4. Redistribution. You may reproduce and distribute copies of the
91
+ Work or Derivative Works thereof in any medium, with or without
92
+ modifications, and in Source or Object form, provided that You
93
+ meet the following conditions:
94
+
95
+ (a) You must give any other recipients of the Work or
96
+ Derivative Works a copy of this License; and
97
+
98
+ (b) You must cause any modified files to carry prominent notices
99
+ stating that You changed the files; and
100
+
101
+ (c) You must retain, in the Source form of any Derivative Works
102
+ that You distribute, all copyright, patent, trademark, and
103
+ attribution notices from the Source form of the Work,
104
+ excluding those notices that do not pertain to any part of
105
+ the Derivative Works; and
106
+
107
+ (d) If the Work includes a "NOTICE" text file as part of its
108
+ distribution, then any Derivative Works that You distribute must
109
+ include a readable copy of the attribution notices contained
110
+ within such NOTICE file, excluding those notices that do not
111
+ pertain to any part of the Derivative Works, in at least one
112
+ of the following places: within a NOTICE text file distributed
113
+ as part of the Derivative Works; within the Source form or
114
+ documentation, if provided along with the Derivative Works; or,
115
+ within a display generated by the Derivative Works, if and
116
+ wherever such third-party notices normally appear. The contents
117
+ of the NOTICE file are for informational purposes only and
118
+ do not modify the License. You may add Your own attribution
119
+ notices within Derivative Works that You distribute, alongside
120
+ or as an addendum to the NOTICE text from the Work, provided
121
+ that such additional attribution notices cannot be construed
122
+ as modifying the License.
123
+
124
+ You may add Your own copyright statement to Your modifications and
125
+ may provide additional or different license terms and conditions
126
+ for use, reproduction, or distribution of Your modifications, or
127
+ for any such Derivative Works as a whole, provided Your use,
128
+ reproduction, and distribution of the Work otherwise complies with
129
+ the conditions stated in this License.
130
+
131
+ 5. Submission of Contributions. Unless You explicitly state otherwise,
132
+ any Contribution intentionally submitted for inclusion in the Work
133
+ by You to the Licensor shall be under the terms and conditions of
134
+ this License, without any additional terms or conditions.
135
+ Notwithstanding the above, nothing herein shall supersede or modify
136
+ the terms of any separate license agreement you may have executed
137
+ with Licensor regarding such Contributions.
138
+
139
+ 6. Trademarks. This License does not grant permission to use the trade
140
+ names, trademarks, service marks, or product names of the Licensor,
141
+ except as required for reasonable and customary use in describing the
142
+ origin of the Work and reproducing the content of the NOTICE file.
143
+
144
+ 7. Disclaimer of Warranty. Unless required by applicable law or
145
+ agreed to in writing, Licensor provides the Work (and each
146
+ Contributor provides its Contributions) on an "AS IS" BASIS,
147
+ WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or
148
+ implied, including, without limitation, any warranties or conditions
149
+ of TITLE, NON-INFRINGEMENT, MERCHANTABILITY, or FITNESS FOR A
150
+ PARTICULAR PURPOSE. You are solely responsible for determining the
151
+ appropriateness of using or redistributing the Work and assume any
152
+ risks associated with Your exercise of permissions under this License.
153
+
154
+ 8. Limitation of Liability. In no event and under no legal theory,
155
+ whether in tort (including negligence), contract, or otherwise,
156
+ unless required by applicable law (such as deliberate and grossly
157
+ negligent acts) or agreed to in writing, shall any Contributor be
158
+ liable to You for damages, including any direct, indirect, special,
159
+ incidental, or consequential damages of any character arising as a
160
+ result of this License or out of the use or inability to use the
161
+ Work (including but not limited to damages for loss of goodwill,
162
+ work stoppage, computer failure or malfunction, or any and all
163
+ other commercial damages or losses), even if such Contributor
164
+ has been advised of the possibility of such damages.
165
+
166
+ 9. Accepting Warranty or Additional Liability. While redistributing
167
+ the Work or Derivative Works thereof, You may choose to offer,
168
+ and charge a fee for, acceptance of support, warranty, indemnity,
169
+ or other liability obligations and/or rights consistent with this
170
+ License. However, in accepting such obligations, You may act only
171
+ on Your own behalf and on Your sole responsibility, not on behalf
172
+ of any other Contributor, and only if You agree to indemnify,
173
+ defend, and hold each Contributor harmless for any liability
174
+ incurred by, or claims asserted against, such Contributor by reason
175
+ of your accepting any such warranty or additional liability.
176
+
177
+ END OF TERMS AND CONDITIONS
178
+
179
+ APPENDIX: How to apply the Apache License to your work.
180
+
181
+ To apply the Apache License to your work, attach the following
182
+ boilerplate notice, with the fields enclosed by brackets "[]"
183
+ replaced with your own identifying information. (Don't include
184
+ the brackets!) The text should be enclosed in the appropriate
185
+ comment syntax for the file format. We also recommend that a
186
+ file or class name and description of purpose be included on the
187
+ same "printed page" as the copyright notice for easier
188
+ identification within third-party archives.
189
+
190
+ Copyright 2026 Alibaba Cloud
191
+
192
+ Licensed under the Apache License, Version 2.0 (the "License");
193
+ you may not use this file except in compliance with the License.
194
+ You may obtain a copy of the License at
195
+
196
+ http://www.apache.org/licenses/LICENSE-2.0
197
+
198
+ Unless required by applicable law or agreed to in writing, software
199
+ distributed under the License is distributed on an "AS IS" BASIS,
200
+ WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
201
+ See the License for the specific language governing permissions and
202
+ limitations under the License.
Modelfile ADDED
@@ -0,0 +1,14 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # ollama create openjev-4b -f Modelfile (change FROM to pick another quant)
2
+ FROM ./APUS-OpenJev-v1-4B-Q8_0.gguf
3
+ TEMPLATE """<|im_start|>user
4
+ {{ .Prompt }}<|im_end|>
5
+ <|im_start|>assistant
6
+ <think>
7
+
8
+ </think>
9
+
10
+ """
11
+ PARAMETER temperature 0
12
+ PARAMETER num_predict 1
13
+ PARAMETER num_ctx 9216
14
+ PARAMETER stop <|im_end|>
README.md ADDED
@@ -0,0 +1,72 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ library_name: gguf
3
+ license: apache-2.0
4
+ base_model: apus-ailab/APUS-OpenJev-v1-4B
5
+ base_model_relation: quantized
6
+ pipeline_tag: text-generation
7
+ language:
8
+ - en
9
+ - zh
10
+ tags:
11
+ - apus-openjev
12
+ - decision-model
13
+ - gguf
14
+ - llama.cpp
15
+ - ollama
16
+ ---
17
+
18
+ # APUS-OpenJev-v1-4B-GGUF
19
+
20
+ [English](README.md) | [中文](README.zh-CN.md) · [Source model](https://huggingface.co/apus-ailab/APUS-OpenJev-v1-4B) · [Collection](https://huggingface.co/collections/apus-ailab/apus-openjev-v1-6ab1ee888eb002fcdd3a2825) · [MLX (Apple Silicon)](https://huggingface.co/apus-ailab/APUS-OpenJev-v1-4B-MLX-8bit)
21
+
22
+ GGUF conversions of [APUS-OpenJev-v1-4B](https://huggingface.co/apus-ailab/APUS-OpenJev-v1-4B) (revision `65797c526c27`) for **Ollama, llama.cpp and LM Studio** on Linux, Windows and macOS (Metal).
23
+
24
+ OpenJev is a **decision model**: each request supplies a state, an instruction and 2–16 candidates; the model answers with one candidate label (A–P) and the application reads the distribution over those labels. It is not a chat model.
25
+
26
+ ## Files and parity
27
+
28
+ Every file was scored on the [Frozen80 panel](https://huggingface.co/datasets/apus-ailab/APUS-OpenJev-Eval-Frozen80) with identical prompt tokens and compared with the HF BF16 release run by its own runtime (full depth, **66/80 · 82.50%**).
29
+
30
+ | File | Size | Frozen80 | Decisions = HF BF16 | Max Δp vs HF BF16 |
31
+ |---|---:|---:|---:|---:|
32
+ | [Q8_0](APUS-OpenJev-v1-4B-Q8_0.gguf) | 4.2 GiB | 67/80 · 83.75% | 79/80 | 0.1288 |
33
+ | [Q4_K_M](APUS-OpenJev-v1-4B-Q4_K_M.gguf) | 2.5 GiB | 67/80 · 83.75% | 77/80 | 0.8879 |
34
+ | [BF16](APUS-OpenJev-v1-4B-BF16.gguf) | 7.8 GiB | 66/80 · 82.50% | 80/80 | 0.0780 |
35
+
36
+ Q8_0 is the recommended default; Q4_K_M uses an importance matrix. Quantized scores are measured separately and do not inherit the BF16 result. Frozen80 is a reused development panel, not a blind benchmark. Per-file details: [evaluation/](evaluation/).
37
+
38
+ ## Ollama
39
+
40
+ ```bash
41
+ ollama run hf.co/apus-ailab/APUS-OpenJev-v1-4B-GGUF:Q8_0
42
+ ```
43
+
44
+ The repository ships the Ollama `template` (Qwen chat turn with thinking disabled) and `params` (`temperature 0`, `num_predict 1`). For a local import use the bundled [Modelfile](Modelfile). To get candidate probabilities, send the rendered prompt through the API:
45
+
46
+ ```bash
47
+ python examples/openjev_local.py --backend ollama --model hf.co/apus-ailab/APUS-OpenJev-v1-4B-GGUF:Q8_0
48
+ ```
49
+
50
+ Ollama returns at most 20 `top_logprobs` and cannot report named tokens, so the distribution is exact only when every candidate label is in the top 20 (Q8_0: **64/80** Frozen80 prompts); otherwise use the selected label or llama-server.
51
+
52
+ ## llama.cpp (exact distribution)
53
+
54
+ ```bash
55
+ llama-server -m APUS-OpenJev-v1-4B-Q8_0.gguf -c 9216 -ngl 999
56
+ python examples/openjev_local.py --backend llama-server --url http://127.0.0.1:8080
57
+ ```
58
+
59
+ `examples/openjev_local.py` renders prompts with [openjev_contracts.py](openjev_contracts.py), the same contract used in training.
60
+
61
+ ## Conversion
62
+
63
+ - llama.cpp `b11118` (`e6ab7c1a4`), `convert_hf_to_gguf.py --no-mtp` (the merged release has no MTP weights).
64
+ - Q4_K_M importance matrix: 448 training-course decisions, 64 per source, disjoint from Frozen80 ([details](evaluation/imatrix-calibration.json), [imatrix.gguf](imatrix.gguf)).
65
+ - All 1-D tensors, including GDN `A_log` / `dt_bias` and norms, stay F32 in every file ([check](evaluation/tensor-check.json)).
66
+ - Scope: full depth only (the 16/20-layer `low` exit is not available), text only (no vision tower), probabilities are **not calibrated**.
67
+
68
+ ## License
69
+
70
+ Apache-2.0, inherited from the source model; see [LICENSE](LICENSE). Base model: [Qwen/Qwen3.5-4B](https://huggingface.co/Qwen/Qwen3.5-4B).
71
+
72
+ **Authors:** gumpcheng ([xDAN2099](https://huggingface.co/xDAN2099)), zhangxu, [APUS AI-LAB](https://github.com/APUS-AI-Lab).
README.zh-CN.md ADDED
@@ -0,0 +1,72 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ library_name: gguf
3
+ license: apache-2.0
4
+ base_model: apus-ailab/APUS-OpenJev-v1-4B
5
+ base_model_relation: quantized
6
+ pipeline_tag: text-generation
7
+ language:
8
+ - en
9
+ - zh
10
+ tags:
11
+ - apus-openjev
12
+ - decision-model
13
+ - gguf
14
+ - llama.cpp
15
+ - ollama
16
+ ---
17
+
18
+ # APUS-OpenJev-v1-4B-GGUF
19
+
20
+ [English](README.md) | [中文](README.zh-CN.md) · [源模型](https://huggingface.co/apus-ailab/APUS-OpenJev-v1-4B) · [Collection](https://huggingface.co/collections/apus-ailab/apus-openjev-v1-6ab1ee888eb002fcdd3a2825) · [MLX(Apple Silicon)](https://huggingface.co/apus-ailab/APUS-OpenJev-v1-4B-MLX-8bit)
21
+
22
+ [APUS-OpenJev-v1-4B](https://huggingface.co/apus-ailab/APUS-OpenJev-v1-4B)(revision `65797c526c27`)的 GGUF 版本,适用于 Linux / Windows / macOS(Metal)上的 **Ollama、llama.cpp、LM Studio**。
23
+
24
+ OpenJev 是**决策模型**:每个请求给出状态、指令和 2–16 个候选,模型回答一个候选标签(A–P),应用读取这些标签上的概率分布。它不是聊天模型。
25
+
26
+ ## 文件与一致性
27
+
28
+ 每个文件都在 [Frozen80](https://huggingface.co/datasets/apus-ailab/APUS-OpenJev-Eval-Frozen80) 上用完全相同的 prompt token 评测,并与 HF BF16 发布版(其自带 runtime、完整深度,**66/80 · 82.50%**)对比。
29
+
30
+ | 文件 | 大小 | Frozen80 | 与 HF BF16 决策一致 | 相对 HF BF16 最大 Δp |
31
+ |---|---:|---:|---:|---:|
32
+ | [Q8_0](APUS-OpenJev-v1-4B-Q8_0.gguf) | 4.2 GiB | 67/80 · 83.75% | 79/80 | 0.1288 |
33
+ | [Q4_K_M](APUS-OpenJev-v1-4B-Q4_K_M.gguf) | 2.5 GiB | 67/80 · 83.75% | 77/80 | 0.8879 |
34
+ | [BF16](APUS-OpenJev-v1-4B-BF16.gguf) | 7.8 GiB | 66/80 · 82.50% | 80/80 | 0.0780 |
35
+
36
+ 默认推荐 Q8_0;Q4_K_M 使用 importance matrix。量化成绩单独测量,不沿用 BF16 成绩。Frozen80 是复用的开发面板,不是盲测。明细见 [evaluation/](evaluation/)。
37
+
38
+ ## Ollama
39
+
40
+ ```bash
41
+ ollama run hf.co/apus-ailab/APUS-OpenJev-v1-4B-GGUF:Q8_0
42
+ ```
43
+
44
+ 仓库自带 Ollama 的 `template`(关闭 thinking 的 Qwen 对话格式)和 `params`(`temperature 0`、`num_predict 1`)。本地导入可用 [Modelfile](Modelfile)。需要候选概率时,通过 API 发送渲染好的 prompt:
45
+
46
+ ```bash
47
+ python examples/openjev_local.py --backend ollama --model hf.co/apus-ailab/APUS-OpenJev-v1-4B-GGUF:Q8_0
48
+ ```
49
+
50
+ Ollama 最多返回 20 个 `top_logprobs`,且不能指定 token,因此只有全部候选标签都在前 20 时分布才是精确的(Q8_0:Frozen80 中 **64/80** 题);否则请使用所选标签,或改用 llama-server。
51
+
52
+ ## llama.cpp(精确分布)
53
+
54
+ ```bash
55
+ llama-server -m APUS-OpenJev-v1-4B-Q8_0.gguf -c 9216 -ngl 999
56
+ python examples/openjev_local.py --backend llama-server --url http://127.0.0.1:8080
57
+ ```
58
+
59
+ `examples/openjev_local.py` 使用 [openjev_contracts.py](openjev_contracts.py) 渲染 prompt,与训练时的格式一致。
60
+
61
+ ## 转换说明
62
+
63
+ - llama.cpp `b11118`(`e6ab7c1a4`),`convert_hf_to_gguf.py --no-mtp`(merged 发布版不含 MTP 权重)。
64
+ - Q4_K_M 的 importance matrix:训练集中 448 条决策,每个来源 64 条,与 Frozen80 无重叠([明细](evaluation/imatrix-calibration.json)、[imatrix.gguf](imatrix.gguf))。
65
+ - 所有 1-D 张量(含 GDN `A_log` / `dt_bias` 和各类 norm)在每个文件中都保持 F32([检查结果](evaluation/tensor-check.json))。
66
+ - 范围:仅完整深度(不提供 16/20 层 `low` 出口),仅文本(不含视觉塔),概率**未经校准**。
67
+
68
+ ## 许可
69
+
70
+ Apache-2.0,继承自源模型,见 [LICENSE](LICENSE)。基座模型:[Qwen/Qwen3.5-4B](https://huggingface.co/Qwen/Qwen3.5-4B)。
71
+
72
+ **作者:** gumpcheng([xDAN2099](https://huggingface.co/xDAN2099))、zhangxu、[APUS AI-LAB](https://github.com/APUS-AI-Lab)。
SHA256SUMS ADDED
@@ -0,0 +1,4 @@
 
 
 
 
 
1
+ 843fbf2825172bd8b89b57752c6d7bcbc19c82fae6100c262aeed1f063bc06d7 ./APUS-OpenJev-v1-4B-BF16.gguf
2
+ 3e77f041b48a6b28081071dc4dca8b047c10acf0a4b9b9b48603cef6b06ee186 ./APUS-OpenJev-v1-4B-Q4_K_M.gguf
3
+ 5e57075a169a76de5150f5c5defd805525ce4df86ae8865f70f128e15e57416c ./APUS-OpenJev-v1-4B-Q8_0.gguf
4
+ 76b68bd161e5b866e00ff7ec184460ef44622efa33e7f61e54d30c603b731a36 ./imatrix.gguf
conversion-provenance.json ADDED
@@ -0,0 +1,23 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "llama_cpp": {
3
+ "tag": "b11118",
4
+ "commit": "e6ab7c1a41054a888ada952eab4c886444c2f5ad"
5
+ },
6
+ "ollama": "0.34.3",
7
+ "mlx": "0.32.2",
8
+ "mlx_lm": "0.31.3",
9
+ "convert_env": {
10
+ "torch": "2.11.0+cpu",
11
+ "transformers": "4.57.6"
12
+ },
13
+ "reference_runtime": {
14
+ "torch": "2.8.0+cu128",
15
+ "transformers": "5.16.1",
16
+ "effort": "high"
17
+ },
18
+ "hardware": "NVIDIA RTX PRO 6000 Blackwell 96GB (Linux); Apple M5 24GB for Mac checks",
19
+ "date": "2026-09-23",
20
+ "source_repository": "apus-ailab/APUS-OpenJev-v1-4B",
21
+ "source_revision": "65797c526c27c4d24f564333779162cd4a64328e",
22
+ "base_model": "Qwen/Qwen3.5-4B"
23
+ }
evaluation/imatrix-calibration.json ADDED
@@ -0,0 +1,19 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "train_file": "/workspace/ms-swift-jev/data/registered-course-v1-20260920/decision.train.jsonl",
3
+ "train_sha256": "172b7427c4dcfa1270750142d2ad9ab6d42ee1350f532137329798a5af7b92c2",
4
+ "seed": 20260923,
5
+ "per_source": 64,
6
+ "records": 448,
7
+ "sources": {
8
+ "osunlp/Mind2Web": 64,
9
+ "oracle_only": 64,
10
+ "mnli": 64,
11
+ "sgd": 64,
12
+ "google-research-datasets/go_emotions": 64,
13
+ "nvidia/HelpSteer3": 64,
14
+ "boolq": 64
15
+ },
16
+ "frozen80_overlap": 0,
17
+ "output_sha256": "b4acbf8ecf6c38b789d809f5c23b44d8ea584a8731c425a497642c40b8782187",
18
+ "characters": 1297640
19
+ }
evaluation/llamacpp-BF16.summary.json ADDED
@@ -0,0 +1,36 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "variant": "4B-BF16",
3
+ "rows": 80,
4
+ "correct": 66,
5
+ "reference_correct": 66,
6
+ "decision_agreement": 80,
7
+ "changed_panel_indexes": [],
8
+ "max_abs_probability_delta": 0.07798188169586251,
9
+ "mean_abs_probability_delta": 0.0036633186839281073,
10
+ "all_candidates_in_top20": 64,
11
+ "reference_choice_in_top20": 80,
12
+ "min_candidate_mass_full_vocab": 0.8304116117153292,
13
+ "seconds": 10.42,
14
+ "by_family": {
15
+ "browser": {
16
+ "rows": 16,
17
+ "correct": 14
18
+ },
19
+ "hs3": {
20
+ "rows": 16,
21
+ "correct": 12
22
+ },
23
+ "boolq": {
24
+ "rows": 16,
25
+ "correct": 16
26
+ },
27
+ "mnli": {
28
+ "rows": 16,
29
+ "correct": 11
30
+ },
31
+ "score": {
32
+ "rows": 16,
33
+ "correct": 13
34
+ }
35
+ }
36
+ }
evaluation/llamacpp-Q4_K_M.summary.json ADDED
@@ -0,0 +1,40 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "variant": "4B-Q4_K_M",
3
+ "rows": 80,
4
+ "correct": 67,
5
+ "reference_correct": 66,
6
+ "decision_agreement": 77,
7
+ "changed_panel_indexes": [
8
+ 59,
9
+ 64,
10
+ 72
11
+ ],
12
+ "max_abs_probability_delta": 0.8879217194689062,
13
+ "mean_abs_probability_delta": 0.03401718616064667,
14
+ "all_candidates_in_top20": 64,
15
+ "reference_choice_in_top20": 80,
16
+ "min_candidate_mass_full_vocab": 0.8305545665216829,
17
+ "seconds": 10.65,
18
+ "by_family": {
19
+ "browser": {
20
+ "rows": 16,
21
+ "correct": 14
22
+ },
23
+ "hs3": {
24
+ "rows": 16,
25
+ "correct": 12
26
+ },
27
+ "boolq": {
28
+ "rows": 16,
29
+ "correct": 16
30
+ },
31
+ "mnli": {
32
+ "rows": 16,
33
+ "correct": 10
34
+ },
35
+ "score": {
36
+ "rows": 16,
37
+ "correct": 15
38
+ }
39
+ }
40
+ }
evaluation/llamacpp-Q8_0.summary.json ADDED
@@ -0,0 +1,38 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "variant": "4B-Q8_0",
3
+ "rows": 80,
4
+ "correct": 67,
5
+ "reference_correct": 66,
6
+ "decision_agreement": 79,
7
+ "changed_panel_indexes": [
8
+ 72
9
+ ],
10
+ "max_abs_probability_delta": 0.12883429532209645,
11
+ "mean_abs_probability_delta": 0.006993976078097623,
12
+ "all_candidates_in_top20": 64,
13
+ "reference_choice_in_top20": 80,
14
+ "min_candidate_mass_full_vocab": 0.7869114107711379,
15
+ "seconds": 12.91,
16
+ "by_family": {
17
+ "browser": {
18
+ "rows": 16,
19
+ "correct": 14
20
+ },
21
+ "hs3": {
22
+ "rows": 16,
23
+ "correct": 12
24
+ },
25
+ "boolq": {
26
+ "rows": 16,
27
+ "correct": 16
28
+ },
29
+ "mnli": {
30
+ "rows": 16,
31
+ "correct": 11
32
+ },
33
+ "score": {
34
+ "rows": 16,
35
+ "correct": 14
36
+ }
37
+ }
38
+ }
evaluation/mlx-8bit.summary.json ADDED
@@ -0,0 +1,35 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "variant": "4B-MLX-8bit",
3
+ "backend": "mlx 0.32.2 on Linux x86_64 (Device(gpu, 0))",
4
+ "rows": 80,
5
+ "correct": 66,
6
+ "reference_correct": 66,
7
+ "decision_agreement": 80,
8
+ "changed_panel_indexes": [],
9
+ "max_abs_probability_delta": 0.1201455295085907,
10
+ "mean_abs_probability_delta": 0.005359249900720897,
11
+ "peak_memory_gb": 6.09,
12
+ "seconds": 105.89,
13
+ "by_family": {
14
+ "browser": {
15
+ "rows": 16,
16
+ "correct": 14
17
+ },
18
+ "hs3": {
19
+ "rows": 16,
20
+ "correct": 12
21
+ },
22
+ "boolq": {
23
+ "rows": 16,
24
+ "correct": 16
25
+ },
26
+ "mnli": {
27
+ "rows": 16,
28
+ "correct": 11
29
+ },
30
+ "score": {
31
+ "rows": 16,
32
+ "correct": 13
33
+ }
34
+ }
35
+ }
evaluation/ppl-crosscheck.summary.json ADDED
@@ -0,0 +1,15 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "hf_bf16": [
3
+ 22.1184,
4
+ 9.8833,
5
+ 7.1177
6
+ ],
7
+ "gguf_bf16_imatrix": [
8
+ 22.0867,
9
+ 9.9745,
10
+ 7.1646
11
+ ],
12
+ "metric": "cumulative PPL over the second half of each 4096-token chunk of the imatrix calibration text (llama-imatrix convention), first 3 chunks",
13
+ "hf_script": "hf_chunk_ppl.py",
14
+ "note": "diagnostic cross-check of GGUF numerics against HF; high 9B values are a property of checkpoint-3000 on raw long text, reproduced by HF"
15
+ }
evaluation/reference.summary.json ADDED
@@ -0,0 +1,6 @@
 
 
 
 
 
 
 
1
+ {
2
+ "model_dir": "src/4B",
3
+ "effort": "high",
4
+ "rows": 80,
5
+ "correct": 66
6
+ }
evaluation/tensor-check.json ADDED
@@ -0,0 +1,45 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "APUS-OpenJev-v1-4B-BF16.gguf": {
3
+ "architecture": "qwen35",
4
+ "tensors": 426,
5
+ "tensor_types": {
6
+ "BF16": 249,
7
+ "F32": 177
8
+ },
9
+ "one_dim_tensors_checked": 153,
10
+ "gdn_decay_tensor_types": {
11
+ "F32": 48
12
+ },
13
+ "one_dim_non_f32": [],
14
+ "bytes": 8424393568
15
+ },
16
+ "APUS-OpenJev-v1-4B-Q8_0.gguf": {
17
+ "architecture": "qwen35",
18
+ "tensors": 426,
19
+ "tensor_types": {
20
+ "F32": 177,
21
+ "Q8_0": 249
22
+ },
23
+ "one_dim_tensors_checked": 153,
24
+ "gdn_decay_tensor_types": {
25
+ "F32": 48
26
+ },
27
+ "one_dim_non_f32": [],
28
+ "bytes": 4482403168
29
+ },
30
+ "APUS-OpenJev-v1-4B-Q4_K_M.gguf": {
31
+ "architecture": "qwen35",
32
+ "tensors": 426,
33
+ "tensor_types": {
34
+ "F32": 177,
35
+ "Q6_K": 33,
36
+ "Q4_K": 216
37
+ },
38
+ "one_dim_tensors_checked": 153,
39
+ "gdn_decay_tensor_types": {
40
+ "F32": 48
41
+ },
42
+ "one_dim_non_f32": [],
43
+ "bytes": 2708804768
44
+ }
45
+ }
examples/openjev_local.py ADDED
@@ -0,0 +1,109 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Call an APUS-OpenJev GGUF through Ollama or llama-server (standard library only).
2
+
3
+ python examples/openjev_local.py --backend ollama --model hf.co/apus-ailab/APUS-OpenJev-v1-4B-GGUF:Q8_0
4
+ python examples/openjev_local.py --backend llama-server --url http://127.0.0.1:8080
5
+
6
+ The model scores caller-supplied candidates; it does not write JSON. The prompt is
7
+ rendered by ``openjev_contracts.py`` (identical to the training contract) and wrapped
8
+ in the no-thinking Qwen chat turn. Candidate labels are A..P, one token each.
9
+
10
+ * llama-server returns the exact candidate distribution (full-vocabulary logprobs).
11
+ * Ollama exposes at most 20 ``top_logprobs``. The distribution is exact when every
12
+ candidate label ranks in the top 20; otherwise only the selected label is reliable
13
+ and ``distribution_complete`` is False.
14
+ """
15
+
16
+ import argparse
17
+ import json
18
+ import math
19
+ import sys
20
+ import urllib.request
21
+ from pathlib import Path
22
+
23
+ sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
24
+ from openjev_contracts import label_mapping, render_prompt # noqa: E402
25
+
26
+ CHAT = "<|im_start|>user\n{prompt}<|im_end|>\n<|im_start|>assistant\n<think>\n\n</think>\n\n"
27
+
28
+
29
+ def _post(url, payload):
30
+ request = urllib.request.Request(url, json.dumps(payload).encode(), {"Content-Type": "application/json"})
31
+ with urllib.request.urlopen(request, timeout=600) as response:
32
+ return json.loads(response.read())
33
+
34
+
35
+ def _normalize(labels, logprobs):
36
+ peak = max(logprobs[label] for label in labels)
37
+ weights = {label: math.exp(logprobs[label] - peak) for label in labels}
38
+ total = sum(weights.values())
39
+ return {label: weight / total for label, weight in weights.items()}
40
+
41
+
42
+ def decide(request, backend="ollama", url=None, model=None):
43
+ mapping = label_mapping(request) # validates the request; label -> candidate id
44
+ labels = list(mapping)
45
+ prompt = CHAT.format(prompt=render_prompt(request))
46
+ if backend == "ollama":
47
+ data = _post(
48
+ (url or "http://127.0.0.1:11434") + "/api/generate",
49
+ {
50
+ "model": model,
51
+ "prompt": prompt,
52
+ "raw": True,
53
+ "stream": False,
54
+ "logprobs": True,
55
+ "top_logprobs": 20,
56
+ "options": {"temperature": 0, "num_predict": 1, "num_ctx": 9216},
57
+ },
58
+ )
59
+ (position,) = data["logprobs"]
60
+ logprobs = {entry["token"]: entry["logprob"] for entry in position["top_logprobs"]}
61
+ selected = data["response"]
62
+ elif backend == "llama-server":
63
+ data = _post(
64
+ (url or "http://127.0.0.1:8080") + "/completion",
65
+ {"prompt": prompt, "n_predict": 1, "n_probs": 1024, "temperature": 0, "cache_prompt": True},
66
+ )
67
+ (position,) = data["completion_probabilities"]
68
+ logprobs = {entry["token"]: entry["logprob"] for entry in position["top_logprobs"]}
69
+ selected = position["token"]
70
+ else:
71
+ raise ValueError("backend must be ollama or llama-server")
72
+
73
+ complete = all(label in logprobs for label in labels)
74
+ result = {"type": request["primitive"], "distribution_complete": complete}
75
+ if complete:
76
+ probabilities = _normalize(labels, logprobs)
77
+ result["probabilities"] = {mapping[label]: p for label, p in probabilities.items()}
78
+ selected = max(probabilities, key=probabilities.get)
79
+ if selected not in mapping:
80
+ raise RuntimeError(f"model produced {selected!r}, not a candidate label")
81
+ if request["primitive"] == "choice":
82
+ result["choice"] = mapping[selected]
83
+ elif complete:
84
+ result["yes_probability"] = result["probabilities"]["yes"]
85
+ else:
86
+ result["answer"] = mapping[selected]
87
+ return result
88
+
89
+
90
+ EXAMPLE = {
91
+ "id": "support-731",
92
+ "group_id": "support-731",
93
+ "primitive": "choice",
94
+ "state": "Order 731 was delivered. The customer confirms that the issue is resolved.",
95
+ "instructions": "Select the next support action.",
96
+ "criteria": [
97
+ {"id": "close_ticket", "description": "Close the ticket as resolved."},
98
+ {"id": "escalate", "description": "Escalate to a human agent."},
99
+ {"id": "refund", "description": "Issue a refund."},
100
+ ],
101
+ }
102
+
103
+ if __name__ == "__main__":
104
+ parser = argparse.ArgumentParser()
105
+ parser.add_argument("--backend", choices=("ollama", "llama-server"), default="ollama")
106
+ parser.add_argument("--url")
107
+ parser.add_argument("--model", help="Ollama model name")
108
+ args = parser.parse_args()
109
+ print(json.dumps(decide(EXAMPLE, args.backend, args.url, args.model), indent=2))
imatrix.gguf ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:76b68bd161e5b866e00ff7ec184460ef44622efa33e7f61e54d30c603b731a36
3
+ size 3626528
openjev_contracts.py ADDED
@@ -0,0 +1,102 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Small shared contract. Prompts use a strict whitelist of input fields."""
2
+
3
+ import json
4
+ import math
5
+
6
+ PROMPT_VERSION = "jev.dynamic.prompt.v2"
7
+ LABELS = tuple("ABCDEFGHIJKLMNOP")
8
+ BINARY_CRITERIA = [
9
+ {"id": "yes", "description": "The stated proposition is true."},
10
+ {"id": "no", "description": "The stated proposition is false."},
11
+ ]
12
+
13
+
14
+ def validate_request(record):
15
+ for key in ("id", "group_id", "state", "instructions"):
16
+ if not isinstance(record.get(key), str) or not record[key].strip():
17
+ raise ValueError(f"{key} must be a nonempty string")
18
+ if record.get("primitive") not in ("choice", "noul", "score_level"):
19
+ raise ValueError("unsupported primitive")
20
+ criteria = record.get("criteria")
21
+ if not isinstance(criteria, list) or not 2 <= len(criteria) <= len(LABELS):
22
+ raise ValueError("criteria must contain 2..16 candidates")
23
+ ids = []
24
+ for candidate in criteria:
25
+ if not isinstance(candidate, dict):
26
+ raise TypeError("candidate must be an object")
27
+ for key in ("id", "description"):
28
+ if not isinstance(candidate.get(key), str) or not candidate[key].strip():
29
+ raise ValueError(f"candidate {key} must be nonempty")
30
+ ids.append(candidate["id"])
31
+ if len(set(ids)) != len(ids):
32
+ raise ValueError("duplicate candidate ids")
33
+ if record["primitive"] != "choice" and criteria != BINARY_CRITERIA:
34
+ raise ValueError("noul and score_level require canonical yes/no criteria")
35
+
36
+
37
+ def validate_record(record):
38
+ validate_request(record)
39
+ if record.get("gold") not in [c["id"] for c in record["criteria"]]:
40
+ raise ValueError("gold must be a candidate id")
41
+ if not isinstance(record.get("provenance"), dict):
42
+ raise TypeError("provenance must be an object")
43
+
44
+
45
+ def label_mapping(record):
46
+ validate_request(record)
47
+ return dict(zip(LABELS, (c["id"] for c in record["criteria"])))
48
+
49
+
50
+ def render_prompt_parts(record):
51
+ """Text prefix/suffix; callers MUST check tokenizer boundary equivalence."""
52
+ validate_request(record)
53
+ prefix = "Shared state:\n" + record["state"] + "\n\n"
54
+ task = {
55
+ "primitive": record["primitive"],
56
+ "instructions": record["instructions"],
57
+ "criteria": [
58
+ {"label": label, "description": candidate["description"]}
59
+ for label, candidate in zip(LABELS, record["criteria"])
60
+ ],
61
+ }
62
+ suffix = json.dumps(task, ensure_ascii=False, sort_keys=True)
63
+ suffix += (
64
+ "\nReturn only the selected letter: "
65
+ + ", ".join(LABELS[: len(record["criteria"])])
66
+ + ".\nAnswer:"
67
+ )
68
+ return prefix, suffix
69
+
70
+
71
+ def render_prompt(record):
72
+ return "".join(render_prompt_parts(record))
73
+
74
+
75
+ def to_messages(record):
76
+ validate_record(record)
77
+ inverse = {candidate: label for label, candidate in label_mapping(record).items()}
78
+ return {
79
+ "messages": [
80
+ {"role": "user", "content": render_prompt(record)},
81
+ {"role": "assistant", "content": inverse[record["gold"]]},
82
+ ]
83
+ }
84
+
85
+
86
+ def format_response(record, probabilities):
87
+ """Map ordered candidate probabilities; score_level is NOT aggregate Score."""
88
+ mapping = label_mapping(record)
89
+ values = list(probabilities)
90
+ if len(values) != len(mapping) or any(
91
+ not math.isfinite(p) or p < 0 or p > 1 for p in values
92
+ ):
93
+ raise ValueError("invalid probabilities")
94
+ if not math.isclose(sum(values), 1, abs_tol=1e-5):
95
+ raise ValueError("probabilities must sum to one")
96
+ distribution = dict(zip(mapping.values(), values))
97
+ result = {"type": record["primitive"], "probabilities": distribution}
98
+ if record["primitive"] == "choice":
99
+ result["choice"] = max(distribution, key=distribution.get)
100
+ else:
101
+ result["yes_probability"] = distribution["yes"]
102
+ return result
params ADDED
@@ -0,0 +1,8 @@
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "temperature": 0,
3
+ "num_predict": 1,
4
+ "num_ctx": 9216,
5
+ "stop": [
6
+ "<|im_end|>"
7
+ ]
8
+ }
template ADDED
@@ -0,0 +1,7 @@
 
 
 
 
 
 
 
 
1
+ <|im_start|>user
2
+ {{ .Prompt }}<|im_end|>
3
+ <|im_start|>assistant
4
+ <think>
5
+
6
+ </think>
7
+