sugarknight commited on
Commit
67c587c
·
verified ·
1 Parent(s): bc73656

Release Decision Mix V2 epoch 3: safetensors and CPU/GPU ONNX

Browse files

Dev-selected experimental checkpoint; private datasets excluded. FP32 top1 parity 311/311, FP16 310/311 with a bounded near-tie difference. Uncalibrated.

README.md CHANGED
@@ -4,61 +4,96 @@ language:
4
  - en
5
  - zh
6
  license: apache-2.0
 
7
  library_name: transformers
8
  pipeline_tag: text-classification
9
- base_model: knowledgator/gliclass-instruct-large-v1.0
10
  tags:
11
  - gliclass
12
  - choice-classification
13
  - experimental
 
14
  ---
15
 
16
- # ERABI Practical V1 (experimental)
17
 
18
- This is an **experimental**, uncalibrated choice-ranking model. It is not an official Jev model, a validated general-purpose reasoner, or an automatic decision-maker. The model ranks 2–16 user-supplied candidate texts for a natural-language context and question and returns all candidate probabilities through the [ERABI code](https://github.com/sugarkwork/erabi). Decisions should be reviewed by a person.
19
 
20
- ## Provenance
21
 
22
- - Base: [knowledgator/gliclass-instruct-large-v1.0](https://huggingface.co/knowledgator/gliclass-instruct-large-v1.0), Apache-2.0, 438,672,897 parameters.
23
- - First fine-tune: one epoch on 2,414 Practical V1 training records, peak learning rate 2.5e-6, 151 optimizer steps, microbatch 2, gradient accumulation 8, fp16 AMP.
24
- - Second, exploratory fine-tune (2026-09-23): one selected epoch on 175 privately held Exam-QA transformations mixed with 175 deterministic Practical V1 replay records, learning rate 1.5e-6, 22 optimizer steps, maximum training length 1,024 tokens.
25
- - Practical V1 data consists of original synthetic Japanese, English, and Simplified Chinese examples in six task families, generated and answer-blind rejudged with DeepSeek V4.1 Flash.
26
- - The Exam-QA source was filtered and transformed with the same DeepSeek model. Symbolic answer labels were mapped to source choice text. Ambiguous, multi-answer, figure-dependent, partial-credit, incomplete, or over-1,024-token items were skipped. Generated distractors were train-only; validation used source-provided choices only.
27
- - Exam-QA source records, transformed JSONL, and API responses are **not published** pending human review and source-by-source redistribution review. They are not claimed as human gold.
28
- - Data and training code: [GitHub repository](https://github.com/sugarkwork/erabi/tree/main/data/practical_v1) and [training script](https://github.com/sugarkwork/erabi/blob/main/scripts/train_practical_v1.py). Labels remain **unreviewed synthetic teacher agreement**, not human gold.
29
 
30
- ## Exploratory evaluation
31
 
32
- | Set | Frozen RC3 before this fine-tune | This checkpoint |
33
- |---|---:|---:|
34
- | Practical V1 dev, 399 cases | 59.90% | 77.19% |
35
- | Practical V1 held-out synthetic eval, 386 cases | 61.66% | 76.17% |
36
- | Existing RC3 Bridge, 480 cases | 88.75% | 88.54% |
37
 
38
- The Practical V1 eval set was used once after selecting by dev and existing-bridge results. Reading inference **regressed** from 54/71 to 50/71 despite aggregate gains. Candidate-order consistency on the existing bridge was 97.50%. These figures are not a benchmark of real-world correctness or Jev parity, because Practical V1 questions and labels come from the same teacher family. There is no independent human-verified final test, temperature calibration, or formal release approval for this checkpoint.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
39
 
40
- The current weights add the private Exam-QA experiment to that checkpoint:
41
 
42
- | Set | Before Exam-QA fine-tune | Current weights |
 
 
 
 
43
  |---|---:|---:|
44
- | Private Exam-QA validation, source choices only, 37 cases | 21.62% (8/37) | **24.32% (9/37)** |
45
- | Practical V1 dev, 399 cases | 77.19% (308/399) | **78.20% (312/399)** |
46
- | Existing RC3 Bridge, 480 cases | 88.54% (425/480) | **88.54% (425/480)** |
47
- | Practical V1 teacher-agreed eval, 386 cases | 76.17% (294/386) | 75.65% (292/386) |
48
 
49
- The Exam-QA gain is only **one additional correct item**, so it is weak exploratory evidence, not a claim of exam competence. The validation set has just nine source groups and has not been independently human-audited. The current safetensors SHA256 is `1ae38ef6c1103f8c021b0d3a974f3b1aedc42d89832d11761e74b95664216251`.
50
 
51
- ## Use
52
 
53
- ```bash
54
- python -m pip install erabi
55
- erabi predict --request request.json
56
- ```
 
 
 
 
57
 
58
- The repository contains three inference formats from the same checkpoint: `model.safetensors` (PyTorch), `onnx/fp32/model.onnx` (CPU), and `onnx/fp16/model.onnx` (NVIDIA GPU). ERABI 0.1.4 pins this World Choice checkpoint by default; 0.1.3 pins the earlier Exam-QA revision, and 0.1.2 pins the Practical V1-only revision. Upgrade with `python -m pip install --upgrade erabi`. `--model-format auto` downloads only the selected variant: FP32 ONNX for CPU with ONNX Runtime, FP16 ONNX for CUDA with CUDA Execution Provider, and otherwise PyTorch safetensors. Install the compatible `onnxruntime` (CPU) or `onnxruntime-gpu` (GPU) separately; do not install both in one environment. You can also select `--model-format pytorch`, `onnx-fp32`, or `onnx-fp16` explicitly.
59
 
60
- The newly exported ONNX FP32 and FP16 variants preserved the PyTorch top-ranked choice on 37/37 private Exam-QA validation cases, up to 853 input tokens. Experimental INT8 variants changed predictions substantially and are not distributed. The first invocation downloads the selected model; later invocations use the Hugging Face cache. Input and output JSON contracts and runtime recommendations are documented in the [ERABI README](https://github.com/sugarkwork/erabi#モデル形式の自動選択とおすすめ). The public ERABI runtime still defaults to a 512-token fail-closed contract. The weights were trained and experimentally checked at up to 1,024 tokens, but using that length requires changing both the runtime limit and preprocessing length while checking the untruncated input. Candidate probabilities are not calibrated confidence guarantees.
61
 
62
- ## License and limitations
63
 
64
- These fine-tuned weights derive from the Apache-2.0-licensed GLiClass base model and are distributed under Apache-2.0; see the [Apache License 2.0](https://www.apache.org/licenses/LICENSE-2.0) and the [base model card](https://huggingface.co/knowledgator/gliclass-instruct-large-v1.0). ERABI source code is separately MIT-licensed. Do not rely on this experimental model for high-stakes or unattended decisions.
 
4
  - en
5
  - zh
6
  license: apache-2.0
7
+ base_model: knowledgator/gliclass-instruct-large-v1.0
8
  library_name: transformers
9
  pipeline_tag: text-classification
 
10
  tags:
11
  - gliclass
12
  - choice-classification
13
  - experimental
14
+ - onnx
15
  ---
16
 
17
+ # ERABI — Decision Mix V2 experimental
18
 
19
+ Jevっぽい「状況を読んで候補を選ぶ」動きを、GLiClassで再現してみたローカル判断エンジンです。Jevの公式版・内部再現版ではありません。状況、判断基準、2〜16候補から全候補の確率分布を返し、ツール選択やNPC判断の実験に使います。文章生成やツール実行はしません。
20
 
21
+ 約438MパラメータのGLiClass追加学習モデルです。今回はWorld Choice checkpointからDecision Mix V2を3 epochs追加学習し、dev選択でepoch 3を採用しました。Hub名は互換性のため旧名のままです。旧モデルはGit revisionを指定して取得できます。
22
 
23
+ [Pythonコード・README](https://github.com/sugarkwork/erabi) / [元GLiClass](https://huggingface.co/knowledgator/gliclass-instruct-large-v1.0)
 
 
 
 
 
 
24
 
25
+ ## すぐ試す
26
 
27
+ Python 3.11以上。Windows PowerShellのCPU ONNX例です。
 
 
 
 
28
 
29
+ ```powershell
30
+ py -3.12 -m venv .venv
31
+ .\.venv\Scripts\Activate.ps1
32
+ python -m pip install --upgrade pip
33
+ python -m pip install torch --index-url https://download.pytorch.org/whl/cpu
34
+ python -m pip install erabi onnxruntime
35
+ ```
36
+
37
+ Linux/macOSは`python3 -m venv .venv`、`source .venv/bin/activate`を使います。GPUは環境に合うGPU用PyTorchを導入し、`onnxruntime`の代わりに`onnxruntime-gpu`を入れてください。両方を同じ環境に入れないでください。
38
+
39
+ 以下を`sample.py`として保存し、`python sample.py`で実行します。
40
+
41
+ ```python
42
+ from time import perf_counter
43
+ from erabi.model_loader import load_engine
44
+ from erabi.schema import ChoiceRequest
45
+
46
+ print("モデルを読み込みます(初回はダウンロード)...", flush=True)
47
+ t = perf_counter()
48
+ engine = load_engine(device="cpu", revision="main")
49
+ print(f"初期化: {perf_counter() - t:.2f}秒 / {engine.model_format}")
50
+ request = ChoiceRequest.from_dict({
51
+ "context": "ユーザーが今日の宮崎の天���を知りたいと言った。",
52
+ "question": "次に使う機能を選んでください。",
53
+ "choices": [
54
+ {"id": "chat", "text": "雑談をする"},
55
+ {"id": "web", "text": "Web検索する"},
56
+ {"id": "image", "text": "イラストを作成する"},
57
+ ],
58
+ })
59
+ print("推論します...", flush=True)
60
+ t = perf_counter()
61
+ result = engine.predict(request)
62
+ print(f"推論: {perf_counter() - t:.3f}秒")
63
+ print("選択:", result.best_candidate_id)
64
+ print("確率:", {c.id: round(c.probability, 4) for c in result.choices})
65
+ ```
66
 
67
+ GPUでは`device="cuda:0"`にします。CPUはONNX FP32、CUDA Providerが使えるGPUはONNX FP16を自動選択し、必要な形式だけ取得します。ONNX Runtimeがない場合はPyTorchです。`model_format="pytorch"`などで明示選択できます。`cache_dir="./model-cache"`または`ERABI_MODEL_CACHE_DIR`で保存先を指定できます。
68
 
69
+ PyPI 0.1.4の無指定モデルは旧World Choice版です。上例の`revision="main"`が今回の最新weightsを選びます。固定運用ではこのページのFiles and versionsからコミットIDを指定してください。
70
+
71
+ ## 同じ独自テスト311件の比較(2026-09-30)
72
+
73
+ | モデル | 正答数 | 正答率 |
74
  |---|---:|---:|
75
+ | ERABI World Choice(学習前) | 183/311 | 58.84% |
76
+ | ERABI Decision Mix V2(今回) | 210/311 | 67.52% |
77
+ | Laya 0.3.21 reviewed | 107/311 | 34.41% |
78
+ | Jev 1.13 remote API | 287/311 | 92.28% |
79
 
80
+ 同じcontext・question・候補順です。ERABIだけが同系列trainで追加学習済みなので、一般的な性能順位ではありません。ラベルはDeepSeek生成・正解非表示再判定の暫定合成教師で、全件の独立人手goldではありません。NPC・platformerは弱点で、platformerは追加学習前40.00%→37.50%へ低下しています。
81
 
82
+ train 3,238 / dev 353 / calibration 200 / final 311をcanonical groupで分割。calibrationは今回未使用。replay trainを混ぜて計5,361件/epoch、LR 1e-6、microbatch 1・勾配蓄積16、FP16 AMPで学習しました。finalはdev選択後に評価し、モデル選択・校正に使用していません。データ本体は非公開です。
83
 
84
+ ## 配布ファイルと制限
85
+
86
+ - `model.safetensors`:PyTorch FP32、約1.75GB
87
+ - `onnx/fp32/model.onnx`:CPU向け、約1.76GB
88
+ - `onnx/fp16/model.onnx`:CUDA向け、約0.88GB
89
+ - tokenizer/config:3形式で共通
90
+
91
+ weights SHA256: `959c7c38ff00c40f39ac5ad0e40344117e8caa14dadc511b2128e6f18ff06934`
92
 
93
+ <!-- RELEASE_VALIDATION -->
94
 
95
+ 311件の変換検証ではCPU FP32はPyTorchとTop-1 311/311一致、GPU FP16は310/311一致でした。近接した1件が丸め差で変わり、正答数はPyTorch/FP32が210、FP16が211です。最大確率差0.00470、元モデルでの該当2候補の差0.00816。FP16の完全一致は保証しません。校正は未実施です。
96
 
97
+ 入力全体は512トークンまで。日本語・英語・中国語を含む実験です。超過を黙って切り詰めません。出力は常にreview対象で、未校正の確率は正解・安全の保証ではありません。コマンド分類を実行許可に使わず、別の権限・パス・許可リスト制御を用意してください。
98
 
99
+ モデルはGLiClassのApache-2.0派生weightsです。ERABIのPythonコードはMITで、コードとweightsのライセンスは別です。用途に応じて元モデル・依存の条件を確認してください。
model.safetensors CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:20385ab5a4a40aefc754a5f5faba0463d1b68831c896fb92d95160c63f44bd4e
3
  size 1754746792
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:959c7c38ff00c40f39ac5ad0e40344117e8caa14dadc511b2128e6f18ff06934
3
  size 1754746792
onnx/fp16/model.onnx CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:56d5a87b79f60299be5f93facac524f8ba6af7e2896399eaa0de210c64ebe801
3
  size 879507317
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:466c2ee5cb9b525d9f5df9b34948bcfdd48770ec05fe533d6e42a528b1722431
3
  size 879507317
onnx/fp32/model.onnx CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:f35ba7019a32b823b116fea5dff839315bb0e982850d66352d531b6b0eaed770
3
  size 1756853324
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:3682944e03e8a1614b3d8766b166365408c3e603bd113ab0db5c0e3ba0dab95b
3
  size 1756853324
tokenizer.json CHANGED
@@ -2,7 +2,7 @@
2
  "version": "1.0",
3
  "truncation": {
4
  "direction": "Right",
5
- "max_length": 1024,
6
  "strategy": "LongestFirst",
7
  "stride": 0
8
  },
 
2
  "version": "1.0",
3
  "truncation": {
4
  "direction": "Right",
5
+ "max_length": 512,
6
  "strategy": "LongestFirst",
7
  "stride": 0
8
  },