Text Classification
Transformers
ONNX
Safetensors
Japanese
English
Chinese
GLiClass
gliclass
choice-classification
experimental
Instructions to use sugarknight/erabi-practical-v1-experimental with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use sugarknight/erabi-practical-v1-experimental with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="sugarknight/erabi-practical-v1-experimental")# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("sugarknight/erabi-practical-v1-experimental", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Release Decision Mix V2 epoch 3: safetensors and CPU/GPU ONNX
Browse filesDev-selected experimental checkpoint; private datasets excluded. FP32 top1 parity 311/311, FP16 310/311 with a bounded near-tie difference. Uncalibrated.
- README.md +69 -34
- model.safetensors +1 -1
- onnx/fp16/model.onnx +1 -1
- onnx/fp32/model.onnx +1 -1
- tokenizer.json +1 -1
README.md
CHANGED
|
@@ -4,61 +4,96 @@ language:
|
|
| 4 |
- en
|
| 5 |
- zh
|
| 6 |
license: apache-2.0
|
|
|
|
| 7 |
library_name: transformers
|
| 8 |
pipeline_tag: text-classification
|
| 9 |
-
base_model: knowledgator/gliclass-instruct-large-v1.0
|
| 10 |
tags:
|
| 11 |
- gliclass
|
| 12 |
- choice-classification
|
| 13 |
- experimental
|
|
|
|
| 14 |
---
|
| 15 |
|
| 16 |
-
# ERABI
|
| 17 |
|
| 18 |
-
|
| 19 |
|
| 20 |
-
|
| 21 |
|
| 22 |
-
|
| 23 |
-
- First fine-tune: one epoch on 2,414 Practical V1 training records, peak learning rate 2.5e-6, 151 optimizer steps, microbatch 2, gradient accumulation 8, fp16 AMP.
|
| 24 |
-
- Second, exploratory fine-tune (2026-09-23): one selected epoch on 175 privately held Exam-QA transformations mixed with 175 deterministic Practical V1 replay records, learning rate 1.5e-6, 22 optimizer steps, maximum training length 1,024 tokens.
|
| 25 |
-
- Practical V1 data consists of original synthetic Japanese, English, and Simplified Chinese examples in six task families, generated and answer-blind rejudged with DeepSeek V4.1 Flash.
|
| 26 |
-
- The Exam-QA source was filtered and transformed with the same DeepSeek model. Symbolic answer labels were mapped to source choice text. Ambiguous, multi-answer, figure-dependent, partial-credit, incomplete, or over-1,024-token items were skipped. Generated distractors were train-only; validation used source-provided choices only.
|
| 27 |
-
- Exam-QA source records, transformed JSONL, and API responses are **not published** pending human review and source-by-source redistribution review. They are not claimed as human gold.
|
| 28 |
-
- Data and training code: [GitHub repository](https://github.com/sugarkwork/erabi/tree/main/data/practical_v1) and [training script](https://github.com/sugarkwork/erabi/blob/main/scripts/train_practical_v1.py). Labels remain **unreviewed synthetic teacher agreement**, not human gold.
|
| 29 |
|
| 30 |
-
##
|
| 31 |
|
| 32 |
-
|
| 33 |
-
|---|---:|---:|
|
| 34 |
-
| Practical V1 dev, 399 cases | 59.90% | 77.19% |
|
| 35 |
-
| Practical V1 held-out synthetic eval, 386 cases | 61.66% | 76.17% |
|
| 36 |
-
| Existing RC3 Bridge, 480 cases | 88.75% | 88.54% |
|
| 37 |
|
| 38 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 39 |
|
| 40 |
-
|
| 41 |
|
| 42 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
| 43 |
|---|---:|---:|
|
| 44 |
-
|
|
| 45 |
-
|
|
| 46 |
-
|
|
| 47 |
-
|
|
| 48 |
|
| 49 |
-
|
| 50 |
|
| 51 |
-
|
| 52 |
|
| 53 |
-
|
| 54 |
-
|
| 55 |
-
|
| 56 |
-
``
|
|
|
|
|
|
|
|
|
|
|
|
|
| 57 |
|
| 58 |
-
|
| 59 |
|
| 60 |
-
|
| 61 |
|
| 62 |
-
|
| 63 |
|
| 64 |
-
|
|
|
|
| 4 |
- en
|
| 5 |
- zh
|
| 6 |
license: apache-2.0
|
| 7 |
+
base_model: knowledgator/gliclass-instruct-large-v1.0
|
| 8 |
library_name: transformers
|
| 9 |
pipeline_tag: text-classification
|
|
|
|
| 10 |
tags:
|
| 11 |
- gliclass
|
| 12 |
- choice-classification
|
| 13 |
- experimental
|
| 14 |
+
- onnx
|
| 15 |
---
|
| 16 |
|
| 17 |
+
# ERABI — Decision Mix V2 experimental
|
| 18 |
|
| 19 |
+
Jevっぽい「状況を読んで候補を選ぶ」動きを、GLiClassで再現してみたローカル判断エンジンです。Jevの公式版・内部再現版ではありません。状況、判断基準、2〜16候補から全候補の確率分布を返し、ツール選択やNPC判断の実験に使います。文章生成やツール実行はしません。
|
| 20 |
|
| 21 |
+
約438MパラメータのGLiClass追加学習モデルです。今回はWorld Choice checkpointからDecision Mix V2を3 epochs追加学習し、dev選択でepoch 3を採用しました。Hub名は互換性のため旧名のままです。旧モデルはGit revisionを指定して取得できます。
|
| 22 |
|
| 23 |
+
[Pythonコード・README](https://github.com/sugarkwork/erabi) / [元GLiClass](https://huggingface.co/knowledgator/gliclass-instruct-large-v1.0)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 24 |
|
| 25 |
+
## すぐ試す
|
| 26 |
|
| 27 |
+
Python 3.11以上。Windows PowerShellのCPU ONNX例です。
|
|
|
|
|
|
|
|
|
|
|
|
|
| 28 |
|
| 29 |
+
```powershell
|
| 30 |
+
py -3.12 -m venv .venv
|
| 31 |
+
.\.venv\Scripts\Activate.ps1
|
| 32 |
+
python -m pip install --upgrade pip
|
| 33 |
+
python -m pip install torch --index-url https://download.pytorch.org/whl/cpu
|
| 34 |
+
python -m pip install erabi onnxruntime
|
| 35 |
+
```
|
| 36 |
+
|
| 37 |
+
Linux/macOSは`python3 -m venv .venv`、`source .venv/bin/activate`を使います。GPUは環境に合うGPU用PyTorchを導入し、`onnxruntime`の代わりに`onnxruntime-gpu`を入れてください。両方を同じ環境に入れないでください。
|
| 38 |
+
|
| 39 |
+
以下を`sample.py`として保存し、`python sample.py`で実行します。
|
| 40 |
+
|
| 41 |
+
```python
|
| 42 |
+
from time import perf_counter
|
| 43 |
+
from erabi.model_loader import load_engine
|
| 44 |
+
from erabi.schema import ChoiceRequest
|
| 45 |
+
|
| 46 |
+
print("モデルを読み込みます(初回はダウンロード)...", flush=True)
|
| 47 |
+
t = perf_counter()
|
| 48 |
+
engine = load_engine(device="cpu", revision="main")
|
| 49 |
+
print(f"初期化: {perf_counter() - t:.2f}秒 / {engine.model_format}")
|
| 50 |
+
request = ChoiceRequest.from_dict({
|
| 51 |
+
"context": "ユーザーが今日の宮崎の天���を知りたいと言った。",
|
| 52 |
+
"question": "次に使う機能を選んでください。",
|
| 53 |
+
"choices": [
|
| 54 |
+
{"id": "chat", "text": "雑談をする"},
|
| 55 |
+
{"id": "web", "text": "Web検索する"},
|
| 56 |
+
{"id": "image", "text": "イラストを作成する"},
|
| 57 |
+
],
|
| 58 |
+
})
|
| 59 |
+
print("推論します...", flush=True)
|
| 60 |
+
t = perf_counter()
|
| 61 |
+
result = engine.predict(request)
|
| 62 |
+
print(f"推論: {perf_counter() - t:.3f}秒")
|
| 63 |
+
print("選択:", result.best_candidate_id)
|
| 64 |
+
print("確率:", {c.id: round(c.probability, 4) for c in result.choices})
|
| 65 |
+
```
|
| 66 |
|
| 67 |
+
GPUでは`device="cuda:0"`にします。CPUはONNX FP32、CUDA Providerが使えるGPUはONNX FP16を自動選択し、必要な形式だけ取得します。ONNX Runtimeがない場合はPyTorchです。`model_format="pytorch"`などで明示選択できます。`cache_dir="./model-cache"`または`ERABI_MODEL_CACHE_DIR`で保存先を指定できます。
|
| 68 |
|
| 69 |
+
PyPI 0.1.4の無指定モデルは旧World Choice版です。上例の`revision="main"`が今回の最新weightsを選びます。固定運用ではこのページのFiles and versionsからコミットIDを指定してください。
|
| 70 |
+
|
| 71 |
+
## 同じ独自テスト311件の比較(2026-09-30)
|
| 72 |
+
|
| 73 |
+
| モデル | 正答数 | 正答率 |
|
| 74 |
|---|---:|---:|
|
| 75 |
+
| ERABI World Choice(学習前) | 183/311 | 58.84% |
|
| 76 |
+
| ERABI Decision Mix V2(今回) | 210/311 | 67.52% |
|
| 77 |
+
| Laya 0.3.21 reviewed | 107/311 | 34.41% |
|
| 78 |
+
| Jev 1.13 remote API | 287/311 | 92.28% |
|
| 79 |
|
| 80 |
+
同じcontext・question・候補順です。ERABIだけが同系列trainで追加学習済みなので、一般的な性能順位ではありません。ラベルはDeepSeek生成・正解非表示再判定の暫定合成教師で、全件の独立人手goldではありません。NPC・platformerは弱点で、platformerは追加学習前40.00%→37.50%へ低下しています。
|
| 81 |
|
| 82 |
+
train 3,238 / dev 353 / calibration 200 / final 311をcanonical groupで分割。calibrationは今回未使用。replay trainを混ぜて計5,361件/epoch、LR 1e-6、microbatch 1・勾配蓄積16、FP16 AMPで学習しました。finalはdev選択後に評価し、モデル選択・校正に使用していません。データ本体は非公開です。
|
| 83 |
|
| 84 |
+
## 配布ファイルと制限
|
| 85 |
+
|
| 86 |
+
- `model.safetensors`:PyTorch FP32、約1.75GB
|
| 87 |
+
- `onnx/fp32/model.onnx`:CPU向け、約1.76GB
|
| 88 |
+
- `onnx/fp16/model.onnx`:CUDA向け、約0.88GB
|
| 89 |
+
- tokenizer/config:3形式で共通
|
| 90 |
+
|
| 91 |
+
weights SHA256: `959c7c38ff00c40f39ac5ad0e40344117e8caa14dadc511b2128e6f18ff06934`
|
| 92 |
|
| 93 |
+
<!-- RELEASE_VALIDATION -->
|
| 94 |
|
| 95 |
+
311件の変換検証ではCPU FP32はPyTorchとTop-1 311/311一致、GPU FP16は310/311一致でした。近接した1件が丸め差で変わり、正答数はPyTorch/FP32が210、FP16が211です。最大確率差0.00470、元モデルでの該当2候補の差0.00816。FP16の完全一致は保証しません。校正は未実施です。
|
| 96 |
|
| 97 |
+
入力全体は512トークンまで。日本語・英語・中国語を含む実験です。超過を黙って切り詰めません。出力は常にreview対象で、未校正の確率は正解・安全の保証ではありません。コマンド分類を実行許可に使わず、別の権限・パス・許可リスト制御を用意してください。
|
| 98 |
|
| 99 |
+
モデルはGLiClassのApache-2.0派生weightsです。ERABIのPythonコードはMITで、コードとweightsのライセンスは別です。用途に応じて元モデル・依存の条件を確認してください。
|
model.safetensors
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
size 1754746792
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:959c7c38ff00c40f39ac5ad0e40344117e8caa14dadc511b2128e6f18ff06934
|
| 3 |
size 1754746792
|
onnx/fp16/model.onnx
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
size 879507317
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:466c2ee5cb9b525d9f5df9b34948bcfdd48770ec05fe533d6e42a528b1722431
|
| 3 |
size 879507317
|
onnx/fp32/model.onnx
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
size 1756853324
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:3682944e03e8a1614b3d8766b166365408c3e603bd113ab0db5c0e3ba0dab95b
|
| 3 |
size 1756853324
|
tokenizer.json
CHANGED
|
@@ -2,7 +2,7 @@
|
|
| 2 |
"version": "1.0",
|
| 3 |
"truncation": {
|
| 4 |
"direction": "Right",
|
| 5 |
-
"max_length":
|
| 6 |
"strategy": "LongestFirst",
|
| 7 |
"stride": 0
|
| 8 |
},
|
|
|
|
| 2 |
"version": "1.0",
|
| 3 |
"truncation": {
|
| 4 |
"direction": "Right",
|
| 5 |
+
"max_length": 512,
|
| 6 |
"strategy": "LongestFirst",
|
| 7 |
"stride": 0
|
| 8 |
},
|