sugarknight commited on
Commit
72ef021
·
verified ·
1 Parent(s): 99d8086

Update to World Choice V1 checkpoint with PyTorch and ONNX variants

Browse files

Selected epoch 3; publish matching safetensors, CPU ONNX FP32, GPU ONNX FP16, and updated limitations. ONNX Top-1 parity 404/404 for both formats.

README.md CHANGED
@@ -11,54 +11,62 @@ tags:
11
  - gliclass
12
  - choice-classification
13
  - experimental
 
14
  ---
15
 
16
- # ERABI Practical V1 (experimental)
17
 
18
- This is an **experimental**, uncalibrated choice-ranking model. It is not an official Jev model, a validated general-purpose reasoner, or an automatic decision-maker. The model ranks 2–16 user-supplied candidate texts for a natural-language context and question and returns all candidate probabilities through the [ERABI code](https://github.com/sugarkwork/erabi). Decisions should be reviewed by a person.
19
 
20
- ## Provenance
21
 
22
- - Base: [knowledgator/gliclass-instruct-large-v1.0](https://huggingface.co/knowledgator/gliclass-instruct-large-v1.0), Apache-2.0, 438,672,897 parameters.
23
- - First fine-tune: one epoch on 2,414 Practical V1 training records, peak learning rate 2.5e-6, 151 optimizer steps, microbatch 2, gradient accumulation 8, fp16 AMP.
24
- - Second, exploratory fine-tune (2026-09-23): one selected epoch on 175 privately held Exam-QA transformations mixed with 175 deterministic Practical V1 replay records, learning rate 1.5e-6, 22 optimizer steps, maximum training length 1,024 tokens.
25
- - Practical V1 data consists of original synthetic Japanese, English, and Simplified Chinese examples in six task families, generated and answer-blind rejudged with DeepSeek V4.1 Flash.
26
- - The Exam-QA source was filtered and transformed with the same DeepSeek model. Symbolic answer labels were mapped to source choice text. Ambiguous, multi-answer, figure-dependent, partial-credit, incomplete, or over-1,024-token items were skipped. Generated distractors were train-only; validation used source-provided choices only.
27
- - Exam-QA source records, transformed JSONL, and API responses are **not published** pending human review and source-by-source redistribution review. They are not claimed as human gold.
28
- - Data and training code: [GitHub repository](https://github.com/sugarkwork/erabi/tree/main/data/practical_v1) and [training script](https://github.com/sugarkwork/erabi/blob/main/scripts/train_practical_v1.py). Labels remain **unreviewed synthetic teacher agreement**, not human gold.
29
 
30
  ## Exploratory evaluation
31
 
32
- | Set | Frozen RC3 before this fine-tune | This checkpoint |
33
  |---|---:|---:|
34
- | Practical V1 dev, 399 cases | 59.90% | 77.19% |
35
- | Practical V1 held-out synthetic eval, 386 cases | 61.66% | 76.17% |
36
- | Existing RC3 Bridge, 480 cases | 88.75% | 88.54% |
 
 
 
37
 
38
- The Practical V1 eval set was used once after selecting by dev and existing-bridge results. Reading inference **regressed** from 54/71 to 50/71 despite aggregate gains. Candidate-order consistency on the existing bridge was 97.50%. These figures are not a benchmark of real-world correctness or Jev parity, because Practical V1 questions and labels come from the same teacher family. There is no independent human-verified final test, temperature calibration, or formal release approval for this checkpoint.
39
 
40
- The current weights add the private Exam-QA experiment to that checkpoint:
41
 
42
- | Set | Before Exam-QA fine-tune | Current weights |
43
- |---|---:|---:|
44
- | Private Exam-QA validation, source choices only, 37 cases | 21.62% (8/37) | **24.32% (9/37)** |
45
- | Practical V1 dev, 399 cases | 77.19% (308/399) | **78.20% (312/399)** |
46
- | Existing RC3 Bridge, 480 cases | 88.54% (425/480) | **88.54% (425/480)** |
47
- | Practical V1 teacher-agreed eval, 386 cases | 76.17% (294/386) | 75.65% (292/386) |
48
 
49
- The Exam-QA gain is only **one additional correct item**, so it is weak exploratory evidence, not a claim of exam competence. The validation set has just nine source groups and has not been independently human-audited. The current safetensors SHA256 is `1ae38ef6c1103f8c021b0d3a974f3b1aedc42d89832d11761e74b95664216251`.
50
 
51
- ## Use
 
 
 
 
52
 
53
  ```bash
54
- python -m pip install erabi
55
- erabi predict --request request.json
 
 
 
56
  ```
57
 
58
- The repository contains three inference formats from the same checkpoint: `model.safetensors` (PyTorch), `onnx/fp32/model.onnx` (CPU), and `onnx/fp16/model.onnx` (NVIDIA GPU). ERABI 0.1.3 pins this updated checkpoint by default; 0.1.2 pins the earlier Practical V1-only revision. Upgrade with `python -m pip install --upgrade erabi`. `--model-format auto` downloads only the selected variant: FP32 ONNX for CPU with ONNX Runtime, FP16 ONNX for CUDA with CUDA Execution Provider, and otherwise PyTorch safetensors. Install the compatible `onnxruntime` (CPU) or `onnxruntime-gpu` (GPU) separately; do not install both in one environment. You can also select `--model-format pytorch`, `onnx-fp32`, or `onnx-fp16` explicitly.
 
 
59
 
60
- The newly exported ONNX FP32 and FP16 variants preserved the PyTorch top-ranked choice on 37/37 private Exam-QA validation cases, up to 853 input tokens. Experimental INT8 variants changed predictions substantially and are not distributed. The first invocation downloads the selected model; later invocations use the Hugging Face cache. Input and output JSON contracts and runtime recommendations are documented in the [ERABI README](https://github.com/sugarkwork/erabi#モデル形式の自動選択とおすすめ). The public ERABI runtime still defaults to a 512-token fail-closed contract. The weights were trained and experimentally checked at up to 1,024 tokens, but using that length requires changing both the runtime limit and preprocessing length while checking the untruncated input. Candidate probabilities are not calibrated confidence guarantees.
61
 
62
  ## License and limitations
63
 
64
- These fine-tuned weights derive from the Apache-2.0-licensed GLiClass base model and are distributed under Apache-2.0; see the [Apache License 2.0](https://www.apache.org/licenses/LICENSE-2.0) and the [base model card](https://huggingface.co/knowledgator/gliclass-instruct-large-v1.0). ERABI source code is separately MIT-licensed. Do not rely on this experimental model for high-stakes or unattended decisions.
 
11
  - gliclass
12
  - choice-classification
13
  - experimental
14
+ - onnx
15
  ---
16
 
17
+ # ERABI World Choice V1 (experimental)
18
 
19
+ This is an **experimental**, uncalibrated choice-ranking model. It is not an official Jev model, a validated general-purpose reasoner, or an unattended decision-maker. Given a natural-language context, a question, and 2–16 user-supplied candidate texts, it returns a score distribution over all candidates through the [ERABI code](https://github.com/sugarkwork/erabi). Important decisions require human review.
20
 
21
+ ## Model lineage and training
22
 
23
+ - Base model: [knowledgator/gliclass-instruct-large-v1.0](https://huggingface.co/knowledgator/gliclass-instruct-large-v1.0), Apache-2.0, approximately 438M parameters.
24
+ - Earlier continual fine-tuning used the private Practical V1 and filtered Exam-QA training splits. Those datasets and labels are provisional and are not claimed as independent human gold.
25
+ - The current 2026-09-27 stage trained for three epochs on 2,825 World Choice V1 training rows, with 2,414 Practical V1 and 175 Exam-QA replay rows (5,414 rows per epoch). Settings were learning rate 1e-6, microbatch 1, gradient accumulation 16, fp16 AMP, and a strict 1,024-token experimental input limit.
26
+ - Epoch 3 was selected using World Choice development accuracy and legacy development retention. Final/reference results did not select the epoch.
27
+ - World Choice V1 contains Japanese, English, and Simplified Chinese synthetic choice tasks covering everyday attributes, animals, quantity, tool routing, conversation memory, robot faults, navigation/resources, RPG survival, NPC interactions, and multi-rule worlds.
28
+
29
+ The World Choice labels were generated and answer-blind rejudged by the same external model family. A fixed review sample found ambiguous and underspecified items, and some categories contain option-length shortcuts. The data therefore remains private and provisional rather than human-verified gold.
30
 
31
  ## Exploratory evaluation
32
 
33
+ | Set | Previous Exam V1 weights | Current weights |
34
  |---|---:|---:|
35
+ | World Choice dev, 575 rows | 47.30% (272/575) | **61.04% (351/575)** |
36
+ | World Choice final reference, 404 rows | 48.76% (197/404) | **63.12% (255/404)** |
37
+ | Practical V1 final, 386 rows | 75.65% (292/386) | **87.56% (338/386)** |
38
+ | Existing RC3 Bridge, 480 rows | 88.54% (425/480) | 88.12% (423/480) |
39
+ | Exam final, 51 rows | 37.25% (19/51) | 41.18% (21/51) |
40
+ | Weakness final, 46 rows | 56.52% (26/46) | 54.35% (25/46) |
41
 
42
+ The World Choice final rows had already been scored before this fine-tune. They are an exposed reference set, not an untouched final acceptance test. The labels are provisional synthetic agreement, and the Weakness set regressed by one item. These results do not establish real-world correctness, exam competence, Jev parity, or production safety.
43
 
44
+ The current PyTorch weights SHA256 is `20385ab5a4a40aefc754a5f5faba0463d1b68831c896fb92d95160c63f44bd4e`.
45
 
46
+ ## Formats and use
 
 
 
 
 
47
 
48
+ The repository provides three formats from the same checkpoint:
49
 
50
+ - `model.safetensors`: PyTorch
51
+ - `onnx/fp32/model.onnx`: CPU-oriented ONNX FP32
52
+ - `onnx/fp16/model.onnx`: NVIDIA GPU-oriented ONNX FP16
53
+
54
+ On all 404 World Choice final-reference rows (maximum observed input 747 tokens), both ONNX formats preserved the PyTorch top-ranked choice: FP32 404/404 and FP16 404/404. Maximum raw-logit differences were 0.0000546 for FP32 and 0.0220 for FP16. Model SHA256 values are `f35ba7019a32b823b116fea5dff839315bb0e982850d66352d531b6b0eaed7770` (FP32) and `56d5a87b79f60299be5f93facac524f8ba6af7e2896399eaa0de210c64ebe8001` (FP16).
55
 
56
  ```bash
57
+ python -m pip install --upgrade erabi
58
+ erabi predict --request request.json \
59
+ --model-id sugarknight/erabi-practical-v1-experimental \
60
+ --revision main \
61
+ --model-format auto
62
  ```
63
 
64
+ Install a compatible `onnxruntime` for CPU or `onnxruntime-gpu` for CUDA separately; do not install both in one environment. With `--model-format auto`, ERABI selects FP32 ONNX for CPU when ONNX Runtime is available, FP16 ONNX for CUDA with the CUDA Execution Provider, and otherwise PyTorch safetensors. Only the selected large model file is downloaded.
65
+
66
+ The currently released ERABI package pins an older verified model revision by default. Until a later package release updates that pin, explicitly pass this model ID with `--revision main` to use these new weights.
67
 
68
+ The public runtime remains fail-closed at 512 tokens by default. Training and parity testing here used up to 1,024 tokens, but longer use requires changing both the runtime limit and preprocessing length while validating the untruncated input. Candidate probabilities are uncalibrated scores, not confidence guarantees.
69
 
70
  ## License and limitations
71
 
72
+ These weights derive from the Apache-2.0-licensed GLiClass base model and are distributed under Apache-2.0. ERABI source code is separately MIT-licensed. Do not use this experimental model for high-stakes or unattended decisions. No private training data, prompts, API responses, or credentials are included in this repository.
model.safetensors CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:1ae38ef6c1103f8c021b0d3a974f3b1aedc42d89832d11761e74b95664216251
3
  size 1754746792
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:20385ab5a4a40aefc754a5f5faba0463d1b68831c896fb92d95160c63f44bd4e
3
  size 1754746792
onnx/fp16/model.onnx CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:4b67d0ec581bb86815d0cebffe24dcdac16b976fedc1ad391b51e9c3eacb60ff
3
  size 879507317
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:56d5a87b79f60299be5f93facac524f8ba6af7e2896399eaa0de210c64ebe801
3
  size 879507317
onnx/fp32/model.onnx CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:4ce73286fc49cc472c7160ac23c4e6a90924f576bc4b075e3053d5245b889350
3
  size 1756853324
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:f35ba7019a32b823b116fea5dff839315bb0e982850d66352d531b6b0eaed770
3
  size 1756853324
tokenizer_config.json CHANGED
@@ -7,7 +7,7 @@
7
  "do_lower_case": false,
8
  "eos_token": "[SEP]",
9
  "is_local": true,
10
- "local_files_only": false,
11
  "mask_token": "[MASK]",
12
  "max_length": 4096,
13
  "model_max_length": 1000000000000000019884624838656,
 
7
  "do_lower_case": false,
8
  "eos_token": "[SEP]",
9
  "is_local": true,
10
+ "local_files_only": true,
11
  "mask_token": "[MASK]",
12
  "max_length": 4096,
13
  "model_max_length": 1000000000000000019884624838656,