Text Generation
MLX
Safetensors
jev-style
English
qwen3_5
decision-model
classification
calibration
qwen3.5
single-prefill
conversational
Instructions to use chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-MLX-bf16 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-MLX-bf16 with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-MLX-bf16") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - jev-style
How to use chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-MLX-bf16 with jev-style:
# Apple silicon pip install "jev-style[mlx]"
from jev_style import JevStyle, noul, choice js = JevStyle.from_pretrained("chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-MLX-bf16") out = js.decide("I was charged twice for one order.", { "billing": noul("This message is about billing."), "team": choice("Which team should handle it?", ["billing", "shipping", "tech"]), }) print(out["answers"]["team"]["choice"]) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- MLX LM
How to use chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-MLX-bf16 with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-MLX-bf16"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-MLX-bf16" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-MLX-bf16", "messages": [ {"role": "user", "content": "Hello"} ] }' - Atomic Chat
Add benchmark, calibration and robustness charts with result-first model cards
Browse files- .gitattributes +3 -0
- README.md +97 -36
- SHA256SUMS.json +30 -2
- evaluation/chart_data.json +183 -0
- figures/benchmark.png +3 -0
- figures/benchmark.svg +578 -0
- figures/calibration.png +3 -0
- figures/calibration.svg +748 -0
- figures/robustness.png +3 -0
- figures/robustness.svg +338 -0
.gitattributes
CHANGED
|
@@ -34,3 +34,6 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
|
|
| 34 |
*.zst filter=lfs diff=lfs merge=lfs -text
|
| 35 |
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
| 36 |
tokenizer.json filter=lfs diff=lfs merge=lfs -text
|
|
|
|
|
|
|
|
|
|
|
|
| 34 |
*.zst filter=lfs diff=lfs merge=lfs -text
|
| 35 |
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
| 36 |
tokenizer.json filter=lfs diff=lfs merge=lfs -text
|
| 37 |
+
figures/benchmark.png filter=lfs diff=lfs merge=lfs -text
|
| 38 |
+
figures/calibration.png filter=lfs diff=lfs merge=lfs -text
|
| 39 |
+
figures/robustness.png filter=lfs diff=lfs merge=lfs -text
|
README.md
CHANGED
|
@@ -14,22 +14,108 @@ tags:
|
|
| 14 |
- jev-style
|
| 15 |
- single-prefill
|
| 16 |
---
|
| 17 |
-
# Jev-Style-Qwen3.5-2B-Decision v2
|
| 18 |
|
| 19 |
-
A
|
| 20 |
|
| 21 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
| 22 |
|
| 23 |
-
|
| 24 |
|
| 25 |
-
|
| 26 |
-
- **99.6% choice agreement** with merged CUDA BF16 on the frozen 500-decision deployment subset.
|
| 27 |
-
- Independent calibration on 3,100 records, automatically applied by the client.
|
| 28 |
-
- BF16 model weights with FP32 normalization gains and FP32 declared-option projection in the decision client.
|
| 29 |
|
| 30 |
-
|
| 31 |
|
| 32 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 33 |
|
| 34 |
```bash
|
| 35 |
python -m pip install -U huggingface_hub
|
|
@@ -55,33 +141,8 @@ print(result["choice"])
|
|
| 55 |
print(result["probabilities"])
|
| 56 |
```
|
| 57 |
|
| 58 |
-
Validated with `mlx==0.32.2` and `mlx-lm==0.31.3`. The companion `calibration.json` is loaded automatically. The
|
| 59 |
-
|
| 60 |
-
## Highlights
|
| 61 |
-
|
| 62 |
-
- **81.27% macro accuracy** for the released merged BF16 model across 11 real-label task groups (3,277 decisions).
|
| 63 |
-
- **9 of 12 task-group accuracy point estimates ahead of English Laya** in the fixed CUDA reference comparison.
|
| 64 |
-
- **+4.53 percentage points over Jev-Style v1** and **+6.12 points over English Laya** in reference macro accuracy on the same evaluation panel.
|
| 65 |
-
- **18.4% lower NLL and 20.0% lower Brier score** than English Laya in the reference comparison.
|
| 66 |
-
- **6.0% option-permutation flip rate**, compared with 9.25% for v1 and 12.0% for English Laya, on 400 Choice/Bool decisions.
|
| 67 |
-
- **One H100 80GB, 36.9 minutes of main training**, with a 2B-class text backbone and rank-32 LoRA.
|
| 68 |
-
|
| 69 |
-
The comparison uses the frozen English task panel and the CUDA reference structure. Deployment variants are measured separately below. The 9/12 count describes task-level point estimates.
|
| 70 |
-
|
| 71 |
-
## Reference evaluation
|
| 72 |
-
|
| 73 |
-
Real-label results are macro-averaged with equal task weights. All three models use the same calibration records and global temperature-fitting objective.
|
| 74 |
-
|
| 75 |
-
| Metric | Jev-Style v1 | English Laya | Jev-Style v2 reference |
|
| 76 |
-
|---|---:|---:|---:|
|
| 77 |
-
| Accuracy ↑ | 76.68% | 75.09% | **81.20%** |
|
| 78 |
-
| Macro-F1 ↑ | 75.42% | 73.45% | **79.78%** |
|
| 79 |
-
| NLL ↓ | 0.5752 | 0.6318 | **0.5154** |
|
| 80 |
-
| Brier ↓ | 0.3290 | 0.3482 | **0.2787** |
|
| 81 |
-
|
| 82 |
-
Accuracy improvements have paired 95% intervals of **+3.58 to +5.52 points vs v1** and **+4.64 to +7.52 points vs English Laya** within this frozen task panel.
|
| 83 |
|
| 84 |
-
The panel covers sentiment, news, natural-language inference, question answering, emotion and email classification. The twelfth task group contains 2,000 teacher-reference typed decisions from 400 states and is reported separately from the real-label macro. Per-task results, all probability metrics, robustness measurements and baseline sensitivity results are supplied in the evaluation files.
|
| 85 |
|
| 86 |
## Decision interface
|
| 87 |
|
|
|
|
| 14 |
- jev-style
|
| 15 |
- single-prefill
|
| 16 |
---
|
| 17 |
+
# Jev-Style-Qwen3.5-2B-Decision v2 (MLX BF16)
|
| 18 |
|
| 19 |
+
A **Jev-style decision model** for classification, routing and typed choices. Give it a state, a question and a list of options; one prefill returns a selected option **with calibrated probabilities**.
|
| 20 |
|
| 21 |
+
| Build | Weight size | Inference |
|
| 22 |
+
|---|---:|---|
|
| 23 |
+
| [HF BF16](https://huggingface.co/chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2) | 3.76 GB | Transformers + decision client |
|
| 24 |
+
| [GGUF Q8_0](https://huggingface.co/chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF) | 2.01 GB | llama.cpp + decision client |
|
| 25 |
+
| **MLX BF16 · this repository** | 3.76 GB | Apple Silicon + native MLX client |
|
| 26 |
|
| 27 |
+
**Download this build:** [model.safetensors](https://huggingface.co/chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-MLX-bf16/resolve/main/model.safetensors?download=true). The repository also includes its calibration, inference client and evaluation records.
|
| 28 |
|
| 29 |
+
## Results
|
|
|
|
|
|
|
|
|
|
| 30 |
|
| 31 |
+
**81.20% macro accuracy on the fixed English reference panel**, compared with 76.68% for v1 and 75.09% for English Laya. The results below use the CUDA reference structure: 11 real-label task groups, 3,277 decisions, equal task weights, and the same 3,100-record calibration split. Results for the released deployment formats appear further below.
|
| 32 |
|
| 33 |
+
| Metric | Jev-Style v1 | English Laya | **Jev-Style v2** |
|
| 34 |
+
|---|---:|---:|---:|
|
| 35 |
+
| Accuracy ↑ | 76.68% | 75.09% | **81.20%** |
|
| 36 |
+
| Macro-F1 ↑ | 75.42% | 73.45% | **79.78%** |
|
| 37 |
+
| Negative log-likelihood ↓ | 0.5752 | 0.6318 | **0.5154** |
|
| 38 |
+
| Brier score ↓ | 0.3290 | 0.3482 | **0.2787** |
|
| 39 |
+
|
| 40 |
+

|
| 41 |
+
|
| 42 |
+
- **Higher accuracy:** +4.53 percentage points over v1 and +6.12 over English Laya; paired 95% intervals are [+3.58, +5.52] and [+4.64, +7.52] points, respectively, within this fixed panel.
|
| 43 |
+
- **Broader task coverage:** accuracy point estimates ahead of English Laya in **9 of 12 task groups**, including the separately scored teacher-reference typed-decisions group.
|
| 44 |
+
- **Better probability quality against English Laya:** **18.4% lower NLL**, **20.0% lower Brier score**, and **26.4% lower task-macro ECE**.
|
| 45 |
+
- **Efficient adaptation:** **36.9 minutes of main training on one H100 80GB**, using rank-32 LoRA on a 2B-class text backbone.
|
| 46 |
+
|
| 47 |
+
## Calibration
|
| 48 |
+
|
| 49 |
+
The reliability diagram plots the v2 model's stated confidence against observed correctness. Every real-label evaluation decision is included; the histogram shows how many predictions fall in each confidence bin. Error bars show Wilson 95% intervals. The accompanying ECE comparison averages per-task calibration errors.
|
| 50 |
+
|
| 51 |
+

|
| 52 |
+
|
| 53 |
+
Temperature is fitted on the calibration split. HF and MLX clients apply the supplied calibration automatically; the calibrated GGUF file incorporates it in the final normalization tensor.
|
| 54 |
+
|
| 55 |
+
## Robustness
|
| 56 |
+
|
| 57 |
+
**Option-order flip rate is halved relative to English Laya**, with **80.00% accuracy after permutation** on the same 400 Choice/Bool decisions. Semantic options are mapped back to their original identities before scoring.
|
| 58 |
+
|
| 59 |
+
| Option-permutation test | Jev-Style v1 | English Laya | **Jev-Style v2** |
|
| 60 |
+
|---|---:|---:|---:|
|
| 61 |
+
| Decision flip rate ↓ | 9.25% | 12.00% | **6.00%** |
|
| 62 |
+
| Accuracy after permutation ↑ | 66.75% | 67.00% | **80.00%** |
|
| 63 |
+
|
| 64 |
+

|
| 65 |
+
|
| 66 |
+
On a separate **200-pair programmatic threshold-policy test**, both decisions in a counterfactual pair are correct in **71.50%** of pairs for v2, compared with 63.00% for v1. This test measures that specific rule family.
|
| 67 |
+
|
| 68 |
+
## Task-level results
|
| 69 |
+
|
| 70 |
+
<details>
|
| 71 |
+
<summary><strong>Per-task accuracy: all 11 real-label tasks and the separate typed-decision group</strong></summary>
|
| 72 |
+
|
| 73 |
+
| Real-label task | Examples | Jev-Style v1 | English Laya | Jev-Style v2 |
|
| 74 |
+
|---|---:|---:|---:|---:|
|
| 75 |
+
| AG News | 300 | 87.67% | 89.00% | 88.00% |
|
| 76 |
+
| ANLI | 300 | 48.00% | 49.67% | 48.67% |
|
| 77 |
+
| BoolQ | 300 | 82.67% | 75.67% | 81.67% |
|
| 78 |
+
| Emotion | 300 | 58.33% | 60.33% | 85.33% |
|
| 79 |
+
| Enron spam | 300 | 77.33% | 96.33% | 97.67% |
|
| 80 |
+
| HANS | 300 | 68.00% | 75.00% | 68.00% |
|
| 81 |
+
| IMDb | 300 | 96.67% | 93.67% | 96.33% |
|
| 82 |
+
| MNLI | 300 | 86.67% | 85.00% | 88.00% |
|
| 83 |
+
| RTE | 277 | 84.48% | 77.98% | 85.92% |
|
| 84 |
+
| SST-2 | 300 | 92.67% | 91.67% | 93.00% |
|
| 85 |
+
| SST-5 | 300 | 61.00% | 31.67% | 60.67% |
|
| 86 |
+
|
| 87 |
+
The separate typed-decisions group contains 2,000 teacher-reference decisions from 400 states. Teacher agreement is 53.35% for v1, 37.55% for English Laya and **73.45% for v2** under the fixed primary interface. This group is excluded from the real-label macro. The comparison here uses the English Laya checkpoint; specialist-checkpoint and rendering sensitivity results are provided in [baseline_sensitivity.json](evaluation/baseline_sensitivity.json).
|
| 88 |
+
|
| 89 |
+
</details>
|
| 90 |
+
|
| 91 |
+
|
| 92 |
+
## Deployment validation
|
| 93 |
+
|
| 94 |
+
| Released format | Weight size | Validated result | Evaluation set |
|
| 95 |
+
|---|---:|---|---|
|
| 96 |
+
| HF BF16 | 3.76 GB | **81.27%** real-label macro accuracy | Full 3,277 real-label decisions |
|
| 97 |
+
| Native MLX BF16 | 3.76 GB | **99.6%** choice agreement with CUDA BF16 | Frozen 500-decision deployment subset |
|
| 98 |
+
| Calibrated GGUF Q8_0 | 2.01 GB | **99.2%** choice agreement with CUDA BF16 | Same 500-decision deployment subset |
|
| 99 |
+
|
| 100 |
+
Each deployment format has its own validation record. Native MLX packaging reproduces the verified MLX client's logits exactly on all 500 deployment cases. The Q8_0 model is approximately **46.7% smaller** than the BF16 GGUF export.
|
| 101 |
+
|
| 102 |
+
On the same 500-case deployment subset, real-label task-macro accuracy is **79.10%** for CUDA BF16, **79.04%** for MLX BF16 and **78.69%** for Q8_0. Full-panel reference results and deployment-subset results use their respective denominators.
|
| 103 |
+
|
| 104 |
+
<details>
|
| 105 |
+
<summary><strong>Evaluation data and downloadable vector charts</strong></summary>
|
| 106 |
+
|
| 107 |
+
- [Reference metrics and paired intervals](evaluation/reference_comparison.json)
|
| 108 |
+
- [Deployment validation](evaluation/deployment.json)
|
| 109 |
+
- [Baseline sensitivity results](evaluation/baseline_sensitivity.json)
|
| 110 |
+
- [Data sources and split manifest](evaluation/data_manifest.json)
|
| 111 |
+
- [Chart data, confidence bins and sample counts](evaluation/chart_data.json)
|
| 112 |
+
- Vector charts: [benchmark](figures/benchmark.svg), [calibration](figures/calibration.svg), [robustness](figures/robustness.svg)
|
| 113 |
+
|
| 114 |
+
The benchmark figures describe the fixed CUDA reference comparison. Reliability pools all real-label examples into confidence bins; task-macro ECE is the mean of 11 separate task ECE values. These are distinct aggregations. The 9/12 figure counts task-level point estimates. Individual prediction probabilities, task summaries, test protocols and calibration records were retained when drawing these charts.
|
| 115 |
+
|
| 116 |
+
</details>
|
| 117 |
+
|
| 118 |
+
## Quick start
|
| 119 |
|
| 120 |
```bash
|
| 121 |
python -m pip install -U huggingface_hub
|
|
|
|
| 141 |
print(result["probabilities"])
|
| 142 |
```
|
| 143 |
|
| 144 |
+
Validated with `mlx==0.32.2` and `mlx-lm==0.31.3`. The companion `calibration.json` is loaded automatically. The reference benchmark and MLX deployment check are reported separately above and in the evaluation files.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 145 |
|
|
|
|
| 146 |
|
| 147 |
## Decision interface
|
| 148 |
|
SHA256SUMS.json
CHANGED
|
@@ -4,8 +4,8 @@
|
|
| 4 |
"sha256": "50cbab8a892c5f2993b8c7351a99182507472def3b1374558308605d99b86b32"
|
| 5 |
},
|
| 6 |
"README.md": {
|
| 7 |
-
"bytes":
|
| 8 |
-
"sha256": "
|
| 9 |
},
|
| 10 |
"calibration.json": {
|
| 11 |
"bytes": 362,
|
|
@@ -78,5 +78,33 @@
|
|
| 78 |
"tokenizer_config.json": {
|
| 79 |
"bytes": 1156,
|
| 80 |
"sha256": "303f948bfd44621a24f1d90b6330d8280076d08b60152f5851d9e9921c0b570e"
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 81 |
}
|
| 82 |
}
|
|
|
|
| 4 |
"sha256": "50cbab8a892c5f2993b8c7351a99182507472def3b1374558308605d99b86b32"
|
| 5 |
},
|
| 6 |
"README.md": {
|
| 7 |
+
"bytes": 10469,
|
| 8 |
+
"sha256": "94cc76f41bf25f0e1958ce7293e1e9c22cf34e7a9195b33f080200d983ec6113"
|
| 9 |
},
|
| 10 |
"calibration.json": {
|
| 11 |
"bytes": 362,
|
|
|
|
| 78 |
"tokenizer_config.json": {
|
| 79 |
"bytes": 1156,
|
| 80 |
"sha256": "303f948bfd44621a24f1d90b6330d8280076d08b60152f5851d9e9921c0b570e"
|
| 81 |
+
},
|
| 82 |
+
"evaluation/chart_data.json": {
|
| 83 |
+
"bytes": 4588,
|
| 84 |
+
"sha256": "0db9855803eb514792a906ff04b1d9ccdcba81babeedcffdd9ae34131f8b04f1"
|
| 85 |
+
},
|
| 86 |
+
"figures/calibration.svg": {
|
| 87 |
+
"bytes": 30706,
|
| 88 |
+
"sha256": "87b305248aef4580849edd523a29c8d7aba0774a879fe737f23c4fd460a861d0"
|
| 89 |
+
},
|
| 90 |
+
"figures/robustness.png": {
|
| 91 |
+
"bytes": 144638,
|
| 92 |
+
"sha256": "9a2032db7052f934d54db70eb490bea4eb160d6fb9c3644d07f5ffa11a91b383"
|
| 93 |
+
},
|
| 94 |
+
"figures/benchmark.png": {
|
| 95 |
+
"bytes": 216104,
|
| 96 |
+
"sha256": "8bc8eecd93cc7ff0727aaef15c644be664fd64abcee7d208e79e27218cb29051"
|
| 97 |
+
},
|
| 98 |
+
"figures/benchmark.svg": {
|
| 99 |
+
"bytes": 23602,
|
| 100 |
+
"sha256": "02c8a57ff3150f88f78cabdd99ce9aac5d9822a858929761748d1d5e679c9969"
|
| 101 |
+
},
|
| 102 |
+
"figures/robustness.svg": {
|
| 103 |
+
"bytes": 13858,
|
| 104 |
+
"sha256": "bccd285dd10dd98cac069928a127d5767cb61704e18d8ac3fa656703e81d572f"
|
| 105 |
+
},
|
| 106 |
+
"figures/calibration.png": {
|
| 107 |
+
"bytes": 219405,
|
| 108 |
+
"sha256": "0555263334858774466ffde686343d6fae3b4ca4b5572e6e41c6044291883b70"
|
| 109 |
}
|
| 110 |
}
|
evaluation/chart_data.json
ADDED
|
@@ -0,0 +1,183 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"source": "evaluation/reference_comparison.json",
|
| 3 |
+
"source_comparison_sha256": "36f56a490fc67e774cbb25b38d4dc9b5e3a42e1f65fcd748af092e127a98eff8",
|
| 4 |
+
"real_label_decisions": 3277,
|
| 5 |
+
"overview": {
|
| 6 |
+
"v1": {
|
| 7 |
+
"accuracy": 0.7667968493600263,
|
| 8 |
+
"macro_f1": 0.7542251333163978,
|
| 9 |
+
"nll": 0.575152117780084,
|
| 10 |
+
"brier": 0.3289652416178414,
|
| 11 |
+
"ece": 0.08512955323369954
|
| 12 |
+
},
|
| 13 |
+
"laya": {
|
| 14 |
+
"accuracy": 0.750889399409255,
|
| 15 |
+
"macro_f1": 0.7344520495576862,
|
| 16 |
+
"nll": 0.6318461242308782,
|
| 17 |
+
"brier": 0.3481931154806123,
|
| 18 |
+
"ece": 0.12362140723612026
|
| 19 |
+
},
|
| 20 |
+
"v2": {
|
| 21 |
+
"accuracy": 0.8120490099551472,
|
| 22 |
+
"macro_f1": 0.797772564877388,
|
| 23 |
+
"nll": 0.5153573778298738,
|
| 24 |
+
"brier": 0.27870374107073076,
|
| 25 |
+
"ece": 0.09100670643474301
|
| 26 |
+
}
|
| 27 |
+
},
|
| 28 |
+
"reliability": {
|
| 29 |
+
"model": "v2 CUDA reference",
|
| 30 |
+
"temperature": 1.0423505400296817,
|
| 31 |
+
"binning": "15 equal-width confidence bins on [0,1]",
|
| 32 |
+
"empty_bins": "omitted",
|
| 33 |
+
"uncertainty": "Wilson 95% intervals within each confidence bin",
|
| 34 |
+
"aggregation": "pooled over real-label examples",
|
| 35 |
+
"pooled_ece": 0.07282541921307938,
|
| 36 |
+
"bins": [
|
| 37 |
+
{
|
| 38 |
+
"lo": 0.3333333333333333,
|
| 39 |
+
"hi": 0.4,
|
| 40 |
+
"n": 11,
|
| 41 |
+
"mean_confidence": 0.38078092576157074,
|
| 42 |
+
"accuracy": 0.45454545454545453,
|
| 43 |
+
"wilson95": [
|
| 44 |
+
0.21271271622459764,
|
| 45 |
+
0.719908462590678
|
| 46 |
+
]
|
| 47 |
+
},
|
| 48 |
+
{
|
| 49 |
+
"lo": 0.4,
|
| 50 |
+
"hi": 0.4666666666666667,
|
| 51 |
+
"n": 31,
|
| 52 |
+
"mean_confidence": 0.44342106018763106,
|
| 53 |
+
"accuracy": 0.3870967741935484,
|
| 54 |
+
"wilson95": [
|
| 55 |
+
0.23733101420380254,
|
| 56 |
+
0.5617589138033927
|
| 57 |
+
]
|
| 58 |
+
},
|
| 59 |
+
{
|
| 60 |
+
"lo": 0.4666666666666667,
|
| 61 |
+
"hi": 0.5333333333333333,
|
| 62 |
+
"n": 130,
|
| 63 |
+
"mean_confidence": 0.5071245091267953,
|
| 64 |
+
"accuracy": 0.4,
|
| 65 |
+
"wilson95": [
|
| 66 |
+
0.3198243104508396,
|
| 67 |
+
0.4859160017880467
|
| 68 |
+
]
|
| 69 |
+
},
|
| 70 |
+
{
|
| 71 |
+
"lo": 0.5333333333333333,
|
| 72 |
+
"hi": 0.6,
|
| 73 |
+
"n": 152,
|
| 74 |
+
"mean_confidence": 0.5683921263707039,
|
| 75 |
+
"accuracy": 0.5197368421052632,
|
| 76 |
+
"wilson95": [
|
| 77 |
+
0.4408087537451127,
|
| 78 |
+
0.5976919125435891
|
| 79 |
+
]
|
| 80 |
+
},
|
| 81 |
+
{
|
| 82 |
+
"lo": 0.6,
|
| 83 |
+
"hi": 0.6666666666666666,
|
| 84 |
+
"n": 158,
|
| 85 |
+
"mean_confidence": 0.6355755895198861,
|
| 86 |
+
"accuracy": 0.5379746835443038,
|
| 87 |
+
"wilson95": [
|
| 88 |
+
0.46025816470988046,
|
| 89 |
+
0.6138884729158649
|
| 90 |
+
]
|
| 91 |
+
},
|
| 92 |
+
{
|
| 93 |
+
"lo": 0.6666666666666666,
|
| 94 |
+
"hi": 0.7333333333333333,
|
| 95 |
+
"n": 145,
|
| 96 |
+
"mean_confidence": 0.7017602376997318,
|
| 97 |
+
"accuracy": 0.6551724137931034,
|
| 98 |
+
"wilson95": [
|
| 99 |
+
0.5747027826751803,
|
| 100 |
+
0.7276323352183527
|
| 101 |
+
]
|
| 102 |
+
},
|
| 103 |
+
{
|
| 104 |
+
"lo": 0.7333333333333333,
|
| 105 |
+
"hi": 0.8,
|
| 106 |
+
"n": 153,
|
| 107 |
+
"mean_confidence": 0.7684716912664443,
|
| 108 |
+
"accuracy": 0.7647058823529411,
|
| 109 |
+
"wilson95": [
|
| 110 |
+
0.6915216339127415,
|
| 111 |
+
0.8249234476941671
|
| 112 |
+
]
|
| 113 |
+
},
|
| 114 |
+
{
|
| 115 |
+
"lo": 0.8,
|
| 116 |
+
"hi": 0.8666666666666667,
|
| 117 |
+
"n": 166,
|
| 118 |
+
"mean_confidence": 0.8361544180666837,
|
| 119 |
+
"accuracy": 0.7048192771084337,
|
| 120 |
+
"wilson95": [
|
| 121 |
+
0.6314328145486011,
|
| 122 |
+
0.7689405717386028
|
| 123 |
+
]
|
| 124 |
+
},
|
| 125 |
+
{
|
| 126 |
+
"lo": 0.8666666666666667,
|
| 127 |
+
"hi": 0.9333333333333333,
|
| 128 |
+
"n": 333,
|
| 129 |
+
"mean_confidence": 0.9044363365371433,
|
| 130 |
+
"accuracy": 0.7627627627627628,
|
| 131 |
+
"wilson95": [
|
| 132 |
+
0.7142396166140905,
|
| 133 |
+
0.8052926304337862
|
| 134 |
+
]
|
| 135 |
+
},
|
| 136 |
+
{
|
| 137 |
+
"lo": 0.9333333333333333,
|
| 138 |
+
"hi": 1.0,
|
| 139 |
+
"n": 1998,
|
| 140 |
+
"mean_confidence": 0.984503687108951,
|
| 141 |
+
"accuracy": 0.9229229229229229,
|
| 142 |
+
"wilson95": [
|
| 143 |
+
0.9103995454324445,
|
| 144 |
+
0.9338231538993982
|
| 145 |
+
]
|
| 146 |
+
}
|
| 147 |
+
]
|
| 148 |
+
},
|
| 149 |
+
"permutation": {
|
| 150 |
+
"v1": {
|
| 151 |
+
"n": 400,
|
| 152 |
+
"semantic_flip_rate": 0.0925,
|
| 153 |
+
"permuted_accuracy": 0.6675
|
| 154 |
+
},
|
| 155 |
+
"laya": {
|
| 156 |
+
"n": 400,
|
| 157 |
+
"semantic_flip_rate": 0.12,
|
| 158 |
+
"permuted_accuracy": 0.67
|
| 159 |
+
},
|
| 160 |
+
"v2": {
|
| 161 |
+
"n": 400,
|
| 162 |
+
"semantic_flip_rate": 0.06,
|
| 163 |
+
"permuted_accuracy": 0.8
|
| 164 |
+
}
|
| 165 |
+
},
|
| 166 |
+
"counterfactual": {
|
| 167 |
+
"v1": {
|
| 168 |
+
"pairs": 200,
|
| 169 |
+
"both_correct": 0.63,
|
| 170 |
+
"scope": "programmatic threshold-policy shift only"
|
| 171 |
+
},
|
| 172 |
+
"laya": {
|
| 173 |
+
"pairs": 200,
|
| 174 |
+
"both_correct": 0.075,
|
| 175 |
+
"scope": "programmatic threshold-policy shift only"
|
| 176 |
+
},
|
| 177 |
+
"v2": {
|
| 178 |
+
"pairs": 200,
|
| 179 |
+
"both_correct": 0.715,
|
| 180 |
+
"scope": "programmatic threshold-policy shift only"
|
| 181 |
+
}
|
| 182 |
+
}
|
| 183 |
+
}
|
figures/benchmark.png
ADDED
|
Git LFS Details
|
figures/benchmark.svg
ADDED
|
|
figures/calibration.png
ADDED
|
Git LFS Details
|
figures/calibration.svg
ADDED
|
|
figures/robustness.png
ADDED
|
Git LFS Details
|
figures/robustness.svg
ADDED
|
|