Ruiruiz30 commited on
Commit
33c5267
·
verified ·
1 Parent(s): 38fccc7

add benchmark and calibration charts to model card

Browse files
README.md CHANGED
@@ -42,6 +42,21 @@ The first request includes MLX graph and memory warm-up. On the same machine, a
42
 
43
  The published Jev-Omni H200 numbers are not transferable to this Mac mini. This model card reports local measurements only.
44
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
45
  ## Validation
46
 
47
  - Upstream unified verification cases: 4/4 argmax decisions matched after 4-bit conversion.
@@ -89,6 +104,12 @@ Omit `--calibration` to inspect the raw quantized probabilities. The included ca
89
  - [`calibration.json`](calibration.json) is the small runtime file consumed by `--calibration`.
90
  - The benchmark runner and raw checkpoints remain in the source project so the long DecisionBench run can be resumed without putting the full benchmark text into this model repository.
91
 
 
 
 
 
 
 
92
  ## Local conversion code
93
 
94
  `omni_mlx/convert.py` contains the conversion path used for this release. The original unquantized checkpoint is not bundled here; it can be obtained from the upstream repository under its own license and terms.
 
42
 
43
  The published Jev-Omni H200 numbers are not transferable to this Mac mini. This model card reports local measurements only.
44
 
45
+ ![Local warm latency](assets/local-latency.svg)
46
+
47
+ ## Results
48
+
49
+ The public JevBench files were evaluated with the same typed-choice mapping used by the runtime. Temperature scaling changes probabilities only; it does not change the selected option.
50
+
51
+ | Benchmark | Accuracy / state macro | Micro accuracy | ECE-10 |
52
+ |---|---:|---:|---:|
53
+ | JevBench public · 231 decisions | **87.88%** | **87.88%** | 0.04497 raw / **0.03069** scaled |
54
+ | DecisionBench Medium · 293 questions | — | — | Full run not published |
55
+
56
+ ![JevBench public accuracy](assets/jevbench-accuracy.svg)
57
+
58
+ The JevBench result is our local public-set measurement, not a claim that the 4-bit MLX conversion reproduces the upstream card's protocol. The upstream [Jev-Omni card](https://huggingface.co/akhilaaa3/Jev-Omni) reports its own merged-model result separately. Dataset revisions, item filtering and scoring splits must match before comparing the numbers.
59
+
60
  ## Validation
61
 
62
  - Upstream unified verification cases: 4/4 argmax decisions matched after 4-bit conversion.
 
104
  - [`calibration.json`](calibration.json) is the small runtime file consumed by `--calibration`.
105
  - The benchmark runner and raw checkpoints remain in the source project so the long DecisionBench run can be resumed without putting the full benchmark text into this model repository.
106
 
107
+ ## Calibration
108
+
109
+ Temperature scaling was fit on 116 even-indexed public JevBench rows and checked on 115 odd-indexed rows. The fitted temperature is `T=1.11516790625`. On the full public set, ECE-10 moves from `0.04497` to `0.03069`; on the held-out split it moves from `0.06261` to `0.06774`. This is a transparent post-hoc calibration file, not a guarantee of calibrated confidence on game footage.
110
+
111
+ ![JevBench calibration](assets/jevbench-calibration.svg)
112
+
113
  ## Local conversion code
114
 
115
  `omni_mlx/convert.py` contains the conversion path used for this release. The original unquantized checkpoint is not bundled here; it can be obtained from the upstream repository under its own license and terms.
assets/jevbench-accuracy.svg ADDED
assets/jevbench-calibration.svg ADDED
assets/local-latency.svg ADDED