Capicua25x commited on
Commit
8dc02b8
·
verified ·
1 Parent(s): facf024

report GSM8K as a range: two runs exist at the same seed and the card published the higher one; align AA-LCR arm B to its score.json (0.800, not the rejudge pass) so every arm uses the same judging pass

Browse files
Files changed (1) hide show
  1. README.md +15 -4
README.md CHANGED
@@ -140,12 +140,12 @@ on this same box, differing only in KV cache dtype.
140
 
141
  | benchmark | n | bf16 ref | FP8 + bf16 KV | FP8 + fp8 KV | **MXFP4 (this)** |
142
  |---|---|---|---|---|---|
143
- | GSM8K, thinking (flex / strict) | 50 | 0.96 / 0.82 | 0.96 / 0.90 | 0.94 / 0.70 | **0.96 / 0.94** |
144
  | GSM8K, no thinking (flex / strict) | 50 | 0.98 / 0.98 | 0.98 / 0.98 | 0.98 / 0.98 | **0.98 / 0.98** |
145
  | IFEval (inst / prompt, strict) | 80 | .9688 / .9500 | .9688 / .9500 | .9688 / .9500 | **.9688 / .9500** |
146
  | GPQA-Diamond (flexible) | 60 | 0.7833 | 0.8333 | 0.8333 | **0.9167** |
147
  | AIME 2025 | 30 | 0.9333 | 1.0000 | 0.9667 | **0.9333** |
148
- | AA-LCR (~107k-token prompts, judge-scored) | 100 | 0.780 | 0.800 ᵃ | 0.810 | **0.780** |
149
  | τ²-bench telecom (Pass^1) | 114 | 0.939 | 0.904 | 0.895 | **0.868** |
150
  | τ²-bench airline (Pass^1) | 50 | 0.760 | — | — | **0.840** |
151
  | HLE | 120 | 0.3083 | — | — | *running* |
@@ -155,6 +155,12 @@ on this same box, differing only in KV cache dtype.
155
  ᵃ Scored on the 90 items it served; 10 were refused because the prompt exceeded that
156
  configuration's 131k window. Blended over the full 100 it reads 0.720.
157
 
 
 
 
 
 
 
158
  **τ² is domain-split, and the split is the finding.** On telecom this build scores 0.868 against
159
  the bf16 reference's 0.939 — eight simulations — and sits four behind the FP8 + bf16 KV arm and
160
  three behind FP8 + fp8 KV. On airline it scores **0.840 against the reference's 0.760**, four items
@@ -165,8 +171,13 @@ On telecom, 113 of 114 simulations ended normally and one hit the harness's erro
165
  (`too_many_errors`, scored 0), so the shortfall is a genuine capability difference rather than
166
  harness noise — but it is one domain and a single-digit item count, not a blanket weakness.
167
 
168
- Everything else is at or above bf16: GSM8K strict-match **+6 items**, GPQA **+8 items**, and
169
- long-context retrieval **identical** to bf16 at ~107k-token prompts.
 
 
 
 
 
170
 
171
  Cells marked *running* / *pending* are genuinely unfinished, not withheld. This card is dated and
172
  will be revised as they land; the commit history is the record of what was known when.
 
140
 
141
  | benchmark | n | bf16 ref | FP8 + bf16 KV | FP8 + fp8 KV | **MXFP4 (this)** |
142
  |---|---|---|---|---|---|
143
+ | GSM8K, thinking (flex / strict) | 50 | 0.96 / 0.82 | 0.96 / 0.90 | 0.94 / 0.70 | **0.94–0.96 / 0.92–0.94** ᶜ |
144
  | GSM8K, no thinking (flex / strict) | 50 | 0.98 / 0.98 | 0.98 / 0.98 | 0.98 / 0.98 | **0.98 / 0.98** |
145
  | IFEval (inst / prompt, strict) | 80 | .9688 / .9500 | .9688 / .9500 | .9688 / .9500 | **.9688 / .9500** |
146
  | GPQA-Diamond (flexible) | 60 | 0.7833 | 0.8333 | 0.8333 | **0.9167** |
147
  | AIME 2025 | 30 | 0.9333 | 1.0000 | 0.9667 | **0.9333** |
148
+ | AA-LCR (~107k-token prompts, judge-scored) | 100 | 0.780 | 0.800 ᵃ | 0.800 | **0.780** |
149
  | τ²-bench telecom (Pass^1) | 114 | 0.939 | 0.904 | 0.895 | **0.868** |
150
  | τ²-bench airline (Pass^1) | 50 | 0.760 | — | — | **0.840** |
151
  | HLE | 120 | 0.3083 | — | — | *running* |
 
155
  ᵃ Scored on the 90 items it served; 10 were refused because the prompt exceeded that
156
  configuration's 131k window. Blended over the full 100 it reads 0.720.
157
 
158
+ ᶜ **Two runs of this build exist at the same seed and identical settings** — 0.94/0.92 and
159
+ 0.96/0.94 — so the honest figure is a range, not a point. The other three columns are single
160
+ runs, which is worth knowing before reading small deltas here as real: on this cell one run's
161
+ difference is one item. Against the reference's 0.82 strict, this build is +5 or +6 items
162
+ depending on which run you take.
163
+
164
  **τ² is domain-split, and the split is the finding.** On telecom this build scores 0.868 against
165
  the bf16 reference's 0.939 — eight simulations — and sits four behind the FP8 + bf16 KV arm and
166
  three behind FP8 + fp8 KV. On airline it scores **0.840 against the reference's 0.760**, four items
 
171
  (`too_many_errors`, scored 0), so the shortfall is a genuine capability difference rather than
172
  harness noise — but it is one domain and a single-digit item count, not a blanket weakness.
173
 
174
+ Everything else is at or above bf16: GSM8K strict-match **+5 to +6 items** (see note ᶜ — two
175
+ runs exist), GPQA **+8 items**, and long-context retrieval **identical** to bf16 at ~107k-token
176
+ prompts.
177
+
178
+ All AA-LCR figures are the runner's own judging pass, taken from each arm's `score.json`. A
179
+ second judging pass over the same generations moves scores by roughly one item in either
180
+ direction; mixing passes between arms would manufacture differences that are not there.
181
 
182
  Cells marked *running* / *pending* are genuinely unfinished, not withheld. This card is dated and
183
  will be revised as they land; the commit history is the record of what was known when.