Capicua25x commited on
Commit
40ecee5
Β·
verified Β·
1 Parent(s): f5e378e

tau2: add airline (0.840 vs ref 0.760) and correct the telecom error count (1 of 114, not zero); drop the single-run quality comparison against AMD's build and state the structural difference instead

Browse files
Files changed (1) hide show
  1. README.md +19 -9
README.md CHANGED
@@ -42,10 +42,15 @@ norm, embeddings, `lm_head`, routers/gates and the **entire vision path** stay b
42
  Verified by tensor inspection: a module counts as quantised only if it carries a real artifact
43
  (`weight_scale`, `weight_packed`, `qweight`, `weight_zero_point`).
44
 
45
- Keeping attention in bf16 is deliberate β€” it is what preserves the strict-format quality below
46
- while still getting the memory win, since the MLP stack is where the parameters actually are.
47
- For contrast, `amd/Qwen3.8-27B-Quark-AWQ-MXFP4` quantises the whole decoder including attention
48
- (W4A4) and measures 0.96/0.78 on GSM8K at 26.8 tok/s on this same hardware.
 
 
 
 
 
49
 
50
  - Format: MXFP4 (E2M1 + E8M0 scale per 32 weights), `pack_method: reorder`, `weight_format: real_quantized`
51
  - Size: **22.3 GB** across 18 shards (bf16 source β‰ˆ 54 GB)
@@ -142,6 +147,7 @@ on this same box, differing only in KV cache dtype.
142
  | AIME 2025 | 30 | 0.9333 | 1.0000 | 0.9667 | **0.9333** |
143
  | AA-LCR (~107k-token prompts, judge-scored) | 100 | 0.780 | 0.800 ᡃ | 0.810 | **0.780** |
144
  | τ²-bench telecom (Pass^1) | 114 | 0.991 | 0.982 ᡇ | 0.939 | **0.868** |
 
145
  | HLE | 120 | 0.3083 | β€” | β€” | *running* |
146
  | SWE-bench Verified | 100 | β€” | β€” | β€” | *pending* |
147
  | Terminal-Bench Hard | 44 | β€” | β€” | β€” | *pending* |
@@ -150,11 +156,15 @@ on this same box, differing only in KV cache dtype.
150
  configuration's 131k window. Blended over the full 100 it reads 0.720.
151
  ᡇ Ran with thinking off by server default, so it is not strictly paired with the other columns.
152
 
153
- **The τ² result is the honest weak spot.** 0.868 against the bf16 reference's 0.991 is roughly
154
- 14 simulations, with **zero agent errors and zero user-simulator errors** β€” so it is a capability
155
- gap on multi-turn tool use, not harness noise. Airline and retail domains are still running and
156
- will show whether telecom is representative. If you are picking a quant for agentic tool-calling
157
- specifically, weigh this cell above the others.
 
 
 
 
158
 
159
  Everything else is at or above bf16: GSM8K strict-match **+6 items**, GPQA **+8 items**, and
160
  long-context retrieval **identical** to bf16 at ~107k-token prompts.
 
42
  Verified by tensor inspection: a module counts as quantised only if it carries a real artifact
43
  (`weight_scale`, `weight_packed`, `qweight`, `weight_zero_point`).
44
 
45
+ Keeping attention in bf16 is deliberate: the MLP stack is where the parameters are, so excluding
46
+ attention costs little size and keeps those layers on the fast bf16 path.
47
+
48
+ For structural comparison, `amd/Qwen3.8-27B-Quark-AWQ-MXFP4` quantises the decoder's attention as
49
+ well β€” 496 quantised modules against 432 here, the difference being exactly the 16 full-attention
50
+ layers' q/k/v/o β€” and is **AWQ-calibrated** (`algo_config.name = awq`) where this build is data-free
51
+ RTN. Those are the two real differences. We have a single unrepeated n=50 GSM8K run against that
52
+ build, which is not enough to publish a quality comparison from: strict-match moves by about
53
+ Β±0.06 across seeds on this hardware, which is wider than any gap it showed.
54
 
55
  - Format: MXFP4 (E2M1 + E8M0 scale per 32 weights), `pack_method: reorder`, `weight_format: real_quantized`
56
  - Size: **22.3 GB** across 18 shards (bf16 source β‰ˆ 54 GB)
 
147
  | AIME 2025 | 30 | 0.9333 | 1.0000 | 0.9667 | **0.9333** |
148
  | AA-LCR (~107k-token prompts, judge-scored) | 100 | 0.780 | 0.800 ᡃ | 0.810 | **0.780** |
149
  | τ²-bench telecom (Pass^1) | 114 | 0.991 | 0.982 ᡇ | 0.939 | **0.868** |
150
+ | τ²-bench airline (Pass^1) | 50 | 0.760 | β€” | β€” | **0.840** |
151
  | HLE | 120 | 0.3083 | β€” | β€” | *running* |
152
  | SWE-bench Verified | 100 | β€” | β€” | β€” | *pending* |
153
  | Terminal-Bench Hard | 44 | β€” | β€” | β€” | *pending* |
 
156
  configuration's 131k window. Blended over the full 100 it reads 0.720.
157
  ᡇ Ran with thinking off by server default, so it is not strictly paired with the other columns.
158
 
159
+ **τ² is domain-split, and the split is the finding.** On telecom this build scores 0.868 against
160
+ the bf16 reference's 0.991 β€” about 14 simulations β€” but on airline it scores **0.840 against the
161
+ reference's 0.760**, i.e. four items *ahead*. So multi-turn tool use is not uniformly degraded;
162
+ telecom specifically is where it loses. Retail is still running and will add a third point.
163
+
164
+ On telecom, 113 of 114 simulations ended normally and one hit the harness's error ceiling
165
+ (`too_many_errors`, scored 0). That single run cannot account for a 14-item gap, so the telecom
166
+ deficit is a real capability gap rather than harness noise β€” but it is one domain, not a blanket
167
+ weakness. If you are choosing specifically for agentic tool-calling, weigh both cells.
168
 
169
  Everything else is at or above bf16: GSM8K strict-match **+6 items**, GPQA **+8 items**, and
170
  long-context retrieval **identical** to bf16 at ~107k-token prompts.