Capicua25x commited on
Commit
9e488f4
·
verified ·
1 Parent(s): eba1212

AA cells to 5-seed medians [range] (window44); fix stale DFlash prompt-strict cell

Browse files
Files changed (1) hide show
  1. README.md +9 -7
README.md CHANGED
@@ -78,20 +78,22 @@ hardware even with a tuned draft config — treat it as experimental.
78
 
79
  ## Quality — AA vs the official FP8
80
 
81
- Same protocol both arms (lm-eval, seed 1234, temp 0.6/top-p 0.95/top-k 20, thinking off);
 
82
  reference = [ornith-ai/Ornith-1.5-35B-A3B-FP8](https://huggingface.co/ornith-ai/Ornith-1.5-35B-A3B-FP8)
83
  measured on the same hardware.
84
 
85
  | suite | this quant (conc) | this quant (DFlash serve) | official FP8 |
86
  |---|---|---|---|
87
- | GSM8K flexible / strict (n=50) | 0.86 / 0.84 | **0.92 / 0.88** | 0.90 / 0.88 |
88
- | IFEval inst-strict (n=80) | **0.875** | 0.875 [0.836–0.883]⁵ | 0.867 |
89
- | IFEval prompt-strict (n=80) | 0.800 | 0.738 | 0.813 |
 
90
  | τ²-bench telecom reward (n=114) | **0.965** (110/114) | same distribution* | not measured |
91
 
92
- ⁵DFlash IFEval cells are 5-seed median [range]. Short-suite single-seed cells (±2 items ≈ noise): the quant sits at or above its official
93
- FP8 reference. \*Speculative decoding is distribution-preserving (rejection sampling), so
94
- τ² quality carries across serve profiles.
95
 
96
  ## Credits
97
 
 
78
 
79
  ## Quality — AA vs the official FP8
80
 
81
+ Same protocol both arms (lm-eval, temp 0.6/top-p 0.95/top-k 20, thinking off); quant cells
82
+ are 5-seed medians [range], seeds 1234–1238; reference cells are single-seed 1234.
83
  reference = [ornith-ai/Ornith-1.5-35B-A3B-FP8](https://huggingface.co/ornith-ai/Ornith-1.5-35B-A3B-FP8)
84
  measured on the same hardware.
85
 
86
  | suite | this quant (conc) | this quant (DFlash serve) | official FP8 |
87
  |---|---|---|---|
88
+ | GSM8K flexible (n=50) | 0.90 [0.84–0.96]⁵ | 0.88 [0.84–0.92]⁵ | 0.90 |
89
+ | GSM8K strict (n=50) | 0.84 [0.80–0.92]⁵ | 0.84 [0.78–0.86]⁵ | 0.88 |
90
+ | IFEval inst-strict (n=80) | 0.859 [0.836–0.883]⁵ | 0.875 [0.836–0.883]⁵ | 0.867 |
91
+ | IFEval prompt-strict (n=80) | 0.800 [0.763–0.825]⁵ | 0.800 [0.763–0.813]⁵ | 0.813 |
92
  | τ²-bench telecom reward (n=114) | **0.965** (110/114) | same distribution* | not measured |
93
 
94
+ ⁵5-seed median [range]. On the short suites the quant tracks its FP8 reference within
95
+ seed noise (±1–2 items per seed); τ² is the deep gate. \*Speculative decoding is
96
+ distribution-preserving (rejection sampling), so τ² quality carries across serve profiles.
97
 
98
  ## Credits
99