sixstringzen commited on
Commit
9efa75b
·
verified ·
1 Parent(s): 8e1cd3d

fix(eval): correct blind-judge attribution and prompt matching

Browse files

Analysis revision 2. Frozen generations and judge records unchanged. Correct A/B decoding, reversed prompt joins, prompt counts, and dependent-rating uncertainty claims.

Files changed (1) hide show
  1. README.md +68 -29
README.md CHANGED
@@ -76,38 +76,77 @@ The finished artifact passed local structural and oMLX load checks on 2026-09-20
76
 
77
  The artifact contains five safetensors shards, 1,876 indexed tensors, and 29 MTP tensors. The index references no missing shards. These checks confirm that oMLX recognizes and loads the files; they do not establish quality parity with the BF16 source model.
78
 
79
- ## Quality evaluation (v1)
80
-
81
- This model is part of the [Hemmingway-1 oMLX quantization collection](https://huggingface.co/collections/sixstringzen/hemmingway-1-omlx-oqe-quantizations-apple-silicon-evidence-6ab06103552bc3c7e9ad8afa) and was compared with a clean BF16 reference in a bounded blind A/B writing study.
82
-
83
- ### Local measurement provenance
84
-
85
- Generations were produced by real MLX/oMLX software on an Apple M5 Max with 128 GB unified memory under a deterministic quant-isolation profile. MTP and mixed-runtime speculative accelerators were disabled for the baseline. Captured local manifests, generation records, and oMLX/macmon telemetry are the measurement evidence. Claude monitored or orchestrated a subset of the local tests; Claude is workflow provenance, not the inference engine or measurement source.
86
-
87
- ### Blind-judge result
88
-
89
- The study used 14 writing tasks, three judge lanes, and normal/swapped response order. Results below pool the candidate comparisons across Claude Opus 5, Gemini 3.8 Flash, and Grok 4.7:
90
-
91
- | Build | Win | Loss | Tie |
92
- | --- | ---: | ---: | ---: |
93
- | oQ2e | 47.8% | 51.1% | 1.1% |
94
- | oQ3.5e | 36.7% | 58.9% | 4.4% |
95
- | oQ3e | 55.6% | 43.3% | 1.1% |
96
- | oQ4e | 32.2% | 48.9% | 18.9% |
97
- | oQ6e | 27.8% | 31.1% | 41.1% |
98
- | oQ8e | 31.1% | 23.3% | 45.6% |
99
-
100
- The hosted reference row in the local report is a cap-matched subset of 11 prompts, not another quantization condition. Raw inter-rater agreement was 83.7% across 606 pairings. Normal-versus-swapped agreement was 39.6% for Claude, 48.5% for Gemini, and 42.6% for Grok, so the results should be read as subjective and order-sensitive.
101
-
102
- ### Interpretation for this build
103
-
104
- High-bit perceptual-reference group in this suite, with a 41.1% tie rate. This is not mathematical equivalence.
105
-
106
- This pass does not establish mathematical equivalence, token-level fidelity, or general benchmark superiority. KLD, top-1/top-k agreement, KV-cache divergence, and broader benchmark task scores remain separate pending measurements. The public task and metadata package is the [companion benchmark dataset](https://huggingface.co/datasets/sixstringzen/hemmingway-1-omlx-quantization-benchmark-v1).
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
107
 
108
  ## Limitations
109
 
110
- Quantization can change word choice, coherence, and instruction following. A controlled BF16 comparison has not been published for this build.
111
 
112
  The sensitivity pass used oMLX's `oqe_code_multilingual` calibration dataset. No prose-specific calibration dataset was used. The MTP tensors are present in the artifact, but MTP-assisted decoding has not been benchmarked separately.
113
 
 
76
 
77
  The artifact contains five safetensors shards, 1,876 indexed tensors, and 29 MTP tensors. The index references no missing shards. These checks confirm that oMLX recognizes and loads the files; they do not establish quality parity with the BF16 source model.
78
 
79
+ ## Quality evaluation (v1, corrected analysis revision 2)
80
+
81
+ Corrected on 2026-09-22 after identifying errors in A/B decoding and normal/swapped prompt matching. The generated responses and judge records are unchanged. See the [correction record](https://huggingface.co/datasets/sixstringzen/hemmingway-1-omlx-quantization-benchmark-v1/blob/main/CORRECTION.md).
82
+
83
+ Each quantization was compared with the local BF16 reference on 15 prompts.
84
+ Three judge lanes evaluated both response orders, producing 90 ratings per quant.
85
+ The hosted comparison used 11 cap-matched prompts and produced 66 ratings.
86
+ The study contains 14 packet files per judge, each holding multiple cases.
87
+
88
+ The ratings describe one frozen output per condition and prompt. Multiple judges
89
+ and swapped orders do not create independent generation samples. We report
90
+ counts, percentages, and mean score differences without confidence intervals or
91
+ statistical significance claims. Percentages can differ from 100% after rounding.
92
+
93
+ Judges scored instruction adherence, task fit, clarity and control, and writing
94
+ judgment on a 0-4 scale. Scores and overall preferences are separate judgments.
95
+ The Claude lane used manual chats, except the final hosted swapped packet,
96
+ which used OpenRouter. Gemini and Grok used OpenRouter. The Grok collection
97
+ includes documented recovery of packets 04, 05, and 06.
98
+
99
+ | Condition | Wins | Losses | Ties | Ratings | Win % | Loss % | Tie % |
100
+ | --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: |
101
+ | oQ2e | 31 | 58 | 1 | 90 | 34.4 | 64.4 | 1.1 |
102
+ | oQ3.5e | 44 | 42 | 4 | 90 | 48.9 | 46.7 | 4.4 |
103
+ | oQ3e | 54 | 35 | 1 | 90 | 60.0 | 38.9 | 1.1 |
104
+ | oQ4e | 30 | 43 | 17 | 90 | 33.3 | 47.8 | 18.9 |
105
+ | oQ6e | 34 | 19 | 37 | 90 | 37.8 | 21.1 | 41.1 |
106
+ | oQ8e | 27 | 22 | 41 | 90 | 30.0 | 24.4 | 45.6 |
107
+
108
+ After matching prompts across orders, agreement was 89.1% for Claude, 85.1% for Gemini, and 83.2% for Grok (101 pairs per judge). Inter-rater agreement was 83.7% across 606 pairings. These rates measure agreement on the frozen outputs; they do not validate the judges' preferences.
109
+
110
+ ### This build
111
+
112
+ Frequent ties with BF16 and more wins than losses in this suite.
113
+
114
+ ### Local generation profile
115
+
116
+ Local generations ran through MLX/oMLX on an Apple M5 Max with 128 GB unified
117
+ memory. Captured server records identify oMLX `0.7.0.dev2`. The generation profile
118
+ used temperature 0, top-p 1, min-p 0, repetition penalty 1, and seed 42. MTP and
119
+ speculative acceleration were disabled for this baseline. The files preserve
120
+ MTP tensors for separate runtime experiments.
121
+
122
+ The quant runs recorded 90 completions at 512 tokens, 24 at 1024, and 18 at 2048.
123
+ The selected BF16 reference recorded 11, one, and three completions at those caps.
124
+ The judge packets selected the final cap for each prompt: 11 at 512, one at
125
+ 1024, and three at 2048. Earlier truncated attempts remain execution records
126
+ and were excluded from the judge packets. Hosted retained a separate 11-prompt
127
+ comparison because four prompts did not have matching output caps.
128
+
129
+ Local manifests, generation records, and oMLX/macmon telemetry carry execution
130
+ evidence. Grafana displays the telemetry. Background tracing remained active,
131
+ so controlled throughput and memory comparisons require a separate run.
132
+
133
+ This small, fixed suite cannot establish a universal ranking or quantify
134
+ token-level fidelity. The higher observed oQ3e win rate does not imply that
135
+ reducing precision improves the source model in general. High tie rates for
136
+ oQ6e and oQ8e do not establish equivalence to BF16. The study includes no direct
137
+ quant-versus-hosted comparison, and the hosted deployment's exact checkpoint,
138
+ precision, and runtime are not verified.
139
+
140
+ KLD, top-1/top-k agreement, KV-cache divergence, broader task scores, and
141
+ controlled runtime comparisons remain pending. Raw generations, provider
142
+ responses, packet mappings, and telemetry remain in the local evidence package.
143
+ The public dataset contains prompts, selected execution metadata, and aggregates.
144
+
145
+ [Companion dataset](https://huggingface.co/datasets/sixstringzen/hemmingway-1-omlx-quantization-benchmark-v1).
146
 
147
  ## Limitations
148
 
149
+ Quantization can change word choice, coherence, and instruction following. The writing comparison above covers a fixed prompt suite; token-level fidelity remains unmeasured.
150
 
151
  The sensitivity pass used oMLX's `oqe_code_multilingual` calibration dataset. No prose-specific calibration dataset was used. The MTP tensors are present in the artifact, but MTP-assisted decoding has not been benchmarked separately.
152