File size: 10,126 Bytes
e94a291
 
a2a51fd
 
 
 
 
 
e94a291
 
a2a51fd
e94a291
a2a51fd
e94a291
a2a51fd
 
6129a87
a2a51fd
e94a291
 
a2a51fd
e94a291
a2a51fd
 
6129a87
 
 
 
1dde698
6129a87
 
 
 
 
 
 
 
 
e94a291
 
 
a2a51fd
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
9efa75b
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
7854bbc
9efa75b
7854bbc
9efa75b
7854bbc
 
 
 
 
 
 
4afc7cb
a2a51fd
 
9efa75b
a2a51fd
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
---
library_name: mlx
license: apache-2.0
language:
- en
base_model: Altworld/Hemmingway-1
base_model_relation: quantized
pipeline_tag: text-generation
tags:
- mlx
- omlx
- oq
- oqe
- quantized
- 6-bit
- mtp
- apple-silicon
- creative-writing
---

# Hemmingway-1 oQ6e with MTP

This repository contains an enhanced oQ6e quantization of [Altworld/Hemmingway-1](https://huggingface.co/Altworld/Hemmingway-1) for MLX and oMLX on Apple silicon. The conversion preserves the model's multi-token prediction (MTP) tensors.

Altworld developed and published the source model. [sixstringzen](https://huggingface.co/sixstringzen) performed this conversion and published the converted weights with their quantization report. The original model, its intended use, and its training details remain documented in the [source model card](https://huggingface.co/Altworld/Hemmingway-1).

## Quantization set

This repository is part of the [Hemmingway-1 oMLX oQe Quantizations](https://huggingface.co/collections/sixstringzen/hemmingway-1-omlx-oqe-quantizations-apple-silicon-evidence-6ab06103552bc3c7e9ad8afa) collection. Every build in the set uses the same source revision, group size, non-quantized dtype, calibration pass, and MTP preservation policy.

| Build | Base precision | Output size |
| --- | ---: | ---: |
| [oQ2e](https://huggingface.co/sixstringzen/Hemmingway-1-oQ2e-mtp) | 2-bit | 10.14 GiB |
| [oQ3e](https://huggingface.co/sixstringzen/Hemmingway-1-oQ3e-mtp) | 3-bit | 12.22 GiB |
| [oQ3.5e](https://huggingface.co/sixstringzen/Hemmingway-1-oQ3.5e-mtp) | 3-bit with additional higher-precision overrides | 13.19 GiB |
| [oQ4e](https://huggingface.co/sixstringzen/Hemmingway-1-oQ4e-mtp) | 4-bit | 15.21 GiB |
| [oQ6e](https://huggingface.co/sixstringzen/Hemmingway-1-oQ6e-mtp) | 6-bit | 21.39 GiB |
| [oQ8e](https://huggingface.co/sixstringzen/Hemmingway-1-oQ8e-mtp) | 8-bit | 27.10 GiB |

## Quantization details

| Item | Value |
| --- | --- |
| Source model | `Altworld/Hemmingway-1` |
| Source revision | `4d711aac0f0043075ae334d2a3de3db3e10135c9` |
| Quantizer | [oMLX](https://github.com/jundot/omlx) `0.7.0.dev2` |
| Method | Enhanced oQ6e mixed-precision affine quantization |
| Base precision | 6-bit |
| Group size | 64 |
| Non-quantized dtype | `bfloat16` |
| Higher-precision tensors | 10 tensors at 8-bit, including `language_model.lm_head` |
| Calibration dataset | `oqe_code_multilingual` |
| Calibration shape | 128 samples at 512 tokens |
| Imatrix entries | 504 |
| Imatrix cache | Reused from the matching source-model sensitivity pass |
| MTP tensors | 29 preserved tensors |
| Output size | 22,964,597,559 bytes (21.39 GiB) |

oQe uses activation importance to assign additional precision to sensitive tensors. This build starts with 6-bit weights, assigns 8-bit precision to nine sensitive or protected tensors, and stores `language_model.lm_head` at 8-bit because that tensor had no matching imatrix entry. The quantization report records no matrix-shape mismatches and no missing weight shards.

The included [`oq_imatrix_report.json`](https://huggingface.co/sixstringzen/Hemmingway-1-oQ6e-mtp/blob/main/oq_imatrix_report.json) records the sensitivity pass, calibration settings, tensor coverage, and fallback. Strict imatrix coverage was disabled for the known `language_model.lm_head` fallback.

## Compatibility

This model was created and tested with oMLX `0.7.0.dev2`. The source model identifies its text architecture as `qwen3_5_text`; the converted artifact uses `qwen3_5`, which matches the architecture name supported by this oMLX build.

The weights use MLX safetensors and are not GGUF files. Compatibility with other MLX runtimes or earlier oMLX releases has not been verified.

## Use with oMLX

Download `sixstringzen/Hemmingway-1-oQ6e-mtp` from the oMLX model browser, then load it as an LLM. Set `enable_thinking` to `false` when you want direct prose without visible planning. Runtime defaults and the registered model identifier can vary with the local oMLX installation.

## Verification

The finished artifact passed local structural and oMLX load checks on 2026-09-20. oMLX loaded the model without reporting a model-format or weight error.

The artifact contains five safetensors shards, 1,876 indexed tensors, and 29 MTP tensors. The index references no missing shards. These checks confirm that oMLX recognizes and loads the files; they do not establish quality parity with the BF16 source model.

## Quality evaluation (v1, corrected analysis revision 2)

Corrected on 2026-09-22 after identifying errors in A/B decoding and normal/swapped prompt matching. The generated responses and judge records are unchanged. See the [correction record](https://huggingface.co/datasets/sixstringzen/hemmingway-1-omlx-quantization-benchmark-v1/blob/main/CORRECTION.md).

Each quantization was compared with the local BF16 reference on 15 prompts.
Three judge lanes evaluated both response orders, producing 90 ratings per quant.
The hosted comparison used 11 cap-matched prompts and produced 66 ratings.
The study contains 14 packet files per judge, each holding multiple cases.

The ratings describe one frozen output per condition and prompt. Multiple judges
and swapped orders do not create independent generation samples. We report
counts, percentages, and mean score differences without confidence intervals or
statistical significance claims. Percentages can differ from 100% after rounding.

Judges scored instruction adherence, task fit, clarity and control, and writing
judgment on a 0-4 scale. Scores and overall preferences are separate judgments.
The Claude lane used manual chats, except the final hosted swapped packet,
which used OpenRouter. Gemini and Grok used OpenRouter. The Grok collection
includes documented recovery of packets 04, 05, and 06.

| Condition | Wins | Losses | Ties | Ratings | Win % | Loss % | Tie % |
| --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: |
| oQ2e | 31 | 58 | 1 | 90 | 34.4 | 64.4 | 1.1 |
| oQ3.5e | 44 | 42 | 4 | 90 | 48.9 | 46.7 | 4.4 |
| oQ3e | 54 | 35 | 1 | 90 | 60.0 | 38.9 | 1.1 |
| oQ4e | 30 | 43 | 17 | 90 | 33.3 | 47.8 | 18.9 |
| oQ6e | 34 | 19 | 37 | 90 | 37.8 | 21.1 | 41.1 |
| oQ8e | 27 | 22 | 41 | 90 | 30.0 | 24.4 | 45.6 |

After matching prompts across orders, agreement was 89.1% for Claude, 85.1% for Gemini, and 83.2% for Grok (101 pairs per judge). Inter-rater agreement was 83.7% across 606 pairings. These rates measure agreement on the frozen outputs; they do not validate the judges' preferences.

### This build

Frequent ties with BF16 and more wins than losses in this suite.

### Local generation profile

Local generations ran through MLX/oMLX on an Apple M5 Max with 128 GB unified
memory. Captured server records identify oMLX `0.7.0.dev2`. The generation profile
used temperature 0, top-p 1, min-p 0, repetition penalty 1, and seed 42. MTP and
speculative acceleration were disabled for this baseline. The files preserve
MTP tensors for separate runtime experiments.

The quant runs recorded 90 completions at 512 tokens, 24 at 1024, and 18 at 2048.
The selected BF16 reference recorded 11, one, and three completions at those caps.
The judge packets selected the final cap for each prompt: 11 at 512, one at
1024, and three at 2048. Earlier truncated attempts remain execution records
and were excluded from the judge packets. Hosted retained a separate 11-prompt
comparison because four prompts did not have matching output caps.

Local manifests, generation records, and oMLX/macmon telemetry carry execution
evidence. Grafana displays the telemetry. Background tracing remained active,
so controlled throughput and memory comparisons require a separate run.

This small, fixed suite cannot establish a universal ranking or quantify
token-level fidelity. The higher observed oQ3e win rate does not imply that
reducing precision improves the source model in general. High tie rates for
oQ6e and oQ8e do not establish equivalence to BF16. The study includes no direct
quant-versus-hosted comparison, and the hosted deployment's exact checkpoint,
precision, and runtime are not verified.

Broader task scores and controlled cross-quant runtime comparisons remain pending. Raw generations, provider
responses, packet mappings, and telemetry remain in the local evidence package.
The public v1 dataset contains prompts, selected execution metadata, and blind-preference aggregates.

[Companion blind-preference dataset](https://huggingface.co/datasets/sixstringzen/hemmingway-1-omlx-quantization-benchmark-v1).

### Evidence v2: direct fidelity and runtime

The companion [Hemmingway-1 Quantization Evidence v2](https://huggingface.co/datasets/sixstringzen/hemmingway-1-omlx-quantization-evidence-v2) reports a local BF16-to-oQ4e teacher-forced comparison and a controlled oQ4e runtime measurement. The fidelity data includes full-vocabulary KLD, top-10 agreement, and matched key/value cache summaries. The runtime data includes aggregate prefill, decode, latency, and telemetry summaries from an Apple M5 Max host.

This evidence is separate from the corrected v1 blind-preference study. It does not establish general quality, benchmark-task performance, or a cross-quant runtime ranking.

## Limitations

Quantization can change word choice, coherence, and instruction following. The writing comparison above covers a fixed prompt suite; token-level fidelity remains unmeasured.

The sensitivity pass used oMLX's `oqe_code_multilingual` calibration dataset. No prose-specific calibration dataset was used. The MTP tensors are present in the artifact, but MTP-assisted decoding has not been benchmarked separately.

The original model's documented limitations and acceptable-use guidance also apply to this quantized release.

## License

The source model is released under the Apache 2.0 license. This quantized derivative uses the same license; refer to the [source repository](https://huggingface.co/Altworld/Hemmingway-1) for the upstream model card and attribution.

## Feedback

Send compatibility reports through this repository's Community tab and include your oMLX version and Apple hardware.