--- library_name: mlx license: apache-2.0 language: - en base_model: Altworld/Hemmingway-1 base_model_relation: quantized pipeline_tag: text-generation tags: - mlx - omlx - oq - oqe - quantized - 2-bit - mtp - apple-silicon - creative-writing --- # Hemmingway-1 oQ2e with MTP This repository contains an enhanced oQ2e quantization of [Altworld/Hemmingway-1](https://huggingface.co/Altworld/Hemmingway-1) for MLX and oMLX on Apple silicon. The conversion preserves the model's multi-token prediction (MTP) tensors. Altworld developed and published the source model. [sixstringzen](https://huggingface.co/sixstringzen) performed this conversion and published the converted weights with their quantization report. The original model, its intended use, and its training details remain documented in the [source model card](https://huggingface.co/Altworld/Hemmingway-1). ## Quantization set This repository is part of the [Hemmingway-1 oMLX oQe Quantizations](https://huggingface.co/collections/sixstringzen/hemmingway-1-omlx-oqe-quantizations-apple-silicon-evidence-6ab06103552bc3c7e9ad8afa) collection. Every build in the set uses the same source revision, group size, non-quantized dtype, calibration pass, and MTP preservation policy. | Build | Base precision | Output size | | --- | ---: | ---: | | [oQ2e](https://huggingface.co/sixstringzen/Hemmingway-1-oQ2e-mtp) | 2-bit | 10.14 GiB | | [oQ3e](https://huggingface.co/sixstringzen/Hemmingway-1-oQ3e-mtp) | 3-bit | 12.22 GiB | | [oQ3.5e](https://huggingface.co/sixstringzen/Hemmingway-1-oQ3.5e-mtp) | 3-bit with additional higher-precision overrides | 13.19 GiB | | [oQ4e](https://huggingface.co/sixstringzen/Hemmingway-1-oQ4e-mtp) | 4-bit | 15.21 GiB | | [oQ6e](https://huggingface.co/sixstringzen/Hemmingway-1-oQ6e-mtp) | 6-bit | 21.39 GiB | | [oQ8e](https://huggingface.co/sixstringzen/Hemmingway-1-oQ8e-mtp) | 8-bit | 27.10 GiB | ## Quantization details | Item | Value | | --- | --- | | Source model | `Altworld/Hemmingway-1` | | Source revision | `4d711aac0f0043075ae334d2a3de3db3e10135c9` | | Quantizer | [oMLX](https://github.com/jundot/omlx) `0.7.0.dev2` | | Method | Enhanced oQ2e mixed-precision affine quantization | | Base precision | 2-bit | | Group size | 64 | | Non-quantized dtype | `bfloat16` | | Higher-precision tensors | 8 tensors at 4-bit, 157 at 5-bit, and `language_model.lm_head` at 8-bit | | Calibration dataset | `oqe_code_multilingual` | | Calibration shape | 128 samples at 512 tokens | | Imatrix entries | 504 | | Imatrix cache | Reused from the matching source-model sensitivity pass | | MTP tensors | 29 preserved tensors | | Output size | 10,884,899,235 bytes (10.14 GiB) | oQe uses activation importance to assign additional precision to sensitive tensors. This build uses 2-bit weights as its base, with mixed-precision overrides ranging from 4 to 8 bits. The quantization report records no matrix-shape mismatches and no missing weight shards. The included [`oq_imatrix_report.json`](https://huggingface.co/sixstringzen/Hemmingway-1-oQ2e-mtp/blob/main/oq_imatrix_report.json) records the sensitivity pass, calibration settings, tensor coverage, and fallback. Strict imatrix coverage was disabled for the known `language_model.lm_head` fallback. ## Compatibility This model was created with oMLX `0.7.0.dev2`. The source model identifies its text architecture as `qwen3_5_text`; the converted artifact uses `qwen3_5`, which matches the architecture name supported by this oMLX build. The weights use MLX safetensors and are not GGUF files. Compatibility with other MLX runtimes or earlier oMLX releases has not been verified. ## Use with oMLX Download `sixstringzen/Hemmingway-1-oQ2e-mtp` from the oMLX model browser, then load it as an LLM. Set `enable_thinking` to `false` when you want direct prose without visible planning. Runtime defaults and the registered model identifier can vary with the local oMLX installation. ## Verification The finished artifact passed local structural checks on 2026-09-20. It contains three safetensors shards, 1,876 indexed tensors, and 29 MTP tensors. The index references no missing shards. These checks confirm that the artifact is complete and internally consistent. A generation smoke test has not been recorded for this quantization, and the checks do not establish quality parity with the BF16 source model. ## Quality evaluation (v1, corrected analysis revision 2) Corrected on 2026-09-22 after identifying errors in A/B decoding and normal/swapped prompt matching. The generated responses and judge records are unchanged. See the [correction record](https://huggingface.co/datasets/sixstringzen/hemmingway-1-omlx-quantization-benchmark-v1/blob/main/CORRECTION.md). Each quantization was compared with the local BF16 reference on 15 prompts. Three judge lanes evaluated both response orders, producing 90 ratings per quant. The hosted comparison used 11 cap-matched prompts and produced 66 ratings. The study contains 14 packet files per judge, each holding multiple cases. The ratings describe one frozen output per condition and prompt. Multiple judges and swapped orders do not create independent generation samples. We report counts, percentages, and mean score differences without confidence intervals or statistical significance claims. Percentages can differ from 100% after rounding. Judges scored instruction adherence, task fit, clarity and control, and writing judgment on a 0-4 scale. Scores and overall preferences are separate judgments. The Claude lane used manual chats, except the final hosted swapped packet, which used OpenRouter. Gemini and Grok used OpenRouter. The Grok collection includes documented recovery of packets 04, 05, and 06. | Condition | Wins | Losses | Ties | Ratings | Win % | Loss % | Tie % | | --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | | oQ2e | 31 | 58 | 1 | 90 | 34.4 | 64.4 | 1.1 | | oQ3.5e | 44 | 42 | 4 | 90 | 48.9 | 46.7 | 4.4 | | oQ3e | 54 | 35 | 1 | 90 | 60.0 | 38.9 | 1.1 | | oQ4e | 30 | 43 | 17 | 90 | 33.3 | 47.8 | 18.9 | | oQ6e | 34 | 19 | 37 | 90 | 37.8 | 21.1 | 41.1 | | oQ8e | 27 | 22 | 41 | 90 | 30.0 | 24.4 | 45.6 | After matching prompts across orders, agreement was 89.1% for Claude, 85.1% for Gemini, and 83.2% for Grok (101 pairs per judge). Inter-rater agreement was 83.7% across 606 pairings. These rates measure agreement on the frozen outputs; they do not validate the judges' preferences. ### This build BF16 was preferred more often: 58 losses, 31 wins, and one tie. ### Local generation profile Local generations ran through MLX/oMLX on an Apple M5 Max with 128 GB unified memory. Captured server records identify oMLX `0.7.0.dev2`. The generation profile used temperature 0, top-p 1, min-p 0, repetition penalty 1, and seed 42. MTP and speculative acceleration were disabled for this baseline. The files preserve MTP tensors for separate runtime experiments. The quant runs recorded 90 completions at 512 tokens, 24 at 1024, and 18 at 2048. The selected BF16 reference recorded 11, one, and three completions at those caps. The judge packets selected the final cap for each prompt: 11 at 512, one at 1024, and three at 2048. Earlier truncated attempts remain execution records and were excluded from the judge packets. Hosted retained a separate 11-prompt comparison because four prompts did not have matching output caps. Local manifests, generation records, and oMLX/macmon telemetry carry execution evidence. Grafana displays the telemetry. Background tracing remained active, so controlled throughput and memory comparisons require a separate run. This small, fixed suite cannot establish a universal ranking or quantify token-level fidelity. The higher observed oQ3e win rate does not imply that reducing precision improves the source model in general. High tie rates for oQ6e and oQ8e do not establish equivalence to BF16. The study includes no direct quant-versus-hosted comparison, and the hosted deployment's exact checkpoint, precision, and runtime are not verified. Broader task scores and controlled cross-quant runtime comparisons remain pending. Raw generations, provider responses, packet mappings, and telemetry remain in the local evidence package. The public v1 dataset contains prompts, selected execution metadata, and blind-preference aggregates. [Companion blind-preference dataset](https://huggingface.co/datasets/sixstringzen/hemmingway-1-omlx-quantization-benchmark-v1). ### Evidence v2: direct fidelity and runtime The companion [Hemmingway-1 Quantization Evidence v2](https://huggingface.co/datasets/sixstringzen/hemmingway-1-omlx-quantization-evidence-v2) reports a local BF16-to-oQ4e teacher-forced comparison and a controlled oQ4e runtime measurement. The fidelity data includes full-vocabulary KLD, top-10 agreement, and matched key/value cache summaries. The runtime data includes aggregate prefill, decode, latency, and telemetry summaries from an Apple M5 Max host. This evidence is separate from the corrected v1 blind-preference study. It does not establish general quality, benchmark-task performance, or a cross-quant runtime ranking. ## Limitations oQ2e is an extreme-compression build. Quantization can change word choice, coherence, and instruction following more visibly at this level. The writing comparison above covers a fixed prompt suite; token-level fidelity remains unmeasured. The sensitivity pass used oMLX's `oqe_code_multilingual` calibration dataset. No prose-specific calibration dataset was used. The MTP tensors are present in the artifact, but MTP-assisted decoding has not been benchmarked separately. The original model's documented limitations and acceptable-use guidance also apply to this quantized release. ## License The source model is released under the Apache 2.0 license. This quantized derivative uses the same license; refer to the [source repository](https://huggingface.co/Altworld/Hemmingway-1) for the upstream model card and attribution. ## Feedback Send compatibility reports through this repository's Community tab and include your oMLX version and Apple hardware.