codelion's picture
Hemmingway-1 OptiQ mixed-precision quant (recipe from Qwen3.5-27B)
2394caf verified
|
Raw History Blame Contribute Delete
4.05 kB
---
library_name: mlx
license: apache-2.0
pipeline_tag: text-generation
base_model: Altworld/Hemmingway-1
base_model_relation: quantized
tags:
- mlx
- quantized
- mixed-precision
- 4bit
- 8bit
- optiq
- apple-silicon
- text-generation
- qwen3.5
- writing
---
# mlx-community/Hemmingway-1-OptiQ-4bit
> **Built with [mlx-optiq](https://mlx-optiq.com)**, the MLX-native toolkit to quantize, fine-tune, and serve LLMs locally on Apple Silicon, no PyTorch and no cloud. [Try the Lab](https://mlx-optiq.com/docs/lab/) · [All OptiQ quants](https://mlx-optiq.com/models) · [Docs](https://mlx-optiq.com/docs/)
A mixed-precision MLX quant of [Altworld/Hemmingway-1](https://huggingface.co/Altworld/Hemmingway-1), a writing-oriented variant of the Qwen3.5 27B architecture. Sensitive layers are kept at 8-bit and robust ones at 4-bit, rather than crushing everything to a uniform width.
## Quantization details
| Property | Value |
|---|---|
| Predominant precision | 4-bit |
| Layers at 8-bit | 219 |
| Layers at 4-bit | 279 |
| Size on disk | 18 GB (from ~50.9 GB bf16) |
| Group size | 64 |
## How the bit-widths were chosen
Stated plainly, because it differs from most OptiQ quants: **the per-layer
allocation was not measured on this model.** It was transferred from
[mlx-community/Qwen3.5-27B-OptiQ-4bit](https://huggingface.co/mlx-community/Qwen3.5-27B-OptiQ-4bit),
whose allocation came from a KL-divergence sensitivity sweep over a six-domain
calibration mix (prose, reasoning, code, agent, tool-call, instructions).
That transfer is sound here because the two share an architecture exactly —
`qwen3_5_text`, 64 layers, 24 attention heads, 4 KV heads, head_dim 256, hidden
5120, vocab 248,320 — so every layer in the recipe has a counterpart with the
same role and shape. All **498 tensors matched with none unmatched**, which is
the check that matters: an unmatched tensor would silently fall back to flat
4-bit and make this a uniform quant wearing a mixed-precision name.
What sensitivity measures is how much a layer's *role in the architecture*
suffers from precision loss. What it cannot know is whether this model's own
training moved that sensitivity around. If you are quantizing your own
fine-tune and want the allocation measured against it, run `optiq convert` and
let the sweep do it.
## What was verified
- 498/498 tensors matched the recipe, 0 unmatched.
- Generation checked for correctness, not just fluency: factual recall, arithmetic
with working shown (240 km in 3 h → 80 km/h), an iterative Fibonacci that runs,
and a technical explanation.
- OptiQ's release contract (artifact layout, metadata, mixed-precision assertions).
**Not run for this model:** the six-metric Capability Score. The published
scores for the Qwen3.5-27B quant describe *that* model, not this one, and are
not claimed here.
## Prose style
The variant is writing-oriented, and it measures that way against the two
signals that actually separate human from machine prose on our labelled set —
em-dashes per 1k words and average sentence length. Same three prompts, same
sampling, against the base Qwen3.5-27B quant:
| | Hemmingway-1 | Qwen3.5-27B | human | AI |
|---|---|---|---|---|
| em-dashes / 1k words | 0.0 | 0.0 | ~0 | 7.0 |
| average sentence | 16.2 words | 20.9 words | 17.4 | 20.4 |
Em-dashes do not separate the two. Sentence length does: the base sits on the
AI median, this one on the human median. A small probe, not a benchmark.
## Use it
```bash
pip install mlx-optiq
optiq serve --model mlx-community/Hemmingway-1-OptiQ-4bit
```
Or with `mlx-lm` directly:
```python
from mlx_lm import generate, load
model, tokenizer = load("mlx-community/Hemmingway-1-OptiQ-4bit")
prompt = tokenizer.apply_chat_template(
[{"role": "user", "content": "Write three sentences about shipping software."}],
add_generation_prompt=True, tokenize=False)
print(generate(model, tokenizer, prompt=prompt, max_tokens=256))
```
The per-layer bit map is in `optiq/metadata.json` and in the `quantization`
block of `config.json`.