How to use from
OpenClaw
Start the MLX server
# Install MLX LM:
uv tool install mlx-lm
# Start a local OpenAI-compatible server:
mlx_lm.server --model "mlx-community/Hemmingway-1-OptiQ-4bit"
Configure OpenClaw
# Install OpenClaw:
npm install -g openclaw@latest
# Register the local server and set it as the default model:
openclaw onboard --non-interactive --mode local \
  --auth-choice custom-api-key \
  --custom-base-url http://127.0.0.1:8080/v1 \
  --custom-model-id "mlx-community/Hemmingway-1-OptiQ-4bit" \
  --custom-provider-id mlx-lm \
  --custom-compatibility openai \
  --custom-text-input \
  --accept-risk \
  --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Quick Links

mlx-community/Hemmingway-1-OptiQ-4bit

Built with mlx-optiq, the MLX-native toolkit to quantize, fine-tune, and serve LLMs locally on Apple Silicon, no PyTorch and no cloud. Try the Lab · All OptiQ quants · Docs

A mixed-precision MLX quant of Altworld/Hemmingway-1, a writing-oriented variant of the Qwen3.5 27B architecture. Sensitive layers are kept at 8-bit and robust ones at 4-bit, rather than crushing everything to a uniform width.

Quantization details

Property Value
Predominant precision 4-bit
Layers at 8-bit 219
Layers at 4-bit 279
Size on disk 18 GB (from ~50.9 GB bf16)
Group size 64

How the bit-widths were chosen

Stated plainly, because it differs from most OptiQ quants: the per-layer allocation was not measured on this model. It was transferred from mlx-community/Qwen3.5-27B-OptiQ-4bit, whose allocation came from a KL-divergence sensitivity sweep over a six-domain calibration mix (prose, reasoning, code, agent, tool-call, instructions).

That transfer is sound here because the two share an architecture exactly — qwen3_5_text, 64 layers, 24 attention heads, 4 KV heads, head_dim 256, hidden 5120, vocab 248,320 — so every layer in the recipe has a counterpart with the same role and shape. All 498 tensors matched with none unmatched, which is the check that matters: an unmatched tensor would silently fall back to flat 4-bit and make this a uniform quant wearing a mixed-precision name.

What sensitivity measures is how much a layer's role in the architecture suffers from precision loss. What it cannot know is whether this model's own training moved that sensitivity around. If you are quantizing your own fine-tune and want the allocation measured against it, run optiq convert and let the sweep do it.

What was verified

  • 498/498 tensors matched the recipe, 0 unmatched.
  • Generation checked for correctness, not just fluency: factual recall, arithmetic with working shown (240 km in 3 h → 80 km/h), an iterative Fibonacci that runs, and a technical explanation.
  • OptiQ's release contract (artifact layout, metadata, mixed-precision assertions).

Not run for this model: the six-metric Capability Score. The published scores for the Qwen3.5-27B quant describe that model, not this one, and are not claimed here.

Prose style

The variant is writing-oriented, and it measures that way against the two signals that actually separate human from machine prose on our labelled set — em-dashes per 1k words and average sentence length. Same three prompts, same sampling, against the base Qwen3.5-27B quant:

Hemmingway-1 Qwen3.5-27B human AI
em-dashes / 1k words 0.0 0.0 ~0 7.0
average sentence 16.2 words 20.9 words 17.4 20.4

Em-dashes do not separate the two. Sentence length does: the base sits on the AI median, this one on the human median. A small probe, not a benchmark.

Use it

pip install mlx-optiq
optiq serve --model mlx-community/Hemmingway-1-OptiQ-4bit

Or with mlx-lm directly:

from mlx_lm import generate, load

model, tokenizer = load("mlx-community/Hemmingway-1-OptiQ-4bit")
prompt = tokenizer.apply_chat_template(
    [{"role": "user", "content": "Write three sentences about shipping software."}],
    add_generation_prompt=True, tokenize=False)
print(generate(model, tokenizer, prompt=prompt, max_tokens=256))

The per-layer bit map is in optiq/metadata.json and in the quantization block of config.json.

Downloads last month
474
Safetensors
Model size
27B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mlx-community/Hemmingway-1-OptiQ-4bit

Base model

Qwen/Qwen3.8-27B
Quantized
(42)
this model