Instructions to use midium-ai/decider-4b-dwq-4bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use midium-ai/decider-4b-dwq-4bit with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] huggingface-cli download --local-dir decider-4b-dwq-4bit midium-ai/decider-4b-dwq-4bit
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
decider-4b · DWQ 4-bit (MLX)
A 4-bit MLX build of Mapika/decider-4b, quantized with DWQ (distilled weight quantization). It's the on-device decision model used by Midium.
Decider is a "System One" model. It doesn't generate text. You give it a state (a message, a document, a reply) and a typed question with an explicit list of options. From one forward pass, it returns a probability for each option: "where should this answer come from?", "did this reply answer the question?", "does this request need the Slack tool?". No decoding, no parsing, and no answers outside the options you defined.
| Size on disk | Peak memory | Accuracy (14 sets pooled) | |
|---|---|---|---|
| Original (BF16) | 8.4 GB | 9.0 GB | 96.3% |
| This repo (DWQ 4-bit) | 2.4 GB | 3.4 GB | 95.5% |
| Plain 4-bit (affine, group 64) | 2.4 GB | 3.4 GB | 94.8% |
DWQ recovers about half of what plain 4-bit quantization loses, at the same size, and stays within about one point of the original on most sets.
How it was made
- Source:
Mapika/decider-4bat revisioneb5fbdf(v2.1), built on Qwen/Qwen3.5-4B-Base. - Method:
mlx_lmDWQ: 4-bit weights, group size 64. The quantization scales and biases were distilled against the BF16 model's outputs on 2,048 calibration samples (MMLU-Pro questions in Decider's own prompt layout, ending in the answer letter, plus WikiText; 513 tokens each), learning rate 1e-6, seed 123. Final validation loss 0.011. - Unchanged:
decider_config.jsonkeeps the original per-type temperatures (choice 1.11, yes/no 1.56, score 1.287), isolated score levels and the plain prompt layout.
Evaluation
Midium's decision sets, each question asked zero-shot with the questions Midium uses in production, scored with the same harness for every build. None of these sets was used for calibration.
| Task (n) | Majority baseline | BF16 | Plain 4-bit | DWQ 4-bit |
|---|---|---|---|---|
| Routing: vault / general / web / chat (140) | 42.9% | 95.0% | 92.9% | 92.9% |
| Routing, held-out (401) | 45.1% | 95.3% | 92.0% | 94.5% |
| Routing, hard (155) | 49.7% | 82.6% | 81.9% | 81.9% |
| Routing, generated in 16 languages (1,168) | 25.9% | 96.0% | 93.0% | 95.6% |
| Reply check: answered / abstained / clarify (180) | 44.4% | 96.1% | 96.1% | 96.1% |
| Reply check, held-out (400) | 69.8% | 97.8% | 95.5% | 98.0% |
| Reply check, hard (150) | 53.3% | 86.7% | 89.3% | 90.0% |
| Reply check, generated (833) | 34.8% | 96.6% | 93.3% | 95.9% |
| Action risk: low / medium / high (153) | 36.0% | 77.8% | 81.7% | 77.8% |
| Tool needed, per tool (3,564) | 95.3% | 97.3% | 96.2% | 96.2% |
| Tool needed, hard (2,700) | 96.5% | 97.0% | 96.7% | 96.3% |
| Model need: quick / reasoning / code / media, hard (118) | 36.4% | 95.8% | 96.6% | 95.8% |
| Model need, generated (1,102) | 27.3% | 97.7% | 95.9% | 97.3% |
| Model need, public benchmarks (1,200) | 25.0% | 95.5% | 92.3% | 94.4% |
Expected calibration error, pooled: BF16 0.089, DWQ 0.074, plain 4-bit 0.030. If you act on cutoffs, choose them on your own data for the build you deploy.
Speed is about 52 ms per question for both 4-bit builds and 55 ms for BF16, on an Apple M5 Max (one question per forward pass, prompt prefilled per question). The gain is memory, not speed.
For comparison, on the same sets the smaller midium-ai/decider-2b-dwq-4bit (1.06 GB) scores 92.9% pooled. This model is 8–16 points better on routing and reply checks, at about twice the latency.
Usage
The prompt format is plain text: a context, then one question, then the option letters. The answer is the model's probability for each option letter at the Answer: ( position.
import mlx.core as mx
from mlx_lm import load
model, tok = load("midium-ai/decider-4b-dwq-4bit", model_config={"model_type": "qwen3_5"})
options = {
"vault": "The company's own internal documents, policies, records or data.",
"general": "General knowledge or reasoning alone; nothing company-specific or current.",
"web": "Live or recent information from the internet.",
"chat": "Nothing: it is conversational, or about the conversation itself.",
}
prompt = (
"Context:\nHow many vacation days do we get each year?\n\n"
"Question: Where should the assistant get the information to answer this message?\nOptions:"
+ "".join(f"\n({'ABCD'[i]}) {k}: {v}" for i, k in enumerate(options))
+ "\nAnswer: ("
)
logits = model(mx.array([tok.encode(prompt, add_special_tokens=False)]))[0, -1]
letters = [tok.encode(c, add_special_tokens=False)[0] for c in "ABCD"]
probs = mx.softmax(logits[mx.array(letters)].astype(mx.float32) / 1.11) # choice temperature
print(dict(zip(options, probs.tolist())))
The model_config override is needed because the checkpoint ships a text-only qwen3_5_text config, and mlx_lm loads it through its qwen3_5 module. For several questions about the same context, prefill the context once and score each question on a copy of the cache. For yes/no questions the options are (A) no and (B) yes at temperature 1.56. See the upstream repo for score questions (asked as one isolated yes/no row per level) and the full question format.
Limitations
- Everything in the original's model card applies here.
- Quantization cost: the largest drops against BF16 are 2 points on the small routing set and about 1 point on tool and model-need questions. Action risk is unchanged.
- Calibration: the DWQ build is less well calibrated than plain 4-bit, and all builds are overconfident on hard, unfamiliar questions. Choose cutoffs on your own data.
- Option wording and order: answers depend on how clearly the options are described. Reordering options changes a few percent of decisions.
- Tested scope: evaluated on MLX (Apple Silicon) only.
License and attribution
Apache-2.0, the same as the original. The quantized weights are a modified version of Mapika/decider-4b (Apache-2.0), itself built on Qwen/Qwen3.5-4B-Base (Apache-2.0). The only change is the DWQ 4-bit quantization described above.
- Downloads last month
- 65
4-bit