---
license: apache-2.0
license_link: https://huggingface.co/IFM/K2-Horizon-3.7B
base_model: IFM/K2-Horizon-3.7B
base_model_relation: quantized
library_name: mlx
pipeline_tag: text-generation
tags:
- mlx
- quantized
- dense
- k2-horizon
---
# K2-Horizon-3.7B (MLX, 4-bit)
Mixed-precision **4-bit** MLX quantization of [IFM/K2-Horizon-3.7B](https://huggingface.co/IFM/K2-Horizon-3.7B), converted from revision `943ce4e`. **6.5 bits/weight effective**, 4.1 GB on disk. For Apple silicon.
K2-Horizon-3.7B is IFM's small dense K2-Horizon model: a 3.7B decoder-only model with a 512K (524,288-token) context window.
## Requirements
mlx-lm doesn't support the `k2_horizon` architecture yet. There's an open request: [mlx-lm#1876](https://github.com/ml-explore/mlx-lm/issues/1876). Until support lands, **this repo ships the MLX model code** ([`k2_horizon.py`](./k2_horizon.py)), which mlx-lm loads through the `model_file` entry in `config.json`. So the released mlx-lm works as-is:
```bash
pip install -U mlx-lm
```
Pass `--trust-remote-code` (or `trust_remote_code=True`). mlx-lm 0.31.3 loads the file without it, but newer versions require it. The file runs on your machine, so read it first if you like; its comments describe how it differs from IFM's PyTorch code.
Once K2-Horizon support lands in a released mlx-lm, this repo will be updated to use it. If something breaks after that, open a discussion here asking for a re-upload.
## How it was quantized
Mixed precision, chosen from measurements rather than a fixed rule (integer affine quantization, bf16 scales and biases):
- **4-bit, group size 32**: the MLP weights (`gate_proj`, `up_proj`, `down_proj`). These hold most of the parameters.
- **8-bit, group size 64**: everything else: attention, the embeddings and `lm_head`.
Plain 4-bit (everything 4-bit, group size 64) loses too much on this model: **+21.5%** WikiText-2 perplexity versus bf16. This recipe loses **+4.4%**. On all three dense K2-Horizon models, plain 4-bit lost 18-22%, far more than on the 36B MoE model, so the dense 4-bit repos use this mixed recipe.
## Memory
Peak **4.3 GB** for a short prompt; fits a **16 GB** Mac. The KV cache adds about 144 KB per token in bf16 (18.0 GB at 128K tokens), so long contexts need more memory; `--max-kv-size` and KV-cache quantization (`--kv-bits 8`, where available) reduce it.
## Conversion check
The MLX implementation was checked against IFM's PyTorch code (`modeling_k2_horizon.py`) in **fp32 on the real weights of this model**:
- Layer by layer, all 36 layers match to a relative error of 4e-6 or better, and the next-token predictions agree at every position.
- Token by token: the full model in fp32 greedily generated with the KV cache, and the PyTorch model picked the same token at every step (256/256 tokens across English, code, math and Chinese prompts). Cached and uncached outputs also agree at every step.
Smoke-tested after conversion with released mlx-lm 0.31.3: `17 * 23` → `391` and "capital of Australia" → `Canberra`, both ending normally. On a **Mac Studio M4 Max 128GB**: 109.6 tok/s generation, peak 4.3 GB (short prompt).
## Benchmarks (all K2-Horizon-3.7B MLX variants)
WikiText-2 test perplexity (128 × 512 tokens, lower is better) and generation speed on an M4 Max 128GB, all measured the same way:
| | [bf16](https://huggingface.co/mlx-community/K2-Horizon-3.7B-bf16) | [8-bit](https://huggingface.co/mlx-community/K2-Horizon-3.7B-8bit) | [4-bit](https://huggingface.co/mlx-community/K2-Horizon-3.7B-4bit) |
|---|---|---|---|
| Bits/weight | 16 | 8.5 | 6.5 |
| Disk | 10.1 GB | 5.4 GB | 4.1 GB |
| Peak memory | 10.2 GB | 5.5 GB | 4.3 GB |
| WikiText-2 perplexity | 17.552 | 17.544 (-0.05%) | 18.324 (+4.4%) |
| Generation | 50.3 tok/s | 84.2 tok/s | 109.6 tok/s |
Perplexity is a coarse signal. Test the versions on your own workload before picking one. Other K2-Horizon sizes: the [K2-Horizon collection](https://huggingface.co/collections/mlx-community/k2-horizon).
## Usage
```bash
mlx_lm.generate --model mlx-community/K2-Horizon-3.7B-4bit --trust-remote-code --prompt "Explain mixture-of-experts in two sentences." --max-tokens 2048
```
```python
from mlx_lm import load, generate
# Newer mlx-lm versions need trust_remote_code=True; on mlx-lm 0.31.3 use load("mlx-community/K2-Horizon-3.7B-4bit").
model, tokenizer = load("mlx-community/K2-Horizon-3.7B-4bit", trust_remote_code=True)
messages = [{"role": "user", "content": "Explain mixture-of-experts in two sentences."}]
prompt = tokenizer.apply_chat_template(messages, add_generation_prompt=True)
print(generate(model, tokenizer, prompt, max_tokens=2048))
```
`mlx_lm.chat` and `mlx_lm.server` take the same `--trust-remote-code` flag.
The model thinks before it answers, inside ` … `. Leave room for that in `max_tokens`. The effort level is set with `reasoning_effort` in the chat template: `"high"` (default), `"medium"` or `"low"`, e.g. `apply_chat_template(messages, add_generation_prompt=True, reasoning_effort="low")`.
**Notes:**
- **Server output:** mlx-lm doesn't recognize the `` tags yet, so `mlx_lm.server` returns the thinking text inside `content`, before ``, rather than in a separate `reasoning` field. K2-Horizon's tool-call format isn't parsed yet either.
- **Chat template change:** the original template raises an error when an earlier assistant message has no thinking field, which is what OpenAI-style clients send. The template here renders empty thinking for those messages instead. Nothing else was changed.
## License
[Apache-2.0](https://huggingface.co/IFM/K2-Horizon-3.7B), inherited from the base model. Refer to the [original model card](https://huggingface.co/IFM/K2-Horizon-3.7B) for architecture, benchmarks and intended use. All credit for the model belongs to IFM.