horizon-draft-0.5b / README.md
nickalessi's picture
Horizon draft 0.5B: weights, tokenizer, model card
200493f verified
|
Raw History Blame Contribute Delete
2.83 kB
---
license: apache-2.0
license_note: derived from Qwen2.5-0.5B (Apache-2.0); base license retained
base_model: Qwen/Qwen2.5-0.5B
tags:
- speculative-decoding
- draft-model
- llama.cpp
- code-generation
language:
- en
---
# Horizon Draft 0.5B
A tiny **standalone draft model** for speculative decoding with
`DeepSeek-R1-Distill-Qwen-7B` as the target, built for consumer CPUs and
`llama.cpp -md`. Part of the [Horizon](https://github.com/Vexp-ai/horizon)
project: a verification-first layer for local LLMs, by the team behind
[vexp](https://vexp.dev).
## What it is
- **Base:** Qwen2.5-0.5B, full-finetuned for one epoch (~2.5 h on a single
24 GB GPU, ~$3 of compute) on 30k reasoning traces distilled from
DeepSeek-R1 (13k math, 13k code, 4k science), formatted with the
**target's chat template**.
- **Vocabulary-aligned with the target:** trained with the target's
tokenizer and the embedding padded to n_vocab 152064, so it passes
llama.cpp's strict speculative vocabulary check and drops straight into
`-md`. A stock Qwen2.5-0.5B does not pair (different special tokens,
151936 vocab), and through permissive paths it reaches only τ≈1.
## Measured results (mainstream DDR4 desktop, Ryzen 9 3900X, 12 threads)
| Config | Code generation | Reasoning segment | RSS |
|---|---|---|---|
| 7B Q4_K_M autoregressive | 8.1 t/s | 8.1 t/s | 7.7 GB |
| + this drafter (Q8_0, γ=8) | **10.7-15.1 t/s (mean 13.3, 1.65×)** | 7.1-9.4 (neutral) | 8.3 GB |
| + same-family 1.5B draft | 7.1-10.9 (no gain) | 5.3 (hurts) | 9.4 GB |
Acceptance on three coding tasks: **36-56%** on code, 20-30% on
chain-of-thought text. Two practical lessons we ship with the numbers:
on CPU the draft must be nearly free (a 1.5B draft with similar acceptance
gains nothing), and speculation pays on code, not on reasoning: enable it
per segment.
## Usage (llama.cpp)
```bash
llama-speculative -m DeepSeek-R1-Distill-Qwen-7B-Q4_K_M.gguf \
-md horizon-draft-0.5b-q8_0.gguf \
--spec-draft-n-max 8 -t 12 -n 512 --temp 0 -p "..."
```
Caveat: some recent `llama-server` builds silently skip speculative
decoding (unified KV cache default) while still loading the draft. Verify
acceptance stats are non-zero; the `llama-speculative` binary is the
reliable path today.
## Limitations
- Target-specific by design: it predicts DeepSeek-R1-Distill-Qwen-7B's
output distribution. Pair it with other targets at your own risk.
- Trained for one epoch on 30k traces: acceptance has headroom (more
epochs, more diverse code styles, DSpark-style parallel drafting are the
declared next steps).
- Not an instruct model: do not use it standalone for generation.
## Reproduce
Training script (`train/train_draft.py`), corpus recipe and the full
measurement logs are in the
[Horizon repository](https://github.com/Vexp-ai/horizon).