hermitdave's picture
Add files using upload-large-folder tool
ecaf5f4 verified
|
Raw
History Blame
2.05 kB

K2-Horizon-MoVA-36B-A4B MLX

MLX conversions of IFM/K2-Horizon-MoVA-36B-A4B, a sparse Mixture-of-Experts model with Mixture-of-Values attention (36B total / 4B active parameters).

Available Formats

Format Size Quality Use Case
oQ4e ~21 GB ~uniform 6-bit quality Best quality-per-GB
6-bit ~28 GB High Quality-focused, fits 40+ GB
8-bit ~40 GB Near-lossless Reference quality, 64 GB+

Quickstart

pip install -U mlx-lm

# Generate
python3 -m mlx_lm.generate \
  --model hermitdave/K2-Horizon-MoVA-36B-A4B-MLX-8bit \
  --prompt "Explain why long-context evaluation is difficult." \
  --max-tokens 512 --temp 1.0 --top-p 0.95

Reasoning

K2-Horizon is a reasoning model. Always use reasoning_effort="high" for best results:

from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
response = client.chat.completions.create(
    model="hermitdave/K2-Horizon-MoVA-36B-A4B-MLX-8bit",
    messages=[{"role": "user", "content": "Explain quantum entanglement."}],
    extra_body={"chat_template_kwargs": {"reasoning_effort": "high"}},
)
print("Reasoning:", getattr(response.choices[0].message, "reasoning_content", None))
print("Answer:", response.choices[0].message.content)

Benchmark Results

Benchmark K2-Horizon-MoVA-36B-A4B
tau3-Banking (Agentic tool use) 26.8
Terminal-Bench 2.1 (Agentic terminal use) 58.6
GPQA Diamond (Graduate-level science QA) 80.8
AA-LCR (Long-context reasoning) 66.3

Scores in %. See model card for full results.

Citation

@misc{k2horizon2026,
  title  = {Introducing K2 Horizon: Frontier Performance, Radically Open},
  author = {{IFM Team}},
  year   = {2026},
  url    = {https://ifm.ai/blog/k2/},
}

License

Apache-2.0 (same as upstream).