--- base_model: IFM/K2-Horizon-MoVA-36B-A4B tags: - mlx - apple-silicon - text-generation - oQ license: apache-2.0 --- # K2-Horizon-MoVA-36B-A4B oQ4e oQ4e (imatrix-enhanced mixed-precision 4-bit) conversion of [IFM/K2-Horizon-MoVA-36B-A4B](https://huggingface.co/IFM/K2-Horizon-MoVA-36B-A4B), a sparse Mixture-of-Experts model with Mixture-of-Values attention (36B total / 4B active parameters, 512K context). **Upstream model:** [IFM/K2-Horizon-MoVA-36B-A4B](https://huggingface.co/IFM/K2-Horizon-MoVA-36B-A4B) by the IFM Team, released under Apache-2.0. **Conversion:** Quantized to MLX format using [Hermes Agent](https://hermes-agent.nousresearch.com) with `mlx-lm` and `oMLX`. ## Quickstart ```bash pip install -U mlx-lm python3 -m mlx_lm.generate \ --model hermitdave/K2-Horizon-MoVA-36B-A4B-oQ4e \ --prompt "Explain why long-context evaluation is difficult." \ --max-tokens 512 --temp 1.0 --top-p 0.95 ``` ## Reasoning K2-Horizon is a reasoning model. Always use `reasoning_effort="high"` for best results: ```python from openai import OpenAI client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY") response = client.chat.completions.create( model="hermitdave/K2-Horizon-MoVA-36B-A4B-oQ4e", messages=[{"role": "user", "content": "Explain quantum entanglement."}], extra_body={"chat_template_kwargs": {"reasoning_effort": "high"}}, ) print("Reasoning:", getattr(response.choices[0].message, "reasoning_content", None)) print("Answer:", response.choices[0].message.content) ``` ## Benchmark Results ### Leaderboard Benchmarks | Benchmark | K2-Horizon-MoVA-36B-A4B | |-----------|------------------------| | tau3-Banking (Agentic tool use) | **26.8** | | Terminal-Bench 2.1 (Agentic terminal use) | **58.6** | | GPQA Diamond (Graduate-level science QA) | 80.8 | | AA-LCR (Long-context reasoning) | 66.3 | Scores in %. See [model card](https://huggingface.co/IFM/K2-Horizon-MoVA-36B-A4B) for full results. ### Head-to-Head Comparison vs Qwen 3.6 35B A3B Independent benchmark comparison — raw scores collected via [BenchLocal](https://github.com/hermitdave/benchlocal), report compiled by Hermes Agent. See [full report](https://hermitdave.github.io/benchmarks/local-moa-report.html). | Benchmark | K2 Score | Qwen Score | Delta | Winner | |-----------|----------|------------|-------|--------| | Bugfix / Coding | 76 | 73 | +3 | K2 | | Information Extraction | 90 | 84 | +6 | K2 | | Formatting & Structured Output | 93 | 92 | +1 | K2 | | Prompt Authority / Safety | 80 | 60 | +20 | K2 | | Reasoning & Maths | 78 | 80 | -2 | Qwen | | Structured Output | 95 | 87 | +8 | K2 | | Tool Call Performance | 93 | 63 | +30 | K2 | | Hermes Agent Capabilities | 54 | 48 | +6 | K2 | | **Average** | **82.4** | **73.4** | **+9.0** | **K2** | **K2 wins 7/8 benchmarks.** Dominates extraction, formatting, tool calling, and safety. Qwen wins only Reasoning & Maths by 2 points. ### Throughput Comparison (oMLX) | Metric | K2 Horizon 36B A4B | Qwen 3.6 35B A3B | Delta | |--------|-------------------|-----------------|-------| | TG TPS (single stream) | 48.1 | 88.4 | +84% Qwen | | TTFT (ms) | 6,300 | 2,371 | -62% Qwen | | TPOT (ms) | 20.9 | 11.4 | -46% Qwen | | Peak Memory (GB) | 22.0 | 21.9 | Tie | **Qwen is ~1.8x faster** due to 3B active parameters (vs 4B) + MTP (multi-token prediction). K2 trades throughput for quality. ### Verdict by Dimension | Dimension | Winner | Notes | |-----------|--------|-------| | **Quality** | K2 Horizon | Wins 7/8 benchmarks, +9.0 avg | | **Throughput** | Qwen 3.6 | ~1.8x faster single-stream and batch | | **Safety** | K2 Horizon | 0 harmful violations vs 1 | | **Agentic Use** | K2 Horizon | Tool call: 93 vs 63, Error recovery: 100 vs 33 | | **Multimodal** | Qwen 3.6 | Native text + image + video | | **Long Context** | Draw | Qwen 1M (ext.) vs K2 512K native | **Bottom line:** For local agent pipelines where tool reliability and structured output matter: **K2 Horizon**. For high-volume batch inference where throughput dominates: **Qwen 3.6**. ## oMLX Patch K2-Horizon requires oMLX v0.6.4+ with the [K2-Horizon support patch (PR #3441)](https://github.com/jundot/omlx/pull/3441). This patch adds: - `k2_horizon` model type support - Reasoning content handling (`` tags) - Tool call parsing (plain text and XML formats) - Multi-turn conversation support Without this patch, oMLX will refuse to load K2-Horizon models with `ValueError: Model type k2_horizon not supported`. ## Chat Template K2-Horizon uses IFM's custom chat template with reasoning and tool calling support. Key tags: | Tag | Purpose | |-----|---------| | ``, ``, `` | Thinking blocks | | `<\|ifm\|im_start|>`, `<\|ifm\|im_end\|>` | Message delimiters | | ``, ``, `` | Tool call structure | All tags are automatically stripped by oMLX before responses reach users. ## Citation ```bibtex @misc{k2horizon2026, title = {Introducing K2 Horizon: Frontier Performance, Radically Open}, author = {{IFM Team}}, year = {2026}, url = {https://ifm.ai/blog/k2/}, } ``` ## License Apache-2.0 (same as upstream).