--- base_model: IFM/K2-Horizon-7B-Uno tags: - mlx - apple-silicon - text-generation - uno - oQ - oQ4e license: apache-2.0 --- # K2-Horizon-7B-Uno oQ4e oQ4e (imatrix-enhanced mixed-precision ~4.5 BPW) quantization of the merged [IFM/K2-Horizon-7B-Uno](https://huggingface.co/IFM/K2-Horizon-7B-Uno) model — a diffusion-augmented LLM based on K2-Horizon-7B. The LoRA adapter is baked into the base weights, so it runs as a standard autoregressive model. **Upstream model:** [IFM/K2-Horizon-7B-Uno](https://huggingface.co/IFM/K2-Horizon-7B-Uno) by Institute of Foundation Models, released under Apache 2.0. **Conversion:** Merged using PyTorch + PEFT, validated with text generation, then quantized to MLX format using [Hermes Agent](https://hermes-agent.nousresearch.com) with `oMLX`. ## Quickstart ```bash pip install -U mlx-lm python3 -m mlx_lm.generate \ --model hermitdave/K2-Horizon-7B-Uno-oQ4e \ --prompt "Explain step by step." \ --max-tokens 512 --temp 1.0 --top-p 0.95 ``` ## Reasoning K2-Horizon-7B is a reasoning model. Always use `reasoning_effort="high"`: ```python from openai import OpenAI client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY") response = client.chat.completions.create( model="hermitdave/K2-Horizon-7B-Uno-oQ4e", messages=[{"role": "user", "content": "Explain step by step."}], extra_body={"chat_template_kwargs": {"reasoning_effort": "high"}}, ) print("Reasoning:", getattr(response.choices[0].message, "reasoning_content", None)) print("Answer:", response.choices[0].message.content) ``` ## oMLX Patch K2-Horizon requires oMLX v0.6.4+ with the [K2-Horizon support patch](https://github.com/jundot/omlx/pull/3441). ## Citation ```bibtex @misc{k2_horizon_7b_uno, title = {K2-Horizon-7B-Uno}, author = {Institute of Foundation Models}, year = {2026}, howpublished = {\url{https://huggingface.co/IFM/K2-Horizon-7B-Uno}}, } ``` ## License Apache 2.0 (same as upstream).