--- library_name: mlx pipeline_tag: text-generation license: apache-2.0 base_model: IFM/K2-Horizon-7B tags: - mlx - safetensors - apple-silicon - conversational - transformers - k2_horizon - text-generation - k2-horizon - 7b - dense - open-weights - ifm - custom_code - en - dataset:IFM/K2-Horizon-Pretrain-Data - dataset:IFM/K2-Horizon-Midtrain-Data - license:apache-2.0 --- # K2-Horizon-7B-mlx-16bit > 16-bit MLX conversion of [IFM/K2-Horizon-7B](https://huggingface.co/IFM/K2-Horizon-7B) optimized for Apple Silicon native GPU inference. **Converted by**: [SirSahOl](https://huggingface.co/SirSahOl) **Source Model**: [IFM/K2-Horizon-7B](https://huggingface.co/IFM/K2-Horizon-7B) **Framework**: [MLX](https://github.com/ml-explore/mlx) by Apple **Quantization**: 16-bit (Average 16.00 (unquantized bfloat16) bits per weight) **Format**: safetensors **License**: apache-2.0 --- ## Model Details * **Architecture**: K2HorizonForCausalLM * **Parameters**: 7B * **Context Length**: 524,288 tokens * **Format**: MLX (Apple Silicon native GPU format) * **Quantization**: 16-bit (Average 16.00 (unquantized bfloat16) bits per weight) * **Active VRAM Footprint**: ~15.2 GB (Minimum recommended: 24 GB – 32 GB Unified Memory) --- ## Quick Start ### Installation ```bash pip install mlx-lm ``` ### CLI Usage ```bash # Chat interactively mlx_lm.chat --model SirSahOl/K2-Horizon-7B-chat-mlx-16bit # Generate text mlx_lm.generate --model SirSahOl/K2-Horizon-7B-chat-mlx-16bit --prompt "Write a short poem about artificial intelligence." ``` ### Python API (with Chat Template) ```python from mlx_lm import load, generate model, tokenizer = load("SirSahOl/K2-Horizon-7B-chat-mlx-16bit") messages = [ {"role": "user", "content": "Explain quantum superposition in simple terms."} ] prompt = tokenizer.apply_chat_template( messages, tokenize=False, add_generation_prompt=True ) response = generate(model, tokenizer, prompt=prompt, verbose=True) print(response) ``` --- ## Performance Benchmarks ### Apple Silicon Hardware Sizing Matrix Estimated decoding throughput, time-to-first-token (TTFT), and active unified memory footprint across Apple Silicon tiers: | Apple Silicon Tier | Unified Memory | Active VRAM | Estimated Speed | Est. TTFT | Recommended Use Case | | :--- | :--- | :--- | :--- | :--- | :--- | | **M1 / M2 / M3 / M4 (Base)** | 24 GB (min. required) | ~15.2 GB | **~12 tokens/sec** | ~240 ms | Full-precision unquantized evaluation on 24GB Macs (M2/M3/M4) | | **M1 / M2 / M3 / M4 Pro** | 36 GB – 48 GB | ~15.2 GB | **~18 tokens/sec** | ~160 ms | Development, prompt engineering, and ground-truth model comparison | | **M1 / M2 / M3 / M4 Max** | 36 GB – 128 GB | ~15.2 GB | **~28 tokens/sec** | ~100 ms | Unquantized reference inference with zero perplexity penalty | | **M1 / M2 / M3 Ultra** | 64 GB – 192 GB | ~15.2 GB | **~42 tokens/sec** | ~65 ms | Enterprise workstation deployment, zero-compromise reference serving | > *Estimates based on Apple Silicon unified memory bandwidth and active parameter footprint. Real-world speeds may vary with context length.* --- ## Multi-Quantization Comparison Evaluate your hardware budget and choose the optimal precision: | Variant | Disk Size | VRAM Footprint | Target Apple Silicon Hardware | Key Advantage | | :--- | :--- | :--- | :--- | :--- | | **[4-bit MLX](https://huggingface.co/SirSahOl/K2-Horizon-7B-chat-mlx-4bit)** | ~4.3 GB | ~4.2 GB | M1 / M2 / M3 / M4 (8GB+ Unified Memory) | Maximum generation speed, lowest memory pressure; ideal for multitasking and everyday local chat alongside IDEs. | | **[8-bit MLX](https://huggingface.co/SirSahOl/K2-Horizon-7B-chat-mlx-8bit)** | ~8.1 GB | ~7.8 GB | M1 / M2 / M3 / M4 Pro/Max (16GB+ Unified Memory) | Near-lossless precision, high-fidelity reasoning, and stable complex instruction following. | | **16-bit MLX** (This Repository) | ~15.2 GB | ~15.2 GB | M2 / M3 / M4 Max/Ultra (32GB+ Unified Memory) | Full unquantized bfloat16 precision; zero perplexity loss, ideal for evaluation and reference output. | --- ## Who Should Use This? | Your Hardware | Recommended Quantization | |---------------|--------------------------| | M1/M2/M3/M4 (8GB – 16GB) | **4-bit** — Best balance of speed, low memory, and multitasking capability | | M1/M2/M3/M4 Pro/Max (18GB – 36GB) | **8-bit** — Higher quality reasoning with comfortable memory headroom | | M1/M2/M3/M4 Max/Ultra (36GB – 192GB) | **16-bit** — Unquantized full precision, zero quality degradation | **General guidance:** - Use **4-bit** if you want to run this model alongside IDEs, browsers, and background development tools. - Use **8-bit** if you have 16GB+ unified memory and require superior reasoning and code accuracy. - Use **16-bit** for research, benchmarking, evaluation, or high-end workstation deployments. --- ## Other Quantization Variants | Variant | Link | |---------|------| | 4-bit | [SirSahOl/K2-Horizon-7B-chat-mlx-4bit](https://huggingface.co/SirSahOl/K2-Horizon-7B-chat-mlx-4bit) | | 8-bit | [SirSahOl/K2-Horizon-7B-chat-mlx-8bit](https://huggingface.co/SirSahOl/K2-Horizon-7B-chat-mlx-8bit) | | 16-bit | [SirSahOl/K2-Horizon-7B-chat-mlx-16bit](https://huggingface.co/SirSahOl/K2-Horizon-7B-chat-mlx-16bit) | --- ## LM Studio & Local Inference Setup Guide To prevent runaway loops and ensure correct conversational turn-taking, configure custom stop strings in your local inference runtime. ### Custom Stop Strings Configuration Ensure the following sequence tokens are configured as strict stop boundaries: 1. `<|im_start|>` 2. `<|im_end|>` 3. `<|endoftext|>` ### Option A: Automatic Preset (Recommended for LM Studio) Create a custom prompt preset JSON file named `ChatML.json` in your LM Studio config directory: - **macOS / Linux**: `~/.lmstudio/config-presets/ChatML.json` - **Windows**: `%USERPROFILE%\.lmstudio\config-presets\ChatML.json` ```json { "name": "ChatML", "inference_params": { "pre_prompt": "You are a helpful, respectful, and honest assistant.", "input_prefix": "<|im_start|>user\n", "input_suffix": "<|im_end|>\n", "pre_prompt_prefix": "<|im_start|>system\n", "pre_prompt_suffix": "<|im_end|>\n", "antiprompt": [ "<|im_start|>", "<|im_end|>", "<|endoftext|>" ], "stopStrings": [ "<|im_start|>", "<|im_end|>", "<|endoftext|>" ], "temperature": 0.7, "max_tokens": 2048 } } ``` Restart LM Studio, load **SirSahOl/K2-Horizon-7B-chat-mlx-16bit**, and select **"ChatML"** from the Prompt Template dropdown. ### Option B: Manual LM Studio Configuration 1. Load **SirSahOl/K2-Horizon-7B-chat-mlx-16bit** in LM Studio. 2. In the right-hand **Advanced Configuration / Inference Parameters** panel: - **Stop Strings (antiprompt / stopStrings)**: Add `<|im_start|>`, `<|im_end|>`, `<|endoftext|>` - **System Prefix**: `<|im_start|>system\n` - **System Suffix**: `<|im_end|>\n` - **User Prefix**: `<|im_start|>user\n` - **Assistant Suffix**: `<|im_end|>\n<|im_start|>assistant\n` ### Option C: Ollama Modelfile Setup Create a `Modelfile` to run this model in Ollama: ```dockerfile FROM SirSahOl/K2-Horizon-7B-chat-mlx-16bit PARAMETER stop "<|im_start|>" PARAMETER stop "<|im_end|>" PARAMETER stop "<|endoftext|>" PARAMETER temperature 0.7 TEMPLATE """{{ if .System }}<|im_start|>system {{ .System }}<|im_end|> {{ end }}{{ if .Prompt }}<|im_start|>user {{ .Prompt }}<|im_end|> {{ end }}<|im_start|>assistant """ ``` Create and run with Ollama: ```bash ollama create k2-horizon-7b-chat-mlx-16bit -f Modelfile ollama run k2-horizon-7b-chat-mlx-16bit ``` --- ## Conversion Details | Property | Value | |----------|-------| | **Source Model** | [IFM/K2-Horizon-7B](https://huggingface.co/IFM/K2-Horizon-7B) | | **Quantization** | 16-bit | | **mlx-lm Version** | 0.31.3 | | **Conversion Time** | 195.06s | | **Output Size** | 16.8 GB | | **Date** | 2026-09-15T00:25:48.108634+00:00 | ### Reproduction To reproduce this conversion: ```bash pip install mlx-lm==0.31.3 python3 -m mlx_lm.convert --hf-path /root/.cache/huggingface/hub/models--IFM--K2-Horizon-7B/snapshots/ff325e226270e05ea081a97fd0c9c62652472fe8 --mlx-path output/K2-Horizon-7B-mlx-16bit ``` --- ## Limitations & Known Issues - 4-bit group-wise quantization introduces minor precision loss compared to unquantized weights; for deep mathematical derivations or precision-critical reasoning, test the 8-bit or 16-bit variants. - High context sequences (>32K tokens) require sufficient unified memory headroom; ensure unified memory is not overcommitted. - This is a weight-only MLX conversion designed specifically for Apple Silicon GPUs (M1/M2/M3/M4 series). --- ## License This model conversion inherits the license of the source model: **apache-2.0**. See the [original model card](https://huggingface.co/IFM/K2-Horizon-7B) for full license details. --- ## Changelog | Version | Date | Changes | |---------|------|---------| | v1.0 | 2026-09-15 | Initial conversion | --- *Converted with [MLX Foundry](https://github.com/SirSahOl/mlx-foundry) — a professional pipeline for converting models to Apple MLX format.*