Text Generation
MLX
Safetensors
qwen3_5
agentic
tool-use
coding
reasoning
instruction-following
bf16
conversational
Instructions to use TokenRhythm/NeoHorse-1-4B-MLX with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use TokenRhythm/NeoHorse-1-4B-MLX with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("TokenRhythm/NeoHorse-1-4B-MLX") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use TokenRhythm/NeoHorse-1-4B-MLX with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "TokenRhythm/NeoHorse-1-4B-MLX"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "TokenRhythm/NeoHorse-1-4B-MLX" } ] } } }Run Pi
# Start Pi in your project directory: pi
- MLX LM
How to use TokenRhythm/NeoHorse-1-4B-MLX with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "TokenRhythm/NeoHorse-1-4B-MLX"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "TokenRhythm/NeoHorse-1-4B-MLX" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "TokenRhythm/NeoHorse-1-4B-MLX", "messages": [ {"role": "user", "content": "Hello"} ] }' - Hermes Agent
How to use TokenRhythm/NeoHorse-1-4B-MLX with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "TokenRhythm/NeoHorse-1-4B-MLX"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default TokenRhythm/NeoHorse-1-4B-MLX
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use TokenRhythm/NeoHorse-1-4B-MLX with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "TokenRhythm/NeoHorse-1-4B-MLX"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "TokenRhythm/NeoHorse-1-4B-MLX" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
|
Download README.md from TokenRhythm/NeoHorse-1-4B-MLX: direct link, hf CLI and curl.
- Browser
- Download file 26.5 kB
-
https://huggingface.co/TokenRhythm/NeoHorse-1-4B-MLX/resolve/main/README.md
- Command line
-
hf download hf://TokenRhythm/NeoHorse-1-4B-MLX/README.md
-
curl -L -o README.md https://huggingface.co/TokenRhythm/NeoHorse-1-4B-MLX/resolve/main/README.md
26.5 kB
| license: apache-2.0 | |
| library_name: mlx | |
| pipeline_tag: text-generation | |
| base_model: TokenRhythm/NeoHorse-1-4B | |
| tags: | |
| - agentic | |
| - tool-use | |
| - coding | |
| - reasoning | |
| - instruction-following | |
| - mlx | |
| - bf16 | |
| ## MLX local inference | |
| This is the **unquantized BF16 MLX** version of [NeoHorse-1-4B](https://huggingface.co/TokenRhythm/NeoHorse-1-4B) for Apple Silicon. Converted from the original BF16 weights with MLX-LM. No weight quantization is applied; MLX-LM adapts tensor names/layouts and normalization representation for its runtime. Benchmark scores below refer to the original model, not a separate evaluation of this MLX version. | |
| ```bash | |
| pip install "mlx-lm>=0.31.3" | |
| mlx_lm.chat --model TokenRhythm/NeoHorse-1-4B-MLX | |
| ``` | |
| The model downloads automatically from Hugging Face. The original chat template is preserved. See [Deployment](#deployment) for local checkpoints, the chat API, and tool calling. | |
| <div align="center"> | |
| <h1>NeoHorse-1-4B</h1> | |
| <p><b>Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness.</b></p> | |
| </div> | |
| <div align="center"> | |
| <a href="https://github.com/TokenRhythm/NeoHorse"><img alt="GitHub" src="https://img.shields.io/badge/GitHub-NeoHorse-181717?logo=github&logoColor=white"></a> | |
| <a href="https://tokenrhythm.ai/"><img alt="Company" src="https://img.shields.io/badge/Company-TokenRhythm-F97316?logo=homeassistant&logoColor=white"></a> | |
| <a href="https://huggingface.co/TokenRhythm"><img alt="Hugging Face" src="https://img.shields.io/badge/Hugging%20Face-Models-FFD21E?logo=huggingface&logoColor=000000"></a> | |
| <a href="https://x.com/opensquilla"><img alt="Twitter / X" src="https://img.shields.io/badge/Twitter%20%2F%20X-OpenSquilla-111827?logo=x&logoColor=white"></a> | |
| <a href="https://www.apache.org/licenses/LICENSE-2.0"><img alt="License: Apache-2.0" src="https://img.shields.io/badge/License-Apache--2.0-64748B"></a> | |
| </div> | |
| <p align="center"> | |
| <a href="https://arxiv.org/abs/2609.08183"><b>Technical Report</b></a> | |
| </p> | |
| <style> | |
| /* Reusable benchmark table architecture. Inline styles remain as a fallback for HF rendering. */ | |
| .vl-table { | |
| width: 100%; | |
| min-width: 100%; | |
| border-collapse: collapse; | |
| table-layout: fixed; | |
| font-size: 15px; | |
| } | |
| .vl-table th { | |
| font-size: 15px !important; | |
| line-height: 1.2; | |
| color: #c2410c; | |
| background: rgba(249,115,22,.10); | |
| } | |
| .vl-table td:not(.benchmark-cell):not([colspan]) { | |
| font-size: 15px; | |
| line-height: 1.2; | |
| vertical-align: middle; | |
| } | |
| .vl-table .benchmark-cell { | |
| padding: 12px 10px 12px 18px !important; | |
| vertical-align: middle; | |
| } | |
| .vl-table .benchmark-capability { | |
| font-size: 15px; | |
| font-weight: 600; | |
| line-height: 1.22; | |
| color: #c2410c; | |
| } | |
| .vl-table .benchmark-name { | |
| margin-top: 4px; | |
| font-size: 11px; | |
| font-weight: 400; | |
| line-height: 1.2; | |
| color: inherit; | |
| } | |
| .vl-table .metric-stack { | |
| display: flex; | |
| flex-direction: column; | |
| gap: 7px; | |
| padding: 3px 0; | |
| } | |
| .vl-table .metric-label { | |
| font-size: 10px; | |
| font-weight: 400; | |
| line-height: 1.1; | |
| color: inherit; | |
| } | |
| .vl-table .metric-value { | |
| margin-top: 2px; | |
| font-size: 15px; | |
| line-height: 1.15; | |
| color: inherit; | |
| } | |
| .model-table td:first-child { | |
| width: 34%; | |
| font-weight: 600; | |
| } | |
| /* HF's theme toggle sets the dark class on an ancestor. */ | |
| .dark .vl-table th, | |
| .dark .vl-table .benchmark-capability { | |
| color: #fdba74 !important; | |
| } | |
| </style> | |
| NeoHorse-1-4B is a 4B causal language model and an initial prototype on the path toward **recursive self-improvement (RSI)**. It is post-trained from Qwen3.5-4B for text-based agent harnesses, tool use, coding, and instruction following. | |
| Derived from [Qwen/Qwen3.5-4B](https://huggingface.co/Qwen/Qwen3.5-4B) and fine-tuned by TokenRhythm. The source checkpoint was repackaged for text-only inference. This repository contains **language-model weights only**, converted to MLX BF16 without weight quantization. | |
| <p align="center"> | |
| <a href="https://huggingface.co/TokenRhythm/NeoHorse-1-4B/resolve/main/4B_head_fig.jpg"> | |
| <img src="https://huggingface.co/TokenRhythm/NeoHorse-1-4B/resolve/main/4B_head_fig.jpg" alt="NeoHorse-1-4B evaluation results" width="100%"> | |
| </a> | |
| </p> | |
| ## Highlights | |
| - **Path toward RSI:** the routing harness assigns tasks to a heterogeneous model pool, records tool interactions and outcomes, estimates capability demand, and uses capability-level feedback to shape the next training mixture. Updated models can return to the harness, closing a prototype evaluation鈥搒election鈥搖pdate loop; extending this loop across successive iterations is the next step toward RSI. | |
| - **Agentic post-training framework:** the associated research explores routing-guided curriculum SFT and routing-guided on-policy distillation to turn execution trajectories into training signal while preserving execution and harness context around each response. | |
| - **Data quality:** exact and near-duplicate removal, evaluation decontamination, structural validation, six-dimensional semantic evaluation, and subscene-level Scene/Goal/Outcome labeling. | |
| - **Broad gains:** 64.87 macro average across ten benchmarks versus 58.94 for Qwen3.5-4B (**+5.93**). | |
| ## Model Details | |
| <div style="width:100%;max-width:none;margin:16px 0;padding:0;overflow-x:auto"> | |
| <table class="vl-table model-table" width="100%" style="display:table;width:100%;min-width:100%;table-layout:fixed;border-collapse:collapse;font-size:13px"> | |
| <thead><tr> | |
| <th style="padding:9px 10px;text-align:left;border-bottom:2px solid #f97316;color:#c2410c;background:rgba(249,115,22,.10)">Property</th> | |
| <th style="padding:9px 10px;text-align:left;border-bottom:2px solid #f97316;color:#c2410c;background:rgba(249,115,22,.10)">Value</th> | |
| </tr></thead><tbody> | |
| <tr> | |
| <td style="padding:9px 10px;border-bottom:1px solid rgba(249,115,22,.16);font-weight:600">Model family</td> | |
| <td style="padding:9px 10px;border-bottom:1px solid rgba(249,115,22,.16)">NeoHorse Agent-Native Causal Language Model</td> | |
| </tr> | |
| <tr> | |
| <td style="padding:9px 10px;border-bottom:1px solid rgba(249,115,22,.16);font-weight:600">Parameters</td> | |
| <td style="padding:9px 10px;border-bottom:1px solid rgba(249,115,22,.16)">Approximately <strong>4B</strong></td> | |
| </tr> | |
| <tr> | |
| <td style="padding:9px 10px;border-bottom:1px solid rgba(249,115,22,.16);font-weight:600">Base model</td> | |
| <td style="padding:9px 10px;border-bottom:1px solid rgba(249,115,22,.16)"><a href="https://huggingface.co/Qwen/Qwen3.5-4B">Qwen3.5-4B</a></td> | |
| </tr> | |
| <tr> | |
| <td style="padding:9px 10px;border-bottom:1px solid rgba(249,115,22,.16);font-weight:600">Post-training</td> | |
| <td style="padding:9px 10px;border-bottom:1px solid rgba(249,115,22,.16)">Routing-guided agentic post-training</td> | |
| </tr> | |
| <tr> | |
| <td style="padding:9px 10px;border-bottom:1px solid rgba(249,115,22,.16);font-weight:600">Interface</td> | |
| <td style="padding:9px 10px;border-bottom:1px solid rgba(249,115,22,.16)">Text input and text output</td> | |
| </tr> | |
| <tr> | |
| <td style="padding:9px 10px;border-bottom:1px solid rgba(249,115,22,.16);font-weight:600">Context length</td> | |
| <td style="padding:9px 10px;border-bottom:1px solid rgba(249,115,22,.16)">262,144 natively and extensible up to 1,010,000 tokens.</td> | |
| </tr> | |
| <tr> | |
| <td style="padding:9px 10px;border-bottom:1px solid rgba(249,115,22,.16);font-weight:600">Weight format / precision</td> | |
| <td style="padding:9px 10px;border-bottom:1px solid rgba(249,115,22,.16)">MLX Safetensors / BF16 (unquantized)</td> | |
| </tr> | |
| </tbody></table> | |
| </div> | |
| ## Evaluation | |
| The 4B track compares NeoHorse-1-4B with five representative open-weight models. Results are grouped by capability in the table below. Higher is better; `螖` is NeoHorse-1-4B minus Qwen3.5-4B. **Bold** marks the best available result; <ins>underlining</ins> marks the second-best. | |
| <div style="overflow-x:auto"> | |
| <table class="vl-table" width="100%" style="display:table;width:100%;min-width:100%;border-collapse:collapse;table-layout:fixed;font-size:13px"> | |
| <thead><tr> | |
| <th style="padding:9px 8px;text-align:center;border-bottom:2px solid #f97316;color:#c2410c;background:rgba(249,115,22,.10);text-align:left">Benchmark</th> | |
| <th style="padding:9px 8px;text-align:center;border-bottom:2px solid #f97316;color:#c2410c;background:rgba(249,115,22,.10)">Qwen3.5-4B</th> | |
| <th style="padding:9px 8px;text-align:center;border-bottom:2px solid #f97316;color:#c2410c;background:rgba(249,115,22,.10)">Gemma-4-E4B-it</th> | |
| <th style="padding:9px 8px;text-align:center;border-bottom:2px solid #f97316;color:#c2410c;background:rgba(249,115,22,.10)">Nanbeige-4.2-3B</th> | |
| <th style="padding:9px 8px;text-align:center;border-bottom:2px solid #f97316;color:#c2410c;background:rgba(249,115,22,.10)">Agents-A1-4B</th> | |
| <th style="padding:9px 8px;text-align:center;border-bottom:2px solid #f97316;color:#c2410c;background:rgba(249,115,22,.10)">Spark-X2.5-4B</th> | |
| <th style="padding:9px 8px;text-align:center;border-bottom:2px solid #f97316;color:#c2410c;background:rgba(249,115,22,.10);background:rgba(249,115,22,.18)">NeoHorse-1-4B</th> | |
| <th style="padding:9px 8px;text-align:center;border-bottom:2px solid #f97316;color:#c2410c;background:rgba(249,115,22,.10);background:rgba(249,115,22,.18)">螖 vs Qwen3.5-4B</th> | |
| </tr></thead><tbody> | |
| <tr><td class="benchmark-capability" colspan="8" style="padding:10px 8px;font-weight:700;color:#c2410c;background:rgba(249,115,22,.10);border-top:2px solid #f97316">馃 Agentic</td></tr> | |
| <tr style="border-bottom:1px solid rgba(128,128,128,.16)"> | |
| <td class="benchmark-cell" style="padding:8px;font-weight:600"><div class="benchmark-name">QwenClawBench</div></td> | |
| <td style="padding:8px;text-align:center"><div class="metric-stack"><span class="metric-value">38.47</span></div></td> | |
| <td style="padding:8px;text-align:center"><div class="metric-stack"><span class="metric-value">22.98</span></div></td> | |
| <td style="padding:8px;text-align:center"><div class="metric-stack"><span class="metric-value">40.66</span></div></td> | |
| <td style="padding:8px;text-align:center"><div class="metric-stack"><span class="metric-value">43.16</span></div></td> | |
| <td style="padding:8px;text-align:center"><div class="metric-stack"><span class="metric-value"><ins>43.52</ins></span></div></td> | |
| <td style="padding:8px;text-align:center;background:rgba(249,115,22,.16);font-weight:700"><div class="metric-stack"><span class="metric-value"><strong>44.68</strong></span></div></td> | |
| <td style="padding:8px;text-align:center;background:rgba(249,115,22,.16);font-weight:700"><div class="metric-stack"><span class="metric-value">+6.21</span></div></td> | |
| </tr> | |
| <tr style="border-bottom:1px solid rgba(128,128,128,.16)"> | |
| <td class="benchmark-cell" style="padding:8px;font-weight:600"><div class="benchmark-name">WorkBuddy Bench</div></td> | |
| <td style="padding:8px;text-align:center"><div class="metric-stack"><span class="metric-value">24.62</span></div></td> | |
| <td style="padding:8px;text-align:center"><div class="metric-stack"><span class="metric-value">11.65</span></div></td> | |
| <td style="padding:8px;text-align:center"><div class="metric-stack"><span class="metric-value">21.03</span></div></td> | |
| <td style="padding:8px;text-align:center"><div class="metric-stack"><span class="metric-value"><ins>33.37</ins></span></div></td> | |
| <td style="padding:8px;text-align:center"><div class="metric-stack"><span class="metric-value">26.47</span></div></td> | |
| <td style="padding:8px;text-align:center;background:rgba(249,115,22,.16);font-weight:700"><div class="metric-stack"><span class="metric-value"><strong>34.41</strong></span></div></td> | |
| <td style="padding:8px;text-align:center;background:rgba(249,115,22,.16);font-weight:700"><div class="metric-stack"><span class="metric-value">+9.79</span></div></td> | |
| </tr> | |
| <tr style="border-bottom:1px solid rgba(128,128,128,.16)"> | |
| <td class="benchmark-cell" style="padding:8px;font-weight:600"><div class="benchmark-name">PinchBench</div></td> | |
| <td style="padding:8px;text-align:center"><div class="metric-stack"><span class="metric-value">71.19</span></div></td> | |
| <td style="padding:8px;text-align:center"><div class="metric-stack"><span class="metric-value">47.60</span></div></td> | |
| <td style="padding:8px;text-align:center"><div class="metric-stack"><span class="metric-value">66.78</span></div></td> | |
| <td style="padding:8px;text-align:center"><div class="metric-stack"><span class="metric-value"><ins>75.07</ins></span></div></td> | |
| <td style="padding:8px;text-align:center"><div class="metric-stack"><span class="metric-value">62.37</span></div></td> | |
| <td style="padding:8px;text-align:center;background:rgba(249,115,22,.16);font-weight:700"><div class="metric-stack"><span class="metric-value"><strong>77.33</strong></span></div></td> | |
| <td style="padding:8px;text-align:center;background:rgba(249,115,22,.16);font-weight:700"><div class="metric-stack"><span class="metric-value">+6.14</span></div></td> | |
| </tr> | |
| <tr style="border-bottom:1px solid rgba(128,128,128,.16)"> | |
| <td class="benchmark-cell" style="padding:8px;font-weight:600"><div class="benchmark-name">VitaBench</div></td> | |
| <td style="padding:8px;text-align:center"><div class="metric-stack"><span class="metric-value">21.50</span></div></td> | |
| <td style="padding:8px;text-align:center"><div class="metric-stack"><span class="metric-value">5.00</span></div></td> | |
| <td style="padding:8px;text-align:center"><div class="metric-stack"><span class="metric-value">31.50</span></div></td> | |
| <td style="padding:8px;text-align:center"><div class="metric-stack"><span class="metric-value"><strong>39.25</strong></span></div></td> | |
| <td style="padding:8px;text-align:center"><div class="metric-stack"><span class="metric-value"><ins>37.00</ins></span></div></td> | |
| <td style="padding:8px;text-align:center;background:rgba(249,115,22,.16);font-weight:700"><div class="metric-stack"><span class="metric-value">32.00</span></div></td> | |
| <td style="padding:8px;text-align:center;background:rgba(249,115,22,.16);font-weight:700"><div class="metric-stack"><span class="metric-value">+10.50</span></div></td> | |
| </tr> | |
| <tr style="border-bottom:1px solid rgba(128,128,128,.16)"> | |
| <td class="benchmark-cell" style="padding:8px;font-weight:600"><div class="benchmark-name">BFCL v4</div></td> | |
| <td style="padding:8px;text-align:center"><div class="metric-stack"><span class="metric-value">61.02</span></div></td> | |
| <td style="padding:8px;text-align:center"><div class="metric-stack"><span class="metric-value">47.18</span></div></td> | |
| <td style="padding:8px;text-align:center"><div class="metric-stack"><span class="metric-value"><strong>67.28</strong></span></div></td> | |
| <td style="padding:8px;text-align:center"><div class="metric-stack"><span class="metric-value">46.60</span></div></td> | |
| <td style="padding:8px;text-align:center"><div class="metric-stack"><span class="metric-value"><ins>63.71</ins></span></div></td> | |
| <td style="padding:8px;text-align:center;background:rgba(249,115,22,.16);font-weight:700"><div class="metric-stack"><span class="metric-value">61.79</span></div></td> | |
| <td style="padding:8px;text-align:center;background:rgba(249,115,22,.16);font-weight:700"><div class="metric-stack"><span class="metric-value">+0.77</span></div></td> | |
| </tr> | |
| <tr style="border-bottom:1px solid rgba(128,128,128,.16)"> | |
| <td class="benchmark-cell" style="padding:8px;font-weight:600"><div class="benchmark-name">tau2-Bench</div></td> | |
| <td style="padding:8px;text-align:center"><div class="metric-stack"><span class="metric-value">84.29</span></div></td> | |
| <td style="padding:8px;text-align:center"><div class="metric-stack"><span class="metric-value">43.60</span></div></td> | |
| <td style="padding:8px;text-align:center"><div class="metric-stack"><span class="metric-value"><ins>85.08</ins></span></div></td> | |
| <td style="padding:8px;text-align:center"><div class="metric-stack"><span class="metric-value">81.00</span></div></td> | |
| <td style="padding:8px;text-align:center"><div class="metric-stack"><span class="metric-value">77.72</span></div></td> | |
| <td style="padding:8px;text-align:center;background:rgba(249,115,22,.16);font-weight:700"><div class="metric-stack"><span class="metric-value"><strong>88.46</strong></span></div></td> | |
| <td style="padding:8px;text-align:center;background:rgba(249,115,22,.16);font-weight:700"><div class="metric-stack"><span class="metric-value">+4.17</span></div></td> | |
| </tr> | |
| <tr><td class="benchmark-capability" colspan="8" style="padding:10px 8px;font-weight:700;color:#c2410c;background:rgba(249,115,22,.10);border-top:2px solid #f97316">馃捇 Coding</td></tr> | |
| <tr style="border-bottom:1px solid rgba(128,128,128,.16)"> | |
| <td class="benchmark-cell" style="padding:8px;font-weight:600"><div class="benchmark-name">HumanEval</div></td> | |
| <td style="padding:8px;text-align:center"><div class="metric-stack"><span class="metric-value">87.20</span></div></td> | |
| <td style="padding:8px;text-align:center"><div class="metric-stack"><span class="metric-value">84.76</span></div></td> | |
| <td style="padding:8px;text-align:center"><div class="metric-stack"><span class="metric-value"><strong>98.78</strong></span></div></td> | |
| <td style="padding:8px;text-align:center"><div class="metric-stack"><span class="metric-value">92.68</span></div></td> | |
| <td style="padding:8px;text-align:center"><div class="metric-stack"><span class="metric-value">92.07</span></div></td> | |
| <td style="padding:8px;text-align:center;background:rgba(249,115,22,.16);font-weight:700"><div class="metric-stack"><span class="metric-value"><ins>96.95</ins></span></div></td> | |
| <td style="padding:8px;text-align:center;background:rgba(249,115,22,.16);font-weight:700"><div class="metric-stack"><span class="metric-value">+9.75</span></div></td> | |
| </tr> | |
| <tr style="border-bottom:1px solid rgba(128,128,128,.16)"> | |
| <td class="benchmark-cell" style="padding:8px;font-weight:600"><div class="benchmark-name">LiveCodeBench v6</div></td> | |
| <td style="padding:8px;text-align:center"><div class="metric-stack"><span class="metric-value">53.71</span></div></td> | |
| <td style="padding:8px;text-align:center"><div class="metric-stack"><span class="metric-value">52.00</span></div></td> | |
| <td style="padding:8px;text-align:center"><div class="metric-stack"><span class="metric-value"><strong>72.50<sup>*</sup></strong></span></div></td> | |
| <td style="padding:8px;text-align:center"><div class="metric-stack"><span class="metric-value">56.57</span></div></td> | |
| <td style="padding:8px;text-align:center"><div class="metric-stack"><span class="metric-value">54.86</span></div></td> | |
| <td style="padding:8px;text-align:center;background:rgba(249,115,22,.16);font-weight:700"><div class="metric-stack"><span class="metric-value"><ins>59.43</ins></span></div></td> | |
| <td style="padding:8px;text-align:center;background:rgba(249,115,22,.16);font-weight:700"><div class="metric-stack"><span class="metric-value">+5.72</span></div></td> | |
| </tr> | |
| <tr><td class="benchmark-capability" colspan="8" style="padding:10px 8px;font-weight:700;color:#c2410c;background:rgba(249,115,22,.10);border-top:2px solid #f97316">馃摎 Instruction Following</td></tr> | |
| <tr style="border-bottom:1px solid rgba(128,128,128,.16)"> | |
| <td class="benchmark-cell" style="padding:8px;font-weight:600"><div class="benchmark-name">IFBench</div></td> | |
| <td style="padding:8px;text-align:center"><div class="metric-stack"><span class="metric-value">60.33</span></div></td> | |
| <td style="padding:8px;text-align:center"><div class="metric-stack"><span class="metric-value">40.00</span></div></td> | |
| <td style="padding:8px;text-align:center"><div class="metric-stack"><span class="metric-value">55.00</span></div></td> | |
| <td style="padding:8px;text-align:center"><div class="metric-stack"><span class="metric-value">63.33</span></div></td> | |
| <td style="padding:8px;text-align:center"><div class="metric-stack"><span class="metric-value"><strong>73.33</strong></span></div></td> | |
| <td style="padding:8px;text-align:center;background:rgba(249,115,22,.16);font-weight:700"><div class="metric-stack"><span class="metric-value"><ins>65.33</ins></span></div></td> | |
| <td style="padding:8px;text-align:center;background:rgba(249,115,22,.16);font-weight:700"><div class="metric-stack"><span class="metric-value">+5.00</span></div></td> | |
| </tr> | |
| <tr style="border-bottom:1px solid rgba(128,128,128,.16)"> | |
| <td class="benchmark-cell" style="padding:8px;font-weight:600"><div class="benchmark-name">IFEval</div></td> | |
| <td style="padding:8px;text-align:center"><div class="metric-stack"><span class="metric-value">87.06</span></div></td> | |
| <td style="padding:8px;text-align:center"><div class="metric-stack"><span class="metric-value">74.68</span></div></td> | |
| <td style="padding:8px;text-align:center"><div class="metric-stack"><span class="metric-value">84.47</span></div></td> | |
| <td style="padding:8px;text-align:center"><div class="metric-stack"><span class="metric-value">83.55</span></div></td> | |
| <td style="padding:8px;text-align:center"><div class="metric-stack"><span class="metric-value"><strong>91.13</strong></span></div></td> | |
| <td style="padding:8px;text-align:center;background:rgba(249,115,22,.16);font-weight:700"><div class="metric-stack"><span class="metric-value"><ins>88.35</ins></span></div></td> | |
| <td style="padding:8px;text-align:center;background:rgba(249,115,22,.16);font-weight:700"><div class="metric-stack"><span class="metric-value">+1.29</span></div></td> | |
| </tr> | |
| <tr><td class="benchmark-capability" colspan="8" style="padding:10px 8px;font-weight:700;color:#c2410c;background:rgba(249,115,22,.10);border-top:2px solid #f97316">馃搳 Overall</td></tr> | |
| <tr style="border-bottom:1px solid rgba(128,128,128,.16)"> | |
| <td class="benchmark-cell" style="padding:8px;font-weight:600"><div class="benchmark-name">Ten-benchmark average</div></td> | |
| <td style="padding:8px;text-align:center"><div class="metric-stack"><span class="metric-value">58.94</span></div></td> | |
| <td style="padding:8px;text-align:center"><div class="metric-stack"><span class="metric-value">42.95</span></div></td> | |
| <td style="padding:8px;text-align:center"><div class="metric-stack"><span class="metric-value"><ins>62.31</ins></span></div></td> | |
| <td style="padding:8px;text-align:center"><div class="metric-stack"><span class="metric-value">61.46</span></div></td> | |
| <td style="padding:8px;text-align:center"><div class="metric-stack"><span class="metric-value">62.22</span></div></td> | |
| <td style="padding:8px;text-align:center;background:rgba(249,115,22,.16);font-weight:700"><div class="metric-stack"><span class="metric-value"><strong>64.87</strong></span></div></td> | |
| <td style="padding:8px;text-align:center;background:rgba(249,115,22,.16);font-weight:700"><div class="metric-stack"><span class="metric-value">+5.93</span></div></td> | |
| </tr> | |
| </tbody></table> | |
| </div> | |
| `*` Nanbeige-4.2-3B LiveCodeBench v6 result is reported in the corresponding model's official blog post or technical report. | |
| > **Reported protocol:** SGLang v0.5.17 路 `temperature=1.0` 路 `top_p=0.95` 路 `top_k=20` 路 `min_p=0.0` 路 `presence_penalty=1.5` 路 `repetition_penalty=1.0` 路 thinking mode enabled with `enable_thinking=true` and `force_nonempty_content=true`. QwenClawBench, WorkBuddy Bench, and tau2-Bench use three runs; PinchBench and VitaBench use one run; the remaining benchmarks follow their official protocols. VitaBench uses the DeepSeek-V4-Flash simulator and judge. | |
| ## Deployment | |
| Use [MLX-LM](https://github.com/ml-explore/mlx-lm) on an Apple Silicon Mac to run this checkpoint. | |
| ### Install and select a local checkpoint | |
| ```bash | |
| pip install "mlx-lm>=0.31.3" | |
| MODEL_PATH="/path/to/NeoHorse-1-4B-MLX" | |
| ``` | |
| Set `MODEL_PATH` to the downloaded MLX directory containing `config.json`, tokenizer files, `chat_template.jinja`, and model weights. You can also use `TokenRhythm/NeoHorse-1-4B-MLX` as the model path to download it automatically from Hugging Face. | |
| ### Chat locally | |
| ```bash | |
| mlx_lm.chat --model "$MODEL_PATH" | |
| ``` | |
| ### Start an API server | |
| ```bash | |
| mlx_lm.server \ | |
| --model "$MODEL_PATH" \ | |
| --host 127.0.0.1 \ | |
| --port 8080 | |
| ``` | |
| The server exposes an OpenAI-compatible `/v1/chat/completions` endpoint. In the requests below, `default_model` refers to the checkpoint selected with `--model`. | |
| ### Basic Usage | |
| After the server starts, run this request in another terminal: | |
| ```bash | |
| curl http://127.0.0.1:8080/v1/chat/completions \ | |
| -H 'Content-Type: application/json' \ | |
| -d '{ | |
| "model": "default_model", | |
| "messages": [ | |
| {"role": "user", "content": "Write a Python function that returns the first n Fibonacci numbers."} | |
| ], | |
| "max_tokens": 2048, | |
| "stream": false | |
| }' | |
| ``` | |
| The generated reply is returned in `choices[0].message.content`. | |
| ### Tool Calling | |
| Pass function definitions in the `tools` field: | |
| ```bash | |
| curl http://127.0.0.1:8080/v1/chat/completions \ | |
| -H 'Content-Type: application/json' \ | |
| -d '{ | |
| "model": "default_model", | |
| "messages": [ | |
| {"role": "user", "content": "Use get_weather to check the current weather in Beijing in celsius."} | |
| ], | |
| "tools": [ | |
| { | |
| "type": "function", | |
| "function": { | |
| "name": "get_weather", | |
| "description": "Get the current weather for a city.", | |
| "parameters": { | |
| "type": "object", | |
| "properties": { | |
| "city": {"type": "string", "description": "City name."}, | |
| "unit": {"type": "string", "enum": ["celsius", "fahrenheit"]} | |
| }, | |
| "required": ["city", "unit"] | |
| } | |
| } | |
| } | |
| ], | |
| "max_tokens": 2048, | |
| "stream": false | |
| }' | |
| ``` | |
| MLX-LM reads the preserved chat template to format tool requests and parse generated calls. When the model chooses to call a tool, the call is returned in `choices[0].message.tool_calls`. Your application executes the function, appends the assistant message and a `role: "tool"` result with the matching `tool_call_id`, then sends the conversation back to the same endpoint for the final answer. | |
| ## License | |
| NeoHorse-1-4B is released under the **Apache License 2.0**. | |
| The upstream model is [Qwen/Qwen3.5-4B](https://huggingface.co/Qwen/Qwen3.5-4B). Its original copyright notice, Copyright 2026 Alibaba Cloud, is retained in the license file. TokenRhythm fine-tuned and repackaged the source checkpoint for text-only inference. This repository provides its MLX BF16 conversion without weight quantization. | |
| ## Citation | |
| ``` | |
| @misc{neohorse2026, | |
| title = {NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness}, | |
| author = {NeoHorse Team}, | |
| year = {2026}, | |
| howpublished = {arXiv preprint}, | |
| eprint = {2609.08183}, | |
| archivePrefix = {arXiv}, | |
| primaryClass = {cs.CL}, | |
| url = {https://arxiv.org/abs/2609.08183} | |
| } | |
| ``` | |
| For questions or issue reports, use the [NeoHorse project repository](https://github.com/TokenRhythm/NeoHorse). | |