---
license: apache-2.0
library_name: mlx
pipeline_tag: text-generation
base_model: TokenRhythm/NeoHorse-1-4B
tags:
- agentic
- tool-use
- coding
- reasoning
- instruction-following
- mlx
- bf16
---
## MLX local inference
This is the **unquantized BF16 MLX** version of [NeoHorse-1-4B](https://huggingface.co/TokenRhythm/NeoHorse-1-4B) for Apple Silicon. Converted from the original BF16 weights with MLX-LM. No weight quantization is applied; MLX-LM adapts tensor names/layouts and normalization representation for its runtime. Benchmark scores below refer to the original model, not a separate evaluation of this MLX version.
```bash
pip install "mlx-lm>=0.31.3"
mlx_lm.chat --model TokenRhythm/NeoHorse-1-4B-MLX
```
The model downloads automatically from Hugging Face. The original chat template is preserved. See [Deployment](#deployment) for local checkpoints, the chat API, and tool calling.
NeoHorse-1-4B
Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness.
Technical Report
NeoHorse-1-4B is a 4B causal language model and an initial prototype on the path toward **recursive self-improvement (RSI)**. It is post-trained from Qwen3.5-4B for text-based agent harnesses, tool use, coding, and instruction following.
Derived from [Qwen/Qwen3.5-4B](https://huggingface.co/Qwen/Qwen3.5-4B) and fine-tuned by TokenRhythm. The source checkpoint was repackaged for text-only inference. This repository contains **language-model weights only**, converted to MLX BF16 without weight quantization.
## Highlights
- **Path toward RSI:** the routing harness assigns tasks to a heterogeneous model pool, records tool interactions and outcomes, estimates capability demand, and uses capability-level feedback to shape the next training mixture. Updated models can return to the harness, closing a prototype evaluation鈥搒election鈥搖pdate loop; extending this loop across successive iterations is the next step toward RSI.
- **Agentic post-training framework:** the associated research explores routing-guided curriculum SFT and routing-guided on-policy distillation to turn execution trajectories into training signal while preserving execution and harness context around each response.
- **Data quality:** exact and near-duplicate removal, evaluation decontamination, structural validation, six-dimensional semantic evaluation, and subscene-level Scene/Goal/Outcome labeling.
- **Broad gains:** 64.87 macro average across ten benchmarks versus 58.94 for Qwen3.5-4B (**+5.93**).
## Model Details
| Property |
Value |
| Model family |
NeoHorse Agent-Native Causal Language Model |
| Parameters |
Approximately 4B |
| Base model |
Qwen3.5-4B |
| Post-training |
Routing-guided agentic post-training |
| Interface |
Text input and text output |
| Context length |
262,144 natively and extensible up to 1,010,000 tokens. |
| Weight format / precision |
MLX Safetensors / BF16 (unquantized) |
## Evaluation
The 4B track compares NeoHorse-1-4B with five representative open-weight models. Results are grouped by capability in the table below. Higher is better; `螖` is NeoHorse-1-4B minus Qwen3.5-4B. **Bold** marks the best available result; underlining marks the second-best.
| Benchmark |
Qwen3.5-4B |
Gemma-4-E4B-it |
Nanbeige-4.2-3B |
Agents-A1-4B |
Spark-X2.5-4B |
NeoHorse-1-4B |
螖 vs Qwen3.5-4B |
| 馃 Agentic |
QwenClawBench |
38.47 |
22.98 |
40.66 |
43.16 |
43.52 |
44.68 |
+6.21 |
WorkBuddy Bench |
24.62 |
11.65 |
21.03 |
33.37 |
26.47 |
34.41 |
+9.79 |
PinchBench |
71.19 |
47.60 |
66.78 |
75.07 |
62.37 |
77.33 |
+6.14 |
VitaBench |
21.50 |
5.00 |
31.50 |
39.25 |
37.00 |
32.00 |
+10.50 |
BFCL v4 |
61.02 |
47.18 |
67.28 |
46.60 |
63.71 |
61.79 |
+0.77 |
tau2-Bench |
84.29 |
43.60 |
85.08 |
81.00 |
77.72 |
88.46 |
+4.17 |
| 馃捇 Coding |
HumanEval |
87.20 |
84.76 |
98.78 |
92.68 |
92.07 |
96.95 |
+9.75 |
LiveCodeBench v6 |
53.71 |
52.00 |
72.50* |
56.57 |
54.86 |
59.43 |
+5.72 |
| 馃摎 Instruction Following |
IFBench |
60.33 |
40.00 |
55.00 |
63.33 |
73.33 |
65.33 |
+5.00 |
IFEval |
87.06 |
74.68 |
84.47 |
83.55 |
91.13 |
88.35 |
+1.29 |
| 馃搳 Overall |
Ten-benchmark average |
58.94 |
42.95 |
62.31 |
61.46 |
62.22 |
64.87 |
+5.93 |
`*` Nanbeige-4.2-3B LiveCodeBench v6 result is reported in the corresponding model's official blog post or technical report.
> **Reported protocol:** SGLang v0.5.17 路 `temperature=1.0` 路 `top_p=0.95` 路 `top_k=20` 路 `min_p=0.0` 路 `presence_penalty=1.5` 路 `repetition_penalty=1.0` 路 thinking mode enabled with `enable_thinking=true` and `force_nonempty_content=true`. QwenClawBench, WorkBuddy Bench, and tau2-Bench use three runs; PinchBench and VitaBench use one run; the remaining benchmarks follow their official protocols. VitaBench uses the DeepSeek-V4-Flash simulator and judge.
## Deployment
Use [MLX-LM](https://github.com/ml-explore/mlx-lm) on an Apple Silicon Mac to run this checkpoint.
### Install and select a local checkpoint
```bash
pip install "mlx-lm>=0.31.3"
MODEL_PATH="/path/to/NeoHorse-1-4B-MLX"
```
Set `MODEL_PATH` to the downloaded MLX directory containing `config.json`, tokenizer files, `chat_template.jinja`, and model weights. You can also use `TokenRhythm/NeoHorse-1-4B-MLX` as the model path to download it automatically from Hugging Face.
### Chat locally
```bash
mlx_lm.chat --model "$MODEL_PATH"
```
### Start an API server
```bash
mlx_lm.server \
--model "$MODEL_PATH" \
--host 127.0.0.1 \
--port 8080
```
The server exposes an OpenAI-compatible `/v1/chat/completions` endpoint. In the requests below, `default_model` refers to the checkpoint selected with `--model`.
### Basic Usage
After the server starts, run this request in another terminal:
```bash
curl http://127.0.0.1:8080/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{
"model": "default_model",
"messages": [
{"role": "user", "content": "Write a Python function that returns the first n Fibonacci numbers."}
],
"max_tokens": 2048,
"stream": false
}'
```
The generated reply is returned in `choices[0].message.content`.
### Tool Calling
Pass function definitions in the `tools` field:
```bash
curl http://127.0.0.1:8080/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{
"model": "default_model",
"messages": [
{"role": "user", "content": "Use get_weather to check the current weather in Beijing in celsius."}
],
"tools": [
{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get the current weather for a city.",
"parameters": {
"type": "object",
"properties": {
"city": {"type": "string", "description": "City name."},
"unit": {"type": "string", "enum": ["celsius", "fahrenheit"]}
},
"required": ["city", "unit"]
}
}
}
],
"max_tokens": 2048,
"stream": false
}'
```
MLX-LM reads the preserved chat template to format tool requests and parse generated calls. When the model chooses to call a tool, the call is returned in `choices[0].message.tool_calls`. Your application executes the function, appends the assistant message and a `role: "tool"` result with the matching `tool_call_id`, then sends the conversation back to the same endpoint for the final answer.
## License
NeoHorse-1-4B is released under the **Apache License 2.0**.
The upstream model is [Qwen/Qwen3.5-4B](https://huggingface.co/Qwen/Qwen3.5-4B). Its original copyright notice, Copyright 2026 Alibaba Cloud, is retained in the license file. TokenRhythm fine-tuned and repackaged the source checkpoint for text-only inference. This repository provides its MLX BF16 conversion without weight quantization.
## Citation
```
@misc{neohorse2026,
title = {NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness},
author = {NeoHorse Team},
year = {2026},
howpublished = {arXiv preprint},
eprint = {2609.08183},
archivePrefix = {arXiv},
primaryClass = {cs.CL},
url = {https://arxiv.org/abs/2609.08183}
}
```
For questions or issue reports, use the [NeoHorse project repository](https://github.com/TokenRhythm/NeoHorse).