Instructions to use BrandeisPatrick/Ouro-1.4B-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use BrandeisPatrick/Ouro-1.4B-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf BrandeisPatrick/Ouro-1.4B-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf BrandeisPatrick/Ouro-1.4B-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf BrandeisPatrick/Ouro-1.4B-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf BrandeisPatrick/Ouro-1.4B-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf BrandeisPatrick/Ouro-1.4B-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf BrandeisPatrick/Ouro-1.4B-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf BrandeisPatrick/Ouro-1.4B-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf BrandeisPatrick/Ouro-1.4B-GGUF:Q4_K_M
Use Docker
docker model run hf.co/BrandeisPatrick/Ouro-1.4B-GGUF:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use BrandeisPatrick/Ouro-1.4B-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "BrandeisPatrick/Ouro-1.4B-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "BrandeisPatrick/Ouro-1.4B-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/BrandeisPatrick/Ouro-1.4B-GGUF:Q4_K_M
- Ollama
How to use BrandeisPatrick/Ouro-1.4B-GGUF with Ollama:
ollama run hf.co/BrandeisPatrick/Ouro-1.4B-GGUF:Q4_K_M
- Unsloth Desktop
- Docker Model Runner
How to use BrandeisPatrick/Ouro-1.4B-GGUF with Docker Model Runner:
docker model run hf.co/BrandeisPatrick/Ouro-1.4B-GGUF:Q4_K_M
- Lemonade
How to use BrandeisPatrick/Ouro-1.4B-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull BrandeisPatrick/Ouro-1.4B-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.Ouro-1.4B-GGUF-Q4_K_M
List all available models
lemonade list
- Atomic Chat
Ouro-1.4B — GGUF
GGUF conversions of ByteDance/Ouro-1.4B, a looped language model: the entire 24-layer decoder stack is applied 4 times per token with shared weights, so a 1.4B-parameter model computes at an effective depth of 96 layers.
What this repository is, and is not
| Model weights | ByteDance Seed's, unchanged. Trained by them, licensed Apache-2.0 by them. Nothing here was fine-tuned. |
| What was added | The ouro architecture for llama.cpp — the graph, the HF→GGUF conversion, and the registration — so the weights can be loaded by GGUF runtimes at all. Plus these conversions and their measured accuracy. |
| Upstream status | Submitted as a llama.cpp pull request from the ouro-arch branch of BrandeisPatrick/loop-transformer. Until it merges, these files need the patched build linked below. |
| Credit | If you use the model, cite ByteDance's paper (below). If you use the port or the evaluation harness, link the GitHub repository. |
These are the first GGUFs of this architecture. llama.cpp had no ouro architecture, so no GGUF
runtime could load Ouro at all — the request on the Ollama tracker
(#14252) has been open since February 2026. The
architecture was written for this release; the patch and the evaluation harness are at
BrandeisPatrick/loop-transformer.
These files need a patched llama.cpp today. The
ouroarchitecture is not yet in upstream llama.cpp, so stock llama.cpp, Ollama and LM Studio cannot load them yet. Build with the patch:git clone https://github.com/BrandeisPatrick/loop-transformer && loop-transformer/llamacpp/build.sh. Once the upstream PR merges and Ollama bumps its llama.cpp pin, these files will run unmodified.
Files
| file | size | GSM8K | note |
|---|---|---|---|
Ouro-1.4B-F16.gguf |
2.7 GB | 80.5% | reference precision |
Ouro-1.4B-Q8_0.gguf |
1.4 GB | 79.8% | recommended — matches F16 within noise |
Ouro-1.4B-Q4_K_M.gguf |
854 MB | 75.0% | smallest; costs ~5 points, see below |
GSM8K is a fixed 200-problem subset (100 for the quants), 3-shot chain of thought, greedy, strict answer match — the protocol the Ouro paper specifies in its Table 16. The published figure for this model is 78.92.
Verification against the reference implementation
The port was validated against numbers measured with the original transformers implementation on the same machine before it existed, so this is a comparison to data rather than to an impression.
| loops | this GGUF (F16) | transformers bf16 | seconds/item |
|---|---|---|---|
| 1 | 26.0 | 23.0 | 2.4 vs 11.1 |
| 2 | 67.0 | 64.0 | ~6 vs 19.3 |
| 4 | 80.5 ± 2.8 | 80.0 ± 2.8 | 11.3 vs 38.6 |
Every depth is within one standard error, greedy output is token-identical on a smoke prompt, and the
model loads as n_layer = 96 at 1.43 B parameters — depth expanded, weights stored once. On an
Apple M4 the GGUF runs about 3.4x faster than the reference does on MPS.
The loop count is a runtime dial
Unusually for a GGUF, the compute/accuracy trade-off is adjustable at load time from a single file, because llama.cpp's loader consults key overrides before the file:
llama-cli -m Ouro-1.4B-Q8_0.gguf --override-kv ouro.num_loops=int:1 # 24 layers, ~36 tok/s
llama-cli -m Ouro-1.4B-Q8_0.gguf --override-kv ouro.num_loops=int:2 # 48 layers
llama-cli -m Ouro-1.4B-Q8_0.gguf # 96 layers, ~9.5 tok/s, default
Accuracy follows depth: 26 / 67 / 80 percent on GSM8K at 1 / 2 / 4 loops. Note the model was trained at 4 loops; the published ablation shows quality degrading beyond that.
Quantization notes
Q4_K_M costs about 5 points, more than a dense model this size usually loses. A natural hypothesis is that a looped model re-applies the same weight error once per loop, so the damage compounds with depth. That was tested and is false: the Q4 penalty is 6.0 points at depth 1 and 5.5 at depth 4, a difference of −0.5 against a combined standard error of 7.9. The cost is flat in depth. Q8_0 is effectively lossless and is the recommended file.
Always quote the loop depth alongside a number from these files — the same file scores 20% or 75% depending only on a load-time flag.
Provenance
Converted from ByteDance/Ouro-1.4B at commit 574fa66cb8bf5abdc979642d01cf2b79b16bfab1 with a
llama.cpp built from upstream 67672dc plus the ouro architecture patch. The early-exit gate is
deliberately not converted: it selects which already-computed loop feeds the LM head rather than
changing what is computed, and at the shipped early_exit_threshold = 1.0 it never fires (measured:
0 of 299 token positions exit early).
Citation
@article{ouro2025,
title = {Scaling Latent Reasoning via Looped Language Models},
author = {ByteDance Seed},
journal= {arXiv:2510.25741},
year = {2025}
}
- Downloads last month
- 102
4-bit
8-bit
16-bit
Model tree for BrandeisPatrick/Ouro-1.4B-GGUF
Base model
ByteDance/Ouro-1.4BPaper for BrandeisPatrick/Ouro-1.4B-GGUF
Evaluation results
- accuracy (F16, 4 loops) on GSM8K (200-problem subset, 3-shot CoT, strict match)test set BrandeisPatrick/loop-transformer80.500
- accuracy (Q8_0, 4 loops) on GSM8K (200-problem subset, 3-shot CoT, strict match)test set BrandeisPatrick/loop-transformer79.800
- accuracy (Q4_K_M, 4 loops) on GSM8K (200-problem subset, 3-shot CoT, strict match)test set BrandeisPatrick/loop-transformer75.000
- accuracy (F16, 1 loop) on GSM8K (200-problem subset, 3-shot CoT, strict match)test set BrandeisPatrick/loop-transformer26.000