File size: 10,017 Bytes
3a750fb
 
 
 
6cdef29
 
3a750fb
 
6cdef29
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
3a750fb
 
 
 
6cdef29
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
3a750fb
6cdef29
 
3a750fb
6cdef29
 
 
 
3a750fb
6cdef29
 
 
 
 
 
 
 
3a750fb
6cdef29
3a750fb
6cdef29
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
3a750fb
 
 
 
 
 
 
 
6cdef29
 
 
 
3a750fb
6cdef29
3a750fb
6cdef29
3a750fb
6cdef29
3a750fb
6cdef29
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
---
license: other
license_name: qwen-community-1.0
license_link: LICENSE
language: [en, ko, zh, ja, multilingual]
library_name: transformers
pipeline_tag: image-text-to-text
tags:
  - darwin
  - darwin-rsi
  - recursive-self-improvement
  - self-improvement
  - vidraft
  - final-bench
  - qwen
  - qwen3.8
  - moe
  - mixture-of-experts
  - sparse-moe
  - 180b
  - hybrid-attention
  - linear-attention
  - long-context
  - 262k-context
  - vision-language
  - multimodal
  - reasoning
  - reasoning-model
  - thinking
  - chain-of-thought
  - math
  - science
  - stem
  - ztc
  - zero-token-confidence
  - confidence-estimation
  - hallucination-detection
  - gpqa
  - gpqa-diamond
  - mmlu-pro
  - mmmu-pro
  - eval-results
  - korean
  - english
  - vllm
  - openai-compatible
  - b200
model-index:
  - name: Darwin-180B-RSI
    results:
      - task: {type: text-generation, name: Graduate-Level Reasoning}
        dataset: {type: Idavidrein/gpqa, name: GPQA Diamond, config: gpqa_diamond, split: train}
        metrics:
          - {type: accuracy, value: 94.44, name: "Accuracy (majority vote, up to 16 samples, 131K thinking)", verified: false}
      - task: {type: text-generation, name: Multi-discipline Knowledge & Reasoning}
        dataset: {type: TIGER-Lab/MMLU-Pro, name: MMLU-Pro, split: test}
        metrics:
          - {type: accuracy, value: 88.12, name: "Accuracy (single sample, 131K thinking)", verified: false}
---

# Darwin-180B-RSI

### 180B Mixture-of-Experts · vision-language · **GPQA Diamond 94.44 % — #1 on the Hugging Face leaderboard** · **self-improving**

`reasoning` · `MoE 512 experts` · `262K long context` · `image + text` · `Korean + English` · `self-improvement` · `ZTC`

<p align="center">
<a href="https://vidraft.net"><img src="https://img.shields.io/badge/🌐_VIDRAFT-vidraft.net-111827?style=for-the-badge"></a>
<a href="https://huggingface.co/datasets/Idavidrein/gpqa"><img src="https://img.shields.io/badge/GPQA_Diamond-94.44%25_%231-gold?style=for-the-badge"></a>
<a href="https://huggingface.co/datasets/TIGER-Lab/MMLU-Pro"><img src="https://img.shields.io/badge/MMLU--Pro-88.12%25-2563eb?style=for-the-badge"></a>
<img src="https://img.shields.io/badge/Self--Improving-RSI-e11d48?style=for-the-badge">
<img src="https://img.shields.io/badge/ZTC-Zero--Token_Confidence-7c3aed?style=for-the-badge">
</p>

**The newest flagship of the Darwin family — #1 on GPQA Diamond,
and a model that gets better by learning from its own verified work.**

---

## 🧬 The Darwin Family

<p align="center">
<a href="https://huggingface.co/FINAL-Bench/Darwin-397B-ZTC"><img src="https://img.shields.io/badge/Darwin--397B--ZTC-GPQA_93.43-16a34a"></a>
<a href="https://huggingface.co/FINAL-Bench/Darwin-28B-REASON"><img src="https://img.shields.io/badge/Darwin--28B--REASON-GPQA_89.39-16a34a"></a>
<a href="https://huggingface.co/FINAL-Bench/Darwin-35B-A3B-Opus"><img src="https://img.shields.io/badge/Darwin--35B--A3B--Opus-♥98-e11d48"></a>
<a href="https://huggingface.co/FINAL-Bench/Darwin-36B-Opus"><img src="https://img.shields.io/badge/Darwin--36B--Opus-♥97-e11d48"></a>
<a href="https://huggingface.co/FINAL-Bench/Darwin-4B-Genesis"><img src="https://img.shields.io/badge/Darwin--4B--Genesis-♥63-e11d48"></a>
</p>
<p align="center">
<a href="https://huggingface.co/FINAL-Bench/Darwin-9B-NEG"><img src="https://img.shields.io/badge/Darwin--9B--NEG-♥57-e11d48"></a>
<a href="https://huggingface.co/FINAL-Bench/POCKET-35B-GGUF"><img src="https://img.shields.io/badge/POCKET--35B-824K_↓-1f6feb"></a>
<a href="https://huggingface.co/FINAL-Bench/POCKET-26B-GGUF"><img src="https://img.shields.io/badge/POCKET--26B-365K_↓-1f6feb"></a>
<a href="https://huggingface.co/FINAL-Bench/POCKET-EN-GGUF"><img src="https://img.shields.io/badge/POCKET--EN-♥43-1f6feb"></a>
<a href="https://huggingface.co/FINAL-Bench/POCKET-KR-GGUF"><img src="https://img.shields.io/badge/POCKET--KR-♥36-1f6feb"></a>
</p>

**Darwin** is [VIDRAFT](https://vidraft.net)'s measurement-driven reasoning model family —
roughly **20 official models**, **400+ community derivatives**, and now **two places in the GPQA Diamond top 3**
(Darwin-180B-RSI #1 · Darwin-397B-ZTC #3).

---

## 🧬 Darwin — evolve the parent, keep what works

Darwin treats a strong open model as a **parent**. It measures where the parent is weak,
and strengthens exactly those parts — instead of re-training everything and risking what already works.

- **Diagnose before you change.** Every Darwin generation starts from a measured weakness map of the parent.
- **Change little, precisely.** Darwin modifies a small, targeted fraction of the network. Knowledge stored in the experts is preserved.
- **Proven capability over new guesses.** Earlier Darwin generations grafted the best-performing expert/FFN blocks from other strong models onto a base backbone; Darwin-180B-RSI adds a new ingredient — **the model's own verified work**.
- **Measured, not claimed.** Every change must beat the parent on held-out tests before it ships.

| Model | Scale | GPQA Diamond |
|:---|:---|:---:|
| Darwin-9B-NEG | 9B | 84.3 |
| Darwin-27B-Opus | 27B dense | 86.9 |
| Darwin-36B-Opus | 36B MoE | 88.4 |
| Darwin-28B-REASON | 28B + DELPHI | 89.39 |
| Darwin-397B-ZTC | 397B MoE (FP8) | 93.43 |
| **Darwin-180B-RSI** | **180B MoE** | **94.44** |

### Lineage

| Role | | |
|:---|:---|:---|
| **Parent** | `Qwen/Qwen3.8-Flash-Next` | 180B MoE vision-language backbone · Qwen Community License 1.0 |
| **Darwin RSI** | self-improvement on verified answers | the parent's own solutions, checked against verifiable answer keys, fed back as training signal |
| **Preserved** | 512 routed experts · router · vision encoder | untouched — the parent's knowledge stays intact |
| **ZTC** | zero-token confidence readout | see below |

---

## 🔁 RSI — a model that improves from its own work

**Recursive self-improvement (RSI)** is the core of this generation.
Instead of distilling a bigger teacher, the model improves by learning from itself:

1. **Solve** — the model works through practice problems it has never seen in evaluation.
2. **Verify** — its answers are checked against verifiable references (answer keys, executable checks). Nothing unverified is learned.
3. **Learn** — it is re-trained on the reasoning that turned out to be correct.
4. **Repeat** — the improved model becomes the next solver.

What it bought in this release:

| | Parent (Qwen3.8-Flash-Next) | **Darwin-180B-RSI** |
|:---|:---:|:---:|
| Average reasoning length (MMLU-Pro) | 4,320 tokens | **3,833 tokens (−11 %)** |
| MMLU-Pro accuracy | 88.04 % | **88.12 %** |

**Same or better accuracy with shorter reasoning** — cheaper and faster to serve.
Practice sets are deduplicated against every evaluation set we report (8-gram overlap filter).

---

## 🏛️ ZTC — it knows before it answers

**Zero-Token Confidence (ZTC)** reads the model's own internal state **once, before generation**,
and returns the probability that the answer it is about to give is correct — **no extra tokens, no second model.**

```json
{"answer": "...", "confidence": 0.93, "ztc_score": 1.84, "truncated": false}
```

Use it to gate actions: when confidence is low, do not call the tool, escalate, or answer "I don't know".
The ZTC readout for this model is being fitted and will ship in `ztc/` (same format as
[Darwin-397B-ZTC](https://huggingface.co/FINAL-Bench/Darwin-397B-ZTC)).

---

## 🏆 Results

| Benchmark | Score | Setting | Leaderboard |
|:---|:---:|:---|:---|
| **GPQA Diamond** (198) | **94.44** | majority vote over up to 16 samples · 131,072-token thinking budget | [**#1**](https://huggingface.co/datasets/Idavidrein/gpqa) |
| **MMLU-Pro** (12,032) | **88.12** | single sample · 131,072-token thinking budget | [leaderboard](https://huggingface.co/datasets/TIGER-Lab/MMLU-Pro) |
| MMMU-Pro (vision, 1,730) | measuring | majority vote | [leaderboard](https://huggingface.co/datasets/MMMU/MMMU_Pro) |

Sampling for all runs: temperature 1.0 · top_p 0.95 · top_k 20 · bf16. All numbers are self-measured and reproducible with the settings above.

**MMLU-Pro by category (single sample)** — strongest in math 95.0 · biology 94.6 · physics 92.5; room to grow in law and history.

---

## ⚙️ Specifications

| | |
|:---|:---|
| Architecture | Mixture-of-Experts, hybrid attention (36 linear-attention + 12 full-attention layers) |
| Layers / hidden | 48 / 2,560 |
| Experts | 512 routed (10 active per token) + shared expert |
| Context | 262,144 tokens |
| Vocabulary | 248,320 |
| Modalities | image + text → text |
| Precision | bf16 (~336 GB) |

---

## 🚀 Quickstart

### Serving with vLLM (8 × B200 or equivalent)

```bash
vllm serve FINAL-Bench/Darwin-180B-RSI \
  --tensor-parallel-size 8 --enable-expert-parallel \
  --max-model-len 135168 --trust-remote-code
```

### Chat Completions (OpenAI-compatible)

```python
from openai import OpenAI
c = OpenAI(base_url="http://localhost:8000/v1", api_key="-")
r = c.chat.completions.create(model="FINAL-Bench/Darwin-180B-RSI",
    messages=[{"role": "user", "content": "Explain why the sky is blue in two sentences."}],
    temperature=1.0, top_p=0.95, extra_body={"top_k": 20})
print(r.choices[0].message.content)
```

### Transformers

```python
from transformers import AutoProcessor, AutoModelForImageTextToText
model_id = "FINAL-Bench/Darwin-180B-RSI"
processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForImageTextToText.from_pretrained(model_id, torch_dtype="auto", device_map="auto")
```

**Tip:** this is a thinking model. Give it room — a thinking budget of 32K–131K tokens is recommended for hard reasoning.
Short budgets truncate the reasoning and cost accuracy.

---

## 📜 License

Darwin-180B-RSI is a derivative of **Qwen3.8-Flash-Next** and is distributed under the **Qwen Community License 1.0** (see `LICENSE`).

## 🏢 About

Built by **[VIDRAFT](https://vidraft.net)** · evaluated with **FINAL-Bench**.