SeaWolf-AI commited on
Commit
6cdef29
·
verified ·
1 Parent(s): 3a750fb

Model card: Darwin concept, RSI, ZTC, results, specs

Browse files
Files changed (1) hide show
  1. README.md +212 -22
README.md CHANGED
@@ -1,35 +1,222 @@
1
  ---
2
- library_name: transformers
3
  license: other
4
  license_name: qwen-community-1.0
5
  license_link: LICENSE
 
 
6
  pipeline_tag: image-text-to-text
7
  tags:
8
- - darwin
9
- - rsi
10
- - recursive-self-improvement
11
- - vidraft
12
- - final-bench
13
- - qwen
14
- - moe
15
- - 180b
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
16
  ---
17
 
18
  # Darwin-180B-RSI
19
 
20
- **Darwin-180B-RSI** is a 180B-class mixture-of-experts vision-language model from **VIDRAFT / FINAL-Bench**, trained with our recursive self-improvement (RSI) pipeline: the model solves practice problems, its own solutions are checked against verifiable answer keys, and it is re-trained on what it got right.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
21
 
22
- Built on [Qwen3.8-Flash-Next](https://huggingface.co/Qwen/Qwen3.8-Flash-Next) (Qwen Community License 1.0). Only a small subset of weights was updated; experts, router and the vision encoder are unchanged.
 
23
 
24
- ## Results
 
 
 
25
 
26
- | Benchmark | Score | Setting |
27
- |---|---|---|
28
- | GPQA Diamond (198) | **94.44** | majority vote over up to 16 samples · thinking budget 131,072 tokens |
 
 
 
 
 
29
 
30
- All numbers are self-measured. Sampling: temperature 1.0, top_p 0.95, top_k 20, bf16. More leaderboard results will be added as they are measured.
31
 
32
- ## Usage
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
33
 
34
  ```python
35
  from transformers import AutoProcessor, AutoModelForImageTextToText
@@ -38,12 +225,15 @@ processor = AutoProcessor.from_pretrained(model_id)
38
  model = AutoModelForImageTextToText.from_pretrained(model_id, torch_dtype="auto", device_map="auto")
39
  ```
40
 
41
- The model is served with vLLM in the same way as the base model (`--tensor-parallel-size 8 --enable-expert-parallel`).
 
 
 
42
 
43
- ## License
44
 
45
- This model is a derivative of Qwen3.8-Flash-Next and is distributed under the **Qwen Community License 1.0** (see `LICENSE`).
46
 
47
- ## Organization
48
 
49
- VIDRAFT · FINAL-Bench
 
1
  ---
 
2
  license: other
3
  license_name: qwen-community-1.0
4
  license_link: LICENSE
5
+ language: [en, ko, zh, ja, multilingual]
6
+ library_name: transformers
7
  pipeline_tag: image-text-to-text
8
  tags:
9
+ - darwin
10
+ - darwin-rsi
11
+ - recursive-self-improvement
12
+ - self-improvement
13
+ - vidraft
14
+ - final-bench
15
+ - qwen
16
+ - qwen3.8
17
+ - moe
18
+ - mixture-of-experts
19
+ - sparse-moe
20
+ - 180b
21
+ - hybrid-attention
22
+ - linear-attention
23
+ - long-context
24
+ - 262k-context
25
+ - vision-language
26
+ - multimodal
27
+ - reasoning
28
+ - reasoning-model
29
+ - thinking
30
+ - chain-of-thought
31
+ - math
32
+ - science
33
+ - stem
34
+ - ztc
35
+ - zero-token-confidence
36
+ - confidence-estimation
37
+ - hallucination-detection
38
+ - gpqa
39
+ - gpqa-diamond
40
+ - mmlu-pro
41
+ - mmmu-pro
42
+ - eval-results
43
+ - korean
44
+ - english
45
+ - vllm
46
+ - openai-compatible
47
+ - b200
48
+ model-index:
49
+ - name: Darwin-180B-RSI
50
+ results:
51
+ - task: {type: text-generation, name: Graduate-Level Reasoning}
52
+ dataset: {type: Idavidrein/gpqa, name: GPQA Diamond, config: gpqa_diamond, split: train}
53
+ metrics:
54
+ - {type: accuracy, value: 94.44, name: "Accuracy (majority vote, up to 16 samples, 131K thinking)", verified: false}
55
+ - task: {type: text-generation, name: Multi-discipline Knowledge & Reasoning}
56
+ dataset: {type: TIGER-Lab/MMLU-Pro, name: MMLU-Pro, split: test}
57
+ metrics:
58
+ - {type: accuracy, value: 88.12, name: "Accuracy (single sample, 131K thinking)", verified: false}
59
  ---
60
 
61
  # Darwin-180B-RSI
62
 
63
+ ### 180B Mixture-of-Experts · vision-language · **GPQA Diamond 94.44 % — #1 on the Hugging Face leaderboard** · **self-improving**
64
+
65
+ `reasoning` · `MoE 512 experts` · `262K long context` · `image + text` · `Korean + English` · `self-improvement` · `ZTC`
66
+
67
+ <p align="center">
68
+ <a href="https://vidraft.net"><img src="https://img.shields.io/badge/🌐_VIDRAFT-vidraft.net-111827?style=for-the-badge"></a>
69
+ <a href="https://huggingface.co/datasets/Idavidrein/gpqa"><img src="https://img.shields.io/badge/GPQA_Diamond-94.44%25_%231-gold?style=for-the-badge"></a>
70
+ <a href="https://huggingface.co/datasets/TIGER-Lab/MMLU-Pro"><img src="https://img.shields.io/badge/MMLU--Pro-88.12%25-2563eb?style=for-the-badge"></a>
71
+ <img src="https://img.shields.io/badge/Self--Improving-RSI-e11d48?style=for-the-badge">
72
+ <img src="https://img.shields.io/badge/ZTC-Zero--Token_Confidence-7c3aed?style=for-the-badge">
73
+ </p>
74
+
75
+ **The newest flagship of the Darwin family — #1 on GPQA Diamond,
76
+ and a model that gets better by learning from its own verified work.**
77
+
78
+ ---
79
+
80
+ ## 🧬 The Darwin Family
81
+
82
+ <p align="center">
83
+ <a href="https://huggingface.co/FINAL-Bench/Darwin-397B-ZTC"><img src="https://img.shields.io/badge/Darwin--397B--ZTC-GPQA_93.43-16a34a"></a>
84
+ <a href="https://huggingface.co/FINAL-Bench/Darwin-28B-REASON"><img src="https://img.shields.io/badge/Darwin--28B--REASON-GPQA_89.39-16a34a"></a>
85
+ <a href="https://huggingface.co/FINAL-Bench/Darwin-35B-A3B-Opus"><img src="https://img.shields.io/badge/Darwin--35B--A3B--Opus-♥98-e11d48"></a>
86
+ <a href="https://huggingface.co/FINAL-Bench/Darwin-36B-Opus"><img src="https://img.shields.io/badge/Darwin--36B--Opus-♥97-e11d48"></a>
87
+ <a href="https://huggingface.co/FINAL-Bench/Darwin-4B-Genesis"><img src="https://img.shields.io/badge/Darwin--4B--Genesis-♥63-e11d48"></a>
88
+ </p>
89
+ <p align="center">
90
+ <a href="https://huggingface.co/FINAL-Bench/Darwin-9B-NEG"><img src="https://img.shields.io/badge/Darwin--9B--NEG-♥57-e11d48"></a>
91
+ <a href="https://huggingface.co/FINAL-Bench/POCKET-35B-GGUF"><img src="https://img.shields.io/badge/POCKET--35B-824K_↓-1f6feb"></a>
92
+ <a href="https://huggingface.co/FINAL-Bench/POCKET-26B-GGUF"><img src="https://img.shields.io/badge/POCKET--26B-365K_↓-1f6feb"></a>
93
+ <a href="https://huggingface.co/FINAL-Bench/POCKET-EN-GGUF"><img src="https://img.shields.io/badge/POCKET--EN-♥43-1f6feb"></a>
94
+ <a href="https://huggingface.co/FINAL-Bench/POCKET-KR-GGUF"><img src="https://img.shields.io/badge/POCKET--KR-♥36-1f6feb"></a>
95
+ </p>
96
+
97
+ **Darwin** is [VIDRAFT](https://vidraft.net)'s measurement-driven reasoning model family —
98
+ roughly **20 official models**, **400+ community derivatives**, and now **two places in the GPQA Diamond top 3**
99
+ (Darwin-180B-RSI #1 · Darwin-397B-ZTC #3).
100
+
101
+ ---
102
+
103
+ ## 🧬 Darwin — evolve the parent, keep what works
104
 
105
+ Darwin treats a strong open model as a **parent**. It measures where the parent is weak,
106
+ and strengthens exactly those parts — instead of re-training everything and risking what already works.
107
 
108
+ - **Diagnose before you change.** Every Darwin generation starts from a measured weakness map of the parent.
109
+ - **Change little, precisely.** Darwin modifies a small, targeted fraction of the network. Knowledge stored in the experts is preserved.
110
+ - **Proven capability over new guesses.** Earlier Darwin generations grafted the best-performing expert/FFN blocks from other strong models onto a base backbone; Darwin-180B-RSI adds a new ingredient — **the model's own verified work**.
111
+ - **Measured, not claimed.** Every change must beat the parent on held-out tests before it ships.
112
 
113
+ | Model | Scale | GPQA Diamond |
114
+ |:---|:---|:---:|
115
+ | Darwin-9B-NEG | 9B | 84.3 |
116
+ | Darwin-27B-Opus | 27B dense | 86.9 |
117
+ | Darwin-36B-Opus | 36B MoE | 88.4 |
118
+ | Darwin-28B-REASON | 28B + DELPHI | 89.39 |
119
+ | Darwin-397B-ZTC | 397B MoE (FP8) | 93.43 |
120
+ | **Darwin-180B-RSI** | **180B MoE** | **94.44** |
121
 
122
+ ### Lineage
123
 
124
+ | Role | | |
125
+ |:---|:---|:---|
126
+ | **Parent** | `Qwen/Qwen3.8-Flash-Next` | 180B MoE vision-language backbone · Qwen Community License 1.0 |
127
+ | **Darwin RSI** | self-improvement on verified answers | the parent's own solutions, checked against verifiable answer keys, fed back as training signal |
128
+ | **Preserved** | 512 routed experts · router · vision encoder | untouched — the parent's knowledge stays intact |
129
+ | **ZTC** | zero-token confidence readout | see below |
130
+
131
+ ---
132
+
133
+ ## 🔁 RSI — a model that improves from its own work
134
+
135
+ **Recursive self-improvement (RSI)** is the core of this generation.
136
+ Instead of distilling a bigger teacher, the model improves by learning from itself:
137
+
138
+ 1. **Solve** — the model works through practice problems it has never seen in evaluation.
139
+ 2. **Verify** — its answers are checked against verifiable references (answer keys, executable checks). Nothing unverified is learned.
140
+ 3. **Learn** — it is re-trained on the reasoning that turned out to be correct.
141
+ 4. **Repeat** — the improved model becomes the next solver.
142
+
143
+ What it bought in this release:
144
+
145
+ | | Parent (Qwen3.8-Flash-Next) | **Darwin-180B-RSI** |
146
+ |:---|:---:|:---:|
147
+ | Average reasoning length (MMLU-Pro) | 4,320 tokens | **3,833 tokens (−11 %)** |
148
+ | MMLU-Pro accuracy | 88.04 % | **88.12 %** |
149
+
150
+ **Same or better accuracy with shorter reasoning** — cheaper and faster to serve.
151
+ Practice sets are deduplicated against every evaluation set we report (8-gram overlap filter).
152
+
153
+ ---
154
+
155
+ ## 🏛️ ZTC — it knows before it answers
156
+
157
+ **Zero-Token Confidence (ZTC)** reads the model's own internal state **once, before generation**,
158
+ and returns the probability that the answer it is about to give is correct — **no extra tokens, no second model.**
159
+
160
+ ```json
161
+ {"answer": "...", "confidence": 0.93, "ztc_score": 1.84, "truncated": false}
162
+ ```
163
+
164
+ Use it to gate actions: when confidence is low, do not call the tool, escalate, or answer "I don't know".
165
+ The ZTC readout for this model is being fitted and will ship in `ztc/` (same format as
166
+ [Darwin-397B-ZTC](https://huggingface.co/FINAL-Bench/Darwin-397B-ZTC)).
167
+
168
+ ---
169
+
170
+ ## 🏆 Results
171
+
172
+ | Benchmark | Score | Setting | Leaderboard |
173
+ |:---|:---:|:---|:---|
174
+ | **GPQA Diamond** (198) | **94.44** | majority vote over up to 16 samples · 131,072-token thinking budget | [**#1**](https://huggingface.co/datasets/Idavidrein/gpqa) |
175
+ | **MMLU-Pro** (12,032) | **88.12** | single sample · 131,072-token thinking budget | [leaderboard](https://huggingface.co/datasets/TIGER-Lab/MMLU-Pro) |
176
+ | MMMU-Pro (vision, 1,730) | measuring | majority vote | [leaderboard](https://huggingface.co/datasets/MMMU/MMMU_Pro) |
177
+
178
+ Sampling for all runs: temperature 1.0 · top_p 0.95 · top_k 20 · bf16. All numbers are self-measured and reproducible with the settings above.
179
+
180
+ **MMLU-Pro by category (single sample)** — strongest in math 95.0 · biology 94.6 · physics 92.5; room to grow in law and history.
181
+
182
+ ---
183
+
184
+ ## ⚙️ Specifications
185
+
186
+ | | |
187
+ |:---|:---|
188
+ | Architecture | Mixture-of-Experts, hybrid attention (36 linear-attention + 12 full-attention layers) |
189
+ | Layers / hidden | 48 / 2,560 |
190
+ | Experts | 512 routed (10 active per token) + shared expert |
191
+ | Context | 262,144 tokens |
192
+ | Vocabulary | 248,320 |
193
+ | Modalities | image + text → text |
194
+ | Precision | bf16 (~336 GB) |
195
+
196
+ ---
197
+
198
+ ## 🚀 Quickstart
199
+
200
+ ### Serving with vLLM (8 × B200 or equivalent)
201
+
202
+ ```bash
203
+ vllm serve FINAL-Bench/Darwin-180B-RSI \
204
+ --tensor-parallel-size 8 --enable-expert-parallel \
205
+ --max-model-len 135168 --trust-remote-code
206
+ ```
207
+
208
+ ### Chat Completions (OpenAI-compatible)
209
+
210
+ ```python
211
+ from openai import OpenAI
212
+ c = OpenAI(base_url="http://localhost:8000/v1", api_key="-")
213
+ r = c.chat.completions.create(model="FINAL-Bench/Darwin-180B-RSI",
214
+ messages=[{"role": "user", "content": "Explain why the sky is blue in two sentences."}],
215
+ temperature=1.0, top_p=0.95, extra_body={"top_k": 20})
216
+ print(r.choices[0].message.content)
217
+ ```
218
+
219
+ ### Transformers
220
 
221
  ```python
222
  from transformers import AutoProcessor, AutoModelForImageTextToText
 
225
  model = AutoModelForImageTextToText.from_pretrained(model_id, torch_dtype="auto", device_map="auto")
226
  ```
227
 
228
+ **Tip:** this is a thinking model. Give it room — a thinking budget of 32K–131K tokens is recommended for hard reasoning.
229
+ Short budgets truncate the reasoning and cost accuracy.
230
+
231
+ ---
232
 
233
+ ## 📜 License
234
 
235
+ Darwin-180B-RSI is a derivative of **Qwen3.8-Flash-Next** and is distributed under the **Qwen Community License 1.0** (see `LICENSE`).
236
 
237
+ ## 🏢 About
238
 
239
+ Built by **[VIDRAFT](https://vidraft.net)** · evaluated with **FINAL-Bench**.