--- license: other license_name: qwen-community-1.0 license_link: LICENSE language: [en, ko, zh, ja, multilingual] library_name: transformers pipeline_tag: image-text-to-text tags: - darwin - darwin-rsi - recursive-self-improvement - self-improvement - vidraft - final-bench - qwen - qwen3.8 - moe - mixture-of-experts - sparse-moe - 180b - hybrid-attention - linear-attention - long-context - 262k-context - vision-language - multimodal - reasoning - reasoning-model - thinking - chain-of-thought - math - science - stem - ztc - zero-token-confidence - confidence-estimation - hallucination-detection - gpqa - gpqa-diamond - mmlu-pro - mmmu-pro - eval-results - korean - english - vllm - openai-compatible - b200 model-index: - name: Darwin-180B-RSI results: - task: {type: text-generation, name: Graduate-Level Reasoning} dataset: {type: Idavidrein/gpqa, name: GPQA Diamond, config: gpqa_diamond, split: train} metrics: - {type: accuracy, value: 94.44, name: "Accuracy (majority vote, up to 16 samples, 131K thinking)", verified: false} - task: {type: text-generation, name: Multi-discipline Knowledge & Reasoning} dataset: {type: TIGER-Lab/MMLU-Pro, name: MMLU-Pro, split: test} metrics: - {type: accuracy, value: 88.12, name: "Accuracy (single sample, 131K thinking)", verified: false} --- # Darwin-180B-RSI ### 180B Mixture-of-Experts · vision-language · **GPQA Diamond 94.44 % — #1 on the Hugging Face leaderboard** · **self-improving** `reasoning` · `MoE 512 experts` · `262K long context` · `image + text` · `Korean + English` · `self-improvement` · `ZTC`
**The newest flagship of the Darwin family — #1 on GPQA Diamond, and a model that gets better by learning from its own verified work.** --- ## 🧬 The Darwin Family **Darwin** is [VIDRAFT](https://vidraft.net)'s measurement-driven reasoning model family — roughly **20 official models**, **400+ community derivatives**, and now **two places in the GPQA Diamond top 3** (Darwin-180B-RSI #1 · Darwin-397B-ZTC #3). --- ## 🧬 Darwin — evolve the parent, keep what works Darwin treats a strong open model as a **parent**. It measures where the parent is weak, and strengthens exactly those parts — instead of re-training everything and risking what already works. - **Diagnose before you change.** Every Darwin generation starts from a measured weakness map of the parent. - **Change little, precisely.** Darwin modifies a small, targeted fraction of the network. Knowledge stored in the experts is preserved. - **Proven capability over new guesses.** Earlier Darwin generations grafted the best-performing expert/FFN blocks from other strong models onto a base backbone; Darwin-180B-RSI adds a new ingredient — **the model's own verified work**. - **Measured, not claimed.** Every change must beat the parent on held-out tests before it ships. | Model | Scale | GPQA Diamond | |:---|:---|:---:| | Darwin-9B-NEG | 9B | 84.3 | | Darwin-27B-Opus | 27B dense | 86.9 | | Darwin-36B-Opus | 36B MoE | 88.4 | | Darwin-28B-REASON | 28B + DELPHI | 89.39 | | Darwin-397B-ZTC | 397B MoE (FP8) | 93.43 | | **Darwin-180B-RSI** | **180B MoE** | **94.44** | ### Lineage | Role | | | |:---|:---|:---| | **Parent** | `Qwen/Qwen3.8-Flash-Next` | 180B MoE vision-language backbone · Qwen Community License 1.0 | | **Darwin RSI** | self-improvement on verified answers | the parent's own solutions, checked against verifiable answer keys, fed back as training signal | | **Preserved** | 512 routed experts · router · vision encoder | untouched — the parent's knowledge stays intact | | **ZTC** | zero-token confidence readout | see below | --- ## 🔁 RSI — a model that improves from its own work **Recursive self-improvement (RSI)** is the core of this generation. Instead of distilling a bigger teacher, the model improves by learning from itself: 1. **Solve** — the model works through practice problems it has never seen in evaluation. 2. **Verify** — its answers are checked against verifiable references (answer keys, executable checks). Nothing unverified is learned. 3. **Learn** — it is re-trained on the reasoning that turned out to be correct. 4. **Repeat** — the improved model becomes the next solver. What it bought in this release: | | Parent (Qwen3.8-Flash-Next) | **Darwin-180B-RSI** | |:---|:---:|:---:| | Average reasoning length (MMLU-Pro) | 4,320 tokens | **3,833 tokens (−11 %)** | | MMLU-Pro accuracy | 88.04 % | **88.12 %** | **Same or better accuracy with shorter reasoning** — cheaper and faster to serve. Practice sets are deduplicated against every evaluation set we report (8-gram overlap filter). --- ## 🏛️ ZTC — it knows before it answers **Zero-Token Confidence (ZTC)** reads the model's own internal state **once, before generation**, and returns the probability that the answer it is about to give is correct — **no extra tokens, no second model.** ```json {"answer": "...", "confidence": 0.93, "ztc_score": 1.84, "truncated": false} ``` Use it to gate actions: when confidence is low, do not call the tool, escalate, or answer "I don't know". The ZTC readout for this model is being fitted and will ship in `ztc/` (same format as [Darwin-397B-ZTC](https://huggingface.co/FINAL-Bench/Darwin-397B-ZTC)). --- ## 🏆 Results | Benchmark | Score | Setting | Leaderboard | |:---|:---:|:---|:---| | **GPQA Diamond** (198) | **94.44** | majority vote over up to 16 samples · 131,072-token thinking budget | [**#1**](https://huggingface.co/datasets/Idavidrein/gpqa) | | **MMLU-Pro** (12,032) | **88.12** | single sample · 131,072-token thinking budget | [leaderboard](https://huggingface.co/datasets/TIGER-Lab/MMLU-Pro) | | MMMU-Pro (vision, 1,730) | measuring | majority vote | [leaderboard](https://huggingface.co/datasets/MMMU/MMMU_Pro) | Sampling for all runs: temperature 1.0 · top_p 0.95 · top_k 20 · bf16. All numbers are self-measured and reproducible with the settings above. **MMLU-Pro by category (single sample)** — strongest in math 95.0 · biology 94.6 · physics 92.5; room to grow in law and history. --- ## ⚙️ Specifications | | | |:---|:---| | Architecture | Mixture-of-Experts, hybrid attention (36 linear-attention + 12 full-attention layers) | | Layers / hidden | 48 / 2,560 | | Experts | 512 routed (10 active per token) + shared expert | | Context | 262,144 tokens | | Vocabulary | 248,320 | | Modalities | image + text → text | | Precision | bf16 (~336 GB) | --- ## 🚀 Quickstart ### Serving with vLLM (8 × B200 or equivalent) ```bash vllm serve FINAL-Bench/Darwin-180B-RSI \ --tensor-parallel-size 8 --enable-expert-parallel \ --max-model-len 135168 --trust-remote-code ``` ### Chat Completions (OpenAI-compatible) ```python from openai import OpenAI c = OpenAI(base_url="http://localhost:8000/v1", api_key="-") r = c.chat.completions.create(model="FINAL-Bench/Darwin-180B-RSI", messages=[{"role": "user", "content": "Explain why the sky is blue in two sentences."}], temperature=1.0, top_p=0.95, extra_body={"top_k": 20}) print(r.choices[0].message.content) ``` ### Transformers ```python from transformers import AutoProcessor, AutoModelForImageTextToText model_id = "FINAL-Bench/Darwin-180B-RSI" processor = AutoProcessor.from_pretrained(model_id) model = AutoModelForImageTextToText.from_pretrained(model_id, torch_dtype="auto", device_map="auto") ``` **Tip:** this is a thinking model. Give it room — a thinking budget of 32K–131K tokens is recommended for hard reasoning. Short budgets truncate the reasoning and cost accuracy. --- ## 📜 License Darwin-180B-RSI is a derivative of **Qwen3.8-Flash-Next** and is distributed under the **Qwen Community License 1.0** (see `LICENSE`). ## 🏢 About Built by **[VIDRAFT](https://vidraft.net)** · evaluated with **FINAL-Bench**.