Darwin-180B-RSI / README.md
SeaWolf-AI's picture
Model card: Darwin concept, RSI, ZTC, results, specs
6cdef29 verified
|
Raw History Blame
10 kB
metadata
license: other
license_name: qwen-community-1.0
license_link: LICENSE
language:
  - en
  - ko
  - zh
  - ja
  - multilingual
library_name: transformers
pipeline_tag: image-text-to-text
tags:
  - darwin
  - darwin-rsi
  - recursive-self-improvement
  - self-improvement
  - vidraft
  - final-bench
  - qwen
  - qwen3.8
  - moe
  - mixture-of-experts
  - sparse-moe
  - 180b
  - hybrid-attention
  - linear-attention
  - long-context
  - 262k-context
  - vision-language
  - multimodal
  - reasoning
  - reasoning-model
  - thinking
  - chain-of-thought
  - math
  - science
  - stem
  - ztc
  - zero-token-confidence
  - confidence-estimation
  - hallucination-detection
  - gpqa
  - gpqa-diamond
  - mmlu-pro
  - mmmu-pro
  - eval-results
  - korean
  - english
  - vllm
  - openai-compatible
  - b200
model-index:
  - name: Darwin-180B-RSI
    results:
      - task:
          type: text-generation
          name: Graduate-Level Reasoning
        dataset:
          type: Idavidrein/gpqa
          name: GPQA Diamond
          config: gpqa_diamond
          split: train
        metrics:
          - type: accuracy
            value: 94.44
            name: Accuracy (majority vote, up to 16 samples, 131K thinking)
            verified: false
      - task:
          type: text-generation
          name: Multi-discipline Knowledge & Reasoning
        dataset:
          type: TIGER-Lab/MMLU-Pro
          name: MMLU-Pro
          split: test
        metrics:
          - type: accuracy
            value: 88.12
            name: Accuracy (single sample, 131K thinking)
            verified: false

Darwin-180B-RSI

180B Mixture-of-Experts · vision-language · GPQA Diamond 94.44 % — #1 on the Hugging Face leaderboard · self-improving

reasoning · MoE 512 experts · 262K long context · image + text · Korean + English · self-improvement · ZTC

The newest flagship of the Darwin family — #1 on GPQA Diamond, and a model that gets better by learning from its own verified work.


🧬 The Darwin Family

Darwin is VIDRAFT's measurement-driven reasoning model family — roughly 20 official models, 400+ community derivatives, and now two places in the GPQA Diamond top 3 (Darwin-180B-RSI #1 · Darwin-397B-ZTC #3).


🧬 Darwin — evolve the parent, keep what works

Darwin treats a strong open model as a parent. It measures where the parent is weak, and strengthens exactly those parts — instead of re-training everything and risking what already works.

  • Diagnose before you change. Every Darwin generation starts from a measured weakness map of the parent.
  • Change little, precisely. Darwin modifies a small, targeted fraction of the network. Knowledge stored in the experts is preserved.
  • Proven capability over new guesses. Earlier Darwin generations grafted the best-performing expert/FFN blocks from other strong models onto a base backbone; Darwin-180B-RSI adds a new ingredient — the model's own verified work.
  • Measured, not claimed. Every change must beat the parent on held-out tests before it ships.
Model Scale GPQA Diamond
Darwin-9B-NEG 9B 84.3
Darwin-27B-Opus 27B dense 86.9
Darwin-36B-Opus 36B MoE 88.4
Darwin-28B-REASON 28B + DELPHI 89.39
Darwin-397B-ZTC 397B MoE (FP8) 93.43
Darwin-180B-RSI 180B MoE 94.44

Lineage

Role
Parent Qwen/Qwen3.8-Flash-Next 180B MoE vision-language backbone · Qwen Community License 1.0
Darwin RSI self-improvement on verified answers the parent's own solutions, checked against verifiable answer keys, fed back as training signal
Preserved 512 routed experts · router · vision encoder untouched — the parent's knowledge stays intact
ZTC zero-token confidence readout see below

🔁 RSI — a model that improves from its own work

Recursive self-improvement (RSI) is the core of this generation. Instead of distilling a bigger teacher, the model improves by learning from itself:

  1. Solve — the model works through practice problems it has never seen in evaluation.
  2. Verify — its answers are checked against verifiable references (answer keys, executable checks). Nothing unverified is learned.
  3. Learn — it is re-trained on the reasoning that turned out to be correct.
  4. Repeat — the improved model becomes the next solver.

What it bought in this release:

Parent (Qwen3.8-Flash-Next) Darwin-180B-RSI
Average reasoning length (MMLU-Pro) 4,320 tokens 3,833 tokens (−11 %)
MMLU-Pro accuracy 88.04 % 88.12 %

Same or better accuracy with shorter reasoning — cheaper and faster to serve. Practice sets are deduplicated against every evaluation set we report (8-gram overlap filter).


🏛️ ZTC — it knows before it answers

Zero-Token Confidence (ZTC) reads the model's own internal state once, before generation, and returns the probability that the answer it is about to give is correct — no extra tokens, no second model.

{"answer": "...", "confidence": 0.93, "ztc_score": 1.84, "truncated": false}

Use it to gate actions: when confidence is low, do not call the tool, escalate, or answer "I don't know". The ZTC readout for this model is being fitted and will ship in ztc/ (same format as Darwin-397B-ZTC).


🏆 Results

Benchmark Score Setting Leaderboard
GPQA Diamond (198) 94.44 majority vote over up to 16 samples · 131,072-token thinking budget #1
MMLU-Pro (12,032) 88.12 single sample · 131,072-token thinking budget leaderboard
MMMU-Pro (vision, 1,730) measuring majority vote leaderboard

Sampling for all runs: temperature 1.0 · top_p 0.95 · top_k 20 · bf16. All numbers are self-measured and reproducible with the settings above.

MMLU-Pro by category (single sample) — strongest in math 95.0 · biology 94.6 · physics 92.5; room to grow in law and history.


⚙️ Specifications

Architecture Mixture-of-Experts, hybrid attention (36 linear-attention + 12 full-attention layers)
Layers / hidden 48 / 2,560
Experts 512 routed (10 active per token) + shared expert
Context 262,144 tokens
Vocabulary 248,320
Modalities image + text → text
Precision bf16 (~336 GB)

🚀 Quickstart

Serving with vLLM (8 × B200 or equivalent)

vllm serve FINAL-Bench/Darwin-180B-RSI \
  --tensor-parallel-size 8 --enable-expert-parallel \
  --max-model-len 135168 --trust-remote-code

Chat Completions (OpenAI-compatible)

from openai import OpenAI
c = OpenAI(base_url="http://localhost:8000/v1", api_key="-")
r = c.chat.completions.create(model="FINAL-Bench/Darwin-180B-RSI",
    messages=[{"role": "user", "content": "Explain why the sky is blue in two sentences."}],
    temperature=1.0, top_p=0.95, extra_body={"top_k": 20})
print(r.choices[0].message.content)

Transformers

from transformers import AutoProcessor, AutoModelForImageTextToText
model_id = "FINAL-Bench/Darwin-180B-RSI"
processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForImageTextToText.from_pretrained(model_id, torch_dtype="auto", device_map="auto")

Tip: this is a thinking model. Give it room — a thinking budget of 32K–131K tokens is recommended for hard reasoning. Short budgets truncate the reasoning and cost accuracy.


📜 License

Darwin-180B-RSI is a derivative of Qwen3.8-Flash-Next and is distributed under the Qwen Community License 1.0 (see LICENSE).

🏢 About

Built by VIDRAFT · evaluated with FINAL-Bench.