Data-center AI, now on a laptop — POCKET-Darwin-180B

Community Article
Published October 2, 2026

We are releasing a 4-bit GGUF build of Darwin-180B-RSI — #1 on seven official Hugging Face leaderboards (self-reported) — that runs without a GPU.

Model FINAL-Bench/POCKET-Darwin-180B-GGUF
Original FINAL-Bench/Darwin-180B-RSI
China mirror ModelScope FINAL-Bench/POCKET-Darwin-180B-GGUF
Base model Qwen3.8-Flash-Next (Alibaba)
License Qwen Community License 1.0

At a glance

  • Size: 360 GB (BF16) → 111 GB (4-bit GGUF, 4 files)
  • No GPU: one server CPU (16 threads) generates 18.4–21.0 tokens/s, peak memory 78.8 GB
  • Laptop: RTX 5060 Laptop (8 GB VRAM) + 32 GB RAM — 4.17 tokens/s
  • Mini PC: with 128 GB RAM the whole model fits in memory, no GPU required
  • Accuracy: MMLU-Pro, 2,000 questions, paired per question — original 87.65% = 4-bit 87.65%

1. From an 8-GPU server to a laptop

Darwin-180B-RSI original (BF16) POCKET-Darwin-180B
Model size 360 GB (131 files) 111 GB (4 files)
Hardware 4–8× B200, or an 8× H100 (80 GB) server Laptop with 8 GB GPU + 32 GB RAM · 128 GB mini PC · CPU-only server · one DGX Spark
Hardware cost (industry estimate) 8× H100 server ≈ US$350K ≈ US$1,500 gaming laptop
MMLU-Pro 87.65% 87.65%

The original weights alone are 360 GB, so they need five or more 80 GB H100s just to load, and in practice a full 8-GPU server. POCKET brings a model of this class to a personal computer.

2. Only 3B of 180B parameters at a time — why it runs on a laptop

Darwin-180B is a Mixture-of-Experts (MoE) model. Of 512 experts, only 10 are used for each token, so only about 3 billion of its 180 billion parameters do work at any moment.

llama.cpp memory-maps the weights instead of loading them all, and reads the experts it needs straight from the SSD. That is how a 111 GB model runs on a laptop with 32 GB of RAM.

3. How we compressed it — graft quantization

POCKET is not re-quantized from scratch.

  1. We take the widely used Unsloth UD-Q4_K_XL build of the base model as a template.
  2. We replace only the 300 tensors that our self-improvement training actually changed, in the same layout (Q8_0).
  3. Every other byte is identical to the base build. All 300 replaced tensors were read back and verified (300/300).

Because only the changed parts are placed on top of a proven quantized build, quality carries over.

4. Quality before and after compression

MMLU-Pro, 2,000 questions (stratified by subject, fixed hash sample), thinking mode, temperature 1.0 · top-p 0.95 · top-k 20, up to 131K tokens.

Version Accuracy
Darwin-180B-RSI-R3 original (BF16) 87.65%
POCKET-Darwin-180B (4-bit) 87.65%

Zero truncated answers. The 4-bit build scores the same as the original.

On the same 2,000 questions, answers are 14.5% shorter on average than the same-format build of the base model, so at the same speed they finish sooner.

5. Measured on real devices

Device Setup Generation speed How it runs
💻 Gaming laptop RTX 5060 Laptop (8 GB) · 32 GB RAM · NVMe SSD 4.17 tok/s Experts streamed from SSD on demand
🖥️ CPU-only server AMD EPYC, 1 socket · 16 threads · no GPU 18.4–21.0 tok/s Whole model in memory, peak 78.8 GB
🧊 Mini PC 128 GB RAM (e.g. Ryzen AI Max+ 395) — Whole model in memory, no GPU
🟩 NVIDIA DGX Spark 128 GB unified memory ≈ 39 tok/s (same-size, same-format reference build) Fully resident

Recommended minimum: 32 GB RAM + 8 GB VRAM + 120 GB free NVMe. More RAM (96–128 GB) keeps more of the model in memory and runs faster.

6. The technology underneath — Model-level Recursive Self-Improvement

  • Base: Qwen3.8-Flash-Next (180B-parameter MoE, 512 experts)
  • The model solves verifiable problems itself, keeps only its own solutions that check out as correct, and trains on them. No human-written solutions or reasoning traces are used.
  • Only the attention paths and shared experts are trained. The 512 routed experts, the router, the per-layer n-gram embeddings and the MTP head are left unchanged, preserving the base model's knowledge.
  • The R3 round compressed here gained +1.03 points over the previous round on 1,000 SuperGPQA questions never used for training or selection (95% CI [+0.05, +2.00], statistically significant).

Original Darwin-180B-RSI on official Hugging Face leaderboards (self-reported, majority voting over multiple samples)

Benchmark Score Rank
AIME 2026 100 #1
HMMT Feb 2026 100 #1
GPQA Diamond 94.44 #1
MMLU-Pro 88.12 #1
MMMU-Pro (vision) 79.48 #1
LEXam (law) 68.94 #1
LEXam-hard (law) 45.72 #1

7. Verified research institution on ModelScope

VIDRAFT's ModelScope organization (FINAL-Bench) holds an official Research Institution verification. Verified organizations are few — including Alibaba's Qwen, Shanghai AI Laboratory, Zhipu AI and OpenBMB — and among the Korean AI organizations we checked, VIDRAFT is the only one verified. POCKET-Darwin-180B is also published on ModelScope with a Chinese model card, so developers in China can download it directly.

8. Who it is for

Organizations that cannot send data to an external cloud — defense, finance and the public sector. A top-tier model can now run with no internet connection, on an in-house server or a single mini PC.

9. How to run it

Requires llama.cpp b11048 or newer (architecture qwen4exp).

# Laptop / desktop (8 GB+ VRAM, 32 GB+ RAM; experts read from SSD on demand)
llama-server -m POCKET-Darwin-180B-UD-Q4_K_XL-00001-of-00004.gguf -ngl 999 --cpu-moe -fa on -c 8192 --jinja

# CPU only, no GPU
llama-server -m POCKET-Darwin-180B-UD-Q4_K_XL-00001-of-00004.gguf -ngl 0 -t 16 -c 8192 --jinja --load-mode none

# DGX Spark / a single large-memory GPU
llama-server -m POCKET-Darwin-180B-UD-Q4_K_XL-00001-of-00004.gguf -ngl 999 -fa on -c 131072 --jinja

Recommended sampling: temperature 1.0 · top-p 0.95 · top-k 20. This is a reasoning model, so allow at least 2,048 output tokens. The reasoning is returned in reasoning_content.


VIDRAFT · vidraft.net · Hugging Face FINAL-Bench · ModelScope FINAL-Bench

Community

Sign up or log in to comment