Data-center AI, now on a laptop — POCKET-Darwin-180B
| Model | FINAL-Bench/POCKET-Darwin-180B-GGUF |
| Original | FINAL-Bench/Darwin-180B-RSI |
| China mirror | ModelScope FINAL-Bench/POCKET-Darwin-180B-GGUF |
| Base model | Qwen3.8-Flash-Next (Alibaba) |
| License | Qwen Community License 1.0 |
At a glance
- Size: 360 GB (BF16) → 111 GB (4-bit GGUF, 4 files)
- No GPU: one server CPU (16 threads) generates 18.4–21.0 tokens/s, peak memory 78.8 GB
- Laptop: RTX 5060 Laptop (8 GB VRAM) + 32 GB RAM — 4.17 tokens/s
- Mini PC: with 128 GB RAM the whole model fits in memory, no GPU required
- Accuracy: MMLU-Pro, 2,000 questions, paired per question — original 87.65% = 4-bit 87.65%
1. From an 8-GPU server to a laptop
| Darwin-180B-RSI original (BF16) | POCKET-Darwin-180B | |
|---|---|---|
| Model size | 360 GB (131 files) | 111 GB (4 files) |
| Hardware | 4–8× B200, or an 8× H100 (80 GB) server | Laptop with 8 GB GPU + 32 GB RAM · 128 GB mini PC · CPU-only server · one DGX Spark |
| Hardware cost (industry estimate) | 8× H100 server ≈ US$350K | ≈ US$1,500 gaming laptop |
| MMLU-Pro | 87.65% | 87.65% |
The original weights alone are 360 GB, so they need five or more 80 GB H100s just to load, and in practice a full 8-GPU server. POCKET brings a model of this class to a personal computer.
2. Only 3B of 180B parameters at a time — why it runs on a laptop
Darwin-180B is a Mixture-of-Experts (MoE) model. Of 512 experts, only 10 are used for each token, so only about 3 billion of its 180 billion parameters do work at any moment.
llama.cpp memory-maps the weights instead of loading them all, and reads the experts it needs straight from the SSD. That is how a 111 GB model runs on a laptop with 32 GB of RAM.
3. How we compressed it — graft quantization
POCKET is not re-quantized from scratch.
- We take the widely used Unsloth UD-Q4_K_XL build of the base model as a template.
- We replace only the 300 tensors that our self-improvement training actually changed, in the same layout (Q8_0).
- Every other byte is identical to the base build. All 300 replaced tensors were read back and verified (300/300).
Because only the changed parts are placed on top of a proven quantized build, quality carries over.
4. Quality before and after compression
MMLU-Pro, 2,000 questions (stratified by subject, fixed hash sample), thinking mode, temperature 1.0 · top-p 0.95 · top-k 20, up to 131K tokens.
| Version | Accuracy |
|---|---|
| Darwin-180B-RSI-R3 original (BF16) | 87.65% |
| POCKET-Darwin-180B (4-bit) | 87.65% |
Zero truncated answers. The 4-bit build scores the same as the original.
On the same 2,000 questions, answers are 14.5% shorter on average than the same-format build of the base model, so at the same speed they finish sooner.
5. Measured on real devices
| Device | Setup | Generation speed | How it runs |
|---|---|---|---|
| 💻 Gaming laptop | RTX 5060 Laptop (8 GB) · 32 GB RAM · NVMe SSD | 4.17 tok/s | Experts streamed from SSD on demand |
| 🖥️ CPU-only server | AMD EPYC, 1 socket · 16 threads · no GPU | 18.4–21.0 tok/s | Whole model in memory, peak 78.8 GB |
| 🧊 Mini PC | 128 GB RAM (e.g. Ryzen AI Max+ 395) | — | Whole model in memory, no GPU |
| 🟩 NVIDIA DGX Spark | 128 GB unified memory | ≈ 39 tok/s (same-size, same-format reference build) | Fully resident |
Recommended minimum: 32 GB RAM + 8 GB VRAM + 120 GB free NVMe. More RAM (96–128 GB) keeps more of the model in memory and runs faster.
6. The technology underneath — Model-level Recursive Self-Improvement
- Base: Qwen3.8-Flash-Next (180B-parameter MoE, 512 experts)
- The model solves verifiable problems itself, keeps only its own solutions that check out as correct, and trains on them. No human-written solutions or reasoning traces are used.
- Only the attention paths and shared experts are trained. The 512 routed experts, the router, the per-layer n-gram embeddings and the MTP head are left unchanged, preserving the base model's knowledge.
- The R3 round compressed here gained +1.03 points over the previous round on 1,000 SuperGPQA questions never used for training or selection (95% CI [+0.05, +2.00], statistically significant).
Original Darwin-180B-RSI on official Hugging Face leaderboards (self-reported, majority voting over multiple samples)
| Benchmark | Score | Rank |
|---|---|---|
| AIME 2026 | 100 | #1 |
| HMMT Feb 2026 | 100 | #1 |
| GPQA Diamond | 94.44 | #1 |
| MMLU-Pro | 88.12 | #1 |
| MMMU-Pro (vision) | 79.48 | #1 |
| LEXam (law) | 68.94 | #1 |
| LEXam-hard (law) | 45.72 | #1 |
7. Verified research institution on ModelScope
VIDRAFT's ModelScope organization (FINAL-Bench) holds an official Research Institution verification. Verified organizations are few — including Alibaba's Qwen, Shanghai AI Laboratory, Zhipu AI and OpenBMB — and among the Korean AI organizations we checked, VIDRAFT is the only one verified. POCKET-Darwin-180B is also published on ModelScope with a Chinese model card, so developers in China can download it directly.
8. Who it is for
Organizations that cannot send data to an external cloud — defense, finance and the public sector. A top-tier model can now run with no internet connection, on an in-house server or a single mini PC.
9. How to run it
Requires llama.cpp b11048 or newer (architecture qwen4exp).
# Laptop / desktop (8 GB+ VRAM, 32 GB+ RAM; experts read from SSD on demand)
llama-server -m POCKET-Darwin-180B-UD-Q4_K_XL-00001-of-00004.gguf -ngl 999 --cpu-moe -fa on -c 8192 --jinja
# CPU only, no GPU
llama-server -m POCKET-Darwin-180B-UD-Q4_K_XL-00001-of-00004.gguf -ngl 0 -t 16 -c 8192 --jinja --load-mode none
# DGX Spark / a single large-memory GPU
llama-server -m POCKET-Darwin-180B-UD-Q4_K_XL-00001-of-00004.gguf -ngl 999 -fa on -c 131072 --jinja
Recommended sampling: temperature 1.0 · top-p 0.95 · top-k 20. This is a reasoning model, so allow at least 2,048 output tokens. The reasoning is returned in reasoning_content.
VIDRAFT · vidraft.net · Hugging Face FINAL-Bench · ModelScope FINAL-Bench

