card: add Long context section (retrieval to 350K, TTFT/decode/memory tables, largest clean max_seq_len)
Browse files
README.md
CHANGED
|
@@ -77,6 +77,7 @@ eval/
|
|
| 77 |
bench/local-*.score.json local-*.meta.json this quant's per-task scores + run metadata
|
| 78 |
bench/ref-*.score.json ref-*.meta.json the unquantized FP8 reference, same harness
|
| 79 |
eval-quant.json eval-overflow-64rows.json dflash-bench-*.json
|
|
|
|
| 80 |
```
|
| 81 |
|
| 82 |
The shards, index, `quantization_config.json` and the tokenizer files are byte-identical to
|
|
@@ -361,6 +362,10 @@ With 86.20 GiB of weights and a 10 GiB reserve, the KV budget is 25.43 GiB:
|
|
| 361 |
and generates coherently on top of these 2.36 bpw weights, but its quality is untested β the
|
| 362 |
default stays FP16. The 39 sliding-window rings are always FP16.
|
| 363 |
|
|
|
|
|
|
|
|
|
|
|
|
|
| 364 |
#### A unified-memory hazard worth knowing about
|
| 365 |
|
| 366 |
On GB10 there is **no separate VRAM**: host and GPU share one pool. exllamav3's shard loader
|
|
@@ -378,6 +383,169 @@ why the numbers above are flat rather than degrading:
|
|
| 378 |
|
| 379 |
---
|
| 380 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 381 |
## How to run
|
| 382 |
|
| 383 |
### Requirements
|
|
@@ -430,8 +598,8 @@ models/
|
|
| 430 |
model:
|
| 431 |
model_dir: models
|
| 432 |
model_name: mimo-2.25bpw-hq
|
| 433 |
-
max_seq_len:
|
| 434 |
-
cache_size:
|
| 435 |
cache_mode: FP16 # default; Q8 paged KV works but is unvalidated for quality
|
| 436 |
chunk_size: 2048
|
| 437 |
max_batch_size: 1 # each extra slot costs another 175.5 MiB ring + paged span + draft cache
|
|
@@ -550,8 +718,11 @@ The drafter loads BF16 in 1.3 s / 2.81 GiB β no quantization needed. Its own K
|
|
| 550 |
perfect 2.34 bpw quant β and the benchmark deltas above say the FP8 model is only ~0.6 pp
|
| 551 |
better at HumanEval+, so most of that is the base model, not the quantization.
|
| 552 |
* **Tensor parallel is unsupported.** The port has only been built and run single-device.
|
| 553 |
-
* **Long context
|
| 554 |
-
|
|
|
|
|
|
|
|
|
|
| 555 |
* **Batch > 1 with DFlash is untested.** Everything above is `max_batch_size: 1`.
|
| 556 |
* **No vision, no audio, no MTP.** The MTP heads in the source repo are not ported, so
|
| 557 |
MTP-based drafting is unavailable; DFlash is the drafting path.
|
|
|
|
| 77 |
bench/local-*.score.json local-*.meta.json this quant's per-task scores + run metadata
|
| 78 |
bench/ref-*.score.json ref-*.meta.json the unquantized FP8 reference, same harness
|
| 79 |
eval-quant.json eval-overflow-64rows.json dflash-bench-*.json
|
| 80 |
+
longctx.json the long-context sweep, every request (see Long context)
|
| 81 |
```
|
| 82 |
|
| 83 |
The shards, index, `quantization_config.json` and the tokenizer files are byte-identical to
|
|
|
|
| 362 |
and generates coherently on top of these 2.36 bpw weights, but its quality is untested β the
|
| 363 |
default stays FP16. The 39 sliding-window rings are always FP16.
|
| 364 |
|
| 365 |
+
Those are the arithmetic limits. What was actually served, loaded and measured β up to
|
| 366 |
+
**350,091 tokens** β is in [Long context](#long-context) below; the cache is **preallocated in
|
| 367 |
+
full at load**, so `max_seq_len` is a memory decision, not a ceiling you pay for on demand.
|
| 368 |
+
|
| 369 |
#### A unified-memory hazard worth knowing about
|
| 370 |
|
| 371 |
On GB10 there is **no separate VRAM**: host and GPU share one pool. exllamav3's shard loader
|
|
|
|
| 383 |
|
| 384 |
---
|
| 385 |
|
| 386 |
+
## Long context
|
| 387 |
+
|
| 388 |
+
Measured on the DGX Spark, batch 1, greedy, thinking off, FP16 KV, needles and haystacks built
|
| 389 |
+
from wikitext-2 paragraphs. **Every retrieval test passed.**
|
| 390 |
+
|
| 391 |
+
### Retrieval
|
| 392 |
+
|
| 393 |
+
A six-digit passcode is hidden in a wikitext haystack at 10 / 50 / 90% depth and asked for at
|
| 394 |
+
the end. Exact match on the digits, fresh city and fresh code per request.
|
| 395 |
+
|
| 396 |
+
| prompt tokens | 10% | 50% | 90% |
|
| 397 |
+
|---|---|---|---|
|
| 398 |
+
| ~8.2K | OK | OK | OK |
|
| 399 |
+
| ~32.8K | OK | OK | OK |
|
| 400 |
+
| ~65.5K | OK | OK | OK |
|
| 401 |
+
| ~131.1K | OK | OK | OK |
|
| 402 |
+
| ~200.1K | OK | OK | OK |
|
| 403 |
+
| ~250.0K | OK | OK | OK |
|
| 404 |
+
| **350,091** | **OK** | β | β |
|
| 405 |
+
|
| 406 |
+
**18/18** at `max_seq_len 262144`, plus **350,091 tokens** answered correctly in a 393,216
|
| 407 |
+
window β the largest prompt this quant has been given. Nine more cells at `max_seq_len 131072`
|
| 408 |
+
with the drafter on also passed, so retrieval is unaffected by speculative decoding.
|
| 409 |
+
|
| 410 |
+
That matters more than it looks: **39 of 48 layers are sliding-window with a 128-token
|
| 411 |
+
window** and run on a fixed 768-token ring, so a needle 25,000 tokens into a 250K prompt has
|
| 412 |
+
to survive ~195 ring rebases and reach the question through the 9 global-attention layers
|
| 413 |
+
alone. The ring had previously only been validated to 2,048 tokens.
|
| 414 |
+
|
| 415 |
+
**Five needles at once, 128K prompt** β all five returned, in order of appearance:
|
| 416 |
+
|
| 417 |
+
```
|
| 418 |
+
Montevideo: 604403
|
| 419 |
+
Bratislava: 991476
|
| 420 |
+
Kathmandu: 472495
|
| 421 |
+
Ulaanbaatar: 379397
|
| 422 |
+
Ljubljana: 361254
|
| 423 |
+
```
|
| 424 |
+
|
| 425 |
+
**A real long document, ~100K tokens** β complete wikitext articles concatenated with *Plain
|
| 426 |
+
maskray* buried in the middle, three questions answerable only from that article:
|
| 427 |
+
|
| 428 |
+
```
|
| 429 |
+
(a) 2008
|
| 430 |
+
(b) ~54 Ma
|
| 431 |
+
(c) 12 and 62 m
|
| 432 |
+
```
|
| 433 |
+
|
| 434 |
+
All three correct against the source ("In 2008, Last and William White elevated the kuhlii
|
| 435 |
+
groupβ¦", "estimated to have occurred ~ 54 Ma", "recorded from between 12 and 62 m"), including
|
| 436 |
+
the source's own "~" hedge.
|
| 437 |
+
|
| 438 |
+
### Speed and memory vs prompt length
|
| 439 |
+
|
| 440 |
+
No drafter, 128 generated tokens, `max_seq_len 262144` (last row 393,216):
|
| 441 |
+
|
| 442 |
+
| prompt | TTFT | prefill | decode | min MemAvailable |
|
| 443 |
+
|---|---|---|---|---|
|
| 444 |
+
| ~8.2K | 9.1β10.6 s | 768β905 tok/s | 31.0 tok/s | 19.85 GiB |
|
| 445 |
+
| ~32.8K | 41.2β42.0 s | 780β799 tok/s | 28.4 tok/s | 19.56 GiB |
|
| 446 |
+
| ~65.5K | 100.1β100.5 s | 652β654 tok/s | 25.3 tok/s | 19.10 GiB |
|
| 447 |
+
| ~131.1K | 270.0β271.5 s | 483β486 tok/s | 21.7 tok/s | 18.98 GiB |
|
| 448 |
+
| ~200.1K | 524.8β525.9 s | 380β381 tok/s | 18.3 tok/s | 18.85 GiB |
|
| 449 |
+
| ~250.0K | 757.6β758.8 s | 329β330 tok/s | 16.7 tok/s | 17.77 GiB |
|
| 450 |
+
| **350,091** | **1,346.8 s** | **260 tok/s** | **13.8 tok/s** | **15.91 GiB** |
|
| 451 |
+
|
| 452 |
+
**Decode falls only 2.3x from 8K to 250K** β the sliding-window ring again: only 9 of 48
|
| 453 |
+
layers grow their KV with context. Prefill falls 2.7x and TTFT is quadratic, as full
|
| 454 |
+
attention on those 9 layers requires. The curve fits
|
| 455 |
+
|
| 456 |
+
```
|
| 457 |
+
TTFT(seconds) β 988Β·L + 8172Β·LΒ² (L = prompt tokens in millions)
|
| 458 |
+
```
|
| 459 |
+
|
| 460 |
+
to better than 1.5% from 64K to 350K β the 350K point was predicted at 1,347 s from a fit to
|
| 461 |
+
the 128K and 250K points alone and came in at 1,346.8 s. Extrapolated: **400K β 28 min,
|
| 462 |
+
500K β 42 min of prefill.** Long context on one GB10 is TTFT-bound, not memory-bound.
|
| 463 |
+
|
| 464 |
+
With the DFlash drafter at `max_seq_len 131072` (real generations, ~380 tokens, unforced):
|
| 465 |
+
|
| 466 |
+
| prompt | TTFT | decode drafted | vs no draft |
|
| 467 |
+
|---|---|---|---|
|
| 468 |
+
| 1,983 | 2.97 s | **36.3 tok/s** | 1.17x |
|
| 469 |
+
| 65,544 | 103.1 s | **29.8 tok/s** | 1.18x |
|
| 470 |
+
| 129,876 | 171.9 s | **27.0 tok/s** | 1.24x |
|
| 471 |
+
|
| 472 |
+
Live acceptance 38β85%. (Drafting speedup on long-context Q&A is smaller than the 1.35β1.78x
|
| 473 |
+
seen on short coding/reasoning prompts, because the answers are short and factual.)
|
| 474 |
+
|
| 475 |
+
### How much context fits
|
| 476 |
+
|
| 477 |
+
The paged cache is **preallocated in full at load** β `cache_size` is paid up front whether or
|
| 478 |
+
not anyone sends a long prompt. Cost per slot:
|
| 479 |
+
|
| 480 |
+
```
|
| 481 |
+
paged KV 27.00 KiB/token (the 9 global-attention layers only)
|
| 482 |
+
SWA ring 175.5 MiB per slot (768-token ring x 39 layers, always FP16)
|
| 483 |
+
draft KV 20.00 KiB/token (5 DFlash layers, sized for the full context)
|
| 484 |
+
drafter weights 2.81 GiB
|
| 485 |
+
```
|
| 486 |
+
|
| 487 |
+
so a slot costs **27.0 KiB/token** without the drafter and **47.0 KiB/token** with it. On this
|
| 488 |
+
box the whole budget reduces to one line that held to within 0.1 GiB across three configurations:
|
| 489 |
+
|
| 490 |
+
```
|
| 491 |
+
MemAvailable after load β 29.2 GiB β (paged KV + ring + drafter weights + draft KV)
|
| 492 |
+
```
|
| 493 |
+
|
| 494 |
+
(121.63 GiB total β 86.15 GiB of weights β ~6 GiB of runtime.)
|
| 495 |
+
|
| 496 |
+
| `max_seq_len` | drafter | preallocated | load | MemAvailable after load | min under load |
|
| 497 |
+
|---|---|---|---|---|---|
|
| 498 |
+
| 393,216 | off | 10.29 GiB | 21.3 s | 18.65 GiB | **15.91 GiB** @ 350K |
|
| 499 |
+
| 262,144 | off | 6.92 GiB | 21.6 s | 22.30 GiB | **17.77 GiB** @ 250K |
|
| 500 |
+
| **131,072** | **DFlash** | 6.13 + 2.81 GiB | 22.6 s | 20.17 GiB | **17.07 GiB** @ 130K |
|
| 501 |
+
|
| 502 |
+
**Largest clean settings: `max_seq_len 393216` without the drafter, `131072` with it.**
|
| 503 |
+
|
| 504 |
+
* **524,288 does not fit at FP16.** Its 13.67 GiB of paged KV leaves ~15.3 GiB after load, and
|
| 505 |
+
the measured 2.7β4.5 GiB transient of a full-length request would land it at ~11β12.5 GiB β
|
| 506 |
+
at or through a 12 GiB safety floor. On unified memory that is not an OOM kill, it is a
|
| 507 |
+
frozen machine. A 500K prefill would also take ~42 minutes.
|
| 508 |
+
* **262,144 with the drafter does not fit either**: 14.73 GiB resident leaves ~14.5 GiB after
|
| 509 |
+
load and ~10 GiB under a full-length request. The drafter's cache is the problem β it is
|
| 510 |
+
sized for the whole context although it never looks back more than 1,024 tokens.
|
| 511 |
+
|
| 512 |
+
So the choice on a 121 GiB box is **3x the window (393K, no drafter)** or **~1.2x the decode
|
| 513 |
+
rate (131K, drafter)**. This repo's reference server runs the latter.
|
| 514 |
+
|
| 515 |
+
### Recommended serving settings
|
| 516 |
+
|
| 517 |
+
```yaml
|
| 518 |
+
model:
|
| 519 |
+
max_seq_len: 131072 # with the drafter; use 393216 and drop draft_model to go longer
|
| 520 |
+
cache_size: 131072
|
| 521 |
+
cache_mode: FP16
|
| 522 |
+
chunk_size: 2048
|
| 523 |
+
max_batch_size: 1 # every extra slot repeats the ring + paged span + draft cache
|
| 524 |
+
draft_model:
|
| 525 |
+
draft_mode: model
|
| 526 |
+
draft_model_name: mimo-dflash-draft
|
| 527 |
+
draft_cache_mode: FP16
|
| 528 |
+
dynamic_draft: true
|
| 529 |
+
```
|
| 530 |
+
|
| 531 |
+
### Caveats
|
| 532 |
+
|
| 533 |
+
* **`max_seq_len` bounds the prompt, not prompt + generation.** A 131,092-token prompt against
|
| 534 |
+
`max_seq_len 131072` returns `400 Bad Request` / `Prompt length 131092 exceeds the β¦` before
|
| 535 |
+
a token is generated. Budget `max_seq_len β₯ prompt + max_tokens`, and remember the chat
|
| 536 |
+
template adds ~20 tokens.
|
| 537 |
+
* **Every number here is `max_batch_size: 1`.** A second slot adds another 175.5 MiB ring, its
|
| 538 |
+
own paged span and its own draft cache; concurrency at 128K is not free.
|
| 539 |
+
* Prefill is chunked at 2,048 tokens. Prefill rates for a prompt sharing a long prefix with the
|
| 540 |
+
previous request are inflated by page reuse (755 vs 473 tok/s at 130K here), so cold numbers
|
| 541 |
+
are the ones quoted above.
|
| 542 |
+
* Q8 paged KV halves the 27 KiB/token term and would make ~786K arithmetically fit, but
|
| 543 |
+
quantized KV quality on top of 2.34 bpw weights is untested and the prefill time would be
|
| 544 |
+
hours. The 39 sliding-window rings stay FP16 regardless.
|
| 545 |
+
* Raw records: [`eval/longctx.json`](eval/longctx.json).
|
| 546 |
+
|
| 547 |
+
---
|
| 548 |
+
|
| 549 |
## How to run
|
| 550 |
|
| 551 |
### Requirements
|
|
|
|
| 598 |
model:
|
| 599 |
model_dir: models
|
| 600 |
model_name: mimo-2.25bpw-hq
|
| 601 |
+
max_seq_len: 131072 # validated; 393216 fits if you drop the draft model
|
| 602 |
+
cache_size: 131072
|
| 603 |
cache_mode: FP16 # default; Q8 paged KV works but is unvalidated for quality
|
| 604 |
chunk_size: 2048
|
| 605 |
max_batch_size: 1 # each extra slot costs another 175.5 MiB ring + paged span + draft cache
|
|
|
|
| 718 |
perfect 2.34 bpw quant β and the benchmark deltas above say the FP8 model is only ~0.6 pp
|
| 719 |
better at HumanEval+, so most of that is the base model, not the quantization.
|
| 720 |
* **Tensor parallel is unsupported.** The port has only been built and run single-device.
|
| 721 |
+
* **Long context is now measured, and it is TTFT-bound rather than memory-bound.** 8K β 350K
|
| 722 |
+
has been served and retrieval is perfect at every length and depth tested, but a 250K prompt
|
| 723 |
+
costs ~12.6 minutes of prefill and a 350K prompt ~22.4 minutes on one GB10. `max_seq_len`
|
| 724 |
+
above 393,216 (or above 131,072 with the drafter) does not fit at FP16 KV. See
|
| 725 |
+
[Long context](#long-context).
|
| 726 |
* **Batch > 1 with DFlash is untested.** Everything above is `max_batch_size: 1`.
|
| 727 |
* **No vision, no audio, no MTP.** The MTP heads in the source repo are not ported, so
|
| 728 |
MTP-based drafting is unavailable; DFlash is the drafting path.
|