benthecarman commited on
Commit
0f8d7e0
Β·
verified Β·
1 Parent(s): 570d176

card: add Long context section (retrieval to 350K, TTFT/decode/memory tables, largest clean max_seq_len)

Browse files
Files changed (1) hide show
  1. README.md +175 -4
README.md CHANGED
@@ -77,6 +77,7 @@ eval/
77
  bench/local-*.score.json local-*.meta.json this quant's per-task scores + run metadata
78
  bench/ref-*.score.json ref-*.meta.json the unquantized FP8 reference, same harness
79
  eval-quant.json eval-overflow-64rows.json dflash-bench-*.json
 
80
  ```
81
 
82
  The shards, index, `quantization_config.json` and the tokenizer files are byte-identical to
@@ -361,6 +362,10 @@ With 86.20 GiB of weights and a 10 GiB reserve, the KV budget is 25.43 GiB:
361
  and generates coherently on top of these 2.36 bpw weights, but its quality is untested β€” the
362
  default stays FP16. The 39 sliding-window rings are always FP16.
363
 
 
 
 
 
364
  #### A unified-memory hazard worth knowing about
365
 
366
  On GB10 there is **no separate VRAM**: host and GPU share one pool. exllamav3's shard loader
@@ -378,6 +383,169 @@ why the numbers above are flat rather than degrading:
378
 
379
  ---
380
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
381
  ## How to run
382
 
383
  ### Requirements
@@ -430,8 +598,8 @@ models/
430
  model:
431
  model_dir: models
432
  model_name: mimo-2.25bpw-hq
433
- max_seq_len: 65536
434
- cache_size: 65536
435
  cache_mode: FP16 # default; Q8 paged KV works but is unvalidated for quality
436
  chunk_size: 2048
437
  max_batch_size: 1 # each extra slot costs another 175.5 MiB ring + paged span + draft cache
@@ -550,8 +718,11 @@ The drafter loads BF16 in 1.3 s / 2.81 GiB β€” no quantization needed. Its own K
550
  perfect 2.34 bpw quant β€” and the benchmark deltas above say the FP8 model is only ~0.6 pp
551
  better at HumanEval+, so most of that is the base model, not the quantization.
552
  * **Tensor parallel is unsupported.** The port has only been built and run single-device.
553
- * **Long context beyond 64K is untested at serving time.** The KV arithmetic says ~980K
554
- tokens fit at 1 slot, and 32K was measured, but nothing between 32K and 1M has been served.
 
 
 
555
  * **Batch > 1 with DFlash is untested.** Everything above is `max_batch_size: 1`.
556
  * **No vision, no audio, no MTP.** The MTP heads in the source repo are not ported, so
557
  MTP-based drafting is unavailable; DFlash is the drafting path.
 
77
  bench/local-*.score.json local-*.meta.json this quant's per-task scores + run metadata
78
  bench/ref-*.score.json ref-*.meta.json the unquantized FP8 reference, same harness
79
  eval-quant.json eval-overflow-64rows.json dflash-bench-*.json
80
+ longctx.json the long-context sweep, every request (see Long context)
81
  ```
82
 
83
  The shards, index, `quantization_config.json` and the tokenizer files are byte-identical to
 
362
  and generates coherently on top of these 2.36 bpw weights, but its quality is untested β€” the
363
  default stays FP16. The 39 sliding-window rings are always FP16.
364
 
365
+ Those are the arithmetic limits. What was actually served, loaded and measured β€” up to
366
+ **350,091 tokens** β€” is in [Long context](#long-context) below; the cache is **preallocated in
367
+ full at load**, so `max_seq_len` is a memory decision, not a ceiling you pay for on demand.
368
+
369
  #### A unified-memory hazard worth knowing about
370
 
371
  On GB10 there is **no separate VRAM**: host and GPU share one pool. exllamav3's shard loader
 
383
 
384
  ---
385
 
386
+ ## Long context
387
+
388
+ Measured on the DGX Spark, batch 1, greedy, thinking off, FP16 KV, needles and haystacks built
389
+ from wikitext-2 paragraphs. **Every retrieval test passed.**
390
+
391
+ ### Retrieval
392
+
393
+ A six-digit passcode is hidden in a wikitext haystack at 10 / 50 / 90% depth and asked for at
394
+ the end. Exact match on the digits, fresh city and fresh code per request.
395
+
396
+ | prompt tokens | 10% | 50% | 90% |
397
+ |---|---|---|---|
398
+ | ~8.2K | OK | OK | OK |
399
+ | ~32.8K | OK | OK | OK |
400
+ | ~65.5K | OK | OK | OK |
401
+ | ~131.1K | OK | OK | OK |
402
+ | ~200.1K | OK | OK | OK |
403
+ | ~250.0K | OK | OK | OK |
404
+ | **350,091** | **OK** | β€” | β€” |
405
+
406
+ **18/18** at `max_seq_len 262144`, plus **350,091 tokens** answered correctly in a 393,216
407
+ window β€” the largest prompt this quant has been given. Nine more cells at `max_seq_len 131072`
408
+ with the drafter on also passed, so retrieval is unaffected by speculative decoding.
409
+
410
+ That matters more than it looks: **39 of 48 layers are sliding-window with a 128-token
411
+ window** and run on a fixed 768-token ring, so a needle 25,000 tokens into a 250K prompt has
412
+ to survive ~195 ring rebases and reach the question through the 9 global-attention layers
413
+ alone. The ring had previously only been validated to 2,048 tokens.
414
+
415
+ **Five needles at once, 128K prompt** β€” all five returned, in order of appearance:
416
+
417
+ ```
418
+ Montevideo: 604403
419
+ Bratislava: 991476
420
+ Kathmandu: 472495
421
+ Ulaanbaatar: 379397
422
+ Ljubljana: 361254
423
+ ```
424
+
425
+ **A real long document, ~100K tokens** β€” complete wikitext articles concatenated with *Plain
426
+ maskray* buried in the middle, three questions answerable only from that article:
427
+
428
+ ```
429
+ (a) 2008
430
+ (b) ~54 Ma
431
+ (c) 12 and 62 m
432
+ ```
433
+
434
+ All three correct against the source ("In 2008, Last and William White elevated the kuhlii
435
+ group…", "estimated to have occurred ~ 54 Ma", "recorded from between 12 and 62 m"), including
436
+ the source's own "~" hedge.
437
+
438
+ ### Speed and memory vs prompt length
439
+
440
+ No drafter, 128 generated tokens, `max_seq_len 262144` (last row 393,216):
441
+
442
+ | prompt | TTFT | prefill | decode | min MemAvailable |
443
+ |---|---|---|---|---|
444
+ | ~8.2K | 9.1–10.6 s | 768–905 tok/s | 31.0 tok/s | 19.85 GiB |
445
+ | ~32.8K | 41.2–42.0 s | 780–799 tok/s | 28.4 tok/s | 19.56 GiB |
446
+ | ~65.5K | 100.1–100.5 s | 652–654 tok/s | 25.3 tok/s | 19.10 GiB |
447
+ | ~131.1K | 270.0–271.5 s | 483–486 tok/s | 21.7 tok/s | 18.98 GiB |
448
+ | ~200.1K | 524.8–525.9 s | 380–381 tok/s | 18.3 tok/s | 18.85 GiB |
449
+ | ~250.0K | 757.6–758.8 s | 329–330 tok/s | 16.7 tok/s | 17.77 GiB |
450
+ | **350,091** | **1,346.8 s** | **260 tok/s** | **13.8 tok/s** | **15.91 GiB** |
451
+
452
+ **Decode falls only 2.3x from 8K to 250K** β€” the sliding-window ring again: only 9 of 48
453
+ layers grow their KV with context. Prefill falls 2.7x and TTFT is quadratic, as full
454
+ attention on those 9 layers requires. The curve fits
455
+
456
+ ```
457
+ TTFT(seconds) β‰ˆ 988Β·L + 8172Β·LΒ² (L = prompt tokens in millions)
458
+ ```
459
+
460
+ to better than 1.5% from 64K to 350K β€” the 350K point was predicted at 1,347 s from a fit to
461
+ the 128K and 250K points alone and came in at 1,346.8 s. Extrapolated: **400K β‰ˆ 28 min,
462
+ 500K β‰ˆ 42 min of prefill.** Long context on one GB10 is TTFT-bound, not memory-bound.
463
+
464
+ With the DFlash drafter at `max_seq_len 131072` (real generations, ~380 tokens, unforced):
465
+
466
+ | prompt | TTFT | decode drafted | vs no draft |
467
+ |---|---|---|---|
468
+ | 1,983 | 2.97 s | **36.3 tok/s** | 1.17x |
469
+ | 65,544 | 103.1 s | **29.8 tok/s** | 1.18x |
470
+ | 129,876 | 171.9 s | **27.0 tok/s** | 1.24x |
471
+
472
+ Live acceptance 38–85%. (Drafting speedup on long-context Q&A is smaller than the 1.35–1.78x
473
+ seen on short coding/reasoning prompts, because the answers are short and factual.)
474
+
475
+ ### How much context fits
476
+
477
+ The paged cache is **preallocated in full at load** β€” `cache_size` is paid up front whether or
478
+ not anyone sends a long prompt. Cost per slot:
479
+
480
+ ```
481
+ paged KV 27.00 KiB/token (the 9 global-attention layers only)
482
+ SWA ring 175.5 MiB per slot (768-token ring x 39 layers, always FP16)
483
+ draft KV 20.00 KiB/token (5 DFlash layers, sized for the full context)
484
+ drafter weights 2.81 GiB
485
+ ```
486
+
487
+ so a slot costs **27.0 KiB/token** without the drafter and **47.0 KiB/token** with it. On this
488
+ box the whole budget reduces to one line that held to within 0.1 GiB across three configurations:
489
+
490
+ ```
491
+ MemAvailable after load β‰ˆ 29.2 GiB βˆ’ (paged KV + ring + drafter weights + draft KV)
492
+ ```
493
+
494
+ (121.63 GiB total βˆ’ 86.15 GiB of weights βˆ’ ~6 GiB of runtime.)
495
+
496
+ | `max_seq_len` | drafter | preallocated | load | MemAvailable after load | min under load |
497
+ |---|---|---|---|---|---|
498
+ | 393,216 | off | 10.29 GiB | 21.3 s | 18.65 GiB | **15.91 GiB** @ 350K |
499
+ | 262,144 | off | 6.92 GiB | 21.6 s | 22.30 GiB | **17.77 GiB** @ 250K |
500
+ | **131,072** | **DFlash** | 6.13 + 2.81 GiB | 22.6 s | 20.17 GiB | **17.07 GiB** @ 130K |
501
+
502
+ **Largest clean settings: `max_seq_len 393216` without the drafter, `131072` with it.**
503
+
504
+ * **524,288 does not fit at FP16.** Its 13.67 GiB of paged KV leaves ~15.3 GiB after load, and
505
+ the measured 2.7–4.5 GiB transient of a full-length request would land it at ~11–12.5 GiB β€”
506
+ at or through a 12 GiB safety floor. On unified memory that is not an OOM kill, it is a
507
+ frozen machine. A 500K prefill would also take ~42 minutes.
508
+ * **262,144 with the drafter does not fit either**: 14.73 GiB resident leaves ~14.5 GiB after
509
+ load and ~10 GiB under a full-length request. The drafter's cache is the problem β€” it is
510
+ sized for the whole context although it never looks back more than 1,024 tokens.
511
+
512
+ So the choice on a 121 GiB box is **3x the window (393K, no drafter)** or **~1.2x the decode
513
+ rate (131K, drafter)**. This repo's reference server runs the latter.
514
+
515
+ ### Recommended serving settings
516
+
517
+ ```yaml
518
+ model:
519
+ max_seq_len: 131072 # with the drafter; use 393216 and drop draft_model to go longer
520
+ cache_size: 131072
521
+ cache_mode: FP16
522
+ chunk_size: 2048
523
+ max_batch_size: 1 # every extra slot repeats the ring + paged span + draft cache
524
+ draft_model:
525
+ draft_mode: model
526
+ draft_model_name: mimo-dflash-draft
527
+ draft_cache_mode: FP16
528
+ dynamic_draft: true
529
+ ```
530
+
531
+ ### Caveats
532
+
533
+ * **`max_seq_len` bounds the prompt, not prompt + generation.** A 131,092-token prompt against
534
+ `max_seq_len 131072` returns `400 Bad Request` / `Prompt length 131092 exceeds the …` before
535
+ a token is generated. Budget `max_seq_len β‰₯ prompt + max_tokens`, and remember the chat
536
+ template adds ~20 tokens.
537
+ * **Every number here is `max_batch_size: 1`.** A second slot adds another 175.5 MiB ring, its
538
+ own paged span and its own draft cache; concurrency at 128K is not free.
539
+ * Prefill is chunked at 2,048 tokens. Prefill rates for a prompt sharing a long prefix with the
540
+ previous request are inflated by page reuse (755 vs 473 tok/s at 130K here), so cold numbers
541
+ are the ones quoted above.
542
+ * Q8 paged KV halves the 27 KiB/token term and would make ~786K arithmetically fit, but
543
+ quantized KV quality on top of 2.34 bpw weights is untested and the prefill time would be
544
+ hours. The 39 sliding-window rings stay FP16 regardless.
545
+ * Raw records: [`eval/longctx.json`](eval/longctx.json).
546
+
547
+ ---
548
+
549
  ## How to run
550
 
551
  ### Requirements
 
598
  model:
599
  model_dir: models
600
  model_name: mimo-2.25bpw-hq
601
+ max_seq_len: 131072 # validated; 393216 fits if you drop the draft model
602
+ cache_size: 131072
603
  cache_mode: FP16 # default; Q8 paged KV works but is unvalidated for quality
604
  chunk_size: 2048
605
  max_batch_size: 1 # each extra slot costs another 175.5 MiB ring + paged span + draft cache
 
718
  perfect 2.34 bpw quant β€” and the benchmark deltas above say the FP8 model is only ~0.6 pp
719
  better at HumanEval+, so most of that is the base model, not the quantization.
720
  * **Tensor parallel is unsupported.** The port has only been built and run single-device.
721
+ * **Long context is now measured, and it is TTFT-bound rather than memory-bound.** 8K β†’ 350K
722
+ has been served and retrieval is perfect at every length and depth tested, but a 250K prompt
723
+ costs ~12.6 minutes of prefill and a 350K prompt ~22.4 minutes on one GB10. `max_seq_len`
724
+ above 393,216 (or above 131,072 with the drafter) does not fit at FP16 KV. See
725
+ [Long context](#long-context).
726
  * **Batch > 1 with DFlash is untested.** Everything above is `max_batch_size: 1`.
727
  * **No vision, no audio, no MTP.** The MTP heads in the source repo are not ported, so
728
  MTP-based drafting is unavailable; DFlash is the drafting path.