wrldsuksgo2mars commited on
Commit
59484d5
·
verified ·
1 Parent(s): 701cd74

docs: add final v0.7.0 runtime qualification

Browse files
Files changed (1) hide show
  1. README.md +18 -8
README.md CHANGED
@@ -43,14 +43,24 @@ recovery contract covers any expert below the 1,024-route floor. Allocation
43
  provenance and the validation report ship with the model; the full error ledger
44
  is retained separately as internal quantization evidence.
45
 
46
- The first weight publication is gated on the exact checkpoint, allocation,
47
- native-byte, manifest, and configuration-parser audits above. The matching
48
- two-GPU B12x/vLLM release is separately gated on loading this materialized
49
- checkpoint with DFlash2 and FP8 MLA, ordinary generation, and five strict
50
- tool-call scenarios. Those exact prompts, expected behavior, actual calls, and
51
- raw reports are added to
52
- `quantization/vllm-prompt-expected-actual.md` and
53
- `quantization/vllm-tool-eval.json` when runtime qualification completes.
 
 
 
 
 
 
 
 
 
 
54
 
55
  ## Serving
56
 
 
43
  provenance and the validation report ship with the model; the full error ledger
44
  is retained separately as internal quantization evidence.
45
 
46
+ ## Qualified two-GPU serving result
47
+
48
+ Release `v0.7.0` of the matching B12x/vLLM recipe qualified the public target
49
+ revision with DFlash2 K5, FP8 MLA, TP2 + EP2 + DCP2, vision up to 16 images,
50
+ and a 1,048,576-token request limit on 2× RTX PRO 6000 Blackwell at a 400 W
51
+ power cap per GPU.
52
+
53
+ - Code-agent decode: **213 tok/s at C1** and **832 tok/s at C16**.
54
+ - 128K cold prefill: **4,572 tok/s**.
55
+ - Available KV cache: **2,758,919 tokens**.
56
+ - Exact 1M multi-needle test: **6/6 needles recovered**.
57
+ - Exact 69-case thinking tool-call suite at parallelism 8: **86/100**
58
+ (55 pass, 8 partial, 6 fail), versus 88/100 for the matched uniform-K3
59
+ control.
60
+
61
+ The full prompts, outputs, scoring, performance curves, and machine-readable
62
+ receipts are in the recipe's
63
+ [`v0.7.0` benchmark report](https://github.com/tpurtell/glm-5.3-flash-ext3-4-bit-2x-rtx/blob/v0.7.0/benchmarks/RESULTS.md).
64
 
65
  ## Serving
66