avlp12's picture
Update 2026-09-08: two-box tensor-parallel speculative decode 76 tok/s (+37%), prefill +4-5% from a bit-identical gather path, batched speculative serving, corrected round-cost convention and ceilings
8d0d47e verified