--- license: other license_name: kimi-k3 license_link: LICENSE base_model: moonshotai/Kimi-K3 base_model_relation: quantized library_name: sglang tags: - kimi-k3 - moe - w4afp8 - int4 - fp8 - sglang - gh200 --- # Kimi-K3-W4AFP8 INT4 (group 128) routed experts with FP8 activations, FP8 attention projections, requantized from the native MXFP4 weights of [moonshotai/Kimi-K3](https://huggingface.co/moonshotai/Kimi-K3). Same tensor format and kernels as [vessl/Kimi-K3-W4AFP8](https://huggingface.co/vessl/Kimi-K3-W4AFP8), with lower error against the native model. ## Two configurations on 32 x GH200 8 nodes x 4 GH200, TP32 / EP32, DSpark speculative decoding (block 2), FP8 KV cache. Real AIME/HMMT prompts with real EOS; numbers are median output tok/s per stream. | | **W4A16** — [moonshotai/Kimi-K3](https://huggingface.co/moonshotai/Kimi-K3) (native MXFP4) | **W4A8** — this repo | |---|---|---| | tok/s per stream, 50 concurrent | 29.6 | 34.1 | | tok/s per stream, 20 concurrent | 40.6 | 43.6 | | aggregate tok/s, 50 concurrent | 883 | 949 | | KL to native, nats/token (95% CI) | +0.0011 (0.0006–0.0016) | +0.0154 (0.0142–0.0167) | | AIME 2026 + HMMT Feb 2026, 8 samples/problem | 82.3% | 81.3% (−1.0 pt, SE 1.6: not significant) | - **Patches in the measured runs:** W4A16 used patches 02 and 04 (no `--max-total-tokens` cap). W4A8 used 01, 02, 03 and 05, i.e. exactly the launch command below. - **Pick W4A16** when you want the reference model. - **Pick W4A8** for about 15% more throughput at 50 streams. - **Scaling out:** two W4A8 replicas (64 GPUs) behind `sglang-router`, 25 streams each, give 35.7 tok/s per stream. - **KL:** teacher-forced on 355k tokens of native-model math/proof continuations. For comparison, vessl/Kimi-K3-W4AFP8 scores +0.0187. ## Setup 1. **Image:** SGLang v0.5.20, arm64, CUDA 13. 2. **Clone the repo:** `git clone https://huggingface.co/Radioheading/Kimi-K3-W4AFP8`. You need its `apply_patch.py` and `patches/` below. 3. **Patches:** apply the ones listed below inside the image's source tree (`cd /sgl-workspace/sglang && git apply `). 4. **W4A8 quantization method:** for W4A8 only, run `python3 apply_patch.py` from this repo. It registers the method (from vessl/Kimi-K3-W4AFP8). 5. **Draft model:** download [RadixArk/Kimi-K3-DSpark](https://huggingface.co/RadixArk/Kimi-K3-DSpark). | patch | needed for | effect | |---|---|---| | `01-sgl-kernel-sm90a.patch` | W4A8 (required) | Builds the SM90 CUTLASS kernels with `sm_90a`. aarch64 wheels lack it, and W4A8/FP8 kernels abort with "Arch conditional MMA instruction". Shortcut: copy the prebuilt `sgl_kernel_sm90a/common_ops.abi3.so` (arm64, CUDA 13, v0.5.20) over `site-packages/sgl_kernel/sm90/common_ops.abi3.so`. Or apply the patch and rebuild the `common_ops_sm90_build` target (~5 min). | | `02-expert-load-filter.patch` | both | Each rank reads only its own experts (`SGLANG_EXPERT_LOAD_FILTER=896`). Load time drops from ~20 to ~5 min. | | `03-fp8-kv-prefill-gather.patch` | both | FP8-KV prefill converts only the tokens it reads. Without it, add `--max-total-tokens 524288` or prefill can OOM. | | `04-marlin-ep-block.patch` | W4A16 | Fixes Marlin MoE tile size under EP (`SGLANG_MARLIN_EP_BLOCK=1`). Without it: 23.6 tok/s at 50 streams. | | `05-w4a8-count-sort.patch` | W4A8 | Bit-identical, faster MoE routing permutation (`K3_W4_SRC2DST=count`). | ## Launch Run on each of the 8 nodes, with `NODE_RANK=0..7` and `HEAD` = node 0's IP: ```bash export NCCL_MIN_NCHANNELS=24 NCCL_MAX_NCHANNELS=24 NCCL_IB_ADAPTIVE_ROUTING=0 NCCL_IB_QPS_PER_CONNECTION=2 export SGLANG_EXPERT_LOAD_FILTER=896 SGLANG_RAGGED_VERIFY_MODE=static SGLANG_JIT_DEEPGEMM_FAST_WARMUP=1 # W4A8 (this repo) export K3_W4_SRC2DST=count MODEL="--model-path Radioheading/Kimi-K3-W4AFP8" # W4A16 (native) instead: # export SGLANG_MARLIN_EP_BLOCK=1 # MODEL="--model-path moonshotai/Kimi-K3 --moe-runner-backend marlin" sglang serve $MODEL --served-model-name moonshotai/Kimi-K3 --trust-remote-code \ --tp-size 32 --ep-size 32 --nnodes 8 --node-rank $NODE_RANK --dist-init-addr $HEAD:20000 \ --kv-cache-dtype fp8_e4m3 --mamba-ssm-dtype float32 --mem-fraction-static 0.82 \ --prefill-attention-backend flashmla --decode-attention-backend flashmla \ --disable-radix-cache --max-running-requests 64 --cuda-graph-max-bs-decode 64 --max-mamba-cache-size 192 \ --chunked-prefill-size 4096 --max-prefill-tokens 8192 \ --speculative-algorithm DSPARK --speculative-draft-model-path RadixArk/Kimi-K3-DSpark \ --speculative-dspark-block-size 2 --speculative-draft-model-quantization unquant \ --speculative-draft-attention-backend flashinfer --speculative-draft-kv-cache-dtype bfloat16 \ --enable-linear-replayssm-spec \ --reasoning-parser kimi_k3 --tool-call-parser kimi_k3 --host 0.0.0.0 --port 30000 ``` The server is OpenAI-compatible: `http://$HEAD:30000/v1`, model `moonshotai/Kimi-K3`. It is ready after 8–15 minutes. ## License Derived from moonshotai/Kimi-K3 and distributed under its license (see `LICENSE`).