RTX 5090 + 64 GiB RAM: Flash Next EXL3 ~42% decode gain, with quality differences disclosed

#1
by Jerrybro - opened

I'm running Qwen3.8 Flash Next Abliterated EXL3 2.50bpw on an RTX 5090, Ryzen 9 9950X3D and 64 GiB RAM. In one matched maintenance-specific 16K-prompt test, static hotset v3 improved decode from 42.80 to 60.76 token/s (+41.98%). The original 50-question direct battery scored v2 42/50 vs maintenance-v3 41/50; that difference is retained.

Architecture and measured conditions

  • All 512 experts per layer remain available across 48 MoE layers. The tested split is 168 GPU-resident / 344 CPU-handled, with 10 routed experts per token per layer.
  • CPU worker threads=8; fused handoff; static SWAP=0; MEMOPS=0; PLE row cache=OFF; band-swizzle=ON; K4/V4 KV cache. MTP is OFF in these measurements. No training, re-quantization or expert pruning.
  • The entire checkpoint is not in VRAM. RAM and disk-streamed PLE are part of the architecture. Sampled card peak minus startup ambient was around 22.8 GiB, not exact model allocation or a new large VRAM reduction over v2.

Each A/B/A phase used three actual 512-token decode samples; the baseline pools six opening/closing samples. The 42% claim is for that selected maintenance prompt, not all workloads or end-to-end latency.

Separate matched matrix Baseline β†’ candidate token/s Gain Baseline drift
Dynamic v2 β†’ maintenance-v3; 16384-token maintenance prompt / 32768 capacity 42.80 β†’ 60.76 +41.98% +3.58%
Dynamic v2 β†’ balanced static; 249965-token engineering prompt / 262144 capacity 42.22 β†’ 58.58 +38.77% +10.24%
Dynamic v2 β†’ balanced static; 249955-token maintenance prompt / 262144 capacity 46.94 β†’ 54.45 +16.00% +2.76%
Balanced static β†’ maintenance-only static; same 249955-token maintenance prompt / 262144 capacity 59.90 β†’ 65.19 +8.83% +3.28%

Different rows cannot be added together. Cold/cached/unknown prefill is separated from decode. Slower or inconclusive experiments are included in the report.

Answer quality and remaining gaps

The original 50-question score differs by one answer. A later graph/binary repeat reproduced the relevant binary error in both runtimes, including v2 baselines; that does not erase the original result or prove equivalence. Eight actual two-step tool tasks and eight short document version/conflict tasks passed all three phases, reported separately. Reasoning truncation and numeric-string failures remain disclosed. Balanced KG has not completed its own full50 battery and cannot inherit the maintenance-v3 score.

The original author labels the checkpoint Abliterated / Uncensored. I reused it, did not perform the abliteration, and do not claim zero refusals or guaranteed correctness.

Similar work and credits

CPU offload and hot-expert placement are existing ideas. r0b0tlab's EXL3 recipe targets a 24 GB RTX 3090 with native 262K context. imp reports 75-80 token/s on a 5090 using NVFP4 and a GPU-managed LRU cache. kutaelee's 5090/128K recipe publishes WSL2/GGUF validation and MTP negatives. These are author-reported results with different hardware/checkpoints/prompts; I am not claiming to be first or fastest.

Credit to Qwen, OrcaRouter, ghost-actual and ExLlamaV3 / CPU-offload engine contributors for their respective work. My contribution is the runtime-placement experiments and measured speed/quality tradeoff on this workstation.

Read and download

Public introduction and full result index Β· Original author's weights Β· Configuration/evidence ZIP

Our file downloads require Hugging Face login and consent to the standard account/email-sharing form, with automatic approval. No discussion reply is required. The ZIP contains configuration references, 15 speed comparisons, quality results and 71 historical evidence snapshots; it has no executable runtime or one-click installer. Complete end-to-end reproduction is still a gap.

I would welcome measurements from similar hardware, particularly comparable prompt depths with MTP stated explicitly, and discussion of how to balance residency, decode speed and answer reliability. Please keep different engines, quantizations and context conditions separate when comparing results.

Sign up or log in to comment