# Validation — RTX 5090, 20 September 2026 ## What passed - Both native safetensors checkpoints loaded through standard ComfyUI loaders. - 48 main renders: 12 prompts × two seeds × BF16/NVFP4, all at 1024×1024. - Eight extended renders: product edit, portrait edit, transparent RGBA and 2K typography, each with BF16 and NVFP4. - Six held-out renders: Chinese lettering, 1536×864 landscape and two-reference image editing, each with BF16 and NVFP4. These prompts were not used to choose the conversion recipe. - Four shipped workflows imported and executed through the actual frontend. - Three subjects tested across all four combinations of BF16/NVFP4 transformer and encoder weights, using the original stock conditioning path. - Real SM120 block-scaled FP4 CUDA kernels observed in both language conditioning and image denoising with the accelerated encoder patch. - All 571 protected tensors checked for exact equality against the source files. All 444 quantized matrices checked for packed shape, marker, scale dtype and finite scales. Source and release file hashes recorded. - Hub-uploaded weights downloaded into a separate directory and matched against the release SHA-256 hashes. The VAE hash matches the official original. - All four API workflows also passed using the Hub-downloaded files on the unpatched base runtime, with custom nodes disabled. Thus stock compatibility and optional accelerated conditioning were tested separately. No calibration, optimization against the test prompts, or fine-tuning was used. The quantization recipe remained unchanged after its first conversion. ## Measured performance 1024×1024 EMBER product prompt from the shipped generation workflow, seed 20260922, 40 steps, Euler/simple, CFG 1, denoise 1. One first-after-unload run, then three warm runs per variant. **Every node executed on every repeat** using `--cache-none`; the histories were checked for an empty cached-node list. Elapsed time comes from server execution-start/success timestamps, not polling latency. Includes conditioning, denoising, VAE decode and PNG save. | Weights / encoder mode | Warm median | Warm range | First after unload | Largest sampled GPU use, warm | |---|---:|---:|---:|---:| | Official BF16 / stock FP32 conditioning | 15.629 s | 15.553–15.650 s | 18.932 s | 31,677 MiB | | Official INT8 ConvRot / stock conditioning | 7.553 s | 7.496–7.591 s | 9.973 s | 18,433 MiB | | Qwen Image 2.1 NVFP4 / patched BF16 + NVFP4 conditioning | **6.746 s** | **6.706–6.749 s** | **8.083 s** | **13,249 MiB** | This NVFP4 workflow delivered **2.32× the BF16 throughput** and took **10.7% less time than INT8** in this one controlled test. These are local measurements, not general speed guarantees. First-after-unload is not a cold filesystem-cache or fresh-boot test. The quality-suite runs used ordinary node caching and are not the basis of this performance table. GPU memory was sampled through `nvidia-smi` every 100 ms. Values are **whole-GPU usage**, including Windows and other idle desktop services, not PyTorch peak allocated memory or a minimum VRAM requirement. The largest sampled NVFP4 warm value is about 12.94 GiB. No other model inference job was running during this benchmark. Hardware other than the RTX 5090 was not benchmarked. Raw measurements: [benchmark.json](validation/benchmark.json). ## Quality observations and limits The main set covers portraits, landscapes, glass products, typography, object counts and spatial relationships, illustration, fashion, night scenes, architecture, macro, astronaut portraits and botanical posters. All 24 matched pairs are published in [COMPARISONS.md](COMPARISONS.md). Visual review found coherent images and preserved useful text, material and editing capabilities. It also found real differences: changed portrait details, object placement, typography layout and small material/color details. In one macro pair the quantized spider is less green. Small astronaut-label spelling was imperfect in the BF16 baseline too. The Chinese probe retained the requested wording, but NVFP4 moved the subtitle above the cup rather than below it. The two-reference test combined the reference woman with the EMBER bottle and removed the violin, while changing hand pose. The product edit preserved the label and scene while changing glass, metal and props. Portrait recoloring can also change sleeve coverage. This is semantic reference preservation, not a pixel-locked edit. The transparent example has genuine RGBA output: alpha minimum 0, maximum 255, 49.3% of pixels below alpha 16 and 49.7% above alpha 240. See the original [RGBA PNG](assets/03_Transparent_RGBA.png) and the [checkerboard composite](assets/transparency-checker.jpg). RGB-only viewers may show the purple RGB values stored underneath transparent pixels. This is a small local evaluation, not a blinded preference study or a benchmark proving no quality loss. There was no LoRA, ControlNet, training or video test. Long text, exhaustive multilingual coverage and every reference-image count were not tested. Stock encoder arithmetic and accelerated encoder arithmetic can produce different images even with identical weights and seeds. ## Runtime and kernel evidence - Windows, Python 3.13.12, RTX 5090 32 GB. - ComfyUI 0.36.0, base commit `99073836d45f66053c45ba8564984e6def9cebba`. - Included Qwen-specific encoder patch; local validation commit `2d324220`. - PyTorch `2.14.0+cu130`; comfy-kitchen `0.2.35`; comfy-aimdo `0.5.5`. - Frontend package `1.53.6`; safetensors `0.8.0`. - Comfy Kitchen attention, normal VRAM mode, async offload, batch size one. - All tests ran with `--disable-all-custom-nodes`. The profiler captured kernels beginning `cutlass3x_sm120_bstensorop_s16864gemm_block_scaled_ue4m3xe2m1_ue4m3xe2m1` for short conditioning sequences and image-denoising sequences. This confirms hardware FP4 execution, rather than relying on checkpoint names or successful loading. Raw observations: [kernel_audit.jsonl](validation/kernel_audit.jsonl). No NVFP4 matmul-fallback warning occurred in the successful audited runs. The first acceleration probe exposed FP32 activations entering a BF16/FP16-only quantizer. The included patch fixes that at the Qwen Image 2.1 encoding boundary: vision preprocessing stays unchanged, the resulting multimodal embeddings become BF16 only for NVFP4 conditioning on supported hardware, and existing ComfyUI matmul dispatch is used. BF16 and INT8 checkpoints retain their original path. ## Precision and provenance The DiT converts 192 matrices and retains its other 73 tensors. The encoder converts 252 matrices and retains its other 498 tensors, including the complete vision tower, embeddings and output head. Per-matrix relative reconstruction RMSE is recorded in the manifests; it is not an image-quality score. Source revision: `Comfy-Org/Qwen-Image-2.1@ace0edeb3791a594ddfa36ed5f41a178a394e921`. The image transformer source SHA-256 is `89f4158d066cc33906a199fca85634f766892dd78f49b6698dabf187ac86c4bc`. The encoder source SHA-256 is `68bdc82bc1b66851162ae656225e7e2068166b603db19bd5d5a3b90eb12669a9`. Full tensor manifests sit beside each checkpoint; `SHA256SUMS` covers the release.