Runtime format audit (2026-10-06)
No quantization-container correction was needed. This family has no n-gram tensors; no n-gram file or declaration was added. This CUDA pack is outside MLX/oMLX/MTPLX export scope.
See runtime_audit.json for pinned config/index/header bindings, architecture, physical-format findings, and applied corrections. This is development evidence; no quality, MTP exactness, speed, or certification claim is added. Historical evidence stays bound to its original revision.
AX-Qwen3-Embedding-8B-CUDA-AXQ-NVFP4-W4A4
AXQuant CUDA NVFP4 W4A4 mixed precision development preview. Converted from the original upstream BF16 source, without AWQ. Native E2M1 FP4 weights and inputs use per-16 E4M3FN scales and FP32 global scales. This is a retrieval embedding checkpoint, with no MTP or generative claim.
All attention projections and the first/last two MLP blocks retain original BF16. Remaining MLP matrices use NVFP4 W4A4.
Saved query instruction, last-token pooling, exactly one <|endoftext|> (151643), and L2 normalization. The chat <|im_end|> (151645) is not appended.
Source and export
- Original source: Qwen/Qwen3-Embedding-8B.
- Immutable source revision:
1d8ad4ca9b3dd8059ad90a75d4983776a23d44af. - Quantized matrices: 96; protected tensors: 302.
- Output weight bytes: 8,188,897,592.
- Full embedding dimension: 4096.
- Reviewed reproduction commit:
3e22743a126a2ecf1c673ab6441041ac37407db8.
Factory conversion uses AXQuant's NumPy reference RTN encoder, independently
checked against preserved BF16 source tensors. Source and pooling assets,
calibration and output payloads are checksum-bound. The committed development
path adds embedding metadata preservation and Qwen embedding protection;
older released AXQuant wheels do not contain these CUDA features. Exports
were produced during development; provenance.json binds the final reviewed
reproduction code and each source/calibration/payload digest. Weight sensitivity
remains explicitly unmeasured. Original license and notices are retained.
Tested native execution
vLLM 0.25.1, Torch 2.11.0+cu130, CUDA 13.0, on RTX 5090 and Jetson Thor. Every allocation is reconciled with actual worker kernels: 64 CutlassNvFp4LinearKernel modules, no MoE tables, and 36 decoder attention layers. Marlin and emulation are rejected. Full-sequence prefill, eager mode, maximum 512 tokens and one sequence were tested.
| GPU | Mean cosine to BF16 | Minimum cosine | Paired top-1 |
|---|---|---|---|
| RTX 5090 | 0.986226 | 0.982638 | 4/4 |
| Jetson Thor | 0.986426 | 0.983609 | 4/4 |
These are eight vectors from four simple query/document pairs, from a calibration-disjoint development corpus also used in candidate selection. This is not an independent MTEB/RTEB benchmark, broad retrieval-quality, long-context, Matryoshka slicing, concurrency or speed certification. Eager/JIT, autotune and original upstream RoPE configuration warnings do not establish performance. Evaluate your own documents before deployment.
Rebuild retrieval indexes with this exact checkpoint. BF16 and quantized vectors are not interchangeable. Prompts, tokenization, pooling and normalization must match when embedding queries and indexed documents.
Reproduce the included development check
Use a compatible CUDA vLLM/PyTorch environment and download the entire repository:
hf download AutomatosX/AX-Qwen3-Embedding-8B-CUDA-AXQ-NVFP4-W4A4 --local-dir model-nvfp4
python model-nvfp4/examples/embedding_smoke.py --model model-nvfp4 --corpus model-nvfp4/evaluation/retrieval-corpus.json --reference model-nvfp4/evaluation/rtx5090-bf16.json --output embedding-runtime.json --require-native-fp4 --memory-fraction .45
For the tested Thor recipe, use --memory-fraction .12 and
--reference model-nvfp4/evaluation/thor-bf16.json. Both example files must
remain together. No model remote code or AXQuant installation is required
for this runtime example. The reference comparison binds the raw original
BF16 configuration digest and the exact corpus, rejects non-finite or
unnormalized vectors, and requires mean cosine at least 0.95 and minimum
cosine at least 0.90 on this small development corpus.
axquant_cuda_plan.json, axquant_cuda_manifest.json, activation_calibration.json,
development_runtime_smoke.json, evaluation/, provenance.json and
SHA256SUMS.txt record the scope. The converter manifest remains
runtime_verified=false and quality_certified=false; runtime development
checks do not become a certificate.
- Downloads last month
- 33