Runtime format audit (2026-10-06)

No quantization-container correction was needed. This family has no n-gram tensors; no n-gram file or declaration was added. This CUDA pack is outside MLX/oMLX/MTPLX export scope.

See runtime_audit.json for pinned config/index/header bindings, architecture, physical-format findings, and applied corrections. This is development evidence; no quality, MTP exactness, speed, or certification claim is added. Historical evidence stays bound to its original revision.

AX-Qwen3-Embedding-8B-CUDA-AXQ-NVFP4-W4A4

AXQuant CUDA NVFP4 W4A4 mixed precision development preview. Converted from the original upstream BF16 source, without AWQ. Native E2M1 FP4 weights and inputs use per-16 E4M3FN scales and FP32 global scales. This is a retrieval embedding checkpoint, with no MTP or generative claim.

All attention projections and the first/last two MLP blocks retain original BF16. Remaining MLP matrices use NVFP4 W4A4.

Saved query instruction, last-token pooling, exactly one <|endoftext|> (151643), and L2 normalization. The chat <|im_end|> (151645) is not appended.

Source and export

  • Original source: Qwen/Qwen3-Embedding-8B.
  • Immutable source revision: 1d8ad4ca9b3dd8059ad90a75d4983776a23d44af.
  • Quantized matrices: 96; protected tensors: 302.
  • Output weight bytes: 8,188,897,592.
  • Full embedding dimension: 4096.
  • Reviewed reproduction commit: 3e22743a126a2ecf1c673ab6441041ac37407db8.

Factory conversion uses AXQuant's NumPy reference RTN encoder, independently checked against preserved BF16 source tensors. Source and pooling assets, calibration and output payloads are checksum-bound. The committed development path adds embedding metadata preservation and Qwen embedding protection; older released AXQuant wheels do not contain these CUDA features. Exports were produced during development; provenance.json binds the final reviewed reproduction code and each source/calibration/payload digest. Weight sensitivity remains explicitly unmeasured. Original license and notices are retained.

Tested native execution

vLLM 0.25.1, Torch 2.11.0+cu130, CUDA 13.0, on RTX 5090 and Jetson Thor. Every allocation is reconciled with actual worker kernels: 64 CutlassNvFp4LinearKernel modules, no MoE tables, and 36 decoder attention layers. Marlin and emulation are rejected. Full-sequence prefill, eager mode, maximum 512 tokens and one sequence were tested.

GPU Mean cosine to BF16 Minimum cosine Paired top-1
RTX 5090 0.986226 0.982638 4/4
Jetson Thor 0.986426 0.983609 4/4

These are eight vectors from four simple query/document pairs, from a calibration-disjoint development corpus also used in candidate selection. This is not an independent MTEB/RTEB benchmark, broad retrieval-quality, long-context, Matryoshka slicing, concurrency or speed certification. Eager/JIT, autotune and original upstream RoPE configuration warnings do not establish performance. Evaluate your own documents before deployment.

Rebuild retrieval indexes with this exact checkpoint. BF16 and quantized vectors are not interchangeable. Prompts, tokenization, pooling and normalization must match when embedding queries and indexed documents.

Reproduce the included development check

Use a compatible CUDA vLLM/PyTorch environment and download the entire repository:

hf download AutomatosX/AX-Qwen3-Embedding-8B-CUDA-AXQ-NVFP4-W4A4 --local-dir model-nvfp4
python model-nvfp4/examples/embedding_smoke.py   --model model-nvfp4 --corpus model-nvfp4/evaluation/retrieval-corpus.json   --reference model-nvfp4/evaluation/rtx5090-bf16.json   --output embedding-runtime.json --require-native-fp4 --memory-fraction .45

For the tested Thor recipe, use --memory-fraction .12 and --reference model-nvfp4/evaluation/thor-bf16.json. Both example files must remain together. No model remote code or AXQuant installation is required for this runtime example. The reference comparison binds the raw original BF16 configuration digest and the exact corpus, rejects non-finite or unnormalized vectors, and requires mean cosine at least 0.95 and minimum cosine at least 0.90 on this small development corpus.

axquant_cuda_plan.json, axquant_cuda_manifest.json, activation_calibration.json, development_runtime_smoke.json, evaluation/, provenance.json and SHA256SUMS.txt record the scope. The converter manifest remains runtime_verified=false and quality_certified=false; runtime development checks do not become a certificate.

Downloads last month
33
Safetensors
Model size
8B params
Tensor type
BF16
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for AutomatosX/AX-Qwen3-Embedding-8B-CUDA-AXQ-NVFP4-W4A4

Quantized
(41)
this model

Collections including AutomatosX/AX-Qwen3-Embedding-8B-CUDA-AXQ-NVFP4-W4A4