How to use from
Docker Model Runner
docker model run hf.co/AutomatosX/AX-Nemotron-3-Super-120B-A12B-CUDA-AXQ-NVFP4-MTP
Quick Links

Runtime format audit (2026-10-06)

No quantization-container correction was needed. This family has no n-gram tensors; no n-gram file or declaration was added. This CUDA pack is outside MLX/oMLX/MTPLX export scope.

See runtime_audit.json for pinned config/index/header bindings, architecture, physical-format findings, and applied corrections. This is development evidence; no quality, MTP exactness, speed, or certification claim is added. Historical evidence stays bound to its original revision.

AX-Nemotron-3-Super-120B-A12B-CUDA-AXQ-NVFP4-MTP

Native AXQuant RTN NVFP4A16 development checkpoint, exported from the NVIDIA BF16 source with no AWQ or third-party quantizer. Weight blocks contain 16 elements, E2M1 weights use E4M3FN block scales and FP32 inverse global scales. Activations remain BF16. The checkpoint uses the public compressed-tensors NVFP4 format.

Source: nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16 at 2dc98e2afe4face0e4ce40972a915c45368bd34a. Serialization backend: numpy-reference. CPU serialization is format evidence only; it does not establish GPU execution or speed.

Precision and MTP

Eligible text expert and MLP matrices use NVFP4. Embeddings, output head, norms, routers, Mamba state and attention projections, shared experts, latent projections, and all integrated mtp.* tensors keep source precision. The original config, tensor names and main-index MTP layout remain available to compatible CUDA runtimes. axquant_nemotron_release.json records byte-preservation checks for every MTP tensor. MTP runtime compatibility is unverified; the name means trained MTP weights are packaged. No runtime, quality, or throughput certificate is claimed.

Consumers

A consumer must support Nemotron-H, compressed-tensors NVFP4A16, and the source integrated MTP layout to execute all components. Runtime selection, kernels and speculative-decoding settings are runtime responsibilities. The AXQuant plan, manifest and SHA256 inventory are included for reproducible artifact inspection.

The source license and available notices are included.

Downloads last month
443
Safetensors
Model size
124B params
Tensor type
BF16
·
U8
·
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for AutomatosX/AX-Nemotron-3-Super-120B-A12B-CUDA-AXQ-NVFP4-MTP

Quantized
(55)
this model

Collections including AutomatosX/AX-Nemotron-3-Super-120B-A12B-CUDA-AXQ-NVFP4-MTP