Instructions to use AutomatosX/AX-Nemotron-3-Super-120B-A12B-CUDA-AXQ-NVFP4-MTP with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use AutomatosX/AX-Nemotron-3-Super-120B-A12B-CUDA-AXQ-NVFP4-MTP with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="AutomatosX/AX-Nemotron-3-Super-120B-A12B-CUDA-AXQ-NVFP4-MTP", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("AutomatosX/AX-Nemotron-3-Super-120B-A12B-CUDA-AXQ-NVFP4-MTP", trust_remote_code=True) model = AutoModelForCausalLM.from_pretrained("AutomatosX/AX-Nemotron-3-Super-120B-A12B-CUDA-AXQ-NVFP4-MTP", trust_remote_code=True, device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use AutomatosX/AX-Nemotron-3-Super-120B-A12B-CUDA-AXQ-NVFP4-MTP with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "AutomatosX/AX-Nemotron-3-Super-120B-A12B-CUDA-AXQ-NVFP4-MTP" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AutomatosX/AX-Nemotron-3-Super-120B-A12B-CUDA-AXQ-NVFP4-MTP", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/AutomatosX/AX-Nemotron-3-Super-120B-A12B-CUDA-AXQ-NVFP4-MTP
- SGLang
How to use AutomatosX/AX-Nemotron-3-Super-120B-A12B-CUDA-AXQ-NVFP4-MTP with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "AutomatosX/AX-Nemotron-3-Super-120B-A12B-CUDA-AXQ-NVFP4-MTP" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AutomatosX/AX-Nemotron-3-Super-120B-A12B-CUDA-AXQ-NVFP4-MTP", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "AutomatosX/AX-Nemotron-3-Super-120B-A12B-CUDA-AXQ-NVFP4-MTP" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AutomatosX/AX-Nemotron-3-Super-120B-A12B-CUDA-AXQ-NVFP4-MTP", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use AutomatosX/AX-Nemotron-3-Super-120B-A12B-CUDA-AXQ-NVFP4-MTP with Docker Model Runner:
docker model run hf.co/AutomatosX/AX-Nemotron-3-Super-120B-A12B-CUDA-AXQ-NVFP4-MTP
Runtime format audit (2026-10-06)
No quantization-container correction was needed. This family has no n-gram tensors; no n-gram file or declaration was added. This CUDA pack is outside MLX/oMLX/MTPLX export scope.
See runtime_audit.json for pinned config/index/header bindings, architecture, physical-format findings, and applied corrections. This is development evidence; no quality, MTP exactness, speed, or certification claim is added. Historical evidence stays bound to its original revision.
AX-Nemotron-3-Super-120B-A12B-CUDA-AXQ-NVFP4-MTP
Native AXQuant RTN NVFP4A16 development checkpoint, exported from the NVIDIA BF16 source with no AWQ or third-party quantizer. Weight blocks contain 16 elements, E2M1 weights use E4M3FN block scales and FP32 inverse global scales. Activations remain BF16. The checkpoint uses the public compressed-tensors NVFP4 format.
Source: nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16 at 2dc98e2afe4face0e4ce40972a915c45368bd34a. Serialization backend: numpy-reference. CPU serialization is format evidence only; it does not establish GPU execution or speed.
Precision and MTP
Eligible text expert and MLP matrices use NVFP4. Embeddings, output head, norms, routers, Mamba state and attention projections, shared experts, latent projections, and all integrated mtp.* tensors keep source precision. The original config, tensor names and main-index MTP layout remain available to compatible CUDA runtimes. axquant_nemotron_release.json records byte-preservation checks for every MTP tensor. MTP runtime compatibility is unverified; the name means trained MTP weights are packaged. No runtime, quality, or throughput certificate is claimed.
Consumers
A consumer must support Nemotron-H, compressed-tensors NVFP4A16, and the source integrated MTP layout to execute all components. Runtime selection, kernels and speculative-decoding settings are runtime responsibilities. The AXQuant plan, manifest and SHA256 inventory are included for reproducible artifact inspection.
The source license and available notices are included.
- Downloads last month
- 443