Text Generation
Transformers
Safetensors
PyTorch
nemotron_h
nvidia
nemotron-3
latent-moe
mtp
conversational
custom_code
8-bit precision
modelopt
Instructions to use nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4", trust_remote_code=True) model = AutoModelForCausalLM.from_pretrained("nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4", trust_remote_code=True, device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4
- SGLang
How to use nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4 with Docker Model Runner:
docker model run hf.co/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4
nemotron120b
#31
by willowoods - opened
README.md
CHANGED
|
@@ -69,7 +69,6 @@ track_downloads: true
|
|
| 69 |
| **Supported Languages** | English, French, German, Italian, Japanese, Spanish, Chinese |
|
| 70 |
| **Best For** | Agentic workflows, long-context reasoning, high-volume workloads (e.g. IT ticket automation), tool use, RAG |
|
| 71 |
| **Reasoning Mode** | Configurable on/off via chat template (`enable_thinking=True/False`) |
|
| 72 |
-
| **Speculative Decoding** | Includes a built-in MTP head, with an updated MTPv2 head available as a separate [checkpoint](https://huggingface.co/nvidia/Nemotron-3-Super-120B-A12B-BF16-MTPv2) |
|
| 73 |
| **License** | [NVIDIA Nemotron Open Model License](https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-nemotron-open-model-license/) |
|
| 74 |
| **Release Date** | March 11, 2026 |
|
| 75 |
|
|
@@ -272,7 +271,7 @@ vllm serve $MODEL_CKPT \
|
|
| 272 |
|
| 273 |
##### vLLM on DGX Spark
|
| 274 |
|
| 275 |
-
To deploy the NVFP4 chekpoint on NVIDIA DGX Spark, make sure that you are using the `vllm/vllm-openai:v0.
|
| 276 |
|
| 277 |
```bash
|
| 278 |
docker run --rm -it --gpus all \
|
|
@@ -284,7 +283,7 @@ docker run --rm -it --gpus all \
|
|
| 284 |
-v ~/.cache/huggingface:/root/.cache/huggingface \
|
| 285 |
-v $(pwd)/super_v3_reasoning_parser.py:/app/super_v3_reasoning_parser.py \
|
| 286 |
-p 8000:8000 \
|
| 287 |
-
vllm/vllm-openai:v0.
|
| 288 |
--model nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4 \
|
| 289 |
--served-model-name nvidia/nemotron-3-super \
|
| 290 |
--host 0.0.0.0 \
|
|
@@ -303,7 +302,7 @@ docker run --rm -it --gpus all \
|
|
| 303 |
--moe-backend marlin \
|
| 304 |
--mamba_ssm_cache_dtype float16 \
|
| 305 |
--quantization fp4 \
|
| 306 |
-
--speculative_config '{"method":"mtp","num_speculative_tokens":3,"
|
| 307 |
--reasoning-parser-plugin /app/super_v3_reasoning_parser.py \
|
| 308 |
--reasoning-parser super_v3 \
|
| 309 |
--enable-auto-tool-choice \
|
|
@@ -617,8 +616,6 @@ Alongside the model, we release our final pre-training and post-training data, a
|
|
| 617 |
|
| 618 |
More details on the datasets and synthetic data generation methods can be found in the technical report _[**_NVIDIA Nemotron 3 Super_**](https://research.nvidia.com/labs/nemotron/files/NVIDIA-Nemotron-3-Super-Technical-Report.pdf)_.
|
| 619 |
|
| 620 |
-
For more information about the datasets used to train this model, please see the [Public Summary of Training Content](https://developer.download.nvidia.com/assets/nemo/docs/public-summary-of-training-content-for-nvidia-nemotron.pdf)
|
| 621 |
-
|
| 622 |
<details>
|
| 623 |
<summary><strong>Click to explore the full dataset catalogue used for training</strong></summary>
|
| 624 |
|
|
|
|
| 69 |
| **Supported Languages** | English, French, German, Italian, Japanese, Spanish, Chinese |
|
| 70 |
| **Best For** | Agentic workflows, long-context reasoning, high-volume workloads (e.g. IT ticket automation), tool use, RAG |
|
| 71 |
| **Reasoning Mode** | Configurable on/off via chat template (`enable_thinking=True/False`) |
|
|
|
|
| 72 |
| **License** | [NVIDIA Nemotron Open Model License](https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-nemotron-open-model-license/) |
|
| 73 |
| **Release Date** | March 11, 2026 |
|
| 74 |
|
|
|
|
| 271 |
|
| 272 |
##### vLLM on DGX Spark
|
| 273 |
|
| 274 |
+
To deploy the NVFP4 chekpoint on NVIDIA DGX Spark, make sure that you are using the `vllm/vllm-openai:v0.20.0` container image and use the following command:
|
| 275 |
|
| 276 |
```bash
|
| 277 |
docker run --rm -it --gpus all \
|
|
|
|
| 283 |
-v ~/.cache/huggingface:/root/.cache/huggingface \
|
| 284 |
-v $(pwd)/super_v3_reasoning_parser.py:/app/super_v3_reasoning_parser.py \
|
| 285 |
-p 8000:8000 \
|
| 286 |
+
vllm/vllm-openai:v0.20.0 \
|
| 287 |
--model nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4 \
|
| 288 |
--served-model-name nvidia/nemotron-3-super \
|
| 289 |
--host 0.0.0.0 \
|
|
|
|
| 302 |
--moe-backend marlin \
|
| 303 |
--mamba_ssm_cache_dtype float16 \
|
| 304 |
--quantization fp4 \
|
| 305 |
+
--speculative_config '{"method":"mtp","num_speculative_tokens":3,"moe_backend":"triton"}' \
|
| 306 |
--reasoning-parser-plugin /app/super_v3_reasoning_parser.py \
|
| 307 |
--reasoning-parser super_v3 \
|
| 308 |
--enable-auto-tool-choice \
|
|
|
|
| 616 |
|
| 617 |
More details on the datasets and synthetic data generation methods can be found in the technical report _[**_NVIDIA Nemotron 3 Super_**](https://research.nvidia.com/labs/nemotron/files/NVIDIA-Nemotron-3-Super-Technical-Report.pdf)_.
|
| 618 |
|
|
|
|
|
|
|
| 619 |
<details>
|
| 620 |
<summary><strong>Click to explore the full dataset catalogue used for training</strong></summary>
|
| 621 |
|