Text Generation
Transformers
Safetensors
PyTorch
nemotron_h
nvidia
nemotron-3
latent-moe
mtp
conversational
custom_code
8-bit precision
modelopt
Files changed (1) hide show
  1. README.md +3 -6
README.md CHANGED
@@ -69,7 +69,6 @@ track_downloads: true
69
  | **Supported Languages** | English, French, German, Italian, Japanese, Spanish, Chinese |
70
  | **Best For** | Agentic workflows, long-context reasoning, high-volume workloads (e.g. IT ticket automation), tool use, RAG |
71
  | **Reasoning Mode** | Configurable on/off via chat template (`enable_thinking=True/False`) |
72
- | **Speculative Decoding** | Includes a built-in MTP head, with an updated MTPv2 head available as a separate [checkpoint](https://huggingface.co/nvidia/Nemotron-3-Super-120B-A12B-BF16-MTPv2) |
73
  | **License** | [NVIDIA Nemotron Open Model License](https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-nemotron-open-model-license/) |
74
  | **Release Date** | March 11, 2026 |
75
 
@@ -272,7 +271,7 @@ vllm serve $MODEL_CKPT \
272
 
273
  ##### vLLM on DGX Spark
274
 
275
- To deploy the NVFP4 chekpoint on NVIDIA DGX Spark, make sure that you are using the `vllm/vllm-openai:v0.27.1` container image and use the following command:
276
 
277
  ```bash
278
  docker run --rm -it --gpus all \
@@ -284,7 +283,7 @@ docker run --rm -it --gpus all \
284
  -v ~/.cache/huggingface:/root/.cache/huggingface \
285
  -v $(pwd)/super_v3_reasoning_parser.py:/app/super_v3_reasoning_parser.py \
286
  -p 8000:8000 \
287
- vllm/vllm-openai:v0.27.1 \
288
  --model nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4 \
289
  --served-model-name nvidia/nemotron-3-super \
290
  --host 0.0.0.0 \
@@ -303,7 +302,7 @@ docker run --rm -it --gpus all \
303
  --moe-backend marlin \
304
  --mamba_ssm_cache_dtype float16 \
305
  --quantization fp4 \
306
- --speculative_config '{"method":"mtp","num_speculative_tokens":3,"model":"nvidia/Nemotron-3-Super-120B-A12B-BF16-MTPv2","moe_backend":"triton"}' \
307
  --reasoning-parser-plugin /app/super_v3_reasoning_parser.py \
308
  --reasoning-parser super_v3 \
309
  --enable-auto-tool-choice \
@@ -617,8 +616,6 @@ Alongside the model, we release our final pre-training and post-training data, a
617
 
618
  More details on the datasets and synthetic data generation methods can be found in the technical report _[**_NVIDIA Nemotron 3 Super_**](https://research.nvidia.com/labs/nemotron/files/NVIDIA-Nemotron-3-Super-Technical-Report.pdf)_.
619
 
620
- For more information about the datasets used to train this model, please see the [Public Summary of Training Content](https://developer.download.nvidia.com/assets/nemo/docs/public-summary-of-training-content-for-nvidia-nemotron.pdf)
621
-
622
  <details>
623
  <summary><strong>Click to explore the full dataset catalogue used for training</strong></summary>
624
 
 
69
  | **Supported Languages** | English, French, German, Italian, Japanese, Spanish, Chinese |
70
  | **Best For** | Agentic workflows, long-context reasoning, high-volume workloads (e.g. IT ticket automation), tool use, RAG |
71
  | **Reasoning Mode** | Configurable on/off via chat template (`enable_thinking=True/False`) |
 
72
  | **License** | [NVIDIA Nemotron Open Model License](https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-nemotron-open-model-license/) |
73
  | **Release Date** | March 11, 2026 |
74
 
 
271
 
272
  ##### vLLM on DGX Spark
273
 
274
+ To deploy the NVFP4 chekpoint on NVIDIA DGX Spark, make sure that you are using the `vllm/vllm-openai:v0.20.0` container image and use the following command:
275
 
276
  ```bash
277
  docker run --rm -it --gpus all \
 
283
  -v ~/.cache/huggingface:/root/.cache/huggingface \
284
  -v $(pwd)/super_v3_reasoning_parser.py:/app/super_v3_reasoning_parser.py \
285
  -p 8000:8000 \
286
+ vllm/vllm-openai:v0.20.0 \
287
  --model nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4 \
288
  --served-model-name nvidia/nemotron-3-super \
289
  --host 0.0.0.0 \
 
302
  --moe-backend marlin \
303
  --mamba_ssm_cache_dtype float16 \
304
  --quantization fp4 \
305
+ --speculative_config '{"method":"mtp","num_speculative_tokens":3,"moe_backend":"triton"}' \
306
  --reasoning-parser-plugin /app/super_v3_reasoning_parser.py \
307
  --reasoning-parser super_v3 \
308
  --enable-auto-tool-choice \
 
616
 
617
  More details on the datasets and synthetic data generation methods can be found in the technical report _[**_NVIDIA Nemotron 3 Super_**](https://research.nvidia.com/labs/nemotron/files/NVIDIA-Nemotron-3-Super-Technical-Report.pdf)_.
618
 
 
 
619
  <details>
620
  <summary><strong>Click to explore the full dataset catalogue used for training</strong></summary>
621