Any-to-Any
Transformers
Safetensors
PyTorch
NemotronH_Nano_Omni_Reasoning_V3
feature-extraction
nvidia
multimodal
custom_code
8-bit precision
modelopt
Instructions to use nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-NVFP4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-NVFP4 with Transformers:
# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-NVFP4", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
Update model card
Browse filesApply BF16 model card updates while preserving base_model metadata.
README.md
CHANGED
|
@@ -1,16 +1,46 @@
|
|
| 1 |
---
|
| 2 |
base_model:
|
| 3 |
- nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16
|
|
|
|
| 4 |
license: other
|
| 5 |
license_name: nvidia-open-model-agreement
|
| 6 |
license_link: >-
|
| 7 |
https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-open-model-agreement/
|
|
|
|
| 8 |
tags:
|
| 9 |
- nvidia
|
|
|
|
| 10 |
- multimodal
|
| 11 |
-
|
|
|
|
|
|
|
| 12 |
---
|
| 13 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 14 |
# Model Overview
|
| 15 |
|
| 16 |
### Description:
|
|
@@ -49,10 +79,9 @@ NGC 04/28/2026 via [URL](https://catalog.ngc.nvidia.com/orgs/nim/teams/nvidia/c
|
|
| 49 |
**Architecture Type:** Mamba2-Transformer Hybrid Mixture of Experts (MoE) <br>
|
| 50 |
|
| 51 |
**Network Architecture:**
|
| 52 |
-
- [Nemotron 3 Nano LLM (30B A3B)](https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16)
|
| 53 |
-
- CRADIO v4-H vision encoder
|
| 54 |
-
- Parakeet speech encoder
|
| 55 |
-
|
| 56 |
|
| 57 |
**Number of model parameters:** 3.1 x 10^10 (31B A3B) <br>
|
| 58 |
|
|
@@ -130,15 +159,6 @@ Nemotron-3-Nano-Omni-30B-A3B-Reasoning <br>
|
|
| 130 |
|
| 131 |
---
|
| 132 |
|
| 133 |
-
## Quick Start Guide
|
| 134 |
-
|
| 135 |
-
### Model Parameters
|
| 136 |
-
|
| 137 |
-
| Mode | temperature | top_p | top_k | max_tokens | reasoning_budget | grace_period |
|
| 138 |
-
|------|-------------|-------|-------|------------|------------------|--------------|
|
| 139 |
-
| **Thinking mode** | 0.6 | 0.95 | — | 20480 | 16384 | 1024 |
|
| 140 |
-
| **Instruct mode** | 0.2 | — | 1 | 1024 | — | — |
|
| 141 |
-
|
| 142 |
### Download Model Weights
|
| 143 |
|
| 144 |
| Precision | Technical Name | HuggingFace URL |
|
|
@@ -159,7 +179,7 @@ hf auth login
|
|
| 159 |
hf auth whoami
|
| 160 |
```
|
| 161 |
|
| 162 |
-
#### Download the weights
|
| 163 |
|
| 164 |
Pick a target directory on a volume with ≥70 GB free (the model is ~62 GB).
|
| 165 |
|
|
@@ -183,7 +203,7 @@ hf download nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16 \
|
|
| 183 |
ls "$WEIGHTS" | head
|
| 184 |
du -sh "$WEIGHTS" # expect ~62 GB
|
| 185 |
test -f "$WEIGHTS/config.json" && echo OK
|
| 186 |
-
```
|
| 187 |
|
| 188 |
---
|
| 189 |
|
|
@@ -225,6 +245,7 @@ vllm serve nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-NVFP4 \
|
|
| 225 |
--tool-call-parser qwen3_coder \
|
| 226 |
--kv-cache-dtype fp8 # Omit this for BF16
|
| 227 |
```
|
|
|
|
| 228 |
|
| 229 |
#### Platform-Specific Notes
|
| 230 |
|
|
@@ -424,7 +445,7 @@ response = client.chat.completions.create(
|
|
| 424 |
}
|
| 425 |
],
|
| 426 |
max_tokens=1024,
|
| 427 |
-
temperature=
|
| 428 |
extra_body={"top_k": 1, "chat_template_kwargs": {"enable_thinking": False}},
|
| 429 |
)
|
| 430 |
print(response.choices[0].message.content)
|
|
@@ -449,7 +470,7 @@ response = client.chat.completions.create(
|
|
| 449 |
}
|
| 450 |
],
|
| 451 |
max_tokens=1024,
|
| 452 |
-
temperature=
|
| 453 |
extra_body={"top_k": 1, "chat_template_kwargs": {"enable_thinking": False}},
|
| 454 |
)
|
| 455 |
print(response.choices[0].message.content)
|
|
@@ -496,7 +517,7 @@ print(response.choices[0].message.content)
|
|
| 496 |
```bash
|
| 497 |
curl -sS http://localhost:8000/v1/chat/completions \
|
| 498 |
-H "Content-Type: application/json" \
|
| 499 |
-
-d '{"model":"nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-NVFP4","messages":[{"role":"user","content":"Hello, what can you do?"}],"temperature":
|
| 500 |
| python3 -c "import sys,json; print(json.load(sys.stdin)['choices'][0]['message']['content'])"
|
| 501 |
```
|
| 502 |
|
|
@@ -556,7 +577,7 @@ def chat(url, model, b64, text, max_tokens):
|
|
| 556 |
]}],
|
| 557 |
"max_tokens": max_tokens,
|
| 558 |
"stream": False,
|
| 559 |
-
"temperature":
|
| 560 |
"chat_template_kwargs": {"enable_thinking": False},
|
| 561 |
}, timeout=120)
|
| 562 |
r.raise_for_status()
|
|
@@ -738,7 +759,6 @@ Higher values improve temporal coverage but increase VRAM and prefill time. Star
|
|
| 738 |
|
| 739 |
### Notes
|
| 740 |
|
| 741 |
-
|
| 742 |
1. **Reasoning default:** Reasoning is on by default. If you omit `chat_template_kwargs`, the model will produce chain-of-thought traces in `content`. This is appropriate for text and image inputs.
|
| 743 |
2. **Video frame sampling:** The default (~32 frames) is too conservative for most real videos. Set `--media-io-kwargs` at server launch.
|
| 744 |
3. **PDF input format:** The API does not accept raw PDF uploads. Render pages to PNG and send as base64 (see PDF Example above).
|
|
@@ -1019,14 +1039,14 @@ We recommend following settings for reaching the optimal performance.
|
|
| 1019 |
### Sampling Parameters
|
| 1020 |
We suggest the following sampling parameters based on the mode and tasks.
|
| 1021 |
* Thinking mode for long document analysis and multimodal reasoning tasks: <br>
|
| 1022 |
-
`temperature=0.
|
| 1023 |
* Instruct mode (non-thinking) for general tasks:<br>
|
| 1024 |
`temperature=0.2`, `top_k=1`<br>
|
| 1025 |
-
* For ASR tasks, we recommend non-thinking mode with
|
| 1026 |
-
`temperature=
|
| 1027 |
|
| 1028 |
### Model output length
|
| 1029 |
-
For most multimodel reasoning tasks, we recommend using output length of at least 20480. For complex reasoning questions especially in math and programing increasing the maximum output length to
|
| 1030 |
|
| 1031 |
## Ethical Considerations:
|
| 1032 |
NVIDIA believes Trustworthy AI is a shared responsibility and we have established policies and practices to enable development for a wide array of AI applications. When downloaded or used in accordance with our terms of service, developers should work with their internal model team to ensure this model meets requirements for the relevant industry and use case and addresses unforeseen product misuse. <br>
|
|
@@ -1040,12 +1060,12 @@ Please report model quality, risk, security vulnerabilities or NVIDIA AI Concern
|
|
| 1040 |
# Citation:
|
| 1041 |
```
|
| 1042 |
@misc{nvidia2026nemotron3nanoomni,
|
| 1043 |
-
|
| 1044 |
-
|
| 1045 |
-
|
| 1046 |
-
|
| 1047 |
-
|
| 1048 |
-
|
| 1049 |
-
|
| 1050 |
}
|
| 1051 |
```
|
|
|
|
| 1 |
---
|
| 2 |
base_model:
|
| 3 |
- nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16
|
| 4 |
+
library_name: transformers
|
| 5 |
license: other
|
| 6 |
license_name: nvidia-open-model-agreement
|
| 7 |
license_link: >-
|
| 8 |
https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-open-model-agreement/
|
| 9 |
+
pipeline_tag: any-to-any
|
| 10 |
tags:
|
| 11 |
- nvidia
|
| 12 |
+
- pytorch
|
| 13 |
- multimodal
|
| 14 |
+
datasets:
|
| 15 |
+
- nvidia/Nemotron-Image-Training-v3
|
| 16 |
+
track_downloads: true
|
| 17 |
---
|
| 18 |
|
| 19 |
+
## At a Glance
|
| 20 |
+
|
| 21 |
+
| | |
|
| 22 |
+
|---|---|
|
| 23 |
+
| **Total parameters** | 31B (Mamba2-Transformer hybrid MoE) |
|
| 24 |
+
| **Active parameters** | ~3B per token |
|
| 25 |
+
| **Max context** | 256k tokens |
|
| 26 |
+
| **Modalities (in)** | Video, Audio, Image, Text |
|
| 27 |
+
| **Modality (out)** | Text |
|
| 28 |
+
| **Reasoning mode** | On by default; toggle via `enable_thinking` |
|
| 29 |
+
| **Best for** | Video+speech analysis, document intelligence (OCR/charts/long docs), GUI/agentic workflows, ASR |
|
| 30 |
+
| **Minimum GPU (BF16)** | 1× H100 80GB (single-GPU); 1× B200 / 1× H200 recommended |
|
| 31 |
+
| **Minimum GPU (FP8)** | 1× L40S 48GB; 1× RTX Pro 6000 / 1× B200 recommended |
|
| 32 |
+
| **Minimum GPU (NVFP4)** | 1× RTX 5090 32GB; 1× DGX Spark / 1× Jetson Thor also supported |
|
| 33 |
+
| **Precisions** | [BF16](https://huggingface.co/nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16) (62 GB) · [FP8](https://huggingface.co/nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-FP8) (33 GB) · [NVFP4](https://huggingface.co/nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-NVFP4) (21 GB) |
|
| 34 |
+
|
| 35 |
+
## Quick Start Guide
|
| 36 |
+
|
| 37 |
+
### Model Parameters
|
| 38 |
+
|
| 39 |
+
| Mode | temperature | top_p | top_k | max_tokens | reasoning_budget | grace_period |
|
| 40 |
+
|------|-------------|-------|-------|------------|------------------|--------------|
|
| 41 |
+
| **Thinking mode** | 0.6 | 0.95 | — | 20480 | 16384 | 1024 |
|
| 42 |
+
| **Instruct mode** | 0.2 | — | 1 | 1024 | — | — |
|
| 43 |
+
|
| 44 |
# Model Overview
|
| 45 |
|
| 46 |
### Description:
|
|
|
|
| 79 |
**Architecture Type:** Mamba2-Transformer Hybrid Mixture of Experts (MoE) <br>
|
| 80 |
|
| 81 |
**Network Architecture:**
|
| 82 |
+
- [Nemotron 3 Nano LLM (30B A3B)](https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16) — 31B-parameter Mamba2-Transformer hybrid MoE backbone with ~3B active parameters per token.
|
| 83 |
+
- [CRADIO v4-H](https://huggingface.co/nvidia/C-RADIOv4-H) — vision encoder for image and video frames.
|
| 84 |
+
- [Parakeet](https://huggingface.co/nvidia/parakeet-tdt-0.6b-v2) — speech encoder for audio inputs.
|
|
|
|
| 85 |
|
| 86 |
**Number of model parameters:** 3.1 x 10^10 (31B A3B) <br>
|
| 87 |
|
|
|
|
| 159 |
|
| 160 |
---
|
| 161 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 162 |
### Download Model Weights
|
| 163 |
|
| 164 |
| Precision | Technical Name | HuggingFace URL |
|
|
|
|
| 179 |
hf auth whoami
|
| 180 |
```
|
| 181 |
|
| 182 |
+
<!-- #### Download the weights
|
| 183 |
|
| 184 |
Pick a target directory on a volume with ≥70 GB free (the model is ~62 GB).
|
| 185 |
|
|
|
|
| 203 |
ls "$WEIGHTS" | head
|
| 204 |
du -sh "$WEIGHTS" # expect ~62 GB
|
| 205 |
test -f "$WEIGHTS/config.json" && echo OK
|
| 206 |
+
``` -->
|
| 207 |
|
| 208 |
---
|
| 209 |
|
|
|
|
| 245 |
--tool-call-parser qwen3_coder \
|
| 246 |
--kv-cache-dtype fp8 # Omit this for BF16
|
| 247 |
```
|
| 248 |
+
Efficient Video Sampling: video-pruning-rate=0.5 drops 50% of redundant video tokens; halves video-prefill VRAM/TTFT.
|
| 249 |
|
| 250 |
#### Platform-Specific Notes
|
| 251 |
|
|
|
|
| 445 |
}
|
| 446 |
],
|
| 447 |
max_tokens=1024,
|
| 448 |
+
temperature=0.2,
|
| 449 |
extra_body={"top_k": 1, "chat_template_kwargs": {"enable_thinking": False}},
|
| 450 |
)
|
| 451 |
print(response.choices[0].message.content)
|
|
|
|
| 470 |
}
|
| 471 |
],
|
| 472 |
max_tokens=1024,
|
| 473 |
+
temperature=0.2,
|
| 474 |
extra_body={"top_k": 1, "chat_template_kwargs": {"enable_thinking": False}},
|
| 475 |
)
|
| 476 |
print(response.choices[0].message.content)
|
|
|
|
| 517 |
```bash
|
| 518 |
curl -sS http://localhost:8000/v1/chat/completions \
|
| 519 |
-H "Content-Type: application/json" \
|
| 520 |
+
-d '{"model":"nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-NVFP4","messages":[{"role":"user","content":"Hello, what can you do?"}],"temperature":0.2,"top_k":1}' \
|
| 521 |
| python3 -c "import sys,json; print(json.load(sys.stdin)['choices'][0]['message']['content'])"
|
| 522 |
```
|
| 523 |
|
|
|
|
| 577 |
]}],
|
| 578 |
"max_tokens": max_tokens,
|
| 579 |
"stream": False,
|
| 580 |
+
"temperature": 0.2,
|
| 581 |
"chat_template_kwargs": {"enable_thinking": False},
|
| 582 |
}, timeout=120)
|
| 583 |
r.raise_for_status()
|
|
|
|
| 759 |
|
| 760 |
### Notes
|
| 761 |
|
|
|
|
| 762 |
1. **Reasoning default:** Reasoning is on by default. If you omit `chat_template_kwargs`, the model will produce chain-of-thought traces in `content`. This is appropriate for text and image inputs.
|
| 763 |
2. **Video frame sampling:** The default (~32 frames) is too conservative for most real videos. Set `--media-io-kwargs` at server launch.
|
| 764 |
3. **PDF input format:** The API does not accept raw PDF uploads. Render pages to PNG and send as base64 (see PDF Example above).
|
|
|
|
| 1039 |
### Sampling Parameters
|
| 1040 |
We suggest the following sampling parameters based on the mode and tasks.
|
| 1041 |
* Thinking mode for long document analysis and multimodal reasoning tasks: <br>
|
| 1042 |
+
`temperature=0.6`, `top_p=0.95`, `grace_period=1024`, `reasoning_budget=16384`, `max_token=20480`, and `max_model_len=210000`<br>
|
| 1043 |
* Instruct mode (non-thinking) for general tasks:<br>
|
| 1044 |
`temperature=0.2`, `top_k=1`<br>
|
| 1045 |
+
* For ASR tasks, we recommend non-thinking mode with <br>
|
| 1046 |
+
`temperature=1.0`, `top_k=1`<br>
|
| 1047 |
|
| 1048 |
### Model output length
|
| 1049 |
+
For most multimodel reasoning tasks, we recommend using output length of at least 20480. For complex reasoning questions especially in math and programing increasing the maximum output length to 210000 tokens can give the model enough room to produce more detailed and correct answers. We also found the proposed Budget-Controlled Reasoning effectiveness in answering complex reasoning questions.
|
| 1050 |
|
| 1051 |
## Ethical Considerations:
|
| 1052 |
NVIDIA believes Trustworthy AI is a shared responsibility and we have established policies and practices to enable development for a wide array of AI applications. When downloaded or used in accordance with our terms of service, developers should work with their internal model team to ensure this model meets requirements for the relevant industry and use case and addresses unforeseen product misuse. <br>
|
|
|
|
| 1060 |
# Citation:
|
| 1061 |
```
|
| 1062 |
@misc{nvidia2026nemotron3nanoomni,
|
| 1063 |
+
title={Nemotron 3 Nano Omni: Efficient and Open Multimodal Intelligence},
|
| 1064 |
+
author={NVIDIA},
|
| 1065 |
+
year={2026},
|
| 1066 |
+
eprint={2604.24954},
|
| 1067 |
+
archivePrefix={arXiv},
|
| 1068 |
+
primaryClass={cs.LG},
|
| 1069 |
+
url={https://arxiv.org/abs/2604.24954},
|
| 1070 |
}
|
| 1071 |
```
|