cvoegele-nv commited on
Commit
75ebf31
·
verified ·
1 Parent(s): 76955e4

Update model card

Browse files

Apply BF16 model card updates while preserving base_model metadata.

Files changed (1) hide show
  1. README.md +52 -32
README.md CHANGED
@@ -1,16 +1,46 @@
1
  ---
2
  base_model:
3
  - nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16
 
4
  license: other
5
  license_name: nvidia-open-model-agreement
6
  license_link: >-
7
  https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-open-model-agreement/
 
8
  tags:
9
  - nvidia
 
10
  - multimodal
11
- pipeline_tag: any-to-any
 
 
12
  ---
13
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
14
  # Model Overview
15
 
16
  ### Description:
@@ -49,10 +79,9 @@ NGC 04/28/2026 via [URL](https://catalog.ngc.nvidia.com/orgs/nim/teams/nvidia/c
49
  **Architecture Type:** Mamba2-Transformer Hybrid Mixture of Experts (MoE) <br>
50
 
51
  **Network Architecture:**
52
- - [Nemotron 3 Nano LLM (30B A3B)](https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16)
53
- - CRADIO v4-H vision encoder
54
- - Parakeet speech encoder
55
-
56
 
57
  **Number of model parameters:** 3.1 x 10^10 (31B A3B) <br>
58
 
@@ -130,15 +159,6 @@ Nemotron-3-Nano-Omni-30B-A3B-Reasoning <br>
130
 
131
  ---
132
 
133
- ## Quick Start Guide
134
-
135
- ### Model Parameters
136
-
137
- | Mode | temperature | top_p | top_k | max_tokens | reasoning_budget | grace_period |
138
- |------|-------------|-------|-------|------------|------------------|--------------|
139
- | **Thinking mode** | 0.6 | 0.95 | — | 20480 | 16384 | 1024 |
140
- | **Instruct mode** | 0.2 | — | 1 | 1024 | — | — |
141
-
142
  ### Download Model Weights
143
 
144
  | Precision | Technical Name | HuggingFace URL |
@@ -159,7 +179,7 @@ hf auth login
159
  hf auth whoami
160
  ```
161
 
162
- #### Download the weights
163
 
164
  Pick a target directory on a volume with ≥70 GB free (the model is ~62 GB).
165
 
@@ -183,7 +203,7 @@ hf download nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16 \
183
  ls "$WEIGHTS" | head
184
  du -sh "$WEIGHTS" # expect ~62 GB
185
  test -f "$WEIGHTS/config.json" && echo OK
186
- ```
187
 
188
  ---
189
 
@@ -225,6 +245,7 @@ vllm serve nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-NVFP4 \
225
  --tool-call-parser qwen3_coder \
226
  --kv-cache-dtype fp8 # Omit this for BF16
227
  ```
 
228
 
229
  #### Platform-Specific Notes
230
 
@@ -424,7 +445,7 @@ response = client.chat.completions.create(
424
  }
425
  ],
426
  max_tokens=1024,
427
- temperature=1.0,
428
  extra_body={"top_k": 1, "chat_template_kwargs": {"enable_thinking": False}},
429
  )
430
  print(response.choices[0].message.content)
@@ -449,7 +470,7 @@ response = client.chat.completions.create(
449
  }
450
  ],
451
  max_tokens=1024,
452
- temperature=1.0,
453
  extra_body={"top_k": 1, "chat_template_kwargs": {"enable_thinking": False}},
454
  )
455
  print(response.choices[0].message.content)
@@ -496,7 +517,7 @@ print(response.choices[0].message.content)
496
  ```bash
497
  curl -sS http://localhost:8000/v1/chat/completions \
498
  -H "Content-Type: application/json" \
499
- -d '{"model":"nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-NVFP4","messages":[{"role":"user","content":"Hello, what can you do?"}],"temperature":1.0,"top_k":1}' \
500
  | python3 -c "import sys,json; print(json.load(sys.stdin)['choices'][0]['message']['content'])"
501
  ```
502
 
@@ -556,7 +577,7 @@ def chat(url, model, b64, text, max_tokens):
556
  ]}],
557
  "max_tokens": max_tokens,
558
  "stream": False,
559
- "temperature": 1.0,
560
  "chat_template_kwargs": {"enable_thinking": False},
561
  }, timeout=120)
562
  r.raise_for_status()
@@ -738,7 +759,6 @@ Higher values improve temporal coverage but increase VRAM and prefill time. Star
738
 
739
  ### Notes
740
 
741
-
742
  1. **Reasoning default:** Reasoning is on by default. If you omit `chat_template_kwargs`, the model will produce chain-of-thought traces in `content`. This is appropriate for text and image inputs.
743
  2. **Video frame sampling:** The default (~32 frames) is too conservative for most real videos. Set `--media-io-kwargs` at server launch.
744
  3. **PDF input format:** The API does not accept raw PDF uploads. Render pages to PNG and send as base64 (see PDF Example above).
@@ -1019,14 +1039,14 @@ We recommend following settings for reaching the optimal performance.
1019
  ### Sampling Parameters
1020
  We suggest the following sampling parameters based on the mode and tasks.
1021
  * Thinking mode for long document analysis and multimodal reasoning tasks: <br>
1022
- `temperature=0.5-0.7`, `top_p=0.95`, `grace_period=1024`, `reasoning_budget=16384`, `max_token=20480`, and `max_model_len=210000`<br>
1023
  * Instruct mode (non-thinking) for general tasks:<br>
1024
  `temperature=0.2`, `top_k=1`<br>
1025
- * For ASR tasks, we recommend non-thinking mode with
1026
- `temperature=0.2`, `top_k=1`<br>
1027
 
1028
  ### Model output length
1029
- For most multimodel reasoning tasks, we recommend using output length of at least 20480. For complex reasoning questions especially in math and programing increasing the maximum output length to 131072 tokens can give the model enough room to produce more detailed and correct answers. We also found the proposed Budget-Controlled Reasoning effectiveness in answering complex reasoning questions.
1030
 
1031
  ## Ethical Considerations:
1032
  NVIDIA believes Trustworthy AI is a shared responsibility and we have established policies and practices to enable development for a wide array of AI applications. When downloaded or used in accordance with our terms of service, developers should work with their internal model team to ensure this model meets requirements for the relevant industry and use case and addresses unforeseen product misuse. <br>
@@ -1040,12 +1060,12 @@ Please report model quality, risk, security vulnerabilities or NVIDIA AI Concern
1040
  # Citation:
1041
  ```
1042
  @misc{nvidia2026nemotron3nanoomni,
1043
-       title={Nemotron 3 Nano Omni: Efficient and Open Multimodal Intelligence},
1044
-       author={NVIDIA},
1045
-       year={2026},
1046
-       eprint={2604.24954},
1047
-       archivePrefix={arXiv},
1048
-       primaryClass={cs.LG},
1049
-       url={https://arxiv.org/abs/2604.24954},
1050
  }
1051
  ```
 
1
  ---
2
  base_model:
3
  - nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16
4
+ library_name: transformers
5
  license: other
6
  license_name: nvidia-open-model-agreement
7
  license_link: >-
8
  https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-open-model-agreement/
9
+ pipeline_tag: any-to-any
10
  tags:
11
  - nvidia
12
+ - pytorch
13
  - multimodal
14
+ datasets:
15
+ - nvidia/Nemotron-Image-Training-v3
16
+ track_downloads: true
17
  ---
18
 
19
+ ## At a Glance
20
+
21
+ | | |
22
+ |---|---|
23
+ | **Total parameters** | 31B (Mamba2-Transformer hybrid MoE) |
24
+ | **Active parameters** | ~3B per token |
25
+ | **Max context** | 256k tokens |
26
+ | **Modalities (in)** | Video, Audio, Image, Text |
27
+ | **Modality (out)** | Text |
28
+ | **Reasoning mode** | On by default; toggle via `enable_thinking` |
29
+ | **Best for** | Video+speech analysis, document intelligence (OCR/charts/long docs), GUI/agentic workflows, ASR |
30
+ | **Minimum GPU (BF16)** | 1× H100 80GB (single-GPU); 1× B200 / 1× H200 recommended |
31
+ | **Minimum GPU (FP8)** | 1× L40S 48GB; 1× RTX Pro 6000 / 1× B200 recommended |
32
+ | **Minimum GPU (NVFP4)** | 1× RTX 5090 32GB; 1× DGX Spark / 1× Jetson Thor also supported |
33
+ | **Precisions** | [BF16](https://huggingface.co/nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16) (62 GB) · [FP8](https://huggingface.co/nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-FP8) (33 GB) · [NVFP4](https://huggingface.co/nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-NVFP4) (21 GB) |
34
+
35
+ ## Quick Start Guide
36
+
37
+ ### Model Parameters
38
+
39
+ | Mode | temperature | top_p | top_k | max_tokens | reasoning_budget | grace_period |
40
+ |------|-------------|-------|-------|------------|------------------|--------------|
41
+ | **Thinking mode** | 0.6 | 0.95 | — | 20480 | 16384 | 1024 |
42
+ | **Instruct mode** | 0.2 | — | 1 | 1024 | — | — |
43
+
44
  # Model Overview
45
 
46
  ### Description:
 
79
  **Architecture Type:** Mamba2-Transformer Hybrid Mixture of Experts (MoE) <br>
80
 
81
  **Network Architecture:**
82
+ - [Nemotron 3 Nano LLM (30B A3B)](https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16) — 31B-parameter Mamba2-Transformer hybrid MoE backbone with ~3B active parameters per token.
83
+ - [CRADIO v4-H](https://huggingface.co/nvidia/C-RADIOv4-H) — vision encoder for image and video frames.
84
+ - [Parakeet](https://huggingface.co/nvidia/parakeet-tdt-0.6b-v2) — speech encoder for audio inputs.
 
85
 
86
  **Number of model parameters:** 3.1 x 10^10 (31B A3B) <br>
87
 
 
159
 
160
  ---
161
 
 
 
 
 
 
 
 
 
 
162
  ### Download Model Weights
163
 
164
  | Precision | Technical Name | HuggingFace URL |
 
179
  hf auth whoami
180
  ```
181
 
182
+ <!-- #### Download the weights
183
 
184
  Pick a target directory on a volume with ≥70 GB free (the model is ~62 GB).
185
 
 
203
  ls "$WEIGHTS" | head
204
  du -sh "$WEIGHTS" # expect ~62 GB
205
  test -f "$WEIGHTS/config.json" && echo OK
206
+ ``` -->
207
 
208
  ---
209
 
 
245
  --tool-call-parser qwen3_coder \
246
  --kv-cache-dtype fp8 # Omit this for BF16
247
  ```
248
+ Efficient Video Sampling: video-pruning-rate=0.5 drops 50% of redundant video tokens; halves video-prefill VRAM/TTFT.
249
 
250
  #### Platform-Specific Notes
251
 
 
445
  }
446
  ],
447
  max_tokens=1024,
448
+ temperature=0.2,
449
  extra_body={"top_k": 1, "chat_template_kwargs": {"enable_thinking": False}},
450
  )
451
  print(response.choices[0].message.content)
 
470
  }
471
  ],
472
  max_tokens=1024,
473
+ temperature=0.2,
474
  extra_body={"top_k": 1, "chat_template_kwargs": {"enable_thinking": False}},
475
  )
476
  print(response.choices[0].message.content)
 
517
  ```bash
518
  curl -sS http://localhost:8000/v1/chat/completions \
519
  -H "Content-Type: application/json" \
520
+ -d '{"model":"nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-NVFP4","messages":[{"role":"user","content":"Hello, what can you do?"}],"temperature":0.2,"top_k":1}' \
521
  | python3 -c "import sys,json; print(json.load(sys.stdin)['choices'][0]['message']['content'])"
522
  ```
523
 
 
577
  ]}],
578
  "max_tokens": max_tokens,
579
  "stream": False,
580
+ "temperature": 0.2,
581
  "chat_template_kwargs": {"enable_thinking": False},
582
  }, timeout=120)
583
  r.raise_for_status()
 
759
 
760
  ### Notes
761
 
 
762
  1. **Reasoning default:** Reasoning is on by default. If you omit `chat_template_kwargs`, the model will produce chain-of-thought traces in `content`. This is appropriate for text and image inputs.
763
  2. **Video frame sampling:** The default (~32 frames) is too conservative for most real videos. Set `--media-io-kwargs` at server launch.
764
  3. **PDF input format:** The API does not accept raw PDF uploads. Render pages to PNG and send as base64 (see PDF Example above).
 
1039
  ### Sampling Parameters
1040
  We suggest the following sampling parameters based on the mode and tasks.
1041
  * Thinking mode for long document analysis and multimodal reasoning tasks: <br>
1042
+ `temperature=0.6`, `top_p=0.95`, `grace_period=1024`, `reasoning_budget=16384`, `max_token=20480`, and `max_model_len=210000`<br>
1043
  * Instruct mode (non-thinking) for general tasks:<br>
1044
  `temperature=0.2`, `top_k=1`<br>
1045
+ * For ASR tasks, we recommend non-thinking mode with <br>
1046
+ `temperature=1.0`, `top_k=1`<br>
1047
 
1048
  ### Model output length
1049
+ For most multimodel reasoning tasks, we recommend using output length of at least 20480. For complex reasoning questions especially in math and programing increasing the maximum output length to 210000 tokens can give the model enough room to produce more detailed and correct answers. We also found the proposed Budget-Controlled Reasoning effectiveness in answering complex reasoning questions.
1050
 
1051
  ## Ethical Considerations:
1052
  NVIDIA believes Trustworthy AI is a shared responsibility and we have established policies and practices to enable development for a wide array of AI applications. When downloaded or used in accordance with our terms of service, developers should work with their internal model team to ensure this model meets requirements for the relevant industry and use case and addresses unforeseen product misuse. <br>
 
1060
  # Citation:
1061
  ```
1062
  @misc{nvidia2026nemotron3nanoomni,
1063
+ title={Nemotron 3 Nano Omni: Efficient and Open Multimodal Intelligence},
1064
+ author={NVIDIA},
1065
+ year={2026},
1066
+ eprint={2604.24954},
1067
+ archivePrefix={arXiv},
1068
+ primaryClass={cs.LG},
1069
+ url={https://arxiv.org/abs/2604.24954},
1070
  }
1071
  ```