Arsh9210 commited on
Commit
432c86c
·
verified ·
1 Parent(s): 12ed9f5

Added README.md

Browse files
Files changed (1) hide show
  1. README.md +556 -0
README.md ADDED
@@ -0,0 +1,556 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: other
3
+ license_name: openmdw-1.1
4
+ license_link: >-
5
+ https://openmdw.ai/license/1-1/
6
+ base_model:
7
+ - nvidia/Nemotron-3-Embed-1B-BF16
8
+ tags:
9
+ - text
10
+ - text-embeddings
11
+ - retrieval
12
+ - semantic-search
13
+ - vllm
14
+ - nvfp4
15
+ - rag
16
+ language:
17
+ - multilingual
18
+ - en
19
+ - ar
20
+ - as
21
+ - bn
22
+ - bg
23
+ - zh
24
+ - da
25
+ - nl
26
+ - fi
27
+ - fr
28
+ - de
29
+ - hi
30
+ - id
31
+ - it
32
+ - ja
33
+ - ko
34
+ - ms
35
+ - mr
36
+ - ne
37
+ - "no"
38
+ - fa
39
+ - pt
40
+ - ro
41
+ - ru
42
+ - es
43
+ - sw
44
+ - sv
45
+ - ta
46
+ - te
47
+ - th
48
+ - uk
49
+ - ur
50
+ - vi
51
+ library_name: vllm
52
+ pipeline_tag: sentence-similarity
53
+ ---
54
+
55
+ <div align="center">
56
+
57
+ **NVIDIA Nemotron 3 Embed**
58
+
59
+ <a href="https://huggingface.co/nvidia/Nemotron-3-Embed-1B-BF16">
60
+ <img alt="Nemotron 3 Embed 1B BF16" src="https://img.shields.io/badge/1B-BF16-76B900?style=for-the-badge&logo=nvidia&logoColor=white">
61
+ </a>
62
+ &nbsp;
63
+ <a href="https://huggingface.co/nvidia/Nemotron-3-Embed-1B-NVFP4">
64
+ <img alt="Nemotron 3 Embed 1B NVFP4" src="https://img.shields.io/badge/1B-NVFP4-76B900?style=for-the-badge&logo=nvidia&logoColor=white">
65
+ </a>
66
+ &nbsp;
67
+ <a href="https://huggingface.co/nvidia/Nemotron-3-Embed-8B-BF16">
68
+ <img alt="Nemotron 3 Embed 8B BF16" src="https://img.shields.io/badge/8B-BF16-76B900?style=for-the-badge&logo=nvidia&logoColor=white">
69
+ </a>
70
+
71
+ </div>
72
+
73
+ # Model Overview
74
+
75
+ ### Description
76
+ **Nemotron-3-Embed-1B-NVFP4** is the quantized version of the [Nemotron-3-Embed-1B-BF16](https://huggingface.co/nvidia/Nemotron-3-Embed-1B-BF16) model, which is developed for text question-answering retrieval. For more information, please check [here](https://huggingface.co/nvidia/Nemotron-3-Embed-1B-BF16). The **Nemotron-3-Embed-1B-NVFP4** model is quantized with [NVIDIA Model Optimizer](https://github.com/NVIDIA/Model-Optimizer), using nvidia-modelopt v0.45.0.
77
+ This model was evaluated on 34 languages: English, Arabic, Assamese, Bengali, Bulgarian, Chinese, Danish, Dutch, Finnish, French, German, Hindi, Hinglish, Indonesian, Italian, Japanese, Korean, Malay, Marathi, Nepali, Norwegian, Persian, Portuguese, Romanian, Russian, Spanish, Swahili, Swedish, Tamil, Telugu, Thai, Ukrainian, Urdu, Vietnamese. Read more details in our [Blog Post](https://huggingface.co/blog/nvidia/nemotron-3-embed-wins-rteb).
78
+
79
+ This model is ready for commercial use.
80
+
81
+ ### License/Terms of Use
82
+
83
+ This model and its associated configuration files are licensed under the [OpenMDW License Agreement, version 1.1 (OpenMDW-1.1)](https://openmdw.ai/license/1-1/). Additional Information: Built with [Ministral-3-3B-Instruct-2512](https://huggingface.co/mistralai/Ministral-3-3B-Instruct-2512), which is released under Apache 2.0.
84
+
85
+ This project will download and install additional third-party open source software projects. Review the license terms of these open source projects before use.
86
+
87
+ ### Deployment Geography
88
+ Global
89
+
90
+ ### Use Case
91
+
92
+ The **Nemotron-3-Embed-1B-NVFP4** is most suitable for users who want to build a multilingual question-and-answer application over a large text corpus, leveraging the latest dense retrieval technologies.
93
+
94
+ Generally [Nemotron-3-Embed-1B-BF16](https://huggingface.co/nvidia/Nemotron-3-Embed-1B-BF16) and **Nemotron-3-Embed-1B-NVFP4** models share the embedding space and can be used interchangeably, but we recommend validating retrieval quality on a representative sample before switching the models.
95
+
96
+ ### Release Date
97
+
98
+ 07/16/2026 via https://huggingface.co/nvidia/Nemotron-3-Embed-1B-NVFP4
99
+
100
+ ## Model Architecture
101
+
102
+ - **Architecture Type:** Transformer
103
+ - **Network Architecture:** [Ministral-3-3B-Instruct-2512](https://huggingface.co/mistralai/Ministral-3-3B-Instruct-2512)-based pruned encoder
104
+ - **Number of model parameters:** The model has 1.14B parameters
105
+ - **Hidden Size:** 2048
106
+
107
+ The **Nemotron-3-Embed-1B-NVFP4** is the quantized version of the [Nemotron-3-Embed-1B-BF16](https://huggingface.co/nvidia/Nemotron-3-Embed-1B-BF16), which is a transformer-based text embedding model trained with bidirectional attention masking, where the final embedding vector is obtained by applying average pooling to the transformer’s token-level representations. It encodes each input text into a dense embedding vector of dimension 2048.
108
+
109
+ ## Input(s)
110
+
111
+ - **Input Type(s):** Text
112
+ - **Input Format(s):**
113
+ - Text: List of strings
114
+ - **Input Parameters:**
115
+ - Text: One-Dimensional (1D)
116
+ - **Other Properties Related to Input:** Text inputs should be tokenized by the model tokenizer. The model’s max sequence length is 32768. Longer inputs should be chunked or truncated.
117
+
118
+ ## Output(s)
119
+
120
+ - **Output Type(s):** Floats
121
+ - **Output Format(s):**
122
+ - List of float arrays
123
+ - **Output Parameters:** One-Dimensional (1D) embedding vector per input text string
124
+ - **Other Properties Related to Output:** The model outputs a 2048-dimensional embedding vector for each input text string. It also supports dynamic embedding sizes by slicing the vector from the start (for example, keeping the first 1024 or 512 dimensions). These sliced embeddings remain highly functional, provided the resulting sub-vector is re-normalized (L2 normalization) after slicing.
125
+
126
+ Our AI models are designed and/or optimized to run on NVIDIA GPU-accelerated systems. By leveraging NVIDIA's hardware (e.g. GPU cores) and software frameworks (e.g., CUDA libraries), the model achieves faster training and inference times compared to CPU-only solutions.
127
+
128
+
129
+ ## Post-Training Quantization and Quantization-Aware Distillation
130
+
131
+ This model is a post-training quantized variant of [nvidia/Nemotron-3-Embed-1B-BF16](https://huggingface.co/nvidia/Nemotron-3-Embed-1B-BF16). Quantization was applied to the weights and activations of linear layers only, targeting the **NVFP4** data type for efficient inference. Quantization-Aware Distillation (QAD) was applied primarily to recover accuracy for long input sequences.
132
+
133
+ ## vLLM Usage
134
+
135
+ This NVFP4 checkpoint is intended for vLLM. The examples below show offline Python use and online serving. Use [nvidia/Nemotron-3-Embed-1B-BF16](https://huggingface.co/nvidia/Nemotron-3-Embed-1B-BF16) if you need Transformers or Sentence Transformers.
136
+
137
+ Use the Hugging Face model ID by default. If you are working from a local checkpoint, replace `MODEL_ID` with that path:
138
+
139
+ ```bash
140
+ MODEL_ID=nvidia/Nemotron-3-Embed-1B-NVFP4
141
+ ```
142
+
143
+ The output tables use `q[i]` for queries and `d[i]` for documents. Scores are rounded to four decimal places, and runtime differences might affect the final decimal places.
144
+
145
+ ### Tested vLLM Versions
146
+
147
+ This checkpoint has been explicitly tested with the following vLLM distributions:
148
+
149
+ | Distribution | Tested version |
150
+ | --- | --- |
151
+ | Python package | `vllm==0.25.0` |
152
+ | Upstream container | `vllm/vllm-openai:v0.22.1`, `vllm/vllm-openai:v0.25.0` |
153
+ | NVIDIA container | `nvcr.io/nvidia/vllm:26.06-py3` |
154
+
155
+ Use vLLM `0.25.0` for the examples below. vLLM `0.23.x` and `0.24.x` have known issues with this NVFP4 checkpoint family. Other versions have not been explicitly validated.
156
+
157
+ To install vLLM without a container image, run the following command:
158
+
159
+ ```bash
160
+ pip install --upgrade "vllm==0.25.0" openai requests numpy
161
+ ```
162
+
163
+
164
+ ### vLLM Offline Python
165
+
166
+ Use the offline Python API for local vLLM inference without an HTTP server. `LLM.embed` accepts formatted strings. Add the `query: ` and `passage: ` prefixes manually. The example uses a 4,096-token limit. Review [CUDA graph sizing](#cuda-graph-sizing) before increasing it.
167
+
168
+ <details>
169
+ <summary>vLLM Offline Python Example</summary>
170
+
171
+ ```python
172
+ import numpy as np
173
+ from vllm import LLM
174
+
175
+ MODEL_ID = "nvidia/Nemotron-3-Embed-1B-NVFP4"
176
+ MAX_MODEL_LEN = 4096
177
+ MAX_BATCHED_TOKENS = 4096
178
+
179
+ QUERIES = [
180
+ "Write a Python function that counts the frequency of each element in a list of lists.",
181
+ "Write a function that orders a dictionary with tuple keys by the product of each key's tuple values.",
182
+ "What symptoms and common triggers help distinguish eczema from other inflammatory skin conditions?",
183
+ "How can someone reduce exposure to pollen during allergy season?",
184
+ ]
185
+
186
+ DOCUMENTS = [
187
+ "def frequency_lists(list1):\n flattened = [item for sublist in list1 for item in sublist]\n counts = {}\n for item in flattened:\n if item in counts:\n counts[item] += 1\n else:\n counts[item] = 1\n return counts",
188
+ "def sort_dict_item(test_dict):\n return {key: test_dict[key] for key in sorted(test_dict.keys(), key=lambda ele: ele[0] * ele[1])}",
189
+ "Eczema commonly causes itchy, dry, inflamed patches of skin. The affected areas may look red, scaly, cracked, or darker than the surrounding skin depending on skin tone. Symptoms can flare after exposure to irritants, allergens, stress, or changes in weather.",
190
+ "People with pollen allergy can reduce exposure by staying indoors on dry, windy days, avoiding early-morning outdoor activity, and going outside after rain when pollen levels are lower. They should check pollen forecasts, close windows and doors when counts are high, and consider starting allergy medication before symptoms begin if high pollen is expected. After being outside, showering, changing clothes, avoiding outdoor laundry drying, and wearing a face mask for yard work can help limit pollen contact.",
191
+ ]
192
+
193
+ def main():
194
+ llm = LLM(
195
+ model=MODEL_ID,
196
+ max_model_len=MAX_MODEL_LEN,
197
+ max_num_batched_tokens=MAX_BATCHED_TOKENS,
198
+ max_cudagraph_capture_size=MAX_BATCHED_TOKENS,
199
+ )
200
+ texts = ["query: " + query for query in QUERIES] + [
201
+ "passage: " + doc for doc in DOCUMENTS
202
+ ]
203
+ outputs = llm.embed(texts, use_tqdm=False)
204
+ embeddings = np.array(
205
+ [output.outputs.embedding for output in outputs],
206
+ dtype=np.float32,
207
+ )
208
+
209
+ query_embeddings = embeddings[: len(QUERIES)]
210
+ document_embeddings = embeddings[len(QUERIES) :]
211
+
212
+ scores = query_embeddings @ document_embeddings.T
213
+ print("Similarity scores:")
214
+ print(f"{'':>8}" + "".join(f"d[{i}] " for i in range(scores.shape[1])))
215
+ for query_index, row in enumerate(scores):
216
+ print(f"q[{query_index}] " + " ".join(f"{score:>7.4f}" for score in row))
217
+
218
+
219
+ if __name__ == "__main__":
220
+ main()
221
+ ```
222
+
223
+ </details>
224
+
225
+ <details>
226
+ <summary>vLLM Offline Python Expected Output</summary>
227
+
228
+ The following output is expected:
229
+
230
+ ```text
231
+ Similarity scores:
232
+ d[0] d[1] d[2] d[3]
233
+ q[0] 0.8064 0.0201 0.0003 -0.0320
234
+ q[1] 0.0445 0.6469 -0.0516 0.0388
235
+ q[2] -0.0083 -0.0402 0.6558 0.1071
236
+ q[3] -0.0222 0.0265 0.1261 0.7677
237
+ ```
238
+
239
+ </details>
240
+
241
+ ### vLLM Online Serving
242
+
243
+ Tested vLLM builds read the checkpoint NVFP4 metadata, so no quantization flag is required. Start the server with the following command:
244
+
245
+ ```bash
246
+ MODEL_ID=nvidia/Nemotron-3-Embed-1B-NVFP4
247
+ MAX_MODEL_LEN=4096
248
+ MAX_BATCHED_TOKENS=4096
249
+
250
+ vllm serve "$MODEL_ID" \
251
+ --max-model-len "$MAX_MODEL_LEN" \
252
+ --max-num-batched-tokens "$MAX_BATCHED_TOKENS" \
253
+ --max-cudagraph-capture-size "$MAX_BATCHED_TOKENS"
254
+ ```
255
+
256
+ #### CUDA Graph Sizing
257
+
258
+ The checkpoint supports sequences up to 32,768 tokens. The examples use 4,096 as a conservative starting point.
259
+
260
+ Use the following guidance to tune CUDA graph capture:
261
+
262
+ - Set `--max-model-len` to the longest request you intend to serve. Tune `--max-num-batched-tokens` for the workload, concurrency, and available GPU memory. When chunked prefill is disabled, the batched-token budget must be at least the model-length limit.
263
+ - For default capture buckets up to 8,192, set `--max-cudagraph-capture-size` equal to `--max-num-batched-tokens`. This setting makes batches up to the scheduler budget eligible for CUDA graph execution. Batches outside the captured range use a slower uncaptured path.
264
+ - Larger capture ranges increase startup time and graph memory. Benchmark representative traffic on the target hardware. Do not capture beyond the batched-token budget.
265
+
266
+ Use a maximum capture size of 4,096 as the conservative default for services that restart or autoscale regularly. A maximum capture size of 8,192 can be reasonable when a longer cold start is acceptable. For maximum capture sizes above 8,192, pass a smaller, workload-aligned set with `--cudagraph-capture-sizes` to keep startup time under control.
267
+
268
+ <details>
269
+ <summary>Example Sparse Capture Sizes for Inputs up to 32,768 Tokens</summary>
270
+
271
+ The following command uses a sparse capture-size list:
272
+
273
+ ```bash
274
+ MODEL_ID=nvidia/Nemotron-3-Embed-1B-NVFP4
275
+ MAX_BATCHED_TOKENS=32768
276
+
277
+ vllm serve "$MODEL_ID" \
278
+ --max-num-batched-tokens "$MAX_BATCHED_TOKENS" \
279
+ --cudagraph-capture-sizes \
280
+ 1 2 4 8 16 24 32 40 48 56 64 72 80 88 96 104 112 120 128 \
281
+ 136 144 152 160 168 176 184 192 200 208 216 224 232 240 248 256 \
282
+ 384 512 768 1024 1536 2048 3072 4096 6144 8192 12288 16384 \
283
+ 24576 32768
284
+ ```
285
+
286
+ Choose sizes from representative batch-token measurements. vLLM pads each execution batch to the next captured size, so denser lists reduce padding but require more startup time and graph memory.
287
+
288
+ In illustrative tests on an NVIDIA GB10 system with vLLM 0.25.0, cold startup using automatic buckets took about 74 seconds at a maximum capture size of 4,096 and 121 seconds at 8,192. At 32,768, startup with automatic buckets was projected to take tens of minutes. With the sparse list of 49 capture sizes above, startup completed in about one minute. Results vary by workload and hardware.
289
+
290
+ </details>
291
+
292
+ Add `--host` or `--port` to the serving command if you need non-default network settings.
293
+
294
+ To serve a local checkpoint, replace `MODEL_ID` with its path. Add `--served-model-name nvidia/Nemotron-3-Embed-1B-NVFP4` if clients should continue using the Hugging Face model ID.
295
+
296
+ #### Recommended Retrieval Endpoint
297
+
298
+ After the server is running, use `/v2/embed` for retrieval. Send raw query and document strings. `input_type` applies the saved query and document prompt metadata.
299
+
300
+ The following example uses the recommended endpoint:
301
+
302
+ ```python
303
+ import numpy as np
304
+ import requests
305
+
306
+ MODEL = "nvidia/Nemotron-3-Embed-1B-NVFP4"
307
+ URL = "http://localhost:8000/v2/embed"
308
+
309
+ QUERIES = [
310
+ "Write a Python function that counts the frequency of each element in a list of lists.",
311
+ "Write a function that orders a dictionary with tuple keys by the product of each key's tuple values.",
312
+ "What symptoms and common triggers help distinguish eczema from other inflammatory skin conditions?",
313
+ "How can someone reduce exposure to pollen during allergy season?",
314
+ ]
315
+
316
+ DOCUMENTS = [
317
+ "def frequency_lists(list1):\n flattened = [item for sublist in list1 for item in sublist]\n counts = {}\n for item in flattened:\n if item in counts:\n counts[item] += 1\n else:\n counts[item] = 1\n return counts",
318
+ "def sort_dict_item(test_dict):\n return {key: test_dict[key] for key in sorted(test_dict.keys(), key=lambda ele: ele[0] * ele[1])}",
319
+ "Eczema commonly causes itchy, dry, inflamed patches of skin. The affected areas may look red, scaly, cracked, or darker than the surrounding skin depending on skin tone. Symptoms can flare after exposure to irritants, allergens, stress, or changes in weather.",
320
+ "People with pollen allergy can reduce exposure by staying indoors on dry, windy days, avoiding early-morning outdoor activity, and going outside after rain when pollen levels are lower. They should check pollen forecasts, close windows and doors when counts are high, and consider starting allergy medication before symptoms begin if high pollen is expected. After being outside, showering, changing clothes, avoiding outdoor laundry drying, and wearing a face mask for yard work can help limit pollen contact.",
321
+ ]
322
+
323
+ def embed(input_type: str, texts: list[str]) -> np.ndarray:
324
+ response = requests.post(
325
+ URL,
326
+ json={
327
+ "model": MODEL,
328
+ "input_type": input_type,
329
+ "texts": texts,
330
+ "embedding_types": ["float"],
331
+ "truncate": "END",
332
+ },
333
+ timeout=120,
334
+ )
335
+ response.raise_for_status()
336
+ return np.array(response.json()["embeddings"]["float"], dtype=np.float32)
337
+
338
+ query_embeddings = embed("query", QUERIES)
339
+ document_embeddings = embed("document", DOCUMENTS)
340
+
341
+ scores = query_embeddings @ document_embeddings.T
342
+ print("Similarity scores:")
343
+ print(f"{'':>8}" + "".join(f"d[{i}] " for i in range(scores.shape[1])))
344
+ for query_index, row in enumerate(scores):
345
+ print(f"q[{query_index}] " + " ".join(f"{score:>7.4f}" for score in row))
346
+ ```
347
+
348
+ <details>
349
+ <summary>Recommended Retrieval Endpoint Expected Output</summary>
350
+
351
+ The following output is expected:
352
+
353
+ ```text
354
+ Similarity scores:
355
+ d[0] d[1] d[2] d[3]
356
+ q[0] 0.8063 0.0201 0.0003 -0.0320
357
+ q[1] 0.0445 0.6469 -0.0516 0.0388
358
+ q[2] -0.0082 -0.0402 0.6558 0.1072
359
+ q[3] -0.0222 0.0265 0.1261 0.7677
360
+ ```
361
+
362
+ </details>
363
+
364
+ You can also use the OpenAI-compatible `/v1/embeddings` endpoint. For those requests, pass strings in `input` and manually prefix them with `query: ` or `passage: `.
365
+
366
+ ### Expected Configuration Warning
367
+
368
+ When vLLM loads this checkpoint, its Transformers configuration parser can emit the following warning.
369
+
370
+ ```text
371
+ [transformers] Unrecognized keys in `rope_parameters` for 'rope_type'='yarn': {'apply_yarn_scaling'}
372
+ ```
373
+
374
+ This warning is expected and does not prevent vLLM from loading the model or running inference. `apply_yarn_scaling` is a temporary vLLM compatibility field that preserves the checkpoint's intended long-context rotary position embedding (RoPE) behavior. Do not remove it from `config.json`. Refer to [vLLM issue #48621](https://github.com/vllm-project/vllm/issues/48621) for upstream compatibility work.
375
+
376
+
377
+ ## Software Integration
378
+
379
+ - **Runtime Engine(s):** vLLM
380
+ - **Supported Hardware Microarchitecture Compatibility:**
381
+ - NVIDIA Ampere
382
+ - NVIDIA Blackwell
383
+ - NVIDIA Hopper
384
+ - NVIDIA Lovelace
385
+ - **Supported Operating System(s):**
386
+ - Linux
387
+
388
+ The integration of foundation and fine-tuned models into AI systems requires additional testing using use-case-specific data to ensure safe and effective deployment. Following the V-model methodology, iterative testing and validation at both unit and system levels are essential to mitigate risks, meet technical and functional requirements, and ensure compliance with safety and ethical standards before deployment.
389
+
390
+ ## Model Version(s)
391
+
392
+ **Nemotron-3-Embed-1B-NVFP4**
393
+
394
+ ## Training, Testing, and Evaluation Datasets
395
+
396
+ ### Dataset Overview
397
+
398
+ This checkpoint is an NVFP4 post-training-quantized derivative of [nvidia/Nemotron-3-Embed-1B-BF16](https://huggingface.co/nvidia/Nemotron-3-Embed-1B-BF16). Quantization-Aware Distillation (QAD) was applied on a small sample of the original BF16 data mix for further accuracy recovery.
399
+
400
+ ### Calibration Dataset
401
+
402
+ **Data Collection Method by dataset:** Automated
403
+
404
+ **Labeling Method by dataset:** Automated
405
+
406
+ **Properties:** A calibration dataset of 512 total samples was used for NVFP4 post-training quantization, consisting of 256 queries and 256 passages from the [abisee/cnn_dailymail](https://huggingface.co/datasets/abisee/cnn_dailymail) dataset, formatted with query and passage prefixes.
407
+
408
+ ### Training Dataset
409
+
410
+ QAD training was performed as part of the NVFP4 model development. The information below describes the QAD training datasets used for this model. For details about the training datasets used to build the underlying base model, [Nemotron-3-Embed-1B-BF16](https://huggingface.co/nvidia/Nemotron-3-Embed-1B-BF16), please refer to its model card.
411
+
412
+ **Total Size:** 20k data samples
413
+
414
+ **Total Number of Datasets:** 5 dataset files
415
+
416
+ **Dataset Partition:** Training [100%], Testing [N/A — evaluation benchmarks used separately], Validation [N/A — evaluation benchmarks used separately].
417
+
418
+ #### Public Datasets:
419
+
420
+ | Dataset name | Reference |
421
+ | --- | --- |
422
+ | MLDR | [https://huggingface.co/datasets/Shitao/MLDR](https://huggingface.co/datasets/Shitao/MLDR) |
423
+
424
+ #### Synthetic Datasets:
425
+
426
+ Synthetic query-document pairs were generated either from scratch or by using seed datasets to generate queries with the models listed below.
427
+
428
+ <table>
429
+ <tr>
430
+ <th>LLMs used to generate synthetic datasets</th>
431
+ </tr>
432
+ <tr>
433
+ <td>nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16<br>nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4</td>
434
+ </tr>
435
+ </table>
436
+
437
+ <table>
438
+ <tr>
439
+ <th colspan="2">Seed Datasets</th>
440
+ </tr>
441
+ <tr>
442
+ <th>Dataset</th>
443
+ <th>Reference</th>
444
+ </tr>
445
+ <tr>
446
+ <td>FinePdfs</td>
447
+ <td><a href="https://huggingface.co/datasets/HuggingFaceFW/finepdfs">https://huggingface.co/datasets/HuggingFaceFW/finepdfs</a></td>
448
+ </tr>
449
+ </table>
450
+
451
+
452
+
453
+ #### Data Modality
454
+
455
+ - Text
456
+
457
+ #### Training Data Size
458
+
459
+ **Text Training Data Size:** 20k
460
+
461
+ **Data Collection Method by dataset:** Hybrid: Human, Automated, Synthetic
462
+
463
+ **Labeling Method by dataset:** Hybrid: Human, Automated, Synthetic
464
+
465
+ **Properties:** Model distillation training of the underlying model was conducted on text datasets using question–passage pairs from publicly available, commercially permissible datasets and synthetically generated datasets. For more information, please visit [this link](https://huggingface.co/nvidia/Nemotron-3-Embed-1B-BF16).
466
+
467
+ ### Testing Dataset
468
+
469
+ **Data Collection Method by dataset:** Not Applicable
470
+
471
+ **Labeling Method by dataset:** Not Applicable
472
+
473
+ **Properties:** Not Applicable. Model quality was assessed using the evaluation benchmark datasets described in the Evaluation Dataset subsection.
474
+
475
+ ### Evaluation Dataset
476
+
477
+ **Data Collection Method by dataset:** Hybrid: Human, Automated, Synthetic
478
+
479
+ **Labeling Method by dataset:** Hybrid: Human, Automated, Synthetic
480
+
481
+ **Properties:** In this section, we compare the performance of quantized model **Nemotron-3-Embed-1B-NVFP4** with baseline implementation **Nemotron-3-Embed-1B-BF16**.
482
+
483
+ This model is evaluated on 16 public tasks on [Retrieval Embedding Benchmark (RTEB)](https://huggingface.co/blog/rteb), a new benchmark designed to reliably evaluate the retrieval accuracy of embedding models for real-world applications. More details on RTEB can be found on their [leaderboard](https://huggingface.co/spaces/mteb/leaderboard?benchmark_name=RTEB%28beta%29).
484
+
485
+ We set the model sequence length to 4096 for the evaluation results below. The `NVFP4` model was evaluated on an NVIDIA GB200 GPU.
486
+
487
+ **Text Retrieval Benchmarks (chunk retrieval) – Avg. NDCG@10**
488
+ | Model Name | Precision | RTEB |
489
+ | --- | --- | --- |
490
+ | Nemotron-3-Embed-1B-BF16 | BF16 | 72.38 |
491
+ | Nemotron-3-Embed-1B-NVFP4 | NVFP4 | 72.00 |
492
+
493
+
494
+ ## Inference
495
+
496
+ - **Acceleration Engine:** vLLM
497
+ - **Test Hardware:**
498
+ - NVIDIA Ampere - A100 PCIe/SXM
499
+ - NVIDIA Blackwell - GB200 and RTX 6000 PRO
500
+ - NVIDIA Hopper - H100 PCIe/SXM
501
+ - NVIDIA Lovelace - L40 and L4
502
+
503
+ ## Ethical Considerations
504
+
505
+ NVIDIA believes Trustworthy AI is a shared responsibility and we have established policies and practices to enable development for a wide array of AI applications. Developers should work with their internal model team to ensure this model meets requirements for the relevant industry and use case and addresses unforeseen product misuse.
506
+
507
+ For more detailed information on ethical considerations for this model, please see the Model Card++ Bias, Explainability, Safety & Security, and Privacy Subcards.
508
+
509
+ Please report model quality, risk, security vulnerabilities or NVIDIA AI concerns [here](https://www.nvidia.com/en-us/support/submit-security-vulnerability/).
510
+
511
+ ## Bias
512
+
513
+ | Field | Response |
514
+ | ----- | ----- |
515
+ | Participation considerations from adversely impacted groups [protected classes](https://www.senate.ca.gov/content/protected-classes) in model design and testing | None |
516
+ | Measures taken to mitigate against unwanted bias | None |
517
+ | Bias Metric (If Measured): | None |
518
+
519
+ ## Explainability
520
+
521
+ | Field | Response |
522
+ | ----- | ----- |
523
+ | Intended Task/Domain: | Passage and query embedding for question and answer retrieval |
524
+ | Model Type: | Transformer encoder |
525
+ | Intended Users: | Generative AI creators working with conversational AI models - users who want to build a multilingual question and answer application over a large text corpus, leveraging the latest dense retrieval technologies. |
526
+ | Output: | Array of float numbers (Dense Vector Representation for the input text) |
527
+ | Describe how the model works: | Model transforms the tokenized input text into a dense vector representation. |
528
+ | Name the adversely impacted groups this has been tested to deliver comparable outcomes regardless of: | Not Applicable |
529
+ | Technical Limitations & Mitigation: | The model’s max sequence length is 32768. Therefore, the longer text inputs should be truncated. |
530
+ | Verified to have met prescribed NVIDIA quality standards: | Yes |
531
+ | Performance Metrics: | Accuracy, Throughput, and Latency |
532
+ | Potential Known Risks: | This model does not always guarantee to retrieve the correct passage(s) for a given query. |
533
+ | Licensing: | This model and its associated configuration files are licensed under the [OpenMDW License Agreement, version 1.1 (OpenMDW-1.1)](https://openmdw.ai/license/1-1/). Additional Information: Built with [Ministral-3-3B-Instruct-2512](https://huggingface.co/mistralai/Ministral-3-3B-Instruct-2512), which is released under Apache 2.0. |
534
+
535
+ ## Privacy
536
+
537
+ | Field | Response |
538
+ | ----- | ----- |
539
+ | Generatable or reverse engineerable personal data? | None |
540
+ | Was consent obtained for any personal data used? | Not Applicable |
541
+ | Personal data used to create this model? | None Known |
542
+ | How often is the dataset reviewed? | Before Every Release |
543
+ | Is there provenance for all datasets used in training? | Yes |
544
+ | Does data labeling (annotation, metadata) comply with privacy laws? | Yes |
545
+ | Was data from user interactions with the AI model (e.g. user input and prompts) used to train the model? | Yes |
546
+ | Is data compliant with data subject requests for data correction or removal, if such a request was made? | Not Applicable |
547
+ | Applicable Privacy Policy | https://www.nvidia.com/en-us/about-nvidia/privacy-policy/ |
548
+
549
+ ## Safety
550
+
551
+ | Field | Response |
552
+ | ----- | ----- |
553
+ | Model Application(s): | Text Embedding for Retrieval |
554
+ | Describe the physical safety impact (if present). | Not Applicable |
555
+ | Use Case Restrictions: | This model and its associated configuration files are licensed under the [OpenMDW License Agreement, version 1.1 (OpenMDW-1.1)](https://openmdw.ai/license/1-1/). Additional Information: Built with [Ministral-3-3B-Instruct-2512](https://huggingface.co/mistralai/Ministral-3-3B-Instruct-2512), which is released under Apache 2.0. |
556
+ | Model and dataset restrictions: | The Principle of least privilege (PoLP) is applied limiting access for dataset generation and model development. Restrictions enforce dataset access during training, and dataset license constraints adhered to. |