Image-Text-to-Text
Transformers
Safetensors
nemotron_parse_tc
feature-extraction
nvidia
VLM
OCR
conversational
custom_code
Instructions to use nvidia/NVIDIA-Nemotron-Parse-v1.1-TC with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use nvidia/NVIDIA-Nemotron-Parse-v1.1-TC with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="nvidia/NVIDIA-Nemotron-Parse-v1.1-TC", trust_remote_code=True) messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("nvidia/NVIDIA-Nemotron-Parse-v1.1-TC", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use nvidia/NVIDIA-Nemotron-Parse-v1.1-TC with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "nvidia/NVIDIA-Nemotron-Parse-v1.1-TC" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "nvidia/NVIDIA-Nemotron-Parse-v1.1-TC", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/nvidia/NVIDIA-Nemotron-Parse-v1.1-TC
- SGLang
How to use nvidia/NVIDIA-Nemotron-Parse-v1.1-TC with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "nvidia/NVIDIA-Nemotron-Parse-v1.1-TC" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "nvidia/NVIDIA-Nemotron-Parse-v1.1-TC", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "nvidia/NVIDIA-Nemotron-Parse-v1.1-TC" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "nvidia/NVIDIA-Nemotron-Parse-v1.1-TC", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use nvidia/NVIDIA-Nemotron-Parse-v1.1-TC with Docker Model Runner:
docker model run hf.co/nvidia/NVIDIA-Nemotron-Parse-v1.1-TC
Update README.md
Browse files
README.md
CHANGED
|
@@ -9,9 +9,9 @@ tags:
|
|
| 9 |
- nvidia
|
| 10 |
- VLM
|
| 11 |
---
|
| 12 |
-
#
|
| 13 |
|
| 14 |
-
nemotron-parse is a general purpose text-extraction model, specifically designed to handle documents. Given an image, nemotron-parse is able to extract formatted-text, with bounding-boxes and the corresponding semantic class. This has downstream benefits for several tasks such as increasing the availability of training-data for Large Language Models (LLMs), improving the accuracy of retriever systems, and enhancing document understanding pipelines.
|
| 15 |
|
| 16 |
This model is ready for commercial use.
|
| 17 |
|
|
@@ -26,8 +26,8 @@ GOVERNING TERMS: The NIM container is governed by the [NVIDIA Software License A
|
|
| 26 |
Global
|
| 27 |
|
| 28 |
## Use Case:
|
| 29 |
-
nemotron-parse will be capable of comprehensive text understanding and document structure understanding. It will be used in retriever and curator solutions. Its text extraction datasets and capabilities will help with LLM and VLM training, as well as improve run-time inference accuracy of VLMs.
|
| 30 |
-
The nemotron-parse model will perform text extraction from PDF and PPT documents. The nemotron-parse can classify the objects (title, section, caption, index, footnote, lists, tables, bibliography, image) in a given document, and provide bounding boxes with coordinates.
|
| 31 |
|
| 32 |
|
| 33 |
## Release Date:
|
|
@@ -62,7 +62,7 @@ Carbon Emissions: 3.21 tCO2e <br>
|
|
| 62 |
* Input Type(s): Red, Green, Blue (RGB) + Prompt (String)
|
| 63 |
* Input Parameters: 2D, 1D
|
| 64 |
- Other Properties Related to Input:
|
| 65 |
-
- Max Input Resolution (Width, Height):
|
| 66 |
- Min Input Resolution (Width, Height): 1024, 1280
|
| 67 |
- Channel Count: 3
|
| 68 |
|
|
@@ -71,8 +71,8 @@ Carbon Emissions: 3.21 tCO2e <br>
|
|
| 71 |
* Output Format: String
|
| 72 |
* Output Parameters: 1D
|
| 73 |
- Other Properties Related to Output:
|
| 74 |
-
- nemotron-parse output format is a string which encodes text content (formatted or not) as well as bounding boxes and class attributes.<br>
|
| 75 |
-
|
| 76 |
Our AI models are designed and/or optimized to run on NVIDIA GPU-accelerated systems. By leveraging NVIDIA’s hardware (e.g. GPU cores) and software frameworks (e.g., CUDA libraries), the model achieves faster training and inference times compared to CPU-only solutions.<br>
|
| 77 |
|
| 78 |
## Software Integration:
|
|
@@ -88,8 +88,8 @@ The integration of foundation and fine-tuned models into AI systems requires add
|
|
| 88 |
|
| 89 |
## Model Version:
|
| 90 |
|
| 91 |
-
V1.1-
|
| 92 |
-
This version
|
| 93 |
|
| 94 |
## Quick Start
|
| 95 |
|
|
@@ -108,7 +108,7 @@ from transformers import AutoModel, AutoProcessor, AutoTokenizer, AutoConfig, Au
|
|
| 108 |
from postprocessing import extract_classes_bboxes, transform_bbox_to_original, postprocess_text
|
| 109 |
|
| 110 |
# Load model and processor
|
| 111 |
-
model_path = "nvidia/NVIDIA-Nemotron-Parse-v1.1-
|
| 112 |
device = "cuda:0"
|
| 113 |
|
| 114 |
model = AutoModel.from_pretrained(
|
|
@@ -166,7 +166,7 @@ for bbox in bboxes:
|
|
| 166 |
|
| 167 |
### Training Dataset
|
| 168 |
|
| 169 |
-
nemotron-parse is first pre-trained on our internal datasets: human, synthetic and automated.
|
| 170 |
Data Modality:
|
| 171 |
*Text
|
| 172 |
*Image<br>
|
|
@@ -175,7 +175,7 @@ Labeling Method by Dataset: Hybrid: Human, Synthetic, Automated
|
|
| 175 |
|
| 176 |
### Testing and Evaluation Dataset:
|
| 177 |
|
| 178 |
-
nemotron-parse is evaluated on multiple datasets for robustness, including public and internal dataset.
|
| 179 |
Data Collection Method by Dataset: Hybrid: Human, Synthetic, Automated
|
| 180 |
Labeling Method by Dataset: Hybrid: Human, Synthetic, Automated
|
| 181 |
|
|
|
|
| 9 |
- nvidia
|
| 10 |
- VLM
|
| 11 |
---
|
| 12 |
+
# Nemotron-Parse-Lite Overview
|
| 13 |
|
| 14 |
+
nemotron-parse-lite is a general purpose text-extraction model, specifically designed to handle documents. Given an image, nemotron-parse-lite is able to extract formatted-text, with bounding-boxes and the corresponding semantic class. This has downstream benefits for several tasks such as increasing the availability of training-data for Large Language Models (LLMs), improving the accuracy of retriever systems, and enhancing document understanding pipelines.
|
| 15 |
|
| 16 |
This model is ready for commercial use.
|
| 17 |
|
|
|
|
| 26 |
Global
|
| 27 |
|
| 28 |
## Use Case:
|
| 29 |
+
nemotron-parse-lite will be capable of comprehensive text understanding and document structure understanding. It will be used in retriever and curator solutions. Its text extraction datasets and capabilities will help with LLM and VLM training, as well as improve run-time inference accuracy of VLMs.
|
| 30 |
+
The nemotron-parse-lite model will perform text extraction from PDF and PPT documents. The nemotron-parse-lite can classify the objects (title, section, caption, index, footnote, lists, tables, bibliography, image) in a given document, and provide bounding boxes with coordinates.
|
| 31 |
|
| 32 |
|
| 33 |
## Release Date:
|
|
|
|
| 62 |
* Input Type(s): Red, Green, Blue (RGB) + Prompt (String)
|
| 63 |
* Input Parameters: 2D, 1D
|
| 64 |
- Other Properties Related to Input:
|
| 65 |
+
- Max Input Resolution (Width, Height): 1664, 2048
|
| 66 |
- Min Input Resolution (Width, Height): 1024, 1280
|
| 67 |
- Channel Count: 3
|
| 68 |
|
|
|
|
| 71 |
* Output Format: String
|
| 72 |
* Output Parameters: 1D
|
| 73 |
- Other Properties Related to Output:
|
| 74 |
+
- nemotron-parse-lite output format is a string which encodes text content (formatted or not) as well as bounding boxes and class attributes.<br>
|
| 75 |
+
In the default prompt setting, text content is represented as markdown, and math expressions as LaTeX, enclosed in \[..\] or \(..\). If a mathematical expression does not require LaTeX formatting to be represented (e.g., consisting only of characters and subscripts/superscripts), it is represented as markdown. Tables are represented as LaTeX.
|
| 76 |
Our AI models are designed and/or optimized to run on NVIDIA GPU-accelerated systems. By leveraging NVIDIA’s hardware (e.g. GPU cores) and software frameworks (e.g., CUDA libraries), the model achieves faster training and inference times compared to CPU-only solutions.<br>
|
| 77 |
|
| 78 |
## Software Integration:
|
|
|
|
| 88 |
|
| 89 |
## Model Version:
|
| 90 |
|
| 91 |
+
V1.1-Lite
|
| 92 |
+
This version offers 20% speed improvement compared to Nemotron-Parse-v1.1. Additionally, unlike Nemotron-Parse-v1.1, it preserves the page order of un-ordered element: Tables, Captions, Pictures, Footnotes.
|
| 93 |
|
| 94 |
## Quick Start
|
| 95 |
|
|
|
|
| 108 |
from postprocessing import extract_classes_bboxes, transform_bbox_to_original, postprocess_text
|
| 109 |
|
| 110 |
# Load model and processor
|
| 111 |
+
model_path = "nvidia/NVIDIA-Nemotron-Parse-v1.1-Lite" # Or use a local path
|
| 112 |
device = "cuda:0"
|
| 113 |
|
| 114 |
model = AutoModel.from_pretrained(
|
|
|
|
| 166 |
|
| 167 |
### Training Dataset
|
| 168 |
|
| 169 |
+
nemotron-parse-lite is first pre-trained on our internal datasets: human, synthetic and automated.
|
| 170 |
Data Modality:
|
| 171 |
*Text
|
| 172 |
*Image<br>
|
|
|
|
| 175 |
|
| 176 |
### Testing and Evaluation Dataset:
|
| 177 |
|
| 178 |
+
nemotron-parse-lite is evaluated on multiple datasets for robustness, including public and internal dataset.
|
| 179 |
Data Collection Method by Dataset: Hybrid: Human, Synthetic, Automated
|
| 180 |
Labeling Method by Dataset: Hybrid: Human, Synthetic, Automated
|
| 181 |
|