Instructions to use nvidia/Llama-3.1-Nemotron-Nano-VL-8B-V1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use nvidia/Llama-3.1-Nemotron-Nano-VL-8B-V1 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="nvidia/Llama-3.1-Nemotron-Nano-VL-8B-V1")# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("nvidia/Llama-3.1-Nemotron-Nano-VL-8B-V1", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use nvidia/Llama-3.1-Nemotron-Nano-VL-8B-V1 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "nvidia/Llama-3.1-Nemotron-Nano-VL-8B-V1" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "nvidia/Llama-3.1-Nemotron-Nano-VL-8B-V1", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/nvidia/Llama-3.1-Nemotron-Nano-VL-8B-V1
- SGLang
How to use nvidia/Llama-3.1-Nemotron-Nano-VL-8B-V1 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "nvidia/Llama-3.1-Nemotron-Nano-VL-8B-V1" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "nvidia/Llama-3.1-Nemotron-Nano-VL-8B-V1", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "nvidia/Llama-3.1-Nemotron-Nano-VL-8B-V1" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "nvidia/Llama-3.1-Nemotron-Nano-VL-8B-V1", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use nvidia/Llama-3.1-Nemotron-Nano-VL-8B-V1 with Docker Model Runner:
docker model run hf.co/nvidia/Llama-3.1-Nemotron-Nano-VL-8B-V1
| Field | Response |
|---|---|
| Intended Application & Domain: | Visual Question Answering |
| Model Type: | Transformer |
| Intended Users: | Generative AI creators working with conversational AI models and image content. |
| Output: | Text (Responds to posed question, stateful - remembers previous answers) |
| Describe how the model works: | Chat based on image/text |
| Name the adversely impacted groups this has been tested to deliver comparable outcomes regardless of: | Not Applicable |
| Technical Limitations: | Context Length: Supports up to 16,000 tokens total (input + output). If exceeded, input is truncated from the start, and generation ends with an EOS token. Longer prompts may risk performance loss. If the model fails (e.g., generates incorrect responses, repeats, or gives poor responses), issues are diagnosed via benchmarks, human review, and internal debugging tools. Only use NVIDIA provided models that use safetensors format. Do not expose the vLLM host to a network where any untrusted connections may reach the host. Only use NVIDIA provided models that use safetensors format. |
| Verified to have met prescribed NVIDIA quality standards: | Yes |
| Performance Metrics: | MMMU Val with chatGPT as a judge, AI2D, ChartQA Test, InfoVQA Val, OCRBench, OCRBenchV2 English, OCRBenchV2 Chinese, DocVQA val, VideoMME (16 frames), SlideQA (F1) |
| Potential Known Risks: | The Model may produce output that is biased, toxic, or incorrect responses. Therefore, the model may amplify those biases and return toxic responses especially when prompted with toxic prompts. The Model may also generate answers that may be inaccurate, omit key information, or include irrelevant or redundant text, producing socially unacceptable or undesirable text, even if the prompt itself does not include anything explicitly offensive. While we have taken safety and security into account and are continuously improving, outputs may still contain political content, misleading information, or unwanted bias beyond our control. |
| Licensing: | Governing Terms: Your use of the software container and model is governed by the NVIDIA Software and Model Evaluation License. Additional Information: Llama 3.1 Community Model License; Built with Llama. |