sanmonga22 commited on
Commit
ce4bfb5
·
verified ·
1 Parent(s): 9fba1f6

Sanitize public model card metadata

Browse files
Files changed (1) hide show
  1. README.md +12 -25
README.md CHANGED
@@ -1,31 +1,18 @@
1
  ---
2
- license: other
3
- base_model: nvidia/Llama-3.1-Nemotron-Nano-VL-8B-V1
4
- tags: [vlm, image-to-text, hexagon, qualcomm, npu, qhexrt, radio, llama]
 
 
 
 
 
5
  ---
6
 
7
- # Llama-3.1-Nemotron-Nano-VL-8B → Qualcomm Hexagon NPU (v81)
8
 
9
- NVIDIA **Llama-3.1-Nemotron-Nano-VL-8B-V1** (RADIO vision + Llama-3.1-8B) ported to the Qualcomm Hexagon NPU
10
- (SM8750 / **v81** / `soc_model 87`) and **device-validated on a Samsung S25**: image → detailed caption, fully
11
- on-device.
12
 
13
- ## Device validation (S25, v81) — image→caption, 3/3 distinct images correct
14
- - **dog** → "…a white dog sits majestically, its gaze directed upwards… lush green grass… tongue…"
15
- - **tabby cat** → "…a tabby cat with a coat of brown and black stripes… eyes wide open, reflecting curiosity…"
16
- - **pizza** → "…a freshly baked pizza… golden crust, sits in a cardboard box…"
17
 
18
- The 8B LLM backbone is independently device-validated (W8 greedy 4/6 exact vs HF gold). Vision graph is export-
19
- faithful (ONNX vs PyTorch cos 1.000000); host preprocess+projector validated cos 1.000000.
20
-
21
- ## Pipeline (all on-device via the `nemotron_vl_generate` host-op)
22
- image(512²) → RADIO input_conditioner + patch_generator (host) → **vision graph `radio_enc.bin`** (32-block
23
- ViT-Huge) → drop 8 prefix → pixel_shuffle + mlp1 projector (host) → **256 image tokens** → spliced at the 256
24
- `<image>` slots → **8-part W8 sharded decode** (Llama-3.1-8B, MAXCTX 512) → lmhead. ~3.7 tok/s decode.
25
-
26
- ## Run (QHexRT)
27
- ```
28
- adb shell "cd /data/local/tmp/wq && export ADSP_LIBRARY_PATH='/data/local/tmp/wq/dsp;/data/local/tmp/wq;/vendor/dsp/cdsp'; \
29
- LD_LIBRARY_PATH=. ./qhx_generate v81/nemotron-vl-8b-vlm.json libQnnHtp.so libQnnSystem.so v81 60 'Describe this image in detail.' my_photo.jpg"
30
- ```
31
- Arch-pinned v81 (`dsp_arch`+`soc_model 87`). Part of the RunAnywhere Hexagon model set.
 
1
  ---
2
+ license: "other"
3
+ tags:
4
+ - "hnpu"
5
+ - "hexagon"
6
+ - "npu"
7
+ - "vlm"
8
+ base_model: "nvidia/Llama-3.1-Nemotron-Nano-VL-8B-V1"
9
+ pipeline_tag: "image-text-to-text"
10
  ---
11
 
12
+ # nemotron nano vl 8b HNPU
13
 
14
+ Prebuilt HNPU artifacts for [nvidia/Llama-3.1-Nemotron-Nano-VL-8B-V1](https://huggingface.co/nvidia/Llama-3.1-Nemotron-Nano-VL-8B-V1), a public vision-language model.
 
 
15
 
16
+ For model behavior, license, intended use, and limitations, see the [upstream model card](https://huggingface.co/nvidia/Llama-3.1-Nemotron-Nano-VL-8B-V1).
 
 
 
17
 
18
+ Artifacts are architecture-pinned. Available artifact directories: `v81/`.