Prune QHexRT HNPU bundle to runtime-minimum artifacts
Browse filesStrip non-runtime _comment metadata from 1 manifest(s) without changing runtime fields.
Audit artifacts: QHexRT/hf_hnpu_audit cleanup plan, manifest reference check, and multi-agent review.
v81/nemotron-vl-8b-vlm.json
CHANGED
|
@@ -1,6 +1,5 @@
|
|
| 1 |
{
|
| 2 |
"schema_version": 1,
|
| 3 |
-
"_comment": "Nemotron-Nano-VL-8B VLM (NVIDIA Llama-3.1-8B backbone + C-RADIOv2-H ViT-Huge vision tower) on Hexagon v81 (SM8850 / soc_model 87). The nemotron_vl_generate host-op runs the whole image->caption pipeline through QHexRT: host RADIO preprocess (input_conditioner CLIP-norm + patch_generator: im_to_patches 16x16 -> Linear 768->1280 + 8 prefix tokens + baked pos_embed) -> NPU vision graph (32 RADIO blocks, 2D I/O embed[1032,1280]->vit_feat[1032,1280], fp32 for vision-cos>=0.999) -> host projector (drop 8 prefix -> pixel_shuffle 0.5 -> [256,5120] -> mlp1 LN/Linear5120->4096/GELU/Linear4096->4096) -> 256 image tokens spliced at the <image>(128256) positions of the llama_3p1 prompt (prompt_pre bakes <img>+<image>x256+</img>) into the SHARED sharded decode. Decode = the SAME 8-chained-part W8 decode as the NB-8 LLM (nemotron-vl-8b-llm.json): the ~7GB W8 exceeds the v81 per-context ceiling so the 32 layers split into 8 parts, LOCAL KV ring, llama3-scaled RoPE (theta 500000, factor 8). MAXCTX 512 (32 q-heads wide-attention rule). Host weight blobs live in fixture_dir 'vlm' (patch_w/cls/pos + proj_* + prompt_pre/post.txt). VLM needs ADSP_LIBRARY_PATH set for the HTP skel.",
|
| 4 |
"model": {
|
| 5 |
"name": "nemotron-vl-8b-vlm",
|
| 6 |
"family": "vlm",
|
|
|
|
| 1 |
{
|
| 2 |
"schema_version": 1,
|
|
|
|
| 3 |
"model": {
|
| 4 |
"name": "nemotron-vl-8b-vlm",
|
| 5 |
"family": "vlm",
|