sanmonga22 commited on
Commit
4cf162b
·
verified ·
1 Parent(s): 72962ad

Prune QHexRT HNPU bundle to runtime-minimum artifacts

Browse files

Strip non-runtime _comment metadata from 1 manifest(s) without changing runtime fields.
Audit artifacts: QHexRT/hf_hnpu_audit cleanup plan, manifest reference check, and multi-agent review.

Files changed (1) hide show
  1. v81/nemotron-vl-8b-vlm.json +0 -1
v81/nemotron-vl-8b-vlm.json CHANGED
@@ -1,6 +1,5 @@
1
  {
2
  "schema_version": 1,
3
- "_comment": "Nemotron-Nano-VL-8B VLM (NVIDIA Llama-3.1-8B backbone + C-RADIOv2-H ViT-Huge vision tower) on Hexagon v81 (SM8850 / soc_model 87). The nemotron_vl_generate host-op runs the whole image->caption pipeline through QHexRT: host RADIO preprocess (input_conditioner CLIP-norm + patch_generator: im_to_patches 16x16 -> Linear 768->1280 + 8 prefix tokens + baked pos_embed) -> NPU vision graph (32 RADIO blocks, 2D I/O embed[1032,1280]->vit_feat[1032,1280], fp32 for vision-cos>=0.999) -> host projector (drop 8 prefix -> pixel_shuffle 0.5 -> [256,5120] -> mlp1 LN/Linear5120->4096/GELU/Linear4096->4096) -> 256 image tokens spliced at the <image>(128256) positions of the llama_3p1 prompt (prompt_pre bakes <img>+<image>x256+</img>) into the SHARED sharded decode. Decode = the SAME 8-chained-part W8 decode as the NB-8 LLM (nemotron-vl-8b-llm.json): the ~7GB W8 exceeds the v81 per-context ceiling so the 32 layers split into 8 parts, LOCAL KV ring, llama3-scaled RoPE (theta 500000, factor 8). MAXCTX 512 (32 q-heads wide-attention rule). Host weight blobs live in fixture_dir 'vlm' (patch_w/cls/pos + proj_* + prompt_pre/post.txt). VLM needs ADSP_LIBRARY_PATH set for the HTP skel.",
4
  "model": {
5
  "name": "nemotron-vl-8b-vlm",
6
  "family": "vlm",
 
1
  {
2
  "schema_version": 1,
 
3
  "model": {
4
  "name": "nemotron-vl-8b-vlm",
5
  "family": "vlm",