alpamayo2-super-p300x2

nvidia/Alpamayo2-Super (34B vision-language-action driving model: Qwen3-VL-32B backbone + flow-matching action expert) running end to end on ONE Tenstorrent Blackhole chip via tt-nn with fused decode kernels; 6 cameras x 4 frames + ego history in, chain-of-causation text and a 6.4 s trajectory out in 3-5 s.

Runs on p150 (mesh P150).

Packaged and published with tt-model-manager 0.1.0 (manifest schema 5.1).

Quickstart

tt-model pull  changh95/Alpamayo2-Super-p300x2 --with-weights
tt-model serve changh95/Alpamayo2-Super-p300x2

pull --with-weights downloads the Docker image and the nvidia/Alpamayo2-Super weights at 00554695e729a6ff0b6281fd2c81b18d06e33dbe (into your HF cache; they are not in the image). serve starts the model's own HTTP server on port 20000 (or the next free port, if that one is busy); the first start compiles kernels for your device, which takes several minutes, and the server is ready when it logs Application startup complete.

Run with tt-model-manager or tt-cli

tt-model serve changh95/Alpamayo2-Super-p300x2      # pulls the image + the NVIDIA weights (73 GB) into your HF cache
python code/alpamayo_tt/server/client_example.py sample.pt --url http://localhost:20000   # or build the JSON below
tt-model stop  changh95/Alpamayo2-Super-p300x2

This is a tt-dit-server package, which needs tt-model-manager from 2026-09 or later (uv tool install git+https://github.com/tenstorrent/tt-model-manager). tt serve changh95/Alpamayo2-Super-p300x2 (tt-cli) works once tt-cli's managed tt-model is at such a version: point it at a current one with TT_TOOL_BIN_TT_MODEL=$(which tt-model) tt serve changh95/Alpamayo2-Super-p300x2, or upgrade the managed copy (uv pip install --python ~/.local/share/tenstorrent/tools/tt-model/bin/python -U git+https://github.com/tenstorrent/tt-model-manager).

The first start compiles kernels and converts the weights to the device layout (a few minutes); later starts are ready in about a minute. One chip of the box is used (any free one); the other chips stay free.

API

  • POST /predict
    • images: images[camera][frame] base64 PNG/JPEG, cameras in camera_ids order, 4 frames per camera, oldest to newest (the last frame is t0). Default camera_ids [0,1,2,3,5,6] = cross-left, front-wide, cross-right, rear-left, rear-right, front-tele (the release "trajectory" profile of the PhysicalAI-AV ring).
    • ego_history_xyz [16,3] and ego_history_rot [16,3,3]: ego poses at t0-1.5 s ... t0 (0.1 s steps).
    • reasoning (default true): generate the chain-of-causation before the trajectory; false prompts for the trajectory only (about 1 s faster, same trajectory quality on the validation clips).
    • sampling: greedy, top_p (0.98), temperature (0.6), seed; num_traj_samples (1-16) draws extra flow-matching noise samples sharing the same prompt (+0.33 s each); diffusion_steps (10).
    • optional ego_future_xyz [64,3] ground truth -> ade_m / minADE_m in the response.
  • Response: trajectory_xyz [64,3] and trajectory_rot [64,3,3] (ego frame at t0, 0.1 s steps, 6.4 s), chain_of_causation, samples, timing_ms, seed.
  • GET /health, GET /info.
curl -s localhost:20000/predict -H 'Content-Type: application/json' -d @req.json | python -m json.tool | head

What runs on the chip

stage implementation time
vision tower, 24 images (17280 patches) Qwen3-VL ViT, per-image attention, deepstack, bf16 0.32 s
prefill of the ~4.6K-token prompt 64 Qwen3 layers, bfp8 attention / bfp4 MLP weights, whole-sequence SDPA 2.1 s
chain-of-causation decode traced step; per layer two fused programs (RMSNorm+QKV+q/k-norm+rope+KV write, o_proj+residual+RMSNorm+SwiGLU+residual) around sdpa_decode 56 ms/token
action expert, 10 Euler steps 64-layer 1536-wide transformer attending to the whole VLM KV cache, traced 0.33 s
unicycle integration (host) reference action_to_traj 1 ms

End to end: 2.9 s (trajectory only) to 3.5-5 s (10-30 token chain of causation) per request; ~25 GB of the chip's 32 GB DRAM. Weights: bfp8 attention, bfp4 MLP, bf16 vision, bfp8 expert, bf16 KV cache.

Accuracy vs the CPU bf16 reference (PhysicalAI-AV validation clips, greedy, seed 42)

clip teacher-forced ADE vs reference this package, ADE vs ground truth reference, ADE vs ground truth
030c760c @5.1s 0.04 m 0.195 m 0.266 m
030c760c, trajectory only 0.07 m 0.183 m 0.239 m
0347d9f9 @5.1s 0.30 m 2.08 m 2.36 m
06b483cf @5.1s 0.19 m 3.17 m 3.21 m
09ad74a2 @5.1s 0.09 m 0.377 m 0.411 m
0a1ef808 @5.1s 0.02 m 1.83 m 1.89 m

Teacher-forced = same chain-of-causation tokens and the same flow-matching noise as the reference, so it isolates the device's numerical error (bfp4 MLP weights dominate). Component PCCs: vision 0.996, text layers 0.9999, expert layers 0.9999, fused vs plain decode logits 0.99998. The reference's first reasoning token is a near-tie that bfp4 noise can flip, so the text can differ from the GPU/CPU run while staying sensible; the flow-matching sampler is multi-modal, so single-sample ADE varies with seed (use num_traj_samples).

Caveats

  • This is the 34B "Super" research/auto-labeling model, not a real-time in-vehicle policy: several seconds per request on one chip (as on an H100).
  • Fixed input profile: 6 cameras x 4 frames (1920x1080 or any size; the processor resizes to ~180 tokens per image), 16 ego-history poses; batch 1; one request at a time.
  • Not bit-exact run to run (vision tower kernels); trajectories vary at the centimetre level.

Licensing

  • Weights: nvidia/Alpamayo2-Super under OpenMDW-1.1, pulled from NVIDIA's repo at serve time, not redistributed here.
  • code/alpamayo2_super: NVlabs/alpamayo2 (Apache-2.0), used for the chat template, trajectory tokenizers and the unicycle action space.
  • code/alpamayo_tt (the Tenstorrent port, fused kernels, server): Apache-2.0.

Links

nvidia/Alpamayo2-Super · NVlabs/alpamayo2 · PhysicalAI-AV dataset · design notes and measurements: code/alpamayo_tt/README.md, code/alpamayo_tt/PLAN.md, code/alpamayo_tt/*_NOTES.md

Provenance

The exact sources the image was built from — code/ in this repo is byte-identical to the model code inside the image:

component built from
tt-metal 54a826b1d931cf5c70f1d933194cfc32ea96b2b6
code/ digest 37c7c852d3b4fd85 (sha256, first 16 hex digits)
built 2026-09-20T05:28:51+00:00 by tt-model 0.1.0
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including changh95/Alpamayo2-Super-p300x2