mast3r-p150
NAVER DUSt3R ViT-L two-view 3D reconstruction (the MASt3R backbone) port on one Tenstorrent Blackhole p150a. Weights: naver/DUSt3R_ViTLarge_BaseDecoder_512_dpt · Paper: arXiv:2312.14132 · Upstream code: naver/mast3r · Port: changh95/tt-mast3r
Runs on p150 (mesh P150).
Packaged and published with tt-model-manager 0.1.0 (manifest schema 5.1).
Quickstart
tt-model pull changh95/mast3r-p150 --with-weights
tt-model serve changh95/mast3r-p150
- Weights
naver/DUSt3R_ViTLarge_BaseDecoder_512_dptgo to your HF cache; the image does not contain them. - Serves on port 20000 (or the next free port); ready when the log says
Application startup complete.
Run with tt-cli
tt serve changh95/mast3r-p150
printf '{"image1":"%s","image2":"%s"}' "$(base64 -w0 media/source_1.png)" "$(base64 -w0 media/source_2.png)" > req.json
curl -s localhost:20000/predict -H 'Content-Type: application/json' -d @req.json
tt model stop changh95/mast3r-p150
POST /predict:image1,image2(base64 PNG/JPEG, view 1 = reference camera; orimages: [b64, b64]); optionaloutput_format(npz|png, defaultnpz),return_pose(false),intrinsics(3×3 K of view 2 in its original pixels),conf_pct(50).GET /health,GET /info.
Response
{"model": "mast3r-p150", "frame": "camera_1", "input_size": 512, "depth_mode": "exp", "conf_mode": "1+exp",
"preprocess": [{"mode": "pad", "orig_w": 512, "orig_h": 512, "offx": 0, "offy": 0, "pad_side": 512, "out_size": 512, "scale": 1.0}, {...}],
"output_format": "npz", "npz_b64": "...",
"npz_keys": {"pts3d1": [[512, 512, 3], "float32"], "conf1": [[512, 512], "float16"],
"pts3d2": [[512, 512, 3], "float32"], "conf2": [[512, 512], "float16"]},
"summary": {"view1": {"finite_fraction": 1.0, "conf_min": 1.0, "conf_mean": 3.70, "conf_max": 26.39,
"z_min": 0.287, "z_median": 0.407, "z_max": 1.102}, "view2": {...}},
"pose": null, "pose_requested": false,
"timing_ms": {"decode": 5.5, "preprocess": 1.9, "forward": 73.6, "encode": 160.3, "total": 241.4}}
npz_b64is a base64.npzholdingpts3d1/pts3d2float32 (512,512,3) andconf1/conf2float16 (512,512); both pointmaps are in the camera-1 frame, depth ispts3d[..., 2]in DUSt3R's own scale (not metres),conf = 1 + exp(c) >= 1.output_format: "png"returns per-view 16-bit depth + 8-bit confidence PNGs instead.preprocessmaps original pixels to the 512 grid,(u, v) = ((x + offx) * scale, (y + offy) * scale)(mode=pad, the default gray pad-to-square, orcrop= centre-crop to square via theMAST3R_PREPROCenv; offsets are negative forcrop). Withreturn_pose: true,poseholdsR,t(view 2 in the view-1 frame,X_cam2 = R @ X_cam1 + t),focal_1,focal_2(512-px units) andused_known_intrinsics;nullwhen PnP fails.- The
timing_msvalues above come from the container image. With the optimizedcode/,forwardis about 28 ms (see below). The hostencodecode did not change.
Demo
Kitchen frames 00 / 03 (VGGT example scene) → both predicted pointmaps in the camera-1 frame, coloured by the source pixels (served npz output, code/make_demo.py).
Demo & Performances
Warm, batch 1, one 512×512 pair, fused traced graph (MAST3R_OPT=all), bench_breakdown.py with 30 iterations per run (median / min). The served path config is worker dispatch, 2 CQ and uint8 input (the tt-model.yaml serve env).
| Metric | Performance |
|---|---|
Model call model(img1, img2) (H2D + trace + D2H of both maps), served path config |
28.1 ms median (27.99–28.23, 2 runs) · 27.55 ms min |
| Device trace (ViT-L encoder ×2, decoder, two DPT heads), served path config | 25.2–25.4 ms median · 24.52 ms min |
| Model call, ETH dispatch, 1 CQ, float input | 29.5–30.0 ms median · 29.27 ms min |
| Device trace, ETH dispatch, 1 CQ | 24.7–25.0 ms median · 24.41 ms min |
| Device span (profiler cycles at 1.35 GHz, ETH dispatch, 1 CQ, 3 replays) | 23.47–23.49 ms · 688 programs |
| Host prep · H2D, served path config (uint8 pixels) | 0.52 ms · 0.60 ms |
| D2H of both maps, ETH dispatch, 1 CQ | 2.95 ms |
return_pose forward on the real pair, served path config: symmetric graph (sym) · two-pass |
38.0 ms median · 55.8 ms median |
The measurement hardware is a Blackhole chip with a 12×10 compute grid and dispatch on ETH cores. The "served path config" rows use the same 12×10 compute grid with worker dispatch and 2 CQ. Accuracy against the torch fp32 reference on the real kitchen pair (served path config): pts3d PCC 0.99021 / 0.99086 (gate 0.989), conf PCC 0.99232 / 0.99497 (gate 0.99); synthetic pair test_mast3r.py end_to_end PCC 0.9985. Details: VERIFICATION_2026-10-03.md.
RTX 5090 reference measurements (2026-09-14) are unchanged. They use the port's torch reference in PyTorch 2.11, batch 1, on the kitchen demo pair. The "incl. H2D/D2H" column compares with our model call (28.1 ms). The "forward only" column compares with our device trace (25.3 ms). Full table: GPU_COMPARISON.md.
| RTX 5090 precision | GPU incl. H2D/D2H | vs current build (28.1 ms) | GPU forward only vs ours (25.3 ms) |
|---|---|---|---|
| fp32 strict, eager | 103.4 ms | Blackhole 3.68× faster | 102.3 ms: Blackhole 4.04× faster |
| tf32, eager | 70.8 ms | Blackhole 2.52× faster | 70.0 ms: Blackhole 2.77× faster |
| bf16 autocast, eager | 63.8 ms | Blackhole 2.27× faster | 62.6 ms: Blackhole 2.47× faster |
| fp16 autocast, eager | 56.6 ms | Blackhole 2.01× faster | 55.5 ms: Blackhole 2.19× faster |
| bf16 weights, eager | 42.3 ms | Blackhole 1.51× faster | 41.3 ms: Blackhole 1.63× faster |
| bf16 weights + SDPA, eager | 35.3 ms | Blackhole 1.26× faster | 34.3 ms: Blackhole 1.36× faster |
bf16 weights + torch.compile (CUDA graphs) |
21.1 ms | GPU 1.33× faster | 20.1 ms: GPU 1.26× faster |
The Blackhole lead comes from one metal trace that runs fused, block-sharded bf16 kernels on all 120 compute cores. The eager GPU rows pay for kernel launches and explicit softmax attention. The autocast rows also cast the fp32 weights again on each forward. torch.compile with CUDA graphs removes these costs, and then the GPU is about 1.3× faster. The served /predict total stays host-bound on both sides, because the npz encode takes about 160 ms.
Caveats
- Does not scale to multiple p150a in a mesh configuration. The current build uses a 12x10 compute grid of Tensix cores. To get this grid, the dispatch functions move from 10 Tensix cores to ETH cores (
patches/tt-metal-eth-dispatch.patch). Thus, this build assumes that you do not need chip-to-chip ethernet communication. - The serve config sets
MAST3R_CQS=2. Two CQs need worker (Tensix) dispatch. With ETH dispatch, setMAST3R_CQS=1. The outputs are bit-identical. - Fixed 512×512 input, exactly one image pair per request (batch 1). Non-square images are gray-padded to square (
MAST3R_PREPROC=cropcentre-crops instead). The estimated focal can be 5-8 % high, so passintrinsicswhen you know them. - Numbers published before 2026-09-14 describe an older network with wrong decoder taps. This card does not repeat them.
- DUSt3R backbone only: no MASt3R matcher / descriptor head and no N-view global alignment. The pose is single-pair PairViewer (estimated-focal PnP-RANSAC). With
return_pose, one symmetric graph (sym) computes the pose maps. Its outputs are bit-identical to the two-pass path. - bf16 on device: on the kitchen pair, pts3d PCC vs fp32 is 0.990 / 0.991 per head and conf PCC is 0.992 / 0.995. Loosen thresholds that you tuned on the reference. Some optimizations change numerics (SDPA chunks, GELU polynomial, HiFi3 head convs, upsample kernel).
OPT_REPORT.mdlists each change. - Dense outputs come base64-encoded (
.npz, or 16-bit/8-bit PNG). - Weights are NAVER's DUSt3R checkpoint under CC-BY-NC-SA-4.0: non-commercial use only, share-alike.
- Not an OpenAI-compatible API;
GET /v1/modelsis a stub so the tt-model ready card does not 404. - p150a power was not measured, so no efficiency comparison is made.
Licensing
- Weights: naver/DUSt3R_ViTLarge_BaseDecoder_512_dpt, CC-BY-NC-SA-4.0 (non-commercial); fetched by
servefrom NAVER's repo, not redistributed here. - Port and serving code (
code/), from changh95/tt-mast3r: the DUSt3R port follows the weights' CC-BY-NC-SA-4.0 (share-alike, non-commercial); the serving layer (code/models/server/) is Apache-2.0 per its SPDX headers; tt-metal / tt-nn are Apache-2.0.
Provenance
These are the exact sources the container image was built from. code/ has since been updated (2026-10-03 optimized build; see OPT_REPORT.md) and is newer than the image. tt-model serve runs the image's code until the image is rebuilt. tt-model.yaml and SERVING.md still describe the image:
| component | built from |
|---|---|
| tt-metal | 8b98410e730bb504fea43a88609756e34821d91d |
code/ digest (image) |
b45444934e1d76bf (sha256, first 16 hex digits; the current code/ differs) |
| built | 2026-09-14T06:35:13+00:00 by tt-model 0.1.0 |
Model tree for changh95/mast3r-p150
Base model
naver/DUSt3R_ViTLarge_BaseDecoder_512_dpt