mast3r-p150

NAVER DUSt3R ViT-L two-view 3D reconstruction (the MASt3R backbone) port on one Tenstorrent Blackhole p150a. Weights: naver/DUSt3R_ViTLarge_BaseDecoder_512_dpt · Paper: arXiv:2312.14132 · Upstream code: naver/mast3r · Port: changh95/tt-mast3r

Runs on p150 (mesh P150).

Packaged and published with tt-model-manager 0.1.0 (manifest schema 5.1).

Quickstart

tt-model pull  changh95/mast3r-p150 --with-weights
tt-model serve changh95/mast3r-p150
  • Weights naver/DUSt3R_ViTLarge_BaseDecoder_512_dpt go to your HF cache; the image does not contain them.
  • Serves on port 20000 (or the next free port); ready when the log says Application startup complete.

Run with tt-cli

tt serve changh95/mast3r-p150
printf '{"image1":"%s","image2":"%s"}' "$(base64 -w0 media/source_1.png)" "$(base64 -w0 media/source_2.png)" > req.json
curl -s localhost:20000/predict -H 'Content-Type: application/json' -d @req.json
tt model stop changh95/mast3r-p150
  • POST /predict: image1, image2 (base64 PNG/JPEG, view 1 = reference camera; or images: [b64, b64]); optional output_format (npz | png, default npz), return_pose (false), intrinsics (3×3 K of view 2 in its original pixels), conf_pct (50).
  • GET /health, GET /info.

Response

{"model": "mast3r-p150", "frame": "camera_1", "input_size": 512, "depth_mode": "exp", "conf_mode": "1+exp",
 "preprocess": [{"mode": "pad", "orig_w": 512, "orig_h": 512, "offx": 0, "offy": 0, "pad_side": 512, "out_size": 512, "scale": 1.0}, {...}],
 "output_format": "npz", "npz_b64": "...",
 "npz_keys": {"pts3d1": [[512, 512, 3], "float32"], "conf1": [[512, 512], "float16"],
              "pts3d2": [[512, 512, 3], "float32"], "conf2": [[512, 512], "float16"]},
 "summary": {"view1": {"finite_fraction": 1.0, "conf_min": 1.0, "conf_mean": 3.70, "conf_max": 26.39,
                       "z_min": 0.287, "z_median": 0.407, "z_max": 1.102}, "view2": {...}},
 "pose": null, "pose_requested": false,
 "timing_ms": {"decode": 5.5, "preprocess": 1.9, "forward": 73.6, "encode": 160.3, "total": 241.4}}
  • npz_b64 is a base64 .npz holding pts3d1/pts3d2 float32 (512,512,3) and conf1/conf2 float16 (512,512); both pointmaps are in the camera-1 frame, depth is pts3d[..., 2] in DUSt3R's own scale (not metres), conf = 1 + exp(c) >= 1. output_format: "png" returns per-view 16-bit depth + 8-bit confidence PNGs instead.
  • preprocess maps original pixels to the 512 grid, (u, v) = ((x + offx) * scale, (y + offy) * scale) (mode = pad, the default gray pad-to-square, or crop = centre-crop to square via the MAST3R_PREPROC env; offsets are negative for crop). With return_pose: true, pose holds R, t (view 2 in the view-1 frame, X_cam2 = R @ X_cam1 + t), focal_1, focal_2 (512-px units) and used_known_intrinsics; null when PnP fails.
  • The timing_ms values above come from the container image. With the optimized code/, forward is about 28 ms (see below). The host encode code did not change.

Demo

Kitchen frames 00 / 03 (VGGT example scene) → both predicted pointmaps in the camera-1 frame, coloured by the source pixels (served npz output, code/make_demo.py).

Demo & Performances

Warm, batch 1, one 512×512 pair, fused traced graph (MAST3R_OPT=all), bench_breakdown.py with 30 iterations per run (median / min). The served path config is worker dispatch, 2 CQ and uint8 input (the tt-model.yaml serve env).

Metric Performance
Model call model(img1, img2) (H2D + trace + D2H of both maps), served path config 28.1 ms median (27.99–28.23, 2 runs) · 27.55 ms min
Device trace (ViT-L encoder ×2, decoder, two DPT heads), served path config 25.2–25.4 ms median · 24.52 ms min
Model call, ETH dispatch, 1 CQ, float input 29.5–30.0 ms median · 29.27 ms min
Device trace, ETH dispatch, 1 CQ 24.7–25.0 ms median · 24.41 ms min
Device span (profiler cycles at 1.35 GHz, ETH dispatch, 1 CQ, 3 replays) 23.47–23.49 ms · 688 programs
Host prep · H2D, served path config (uint8 pixels) 0.52 ms · 0.60 ms
D2H of both maps, ETH dispatch, 1 CQ 2.95 ms
return_pose forward on the real pair, served path config: symmetric graph (sym) · two-pass 38.0 ms median · 55.8 ms median

The measurement hardware is a Blackhole chip with a 12×10 compute grid and dispatch on ETH cores. The "served path config" rows use the same 12×10 compute grid with worker dispatch and 2 CQ. Accuracy against the torch fp32 reference on the real kitchen pair (served path config): pts3d PCC 0.99021 / 0.99086 (gate 0.989), conf PCC 0.99232 / 0.99497 (gate 0.99); synthetic pair test_mast3r.py end_to_end PCC 0.9985. Details: VERIFICATION_2026-10-03.md.

RTX 5090 reference measurements (2026-09-14) are unchanged. They use the port's torch reference in PyTorch 2.11, batch 1, on the kitchen demo pair. The "incl. H2D/D2H" column compares with our model call (28.1 ms). The "forward only" column compares with our device trace (25.3 ms). Full table: GPU_COMPARISON.md.

RTX 5090 precision GPU incl. H2D/D2H vs current build (28.1 ms) GPU forward only vs ours (25.3 ms)
fp32 strict, eager 103.4 ms Blackhole 3.68× faster 102.3 ms: Blackhole 4.04× faster
tf32, eager 70.8 ms Blackhole 2.52× faster 70.0 ms: Blackhole 2.77× faster
bf16 autocast, eager 63.8 ms Blackhole 2.27× faster 62.6 ms: Blackhole 2.47× faster
fp16 autocast, eager 56.6 ms Blackhole 2.01× faster 55.5 ms: Blackhole 2.19× faster
bf16 weights, eager 42.3 ms Blackhole 1.51× faster 41.3 ms: Blackhole 1.63× faster
bf16 weights + SDPA, eager 35.3 ms Blackhole 1.26× faster 34.3 ms: Blackhole 1.36× faster
bf16 weights + torch.compile (CUDA graphs) 21.1 ms GPU 1.33× faster 20.1 ms: GPU 1.26× faster

The Blackhole lead comes from one metal trace that runs fused, block-sharded bf16 kernels on all 120 compute cores. The eager GPU rows pay for kernel launches and explicit softmax attention. The autocast rows also cast the fp32 weights again on each forward. torch.compile with CUDA graphs removes these costs, and then the GPU is about 1.3× faster. The served /predict total stays host-bound on both sides, because the npz encode takes about 160 ms.

Caveats

  • Does not scale to multiple p150a in a mesh configuration. The current build uses a 12x10 compute grid of Tensix cores. To get this grid, the dispatch functions move from 10 Tensix cores to ETH cores (patches/tt-metal-eth-dispatch.patch). Thus, this build assumes that you do not need chip-to-chip ethernet communication.
  • The serve config sets MAST3R_CQS=2. Two CQs need worker (Tensix) dispatch. With ETH dispatch, set MAST3R_CQS=1. The outputs are bit-identical.
  • Fixed 512×512 input, exactly one image pair per request (batch 1). Non-square images are gray-padded to square (MAST3R_PREPROC=crop centre-crops instead). The estimated focal can be 5-8 % high, so pass intrinsics when you know them.
  • Numbers published before 2026-09-14 describe an older network with wrong decoder taps. This card does not repeat them.
  • DUSt3R backbone only: no MASt3R matcher / descriptor head and no N-view global alignment. The pose is single-pair PairViewer (estimated-focal PnP-RANSAC). With return_pose, one symmetric graph (sym) computes the pose maps. Its outputs are bit-identical to the two-pass path.
  • bf16 on device: on the kitchen pair, pts3d PCC vs fp32 is 0.990 / 0.991 per head and conf PCC is 0.992 / 0.995. Loosen thresholds that you tuned on the reference. Some optimizations change numerics (SDPA chunks, GELU polynomial, HiFi3 head convs, upsample kernel). OPT_REPORT.md lists each change.
  • Dense outputs come base64-encoded (.npz, or 16-bit/8-bit PNG).
  • Weights are NAVER's DUSt3R checkpoint under CC-BY-NC-SA-4.0: non-commercial use only, share-alike.
  • Not an OpenAI-compatible API; GET /v1/models is a stub so the tt-model ready card does not 404.
  • p150a power was not measured, so no efficiency comparison is made.

Licensing

  • Weights: naver/DUSt3R_ViTLarge_BaseDecoder_512_dpt, CC-BY-NC-SA-4.0 (non-commercial); fetched by serve from NAVER's repo, not redistributed here.
  • Port and serving code (code/), from changh95/tt-mast3r: the DUSt3R port follows the weights' CC-BY-NC-SA-4.0 (share-alike, non-commercial); the serving layer (code/models/server/) is Apache-2.0 per its SPDX headers; tt-metal / tt-nn are Apache-2.0.

Provenance

These are the exact sources the container image was built from. code/ has since been updated (2026-10-03 optimized build; see OPT_REPORT.md) and is newer than the image. tt-model serve runs the image's code until the image is rebuilt. tt-model.yaml and SERVING.md still describe the image:

component built from
tt-metal 8b98410e730bb504fea43a88609756e34821d91d
code/ digest (image) b45444934e1d76bf (sha256, first 16 hex digits; the current code/ differs)
built 2026-09-14T06:35:13+00:00 by tt-model 0.1.0
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for changh95/mast3r-p150

Finetuned
(3)
this model

Collection including changh95/mast3r-p150

Paper for changh95/mast3r-p150