p150 ETH-dispatch compliance (2026-10-05): default ETH dispatch, 1 CQ, 12x10 in Python API and server; numbers re-measured
Browse filesThe Python API and the HTTP server now open the device with dispatch on ETH cores, 1 command queue and a 12x10 compute grid by default (the p150 target). Worker dispatch stays only as an explicit opt-in (11x10 on a p150). Card numbers re-measured in this configuration and independently verified; outputs unchanged. tt-model.yaml and SERVING.md are unchanged; they still describe the container image.
- GPU_COMPARISON.md +25 -0
- OPT_REPORT.md +91 -4
- PYTHON.md +8 -0
- README.md +28 -25
- VERIFICATION_2026-10-03.md +74 -0
- code/bench_breakdown.py +4 -4
- code/eval_eth3d.py +4 -1
- code/eval_mast3r.py +4 -2
- code/make_demo.py +3 -2
- code/mast3r_p150/device.py +15 -5
- code/models/server/app.py +8 -1
- code/models/tests/test_fused_host.py +4 -2
- code/test_mast3r.py +4 -2
- code/tools_prof/eth_validate.sh +25 -0
- code/tools_prof/real_pair_acc.py +3 -3
- code/tools_prof/sym_check.py +3 -2
GPU_COMPARISON.md
CHANGED
|
@@ -251,3 +251,28 @@ The 2026-10-04 build runs the encoder and decoder MLPs at LoFi (`MAST3R_OPT` kno
|
|
| 251 |
- With `return_pose`, the `sym` graph now takes 36.5 ms (`sym_check.py`). Two GPU forwards with bf16 weights and SDPA take 70.6 ms.
|
| 252 |
- The newer `code/` writes a stored npz. The npz encode decreases from about 220 ms to about 17 ms. The served `/predict` total decreases from about 285 ms to about 79 ms. The GPU served-like total still uses the deflated npz, so this section does not compare served totals.
|
| 253 |
- The Python API (`mast3r_p150`) measures 26.8 ms median for the device call (ETH dispatch, 1 CQ).
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 251 |
- With `return_pose`, the `sym` graph now takes 36.5 ms (`sym_check.py`). Two GPU forwards with bf16 weights and SDPA take 70.6 ms.
|
| 252 |
- The newer `code/` writes a stored npz. The npz encode decreases from about 220 ms to about 17 ms. The served `/predict` total decreases from about 285 ms to about 79 ms. The GPU served-like total still uses the deflated npz, so this section does not compare served totals.
|
| 253 |
- The Python API (`mast3r_p150`) measures 26.8 ms median for the device call (ETH dispatch, 1 CQ).
|
| 254 |
+
|
| 255 |
+
## Update 2026-10-05: p150 configuration (ETH dispatch, 1 CQ, 12x10)
|
| 256 |
+
|
| 257 |
+
The GPU numbers did not change. The Blackhole numbers in the 2026-10-03 and 2026-10-04 sections used the served path config with worker (Tensix) dispatch and 2 CQ. On a single p150, worker dispatch gives an 11x10 grid, so those numbers do not represent a p150. The newer `code/` uses dispatch on ETH cores, 1 CQ and the 12x10 grid by default. An independent verifier measured the Blackhole numbers again in this configuration (`bench_breakdown.py`, uint8 input, 30 iterations, 2 runs). The outputs are bit-identical to the 2-CQ outputs.
|
| 258 |
+
|
| 259 |
+
- Model call: 27.7 ms (run medians 27.83 / 27.62 ms). The previous value was 26.5 ms (2 CQ).
|
| 260 |
+
- Device trace: 23.8 ms (run medians 23.94 / 23.74 ms). The previous value was 23.7 ms (2 CQ).
|
| 261 |
+
- With 1 CQ, the head-1 readback does not overlap the head-2 compute. Thus the model call is about 1.2 ms slower. The device trace is the same within noise.
|
| 262 |
+
|
| 263 |
+
| GPU variant / precision | GPU incl_h2d | ratio vs 27.7 ms | GPU excl_h2d | ratio vs 23.8 ms |
|
| 264 |
+
|---|---:|---:|---:|---:|
|
| 265 |
+
| ref, fp32 strict | 103.409 | 0.27 (Blackhole 3.73x faster) | 102.278 | 0.23 (Blackhole 4.30x faster) |
|
| 266 |
+
| ref, tf32 | 70.753 | 0.39 (Blackhole 2.55x faster) | 69.983 | 0.34 (Blackhole 2.94x faster) |
|
| 267 |
+
| ref, bf16 autocast | 63.753 | 0.43 (Blackhole 2.30x faster) | 62.613 | 0.38 (Blackhole 2.63x faster) |
|
| 268 |
+
| ref, fp16 autocast | 56.552 | 0.49 (Blackhole 2.04x faster) | 55.505 | 0.43 (Blackhole 2.33x faster) |
|
| 269 |
+
| bf16_weights, eager | 42.329 | 0.65 (Blackhole 1.53x faster) | 41.277 | 0.58 (Blackhole 1.73x faster) |
|
| 270 |
+
| fp16_weights, eager | 39.459 | 0.70 (Blackhole 1.42x faster) | 38.308 | 0.62 (Blackhole 1.61x faster) |
|
| 271 |
+
| SDPA, bf16_weights, eager | 35.327 | 0.78 (Blackhole 1.28x faster) | 34.285 | 0.69 (Blackhole 1.44x faster) |
|
| 272 |
+
| fp16 autocast + compile reduce-overhead | 23.214 | 1.19 (GPU 1.19x faster) | 22.149 | 1.07 (GPU 1.07x faster) |
|
| 273 |
+
| bf16_weights + compile reduce-overhead | 21.106 | 1.31 (GPU 1.31x faster) | 20.109 | 1.18 (GPU 1.18x faster) |
|
| 274 |
+
|
| 275 |
+
- `torch.compile` with CUDA graphs is still faster on the GPU: 1.19-1.31x on the model call and 1.07-1.18x on the device forward.
|
| 276 |
+
- With `return_pose`, the `sym` graph takes 38.0-38.9 ms (`sym_check.py`). Two GPU forwards with bf16 weights and SDPA take 70.6 ms.
|
| 277 |
+
- The served `/predict` npz total is 80-83 ms (stored npz). The GPU served-like total still uses the deflated npz, so this section does not compare served totals.
|
| 278 |
+
- The Python API (`mast3r_p150`) measures 26.9-27.0 ms median for the device call.
|
OPT_REPORT.md
CHANGED
|
@@ -1,6 +1,10 @@
|
|
| 1 |
# mast3r-p150 — 12x10 optimisation report (Galaxy BH chips 9, 3, 18)
|
| 2 |
|
| 3 |
-
**Status (2026-10-
|
|
|
|
|
|
|
|
|
|
|
|
|
| 4 |
-72 % server time / -60 % client wall (stored npz), png -34 %, new binary route /predict_npz (-70 % client wall),
|
| 5 |
and LoFi-RNE MLP linears (served e2e -7.2 %, all gates pass, two ungated metrics slightly worse, disclosed).**
|
| 6 |
|
|
@@ -21,9 +25,92 @@ Baseline and its profile: `OPT_BASELINE.md`. All knobs are `MAST3R_OPT` entries
|
|
| 21 |
Default set now: out, mm, dpt, l1, dptf, gelut, lnfold, phase, hostcol, phase0, decb, resl1, ropes, sym, tsplit,
|
| 22 |
sdpa2, hrope, hcat, mln, mc2d, mc2dr, resl1d, addln, cdbl, pil, til, mm32, dmm, dadd, addln2, rchain, dlin, dfront, dshard,
|
| 23 |
ups1, tailf, pemm, pcat, p1hs, hf3, gpoly, qkx, sdbl, pilhs, tups, sgat, tapm, obs, obsd, resmm, resmmd, esplit, demb, hrqk, lofie, lofid
|
| 24 |
-
(gelua and lofi are off; lofie / lofid since the 2026-10-04 audit integration, below).
|
| 25 |
-
|
| 26 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 27 |
|
| 28 |
## Audit integration 2026-10-04 (chip 1, start = ad39150, final = this commit)
|
| 29 |
|
|
|
|
| 1 |
# mast3r-p150 — 12x10 optimisation report (Galaxy BH chips 9, 3, 18)
|
| 2 |
|
| 3 |
+
**Status (2026-10-05): p150 ETH-dispatch compliance done (section "p150 ETH-dispatch compliance 2026-10-05"): the
|
| 4 |
+
server, the harnesses and the Python API default to ETH dispatch + 1 CQ + 12x10; outputs bit-identical to the old
|
| 5 |
+
Tensix 2-CQ served default; served forward 25.6 -> 26.8 ms (lost CQ1 readback overlap), all gates pass.**
|
| 6 |
+
|
| 7 |
+
Previous status (2026-10-04): audit integration done (section "Audit integration 2026-10-04"): served npz request
|
| 8 |
-72 % server time / -60 % client wall (stored npz), png -34 %, new binary route /predict_npz (-70 % client wall),
|
| 9 |
and LoFi-RNE MLP linears (served e2e -7.2 %, all gates pass, two ungated metrics slightly worse, disclosed).**
|
| 10 |
|
|
|
|
| 25 |
Default set now: out, mm, dpt, l1, dptf, gelut, lnfold, phase, hostcol, phase0, decb, resl1, ropes, sym, tsplit,
|
| 26 |
sdpa2, hrope, hcat, mln, mc2d, mc2dr, resl1d, addln, cdbl, pil, til, mm32, dmm, dadd, addln2, rchain, dlin, dfront, dshard,
|
| 27 |
ups1, tailf, pemm, pcat, p1hs, hf3, gpoly, qkx, sdbl, pilhs, tups, sgat, tapm, obs, obsd, resmm, resmmd, esplit, demb, hrqk, lofie, lofid
|
| 28 |
+
(gelua and lofi are off; lofie / lofid since the 2026-10-04 audit integration, below). Since 2026-10-05 the serve
|
| 29 |
+
config pins `MAST3R_DISPATCH=auto` + `MAST3R_CQS=1` (ETH dispatch, 1 CQ, 12x10 = the p150 configuration; section
|
| 30 |
+
"p150 ETH-dispatch compliance 2026-10-05"). Before that it pinned `MAST3R_CQS=2` (split trace, head-1 readback on CQ1),
|
| 31 |
+
which needs Tensix dispatch: 12x10 on the Galaxy, 11x10 on a p150. `MAST3R_CQS=2` stays as a Galaxy-only opt-in.
|
| 32 |
+
With `hostcol` the server uploads the views as uint8.
|
| 33 |
+
|
| 34 |
+
## p150 ETH-dispatch compliance 2026-10-05 (chip 16, start = d02a5c2, final = this commit)
|
| 35 |
+
|
| 36 |
+
Requirement: on a single p150 the model must use the 12x10 grid, so dispatch must run on ETH cores (stock Tensix
|
| 37 |
+
dispatch leaves 11x10). ETH dispatch in the patched tt-metal has 1 CQ. On this Galaxy, Tensix dispatch still gives
|
| 38 |
+
12x10 (13 Tensix columns), so numbers measured with it do not represent a p150. p150-equivalent here = ETH dispatch,
|
| 39 |
+
`num_command_queues=1`, grid capped at 12x10. (The first attempt was cut short by host crash #5; its WIP commit
|
| 40 |
+
2cf38a7 was reviewed, its host test fixed and every number below measured again after the reboot.)
|
| 41 |
+
|
| 42 |
+
### Audit (default of every device-open path)
|
| 43 |
+
|
| 44 |
+
| path | before (d02a5c2) | after |
|
| 45 |
+
|---|---|---|
|
| 46 |
+
| Python API `Mast3rP150.from_pretrained` (`mast3r_p150/device.py`) | ETH, 1 CQ, 12x10 (`dispatch="auto"`; `MAST3R_CQS` unset = 1) | unchanged; `MAST3R_CQS=2` now warns (Galaxy-only, 11x10 on p150) |
|
| 47 |
+
| HTTP server `models/server/app.py` | `ttnn.open_device(**open_kwargs)` = Tensix dispatch; CQs from `MAST3R_CQS` (code default 1) | `mast3r_p150.device.open_device`: ETH, 1 CQ, 12x10; `/info` -> `device_config` |
|
| 48 |
+
| `tt-model.yaml` `serve.env` (+ SERVING.md) | `MAST3R_CQS: "2"` -> **Tensix dispatch, 2 CQ**, 12x10 on the Galaxy / 11x10 on a p150 | `MAST3R_DISPATCH: "auto"`, `MAST3R_CQS: "1"` -> ETH, 1 CQ, 12x10 |
|
| 49 |
+
| `test_mast3r.py` (card PCC + latency row) | Tensix; validation ran it with `MAST3R_CQS=2` | shared opener: ETH, 1 CQ, 12x10 |
|
| 50 |
+
| `eval_mast3r.py`, `make_demo.py --local`, `eval_eth3d.py` | Tensix, 1 CQ (eval_eth3d: plain open) | shared opener: ETH, 1 CQ, 12x10 |
|
| 51 |
+
| `bench_breakdown.py` (OPT_REPORT trace / e2e rows) | default `--mode worker11` (Tensix, 11x10); eth12 rows used `--mode eth12`; served rows `MAST3R_CQS=2 --mode worker` | default `--mode eth12` (ETH, 1 CQ, 12x10); worker modes stay as explicit A/B |
|
| 52 |
+
| `tools_prof/real_pair_acc.py` (real-pair gate) | Tensix unless `MAST3R_ETH=1` | ETH unless `MAST3R_ETH=0` |
|
| 53 |
+
| `tools_prof/sym_check.py` (pose path) | Tensix; validation ran it with `MAST3R_CQS=2` | shared opener: ETH, 1 CQ |
|
| 54 |
+
| `test_api_device.py`, `test_warmup_device.py`, `first_call_bench.py` | via `from_pretrained`: ETH, 1 CQ, 12x10 | unchanged |
|
| 55 |
+
| `profile_eager.py` | default `--mode eth12` | unchanged |
|
| 56 |
+
|
| 57 |
+
1-CQ path: the model already had one (`MAST3R_CQS=1`, one trace on CQ0, both heads read on CQ0 after it); only the
|
| 58 |
+
served default used CQ1. No new device code. Fallback: without the patch (marker in `tt_metal/impl/dispatch/topology.cpp`)
|
| 59 |
+
`auto` uses Tensix dispatch with a `RuntimeWarning` + log line naming the 11x10 grid; an ETH open that raises also
|
| 60 |
+
falls back with a warning (unless `MAST3R_DISPATCH=eth`). `MAST3R_DISPATCH=eth` with `MAST3R_CQS=2` is an error.
|
| 61 |
+
|
| 62 |
+
### Bit identity (ETH 1 CQ vs the previous served default, Tensix 2 CQ; same chip, same session)
|
| 63 |
+
|
| 64 |
+
- `bench_breakdown.py --dump` (ETH) / `--cmp` (`MAST3R_CQS=2 --mode worker`): float randn pair and uint8 pair, both
|
| 65 |
+
heads bit-identical (n_diff = 0).
|
| 66 |
+
- HTTP: two ETH servers vs one Tensix 2-CQ server: npz arrays (`/predict` and `/predict_npz`) and PNG bytes identical.
|
| 67 |
+
- `test_api_device.py` (ETH): `model(...)` == the server pipeline bit for bit, with and without pose.
|
| 68 |
+
|
| 69 |
+
### Accuracy (`code/tools_prof/eth_validate.sh 16`; all in ETH, 1 CQ, 12x10; gate code unchanged; all gates pass)
|
| 70 |
+
|
| 71 |
+
| metric | published (audit integration 2026-10-04) | now (ETH, 1 CQ) | gate |
|
| 72 |
+
|---|---|---|---|
|
| 73 |
+
| randn head1 / head2 (all), float | 0.99862 / 0.99869 (0.99891) | 0.99862 / 0.99869 (0.99891) | 0.998 |
|
| 74 |
+
| `test_mast3r.py --layer end_to_end` | 0.9985 PASS (2 CQ) | 0.9985 PASS (head1 0.9985, head2 0.9982) | 0.99 |
|
| 75 |
+
| real pair float pts3d h1 / h2, conf h1 / h2 | 0.99016 / 0.99088, 0.99269 / 0.99486 | same digits | 0.989 / 0.99 |
|
| 76 |
+
| served-path u8 real pair pts3d h1 / h2, conf h1 / h2 | 0.99015 / 0.99097, 0.99246 / 0.99519 (2 CQ) | same digits | 0.989 / 0.99 |
|
| 77 |
+
| sym out_ii / out_ji / out_jj vs two-pass | bit-identical | bit-identical | identical |
|
| 78 |
+
| synthetic u8 head1 / head2 (not gated) | 0.99593 / 0.99927 | 0.99593 / 0.99927 | |
|
| 79 |
+
|
| 80 |
+
### Performance before / after (chip 16; medians, min in brackets)
|
| 81 |
+
|
| 82 |
+
| metric (method) | published | p150 config now (ETH, 1 CQ, 12x10) | Tensix 2 CQ, same session (Galaxy-only) |
|
| 83 |
+
|---|---|---|---|
|
| 84 |
+
| trace, float (`bench_breakdown`, 30 it) | 23.32 ms eth12 | 23.42 (23.07) | 23.67 (23.15) |
|
| 85 |
+
| e2e `model(img1, img2)`, float | 28.53 eth12 | 28.38 (27.86) | 27.55 (26.89) |
|
| 86 |
+
| e2e, uint8 views (served input path) | 26.27 / 26.36 (worker 2 CQ) | 27.83 (27.37) | 26.64 (26.36) |
|
| 87 |
+
| `test_mast3r.py --layer end_to_end --runs 25` | 27.51 ms (2 CQ, round 9) | 27.77 ms | |
|
| 88 |
+
| pose: single pair / `sym` / two-pass e2e (`sym_check`) | 26.40 / 36.98 / 52.48 (2 CQ) | 27.22 / 38.89 / 54.90 | |
|
| 89 |
+
| server `timing_ms.forward`, npz (2 x 30 req) | 25.5 / 25.57 (2 CQ) | 26.55-26.98 (2 servers x 2 runs) | 25.63 |
|
| 90 |
+
| `/predict` npz total / client wall | 81.5-83.3 / 130-140 | 80.3-82.5 / 123-136 | 79.4-80.0 / 127-130 |
|
| 91 |
+
| png total (2 x 15 req) | 124.5-139.6 | 131.8-143.7 | 123.9-124.7 |
|
| 92 |
+
| `/predict_npz` total / client wall | 72-76 / 96.8-102.0 | 72.7-74.0 / 97.5-104.5 | 71.8-72.0 / 96.8-100.4 |
|
| 93 |
+
| Python API device call `timing_ms['forward']` (30 calls) | 26.7 (26.4) | 26.99 (26.54) | |
|
| 94 |
+
| `model(PIL, PIL)` / `model(path, path)` whole call | 41.9 / 57.1 | 46.75 / 58.65 | |
|
| 95 |
+
| `predict_pairs` per pair | 27.9 | 28.09 | |
|
| 96 |
+
| startup (`from_pretrained`, warm cache) | ~9 s | 7.5 s (device warm-up 5.8 s, host 1.1 s) | |
|
| 97 |
+
| first call / warm same input (`test_warmup_device`): pair, pose, batch | 0.92x, 1.08x, 0.99x | 0.97x, 1.09x, 1.00x (gate 10 % + 5 ms: pass) | |
|
| 98 |
+
| smoke `--require-pose --npz-route` | PASS | PASS x2 | PASS |
|
| 99 |
+
|
| 100 |
+
Reading: on the device the configurations are the same (trace 23.4 vs 23.7 ms, within noise). The p150 config loses
|
| 101 |
+
the CQ1 overlap of the head-1 readback: the served forward is about 1.1 ms (4 %) slower than the Galaxy-only 2-CQ
|
| 102 |
+
mode, and the uint8 e2e about 1.2 ms. Request totals move by 1-3 ms, inside the host noise. The 2-CQ numbers are not
|
| 103 |
+
reachable on a p150 at 12x10, so the ETH column is the honest p150 number. Logs:
|
| 104 |
+
`<scratchpad>/mast3r-p150-eth-reval/` (`val/*.log`, `http/ab_*.txt`, `api.json`, `warmup.json`).
|
| 105 |
+
|
| 106 |
+
### Incident during this task
|
| 107 |
+
|
| 108 |
+
Before the chip was claimed, a host-test command (`pytest models/tests`) also collected the two device tests
|
| 109 |
+
and ran them for about 2 minutes (04:53-04:56 UTC) without `chipenv.sh`, so with every Galaxy chip visible
|
| 110 |
+
(device 0 = UMD chip 0, not in the pool). The run reached its summary (62 passed, 2 failed) and was interrupted
|
| 111 |
+
(SIGINT) while it closed the device. A chip-2 hang reported
|
| 112 |
+
by another agent at 04:55 overlaps this window; a link was not proven. Host tests are now run with
|
| 113 |
+
`TT_VISIBLE_DEVICES=none` and an explicit file list.
|
| 114 |
|
| 115 |
## Audit integration 2026-10-04 (chip 1, start = ad39150, final = this commit)
|
| 116 |
|
PYTHON.md
CHANGED
|
@@ -77,6 +77,9 @@ Startup takes about 67 s with an empty kernel cache and about 9 s with a warm ca
|
|
| 77 |
Of the 9 s, the default warm-up takes about 7 s (device graphs about 6 s, host about 1 s).
|
| 78 |
The device-graph knobs (`TT_FUSED`, `MAST3R_OPT`, `MAST3R_CQS`, ...) come from the environment, as in the server.
|
| 79 |
Their defaults are the published configuration.
|
|
|
|
|
|
|
|
|
|
| 80 |
You can open only one model at a time in a process. `model.info` shows the dispatch mode, the grid and the weights path.
|
| 81 |
|
| 82 |
## Calling the model
|
|
@@ -256,5 +259,10 @@ The object is thread-safe because a lock serialises the device calls.
|
|
| 256 |
| `predict_pairs`, per pair | 27.9 |
|
| 257 |
| `model(..., return_pose=True)` whole call | 206 / 200 |
|
| 258 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 259 |
The host time outside the device call is mostly the PIL bicubic resize of the 779x520 photos (about 6 ms per view).
|
| 260 |
The resize is the same as the server's, so the outputs stay bit-identical.
|
|
|
|
| 77 |
Of the 9 s, the default warm-up takes about 7 s (device graphs about 6 s, host about 1 s).
|
| 78 |
The device-graph knobs (`TT_FUSED`, `MAST3R_OPT`, `MAST3R_CQS`, ...) come from the environment, as in the server.
|
| 79 |
Their defaults are the published configuration.
|
| 80 |
+
The default is the p150 configuration: ETH dispatch, 1 command queue (`MAST3R_CQS=1`), 12x10 grid.
|
| 81 |
+
`MAST3R_CQS=2` is a Galaxy-only opt-in. It needs Tensix dispatch (on a single p150 that gives 11x10), and `from_pretrained` shows a warning.
|
| 82 |
+
The server and the harnesses (`test_mast3r.py`, `eval_mast3r.py`, `make_demo.py`) open the chip with the same function (`mast3r_p150.device.open_device`).
|
| 83 |
You can open only one model at a time in a process. `model.info` shows the dispatch mode, the grid and the weights path.
|
| 84 |
|
| 85 |
## Calling the model
|
|
|
|
| 259 |
| `predict_pairs`, per pair | 27.9 |
|
| 260 |
| `model(..., return_pose=True)` whole call | 206 / 200 |
|
| 261 |
|
| 262 |
+
Re-checked on 2026-10-05 (ETH dispatch, 1 CQ, 12x10, `test_api_device.py`, another chip): device call 27.0 / 26.5 ms,
|
| 263 |
+
`model(PIL, PIL)` 46.8 / 36.5 ms, `model(path, path)` 58.7 / 56.9 ms, `predict_pairs` 28.1 ms per pair, pose 210 / 198 ms.
|
| 264 |
+
The outputs are bit-identical to the server pipeline. Startup 7.5 s (warm-up: device 5.8 s, host 1.1 s).
|
| 265 |
+
`test_warmup_device.py`: the first call is 0.97x (`pair`), 1.09x (`pose`) and 1.00x (`batch`) of the warm calls on the same input.
|
| 266 |
+
|
| 267 |
The host time outside the device call is mostly the PIL bicubic resize of the 779x520 photos (about 6 ms per view).
|
| 268 |
The resize is the same as the server's, so the outputs stay bit-identical.
|
README.md
CHANGED
|
@@ -56,7 +56,7 @@ Run the snippet from the repo root, or give absolute paths.
|
|
| 56 |
|
| 57 |
The first call of `from_pretrained` downloads the NAVER weights to your HF cache. It also compiles the kernels and captures the metal trace (about 65 s). The package does not install ttnn. You do not need the HTTP server.
|
| 58 |
|
| 59 |
-
`Mast3rP150.from_pretrained()` prepares the plain call, the pose call (`return_pose=True`) and `predict_pairs` before it returns. Thus the first call is as fast as the later calls (within 10 %). Start-up takes about
|
| 60 |
|
| 61 |
| | name | type / default | meaning |
|
| 62 |
|---|---|---|---|
|
|
@@ -103,42 +103,45 @@ Kitchen frames 00 / 03 (VGGT example scene) → both predicted pointmaps in the
|
|
| 103 |
|
| 104 |
## Demo & Performances
|
| 105 |
|
| 106 |
-
Warm, batch 1, one 512×512 pair, fused traced graph (`MAST3R_OPT=all`), `bench_breakdown.py` with 30 iterations per run (median / min).
|
| 107 |
|
| 108 |
| Metric | Performance |
|
| 109 |
|---|---:|
|
| 110 |
-
| Model call `model(img1, img2)` (H2D + trace + D2H of both maps), served path config | **
|
| 111 |
-
| Device trace (ViT-L encoder ×2, decoder, two DPT heads), served path config | **23.
|
| 112 |
-
| Model call,
|
| 113 |
-
| Device trace,
|
| 114 |
-
| Device span (profiler cycles at 1.35 GHz,
|
| 115 |
-
| Host prep · H2D, served path config (uint8 pixels) | 0.
|
| 116 |
-
| D2H of both maps,
|
| 117 |
-
| `return_pose` forward on the real pair (`sym_check.py`): symmetric graph (`sym`) · two-pass | **
|
| 118 |
-
| Python model() call (`mast3r_p150`, `timing_ms['forward']`,
|
|
|
|
| 119 |
|
| 120 |
-
The measurement hardware is a Blackhole chip with a 12×10 compute grid and dispatch on ETH cores.
|
| 121 |
|
| 122 |
-
|
| 123 |
|
| 124 |
-
|
| 125 |
|
| 126 |
-
|
|
|
|
|
|
|
| 127 |
|---|---:|---:|---:|
|
| 128 |
-
| fp32 strict, eager | 103.4 ms | **Blackhole 3.
|
| 129 |
-
| tf32, eager | 70.8 ms | **Blackhole 2.
|
| 130 |
-
| bf16 autocast, eager | 63.8 ms | **Blackhole 2.
|
| 131 |
-
| fp16 autocast, eager | 56.6 ms | **Blackhole 2.
|
| 132 |
-
| bf16 weights, eager | 42.3 ms | **Blackhole 1.
|
| 133 |
-
| bf16 weights + SDPA, eager | 35.3 ms | **Blackhole 1.
|
| 134 |
-
| bf16 weights + `torch.compile` (CUDA graphs) | 21.1 ms | GPU 1.
|
| 135 |
|
| 136 |
-
The Blackhole lead comes from one metal trace that runs fused, block-sharded bf16 kernels on all 120 compute cores. The eager GPU rows pay for kernel launches and explicit softmax attention. The autocast rows also cast the fp32 weights again on each forward. `torch.compile` with CUDA graphs removes these costs, and then the GPU is about 1.2× faster. The newer `code/` writes a stored npz. This change decreases the npz encode from about 220 ms to about 17 ms and the served `/predict` total from about 285 ms to about
|
| 137 |
|
| 138 |
## Caveats
|
| 139 |
|
| 140 |
- Does not scale to multiple p150a in a mesh configuration. The current build uses a 12x10 compute grid of Tensix cores. To get this grid, the dispatch functions move from 10 Tensix cores to ETH cores (`patches/tt-metal-eth-dispatch.patch`). Thus, this build assumes that you do not need chip-to-chip ethernet communication.
|
| 141 |
-
- The
|
| 142 |
- Fixed 512×512 input, exactly one image pair per request (batch 1). Non-square images are gray-padded to square (`MAST3R_PREPROC=crop` centre-crops instead). The estimated focal can be 5-8 % high, so pass `intrinsics` when you know them.
|
| 143 |
- Numbers published before 2026-09-14 describe an older network with wrong decoder taps. This card does not repeat them.
|
| 144 |
- DUSt3R backbone only: no MASt3R matcher / descriptor head and no N-view global alignment. The pose is single-pair PairViewer (estimated-focal PnP-RANSAC). With `return_pose`, one symmetric graph (`sym`) computes the pose maps. Its outputs are bit-identical to the two-pass path.
|
|
@@ -156,7 +159,7 @@ The Blackhole lead comes from one metal trace that runs fused, block-sharded bf1
|
|
| 156 |
|
| 157 |
## Provenance
|
| 158 |
|
| 159 |
-
These are the exact sources the container image was built from. **`code/` has since been updated (2026-10-04 build with the LoFi MLP, the faster npz encode and the `mast3r_p150` Python package; see [`OPT_REPORT.md`](OPT_REPORT.md) and [`PYTHON.md`](PYTHON.md)) and is newer than the image.** `tt-model serve` runs the image's code until the image is rebuilt. `tt-model.yaml` and `SERVING.md` still describe the image:
|
| 160 |
|
| 161 |
| component | built from |
|
| 162 |
| --- | --- |
|
|
|
|
| 56 |
|
| 57 |
The first call of `from_pretrained` downloads the NAVER weights to your HF cache. It also compiles the kernels and captures the metal trace (about 65 s). The package does not install ttnn. You do not need the HTTP server.
|
| 58 |
|
| 59 |
+
`Mast3rP150.from_pretrained()` prepares the plain call, the pose call (`return_pose=True`) and `predict_pairs` before it returns. Thus the first call is as fast as the later calls (within 10 %). Start-up takes about 8 s with a warm kernel cache. Use `warmup_variants=["pair"]` for a faster start-up, or `model.warmup("crop")` to prepare more variants later.
|
| 60 |
|
| 61 |
| | name | type / default | meaning |
|
| 62 |
|---|---|---|---|
|
|
|
|
| 103 |
|
| 104 |
## Demo & Performances
|
| 105 |
|
| 106 |
+
Warm, batch 1, one 512×512 pair, fused traced graph (`MAST3R_OPT=all`), `bench_breakdown.py` with 30 iterations per run (median / min). All rows use the p150 configuration: dispatch on ETH cores, 1 command queue (CQ), 12×10 compute grid. The served path config is this configuration with uint8 input (the serve env of the newer `code/`).
|
| 107 |
|
| 108 |
| Metric | Performance |
|
| 109 |
|---|---:|
|
| 110 |
+
| Model call `model(img1, img2)` (H2D + trace + D2H of both maps), served path config | **27.6–27.8 ms median** (2 runs) · 26.9 ms min |
|
| 111 |
+
| Device trace (ViT-L encoder ×2, decoder, two DPT heads), served path config | **23.7–23.9 ms median** · 23.1 ms min |
|
| 112 |
+
| Model call, float input | 28.4 ms median (2 runs) · 27.9 ms min |
|
| 113 |
+
| Device trace, float input | **23.3–23.4 ms median** · 23.0 ms min |
|
| 114 |
+
| Device span (profiler cycles at 1.35 GHz, 3 replays), measured before the LoFi MLP change | 23.47–23.49 ms · 688 programs |
|
| 115 |
+
| Host prep · H2D, served path config (uint8 pixels) | 0.43 ms · 0.54 ms |
|
| 116 |
+
| D2H of both maps, served path config · float input | 2.75 ms · 3.1–3.4 ms |
|
| 117 |
+
| `return_pose` forward on the real pair (`sym_check.py`): symmetric graph (`sym`) · two-pass | **38.0–38.9 ms median** · 53.9–54.9 ms median |
|
| 118 |
+
| Python model() call (`mast3r_p150`, `timing_ms['forward']`, 30 warm calls) | **26.9–27.0 ms median** · 26.5 ms min |
|
| 119 |
+
| HTTP server, newer `code/` (`timing_ms.forward` · `/predict` npz total · `/predict_npz` total, 30 requests) | **26.6–27.0 ms median** · 80–83 ms · 73–75 ms |
|
| 120 |
|
| 121 |
+
The measurement hardware is a Blackhole chip with a 12×10 compute grid and dispatch on ETH cores (`patches/tt-metal-eth-dispatch.patch`), with 1 command queue. This is the configuration of a single p150. The 2026-10-04 build runs the encoder and decoder MLPs at LoFi (`MAST3R_OPT` knobs `lofie`, `lofid`). An independent verifier measured the rows above again on 2026-10-05 in this configuration. The device span row is from the previous build. Accuracy against the torch fp32 reference on the real kitchen pair (served path config): pts3d PCC 0.99015 / 0.99097 (gate 0.989), conf PCC 0.99246 / 0.99519 (gate 0.99); synthetic pair `test_mast3r.py` end_to_end PCC 0.9985. Details: [`VERIFICATION_2026-10-03.md`](VERIFICATION_2026-10-03.md).
|
| 122 |
|
| 123 |
+
Before 2026-10-05, this card showed a served model call of 26.5 ms. That number used worker (Tensix) dispatch with 2 CQ, which does not give a 12×10 grid on a p150. With 1 CQ, the head-1 readback does not overlap the head-2 compute, so the model call is about 1.2 ms slower. The outputs are bit-identical.
|
| 124 |
|
| 125 |
+
A whole `model(path, path)` call takes 50–59 ms median, because it includes PNG decode and resize on the host. `predict_pairs` takes 27.3–28.1 ms per pair. The Python API gives arrays that are bit-identical to the server `/predict` and `/predict_npz` outputs, including the pose.
|
| 126 |
|
| 127 |
+
RTX 5090 reference measurements (2026-09-14) are unchanged. They use the port's torch reference in PyTorch 2.11, batch 1, on the kitchen demo pair. The "incl. H2D/D2H" column compares with our model call (27.7 ms). The "forward only" column compares with our device trace (23.8 ms). Full table: [`GPU_COMPARISON.md`](GPU_COMPARISON.md).
|
| 128 |
+
|
| 129 |
+
| RTX 5090 precision | GPU incl. H2D/D2H | vs current build (27.7 ms) | GPU forward only vs ours (23.8 ms) |
|
| 130 |
|---|---:|---:|---:|
|
| 131 |
+
| fp32 strict, eager | 103.4 ms | **Blackhole 3.73× faster** | 102.3 ms: Blackhole 4.30× faster |
|
| 132 |
+
| tf32, eager | 70.8 ms | **Blackhole 2.55× faster** | 70.0 ms: Blackhole 2.94× faster |
|
| 133 |
+
| bf16 autocast, eager | 63.8 ms | **Blackhole 2.30× faster** | 62.6 ms: Blackhole 2.63× faster |
|
| 134 |
+
| fp16 autocast, eager | 56.6 ms | **Blackhole 2.04× faster** | 55.5 ms: Blackhole 2.33× faster |
|
| 135 |
+
| bf16 weights, eager | 42.3 ms | **Blackhole 1.53× faster** | 41.3 ms: Blackhole 1.73× faster |
|
| 136 |
+
| bf16 weights + SDPA, eager | 35.3 ms | **Blackhole 1.28× faster** | 34.3 ms: Blackhole 1.44× faster |
|
| 137 |
+
| bf16 weights + `torch.compile` (CUDA graphs) | 21.1 ms | GPU 1.31× faster | 20.1 ms: GPU 1.18× faster |
|
| 138 |
|
| 139 |
+
The Blackhole lead comes from one metal trace that runs fused, block-sharded bf16 kernels on all 120 compute cores. The eager GPU rows pay for kernel launches and explicit softmax attention. The autocast rows also cast the fp32 weights again on each forward. `torch.compile` with CUDA graphs removes these costs, and then the GPU is about 1.2-1.3× faster. The newer `code/` writes a stored npz. This change decreases the npz encode from about 220 ms to about 17 ms and the served `/predict` total from about 285 ms to about 80 ms.
|
| 140 |
|
| 141 |
## Caveats
|
| 142 |
|
| 143 |
- Does not scale to multiple p150a in a mesh configuration. The current build uses a 12x10 compute grid of Tensix cores. To get this grid, the dispatch functions move from 10 Tensix cores to ETH cores (`patches/tt-metal-eth-dispatch.patch`). Thus, this build assumes that you do not need chip-to-chip ethernet communication.
|
| 144 |
+
- The newer `code/` uses ETH dispatch and 1 CQ by default (`MAST3R_DISPATCH=auto`, `MAST3R_CQS=1`). `MAST3R_CQS=2` is an opt-in for Galaxy only. It needs worker (Tensix) dispatch, which gives an 11×10 grid on a p150. The container image and its `tt-model.yaml` still set `MAST3R_CQS=2` until the image is rebuilt. Without the ETH-dispatch patch, `auto` uses Tensix dispatch (11×10 on a p150) and logs a warning.
|
| 145 |
- Fixed 512×512 input, exactly one image pair per request (batch 1). Non-square images are gray-padded to square (`MAST3R_PREPROC=crop` centre-crops instead). The estimated focal can be 5-8 % high, so pass `intrinsics` when you know them.
|
| 146 |
- Numbers published before 2026-09-14 describe an older network with wrong decoder taps. This card does not repeat them.
|
| 147 |
- DUSt3R backbone only: no MASt3R matcher / descriptor head and no N-view global alignment. The pose is single-pair PairViewer (estimated-focal PnP-RANSAC). With `return_pose`, one symmetric graph (`sym`) computes the pose maps. Its outputs are bit-identical to the two-pass path.
|
|
|
|
| 159 |
|
| 160 |
## Provenance
|
| 161 |
|
| 162 |
+
These are the exact sources the container image was built from. **`code/` has since been updated (2026-10-04 build with the LoFi MLP, the faster npz encode and the `mast3r_p150` Python package; 2026-10-05: ETH dispatch and 1 CQ are the default for the server and the harnesses; see [`OPT_REPORT.md`](OPT_REPORT.md) and [`PYTHON.md`](PYTHON.md)) and is newer than the image.** `tt-model serve` runs the image's code until the image is rebuilt. `tt-model.yaml` and `SERVING.md` still describe the image:
|
| 163 |
|
| 164 |
| component | built from |
|
| 165 |
| --- | --- |
|
VERIFICATION_2026-10-03.md
CHANGED
|
@@ -120,3 +120,77 @@ Python API (`09c7d11`, ETH dispatch, 12×10 grid, 1 CQ, 30 warm calls on the dem
|
|
| 120 |
- `predict_pairs` gives 30 results. Each result is equal to one `model()` call.
|
| 121 |
- `timing_ms['forward']`: 26.82 ms median, 26.52 ms min. A whole `model(path, path)` call: 56.9 ms median. `predict_pairs`: 27.57 ms per pair.
|
| 122 |
- Startup with a warm cache: 7.4 s. Host tests (`test_api_host.py`, `test_server_host.py`): 21 passed.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 120 |
- `predict_pairs` gives 30 results. Each result is equal to one `model()` call.
|
| 121 |
- `timing_ms['forward']`: 26.82 ms median, 26.52 ms min. A whole `model(path, path)` call: 56.9 ms median. `predict_pairs`: 27.57 ms per pair.
|
| 122 |
- Startup with a warm cache: 7.4 s. Host tests (`test_api_host.py`, `test_server_host.py`): 21 passed.
|
| 123 |
+
|
| 124 |
+
## p150 ETH-dispatch compliance (2026-10-05)
|
| 125 |
+
|
| 126 |
+
On a single p150, the model must use the 12×10 compute grid. Thus dispatch must run on ETH cores, and the patched tt-metal then has 1 command queue (CQ). Worker (Tensix) dispatch gives an 11×10 grid on a p150. The measurement Galaxy chips give 12×10 also with worker dispatch, so the earlier "served path config" rows (worker dispatch, 2 CQ) do not represent a p150. Commit `c71e46d` makes ETH dispatch, 1 CQ and 12×10 the default for every device-open path. An independent verifier measured `c71e46d` against `d02a5c2` on one chip.
|
| 127 |
+
|
| 128 |
+
### Audit (default of each device-open path)
|
| 129 |
+
|
| 130 |
+
| path | before (`d02a5c2`) | after (`c71e46d`) |
|
| 131 |
+
|---|---|---|
|
| 132 |
+
| Python API `Mast3rP150.from_pretrained` | ETH, 1 CQ, 12×10 | unchanged; `MAST3R_CQS=2` gives a warning (Galaxy only) |
|
| 133 |
+
| HTTP server `models/server/app.py` | `ttnn.open_device`: worker dispatch; CQs from `MAST3R_CQS` | `mast3r_p150.device.open_device`: ETH, 1 CQ, 12×10; `/info` has `device_config` |
|
| 134 |
+
| `tt-model.yaml` serve env | `MAST3R_CQS=2`: worker dispatch, 2 CQ | `MAST3R_DISPATCH=auto`, `MAST3R_CQS=1`: ETH, 1 CQ |
|
| 135 |
+
| `test_mast3r.py` | worker dispatch (validation ran with `MAST3R_CQS=2`) | ETH, 1 CQ, 12×10 |
|
| 136 |
+
| `eval_mast3r.py`, `make_demo.py --local`, `eval_eth3d.py` | worker dispatch, 1 CQ | ETH, 1 CQ, 12×10 |
|
| 137 |
+
| `bench_breakdown.py` | default `--mode worker11` | default `--mode eth12`; worker modes stay as explicit A/B options |
|
| 138 |
+
| `tools_prof/real_pair_acc.py` | worker unless `MAST3R_ETH=1` | ETH unless `MAST3R_ETH=0` |
|
| 139 |
+
| `tools_prof/sym_check.py` | worker dispatch (run with `MAST3R_CQS=2`) | ETH, 1 CQ |
|
| 140 |
+
| `test_api_device.py`, `test_warmup_device.py`, `first_call_bench.py`, `profile_eager.py` | ETH, 1 CQ, 12×10 | unchanged |
|
| 141 |
+
|
| 142 |
+
`MAST3R_DISPATCH=eth` together with `MAST3R_CQS=2` raises an error. Without the ETH-dispatch patch, `auto` uses worker dispatch and logs a warning that names the 11×10 grid.
|
| 143 |
+
|
| 144 |
+
### Configuration check
|
| 145 |
+
|
| 146 |
+
- Python API with defaults: `ttnn.open_device` received `dispatch_core_config=ETH` and no `num_command_queues` (default 1). The device reports a 12×10 grid. A command on CQ 1 fails, so the device has 1 CQ.
|
| 147 |
+
- HTTP server with the serve env of `c71e46d`: the log shows `Device open: dispatch eth, grid 12x10, 1 CQ`. `/info` gives the same `device_config`.
|
| 148 |
+
- The `d02a5c2` server with its own serve env opened with worker dispatch and 2 CQ. This confirms the audit.
|
| 149 |
+
|
| 150 |
+
### Measured (ETH dispatch, 1 CQ, 12×10; median, min in brackets)
|
| 151 |
+
|
| 152 |
+
| measurement | claim | verifier | previous published value |
|
| 153 |
+
|---|---:|---:|---:|
|
| 154 |
+
| device trace, float (`bench_breakdown`, 30 iterations) | 23.42 (23.07) | 23.29 (23.00) | 23.32 (ETH, 2026-10-04) |
|
| 155 |
+
| model call, float | 28.38 (27.86) | 28.36 (28.09) | 28.53 (ETH, 2026-10-04) |
|
| 156 |
+
| device trace, uint8 (served input) | 23.94 (23.24) | 23.74 (23.12) | 23.57-23.84 (worker, 2 CQ) |
|
| 157 |
+
| model call, uint8 (served input) | 27.83 (27.37) | 27.62 (26.92) | 26.47-26.52 (worker, 2 CQ) |
|
| 158 |
+
| `test_mast3r.py --layer end_to_end --runs 25` | 27.77 | 26.86 | 27.51 (2 CQ) |
|
| 159 |
+
| `sym_check`: single pair / `sym` / two-pass | 27.22 / 38.89 / 54.90 | 27.00 / 37.99 / 53.93 | 26.40 / 36.98 / 52.48 (2 CQ) |
|
| 160 |
+
| HTTP `timing_ms.forward` (npz, 30 requests per run) | 26.55-26.98 | 26.57-26.85 | 25.50-25.61 (worker, 2 CQ, same session) |
|
| 161 |
+
| HTTP `/predict` npz total / client wall | 80.3-82.5 / 123-136 | 80.3-83.1 / 125-140 | 81.5-83.3 / 130-140 (2 CQ) |
|
| 162 |
+
| HTTP png total (15 requests) | 131.8-143.7 | 132.6-137.0 | 124.5-139.6 (2 CQ) |
|
| 163 |
+
| HTTP `/predict_npz` total / client wall | 72.7-74.0 / 97.5-104.5 | 72.8-75.1 / 98.4-100.8 | 72-76 / 96.8-102 (2 CQ) |
|
| 164 |
+
| Python API `timing_ms['forward']` (30 calls) | 26.99 (26.54) | 26.85 (26.47) | 26.82 (26.52) (ETH, 1 CQ) |
|
| 165 |
+
| Python API `predict_pairs`, per pair | 28.09 | 27.33 | 27.57 (ETH, 1 CQ) |
|
| 166 |
+
| Python API startup, warm cache | 7.5 s | 8.1 s | 7.4 s (ETH, 1 CQ) |
|
| 167 |
+
|
| 168 |
+
- The device trace does not change within noise. With 1 CQ, the head-1 readback does not overlap the head-2 compute. Thus the served forward is about 1.1 ms (4 %) slower than the worker 2-CQ mode, and the uint8 model call is about 1.2 ms slower. Request totals change by 1-3 ms, which is in the host noise.
|
| 169 |
+
- First call after `from_pretrained` (`test_warmup_device.py`, 3 runs): `pair` 0.98-1.04x, `pose` 0.95-1.06x, `batch` 0.89-1.13x of the warm calls. All runs pass the gate (10 % + 5 ms). The `batch` call takes about 74 ms and is host-bound, so its spread is host noise.
|
| 170 |
+
- The server smoke test (`--require-pose --npz-route`) passes.
|
| 171 |
+
|
| 172 |
+
### Accuracy (ETH dispatch, 1 CQ, 12×10; gate code did not change; all gates pass)
|
| 173 |
+
|
| 174 |
+
| metric | gate | `c71e46d` |
|
| 175 |
+
|---|---:|---:|
|
| 176 |
+
| synthetic randn pair, float, PCC head 1 / head 2 / all | ≥ 0.998 | 0.99862 / 0.99869 / 0.99891 |
|
| 177 |
+
| `test_mast3r.py --layer end_to_end` PCC | ≥ 0.998 | 0.9985 (PASS; head 1 0.9985, head 2 0.9982) |
|
| 178 |
+
| real pair, float: pts3d / conf PCC | ≥ 0.989 / ≥ 0.99 | 0.99016 / 0.99088 · 0.99269 / 0.99486 |
|
| 179 |
+
| real pair, uint8: pts3d / conf PCC | ≥ 0.989 / ≥ 0.99 | 0.99015 / 0.99097 · 0.99246 / 0.99519 |
|
| 180 |
+
| synthetic uint8 pair (not gated) | | 0.99593 / 0.99927 |
|
| 181 |
+
|
| 182 |
+
Every value is equal to the 2026-10-04 value to the printed digit.
|
| 183 |
+
|
| 184 |
+
Bit identity against the previous served default (worker dispatch, 2 CQ), same chip and session:
|
| 185 |
+
|
| 186 |
+
- `bench_breakdown.py`, float and uint8 pairs: head 1 and head 2 are bit-identical (n_diff = 0).
|
| 187 |
+
- HTTP: the npz arrays from `/predict` and `/predict_npz` are identical (n_diff = 0 for `pts3d1`, `conf1`, `pts3d2`, `conf2`). The PNG bytes are identical. The pose rotation is the same (43.3 deg).
|
| 188 |
+
- `sym_check`: `out_ii`, `out_ji` and `out_jj` of the `sym` graph are bit-identical to the two-pass path.
|
| 189 |
+
- `test_api_device.py`: `model()` is bit-identical to the server pipeline, with and without pose.
|
| 190 |
+
- Host tests (`test_fused_host.py`, `test_api_host.py`, `test_server_host.py`): 59 passed.
|
| 191 |
+
|
| 192 |
+
### Review notes
|
| 193 |
+
|
| 194 |
+
- The container image was built before this change. Its `tt-model.yaml` still sets `MAST3R_CQS=2` (worker dispatch) until the image is rebuilt.
|
| 195 |
+
- `auto` selects ETH dispatch only when the tt-metal source tree has the patch marker. If the tree does not have the patch, the server uses worker dispatch (11×10 on a p150) and logs a warning. It does not stop.
|
| 196 |
+
- `tools_prof/real_pair_acc.py` opens the device through its own helper. To use `MAST3R_CQS=2` with it, also set `MAST3R_ETH=0`. This affects only the tooling.
|
code/bench_breakdown.py
CHANGED
|
@@ -9,9 +9,9 @@ Phases per call (same work as TtDust3r.__call__ / test_mast3r.py end_to_end):
|
|
| 9 |
e2e : the public model call (host_prep + h2d + trace + d2h, non-blocking trace)
|
| 10 |
|
| 11 |
Usage (after chipenv.sh):
|
| 12 |
-
python bench_breakdown.py --mode
|
| 13 |
-
python bench_breakdown.py --mode
|
| 14 |
-
python bench_breakdown.py --mode
|
| 15 |
MAST3R_CQS=2 python bench_breakdown.py --mode worker # Tensix dispatch 12x10, 2 CQs (split trace, overlapped head-1 read)
|
| 16 |
add --u8 for uint8 pixel views (HWC-backed, as the server's PIL preprocess produces them)
|
| 17 |
|
|
@@ -31,7 +31,7 @@ sys.path.insert(0, _CODE_ROOT)
|
|
| 31 |
sys.path.insert(0, "/home/ttuser/experiments/tt-models/tools")
|
| 32 |
|
| 33 |
ap = argparse.ArgumentParser()
|
| 34 |
-
ap.add_argument("--mode", default="
|
| 35 |
ap.add_argument("--iters", type=int, default=30)
|
| 36 |
ap.add_argument("--no-ref", action="store_true")
|
| 37 |
ap.add_argument("--u8", action="store_true", help="uint8 pixel views (exact reference input (v/255-0.5)/0.5)")
|
|
|
|
| 9 |
e2e : the public model call (host_prep + h2d + trace + d2h, non-blocking trace)
|
| 10 |
|
| 11 |
Usage (after chipenv.sh):
|
| 12 |
+
python bench_breakdown.py # default --mode eth12: ETH dispatch, 1 CQ, 12x10 (= p150 + ETH dispatch)
|
| 13 |
+
python bench_breakdown.py --mode worker11 # stock p150 Tensix dispatch emulated: grid capped 11x10
|
| 14 |
+
python bench_breakdown.py --mode worker # Galaxy Tensix dispatch (12x10 here, Galaxy-only A/B)
|
| 15 |
MAST3R_CQS=2 python bench_breakdown.py --mode worker # Tensix dispatch 12x10, 2 CQs (split trace, overlapped head-1 read)
|
| 16 |
add --u8 for uint8 pixel views (HWC-backed, as the server's PIL preprocess produces them)
|
| 17 |
|
|
|
|
| 31 |
sys.path.insert(0, "/home/ttuser/experiments/tt-models/tools")
|
| 32 |
|
| 33 |
ap = argparse.ArgumentParser()
|
| 34 |
+
ap.add_argument("--mode", default="eth12", choices=["worker11", "worker", "eth12"])
|
| 35 |
ap.add_argument("--iters", type=int, default=30)
|
| 36 |
ap.add_argument("--no-ref", action="store_true")
|
| 37 |
ap.add_argument("--u8", action="store_true", help="uint8 pixel views (exact reference input (v/255-0.5)/0.5)")
|
code/eval_eth3d.py
CHANGED
|
@@ -78,7 +78,10 @@ def main():
|
|
| 78 |
pairs = pick_pairs(records, args.pairs, max_rot_deg=args.max_rot_deg)
|
| 79 |
print(f"# picked {len(pairs)} pairs with GT rot < {args.max_rot_deg}°")
|
| 80 |
|
| 81 |
-
|
|
|
|
|
|
|
|
|
|
| 82 |
if hasattr(device, "enable_program_cache"):
|
| 83 |
device.enable_program_cache()
|
| 84 |
try:
|
|
|
|
| 78 |
pairs = pick_pairs(records, args.pairs, max_rot_deg=args.max_rot_deg)
|
| 79 |
print(f"# picked {len(pairs)} pairs with GT rot < {args.max_rot_deg}°")
|
| 80 |
|
| 81 |
+
# Same opener as the Python API / server: ETH dispatch + 1 CQ + 12x10 by default (MAST3R_DISPATCH).
|
| 82 |
+
from mast3r_p150.device import open_device as _open_port_device
|
| 83 |
+
device, _dinfo = _open_port_device(args.device_id, None, l1_small_size=32 * 1024)
|
| 84 |
+
print(f"# device: dispatch {_dinfo['dispatch']}, grid {_dinfo['grid']}, {_dinfo['num_cqs']} CQ")
|
| 85 |
if hasattr(device, "enable_program_cache"):
|
| 86 |
device.enable_program_cache()
|
| 87 |
try:
|
code/eval_mast3r.py
CHANGED
|
@@ -297,8 +297,10 @@ def main():
|
|
| 297 |
# of this port runs in); legacy / MAST3R_TRACE=0 keep the plain open.
|
| 298 |
fused = fused_config()
|
| 299 |
print(f"# port path: {'fused ' + str(fused.summary()) if fused.enabled else 'legacy (TT_FUSED=0)'}")
|
| 300 |
-
|
| 301 |
-
|
|
|
|
|
|
|
| 302 |
if hasattr(device, "enable_program_cache"):
|
| 303 |
device.enable_program_cache()
|
| 304 |
try:
|
|
|
|
| 297 |
# of this port runs in); legacy / MAST3R_TRACE=0 keep the plain open.
|
| 298 |
fused = fused_config()
|
| 299 |
print(f"# port path: {'fused ' + str(fused.summary()) if fused.enabled else 'legacy (TT_FUSED=0)'}")
|
| 300 |
+
# Same opener as the Python API / server: ETH dispatch + 1 CQ + 12x10 by default (MAST3R_DISPATCH).
|
| 301 |
+
from mast3r_p150.device import open_device as _open_port_device
|
| 302 |
+
device, _dinfo = _open_port_device(args.device_id, None, fused=fused, l1_small_size=32 * 1024)
|
| 303 |
+
print(f"# device: dispatch {_dinfo['dispatch']}, grid {_dinfo['grid']}, {_dinfo['num_cqs']} CQ")
|
| 304 |
if hasattr(device, "enable_program_cache"):
|
| 305 |
device.enable_program_cache()
|
| 306 |
try:
|
code/make_demo.py
CHANGED
|
@@ -118,8 +118,9 @@ def predict_local(paths, return_pose: bool = False):
|
|
| 118 |
mode = preprocess_mode()
|
| 119 |
imgs, recs = zip(*[preprocess_image(Image.open(p), IMG_SIZE, mode) for p in paths])
|
| 120 |
fused = port.fused_config()
|
| 121 |
-
device
|
| 122 |
-
|
|
|
|
| 123 |
try:
|
| 124 |
state = load_checkpoint()
|
| 125 |
with torch.no_grad():
|
|
|
|
| 118 |
mode = preprocess_mode()
|
| 119 |
imgs, recs = zip(*[preprocess_image(Image.open(p), IMG_SIZE, mode) for p in paths])
|
| 120 |
fused = port.fused_config()
|
| 121 |
+
from mast3r_p150.device import open_device as _open_port_device
|
| 122 |
+
device, _ = _open_port_device(int(os.environ.get("TT_DEVICE_ID", "0")), None, fused=fused,
|
| 123 |
+
l1_small_size=32 * 1024)
|
| 124 |
try:
|
| 125 |
state = load_checkpoint()
|
| 126 |
with torch.no_grad():
|
code/mast3r_p150/device.py
CHANGED
|
@@ -10,6 +10,11 @@
|
|
| 10 |
* otherwise -> tt-metal's default (Tensix) dispatch, with a warning.
|
| 11 |
|
| 12 |
``dispatch="eth"`` / ``"worker"`` force one mode (``$MAST3R_DISPATCH`` sets the default).
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 13 |
"""
|
| 14 |
from __future__ import annotations
|
| 15 |
|
|
@@ -66,14 +71,18 @@ def resolve_dispatch(dispatch: Optional[str], num_cqs: int) -> Tuple[str, Option
|
|
| 66 |
if d == "worker":
|
| 67 |
return "worker", None
|
| 68 |
if num_cqs != 1:
|
| 69 |
-
|
|
|
|
|
|
|
|
|
|
| 70 |
patched = eth_dispatch_patch_present()
|
| 71 |
if patched:
|
| 72 |
return "eth", None
|
| 73 |
why = ("the tt-metal tree has no ETH-dispatch patch" if patched is False
|
| 74 |
else "the tt-metal source tree was not found, so ETH-dispatch support is unknown")
|
| 75 |
-
return "worker", (f"mast3r_p150: {why}; using tt-metal's default (Tensix) dispatch.
|
| 76 |
-
"
|
|
|
|
| 77 |
|
| 78 |
|
| 79 |
def open_device(device_id: int = 0, dispatch: Optional[str] = None, *, fused=None, **extra):
|
|
@@ -82,7 +91,7 @@ def open_device(device_id: int = 0, dispatch: Optional[str] = None, *, fused=Non
|
|
| 82 |
(adds ``trace_region_size`` / ``num_command_queues``)."""
|
| 83 |
import ttnn
|
| 84 |
|
| 85 |
-
base =
|
| 86 |
kw = fused.open_device_kwargs(**base) if fused is not None else base
|
| 87 |
mode, warn = resolve_dispatch(dispatch, int(kw.get("num_command_queues", 1)))
|
| 88 |
if warn:
|
|
@@ -112,4 +121,5 @@ def open_device(device_id: int = 0, dispatch: Optional[str] = None, *, fused=Non
|
|
| 112 |
"before the first device open in this process)")
|
| 113 |
warnings.warn(msg, RuntimeWarning, stacklevel=3)
|
| 114 |
log.warning(msg)
|
| 115 |
-
return device, {"dispatch": mode, "grid": f"{grid[0]}x{grid[1]}",
|
|
|
|
|
|
| 10 |
* otherwise -> tt-metal's default (Tensix) dispatch, with a warning.
|
| 11 |
|
| 12 |
``dispatch="eth"`` / ``"worker"`` force one mode (``$MAST3R_DISPATCH`` sets the default).
|
| 13 |
+
``MAST3R_CQS=2`` (2 command queues) is an explicit opt-in that needs Tensix dispatch: on the
|
| 14 |
+
Galaxy that still gives 12x10, on a single p150 only 11x10; a warning says so.
|
| 15 |
+
|
| 16 |
+
The Python API (``Mast3rP150.from_pretrained``), the HTTP server (``models.server.app``) and the
|
| 17 |
+
harnesses (``test_mast3r.py``, ``eval_mast3r.py``, ``make_demo.py``) all open the chip here.
|
| 18 |
"""
|
| 19 |
from __future__ import annotations
|
| 20 |
|
|
|
|
| 71 |
if d == "worker":
|
| 72 |
return "worker", None
|
| 73 |
if num_cqs != 1:
|
| 74 |
+
# MAST3R_CQS=2 is an explicit opt-in: ETH dispatch has 1 CQ, so it needs Tensix dispatch.
|
| 75 |
+
return "worker", ("mast3r_p150: MAST3R_CQS=2 needs Tensix (worker) dispatch. This is a Galaxy-only "
|
| 76 |
+
"opt-in: on a single p150 Tensix dispatch leaves an 11x10 grid, not the tuned 12x10. "
|
| 77 |
+
"Unset MAST3R_CQS (or set it to 1) for the p150 configuration (ETH dispatch, 1 CQ, 12x10).")
|
| 78 |
patched = eth_dispatch_patch_present()
|
| 79 |
if patched:
|
| 80 |
return "eth", None
|
| 81 |
why = ("the tt-metal tree has no ETH-dispatch patch" if patched is False
|
| 82 |
else "the tt-metal source tree was not found, so ETH-dispatch support is unknown")
|
| 83 |
+
return "worker", (f"mast3r_p150: {why}; using tt-metal's default (Tensix) dispatch. One Tensix column then "
|
| 84 |
+
"runs dispatch, so a single p150 has an 11x10 compute grid, not the tuned 12x10. The published "
|
| 85 |
+
"numbers use ETH dispatch, 1 command queue and a 12x10 grid; expect lower throughput.")
|
| 86 |
|
| 87 |
|
| 88 |
def open_device(device_id: int = 0, dispatch: Optional[str] = None, *, fused=None, **extra):
|
|
|
|
| 91 |
(adds ``trace_region_size`` / ``num_command_queues``)."""
|
| 92 |
import ttnn
|
| 93 |
|
| 94 |
+
base = {"device_id": device_id, "l1_small_size": L1_SMALL_SIZE, **extra}
|
| 95 |
kw = fused.open_device_kwargs(**base) if fused is not None else base
|
| 96 |
mode, warn = resolve_dispatch(dispatch, int(kw.get("num_command_queues", 1)))
|
| 97 |
if warn:
|
|
|
|
| 121 |
"before the first device open in this process)")
|
| 122 |
warnings.warn(msg, RuntimeWarning, stacklevel=3)
|
| 123 |
log.warning(msg)
|
| 124 |
+
return device, {"dispatch": mode, "grid": f"{grid[0]}x{grid[1]}",
|
| 125 |
+
"num_cqs": int(kw.get("num_command_queues", 1)), "open_kwargs": {k: v for k, v in kw.items()}}
|
code/models/server/app.py
CHANGED
|
@@ -289,7 +289,13 @@ async def lifespan(_app: FastAPI):
|
|
| 289 |
log.info("Port path: %s", "fused %s" % fused.summary() if fused.enabled else "legacy (TT_FUSED=0)")
|
| 290 |
log.info("Preprocessing: MAST3R_PREPROC=%s (%s to square, bicubic %dx%d)", cfg.preproc,
|
| 291 |
"gray-pad" if cfg.preproc == "pad" else "centre-crop", IMG_SIZE, IMG_SIZE)
|
| 292 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 293 |
if hasattr(device, "enable_program_cache"): # already on by default; harmless
|
| 294 |
device.enable_program_cache()
|
| 295 |
STATE["device"] = device
|
|
@@ -534,6 +540,7 @@ def info() -> dict:
|
|
| 534 |
"estimated- or known-focal PnP-RANSAC",
|
| 535 |
},
|
| 536 |
"limits": {"max_images": 2, "batch": 1, "fixed_input": f"{IMG_SIZE}x{IMG_SIZE}"},
|
|
|
|
| 537 |
"fused": STATE.get("fused"), # TT_FUSED configuration the port was built with (None before startup)
|
| 538 |
"reference_numbers": {"latency_ms_per_pair_b1": 73, "pcc_vs_torch_reference": 0.9970,
|
| 539 |
"served_forward_ms_median": 73.6, "legacy_latency_ms_per_pair_b1": 234,
|
|
|
|
| 289 |
log.info("Port path: %s", "fused %s" % fused.summary() if fused.enabled else "legacy (TT_FUSED=0)")
|
| 290 |
log.info("Preprocessing: MAST3R_PREPROC=%s (%s to square, bicubic %dx%d)", cfg.preproc,
|
| 291 |
"gray-pad" if cfg.preproc == "pad" else "centre-crop", IMG_SIZE, IMG_SIZE)
|
| 292 |
+
# Dispatch (2026-10-05): the same opener as the Python API. MAST3R_DISPATCH=auto (default): ETH
|
| 293 |
+
# dispatch + 1 CQ + 12x10 grid when tt-metal has the ETH-dispatch patch (the p150 configuration),
|
| 294 |
+
# else Tensix dispatch with a warning. MAST3R_CQS=2 is a Galaxy-only opt-in (Tensix dispatch).
|
| 295 |
+
from mast3r_p150.device import open_device as _open_port_device
|
| 296 |
+
device, dinfo = _open_port_device(cfg.device_id, None, fused=fused, l1_small_size=L1_SMALL_SIZE)
|
| 297 |
+
STATE["dispatch"] = {"dispatch": dinfo["dispatch"], "grid": dinfo["grid"], "num_cqs": dinfo["num_cqs"]}
|
| 298 |
+
log.info("Device open: dispatch %s, grid %s, %d CQ", dinfo["dispatch"], dinfo["grid"], dinfo["num_cqs"])
|
| 299 |
if hasattr(device, "enable_program_cache"): # already on by default; harmless
|
| 300 |
device.enable_program_cache()
|
| 301 |
STATE["device"] = device
|
|
|
|
| 540 |
"estimated- or known-focal PnP-RANSAC",
|
| 541 |
},
|
| 542 |
"limits": {"max_images": 2, "batch": 1, "fixed_input": f"{IMG_SIZE}x{IMG_SIZE}"},
|
| 543 |
+
"device_config": STATE.get("dispatch"), # dispatch (eth | worker), compute grid, command queues
|
| 544 |
"fused": STATE.get("fused"), # TT_FUSED configuration the port was built with (None before startup)
|
| 545 |
"reference_numbers": {"latency_ms_per_pair_b1": 73, "pcc_vs_torch_reference": 0.9970,
|
| 546 |
"served_forward_ms_median": 73.6, "legacy_latency_ms_per_pair_b1": 234,
|
code/models/tests/test_fused_host.py
CHANGED
|
@@ -532,7 +532,8 @@ def test_open_device_kwargs_select_execution_mode():
|
|
| 532 |
def test_harness_and_eval_open_the_device_in_the_reported_mode():
|
| 533 |
"""``test_mast3r.py`` exposes the on-device DPT head as its own layer (the ``dpt_head``
|
| 534 |
layer is the torch bf16 host reference and runs no ttnn op) and, like ``eval_mast3r.py``
|
| 535 |
-
and the server, opens the device via ``
|
|
|
|
| 536 |
device caches before ``close_device``. Source-level checks: the scripts import ttnn only
|
| 537 |
inside ``main`` and are never executed here."""
|
| 538 |
import importlib.util
|
|
@@ -545,7 +546,8 @@ def test_harness_and_eval_open_the_device_in_the_reported_mode():
|
|
| 545 |
|
| 546 |
for script in ("test_mast3r.py", "eval_mast3r.py", os.path.join("models", "server", "app.py")):
|
| 547 |
src = open(os.path.join(_CODE_ROOT, script)).read()
|
| 548 |
-
assert "
|
|
|
|
| 549 |
assert "release_device_caches()" in src and "close_device(" in src, script
|
| 550 |
assert src.index("release_device_caches()") < src.rindex("close_device("), script
|
| 551 |
harness_src = open(os.path.join(_CODE_ROOT, "test_mast3r.py")).read()
|
|
|
|
| 532 |
def test_harness_and_eval_open_the_device_in_the_reported_mode():
|
| 533 |
"""``test_mast3r.py`` exposes the on-device DPT head as its own layer (the ``dpt_head``
|
| 534 |
layer is the torch bf16 host reference and runs no ttnn op) and, like ``eval_mast3r.py``
|
| 535 |
+
and the server, opens the device via ``mast3r_p150.device.open_device`` (which applies
|
| 536 |
+
``open_device_kwargs`` and the ETH-dispatch default) and releases the port's
|
| 537 |
device caches before ``close_device``. Source-level checks: the scripts import ttnn only
|
| 538 |
inside ``main`` and are never executed here."""
|
| 539 |
import importlib.util
|
|
|
|
| 546 |
|
| 547 |
for script in ("test_mast3r.py", "eval_mast3r.py", os.path.join("models", "server", "app.py")):
|
| 548 |
src = open(os.path.join(_CODE_ROOT, script)).read()
|
| 549 |
+
assert "_open_port_device(" in src and "fused=fused" in src, script
|
| 550 |
+
assert "ttnn.open_device(" not in src, script
|
| 551 |
assert "release_device_caches()" in src and "close_device(" in src, script
|
| 552 |
assert src.index("release_device_caches()") < src.rindex("close_device("), script
|
| 553 |
harness_src = open(os.path.join(_CODE_ROOT, "test_mast3r.py")).read()
|
code/test_mast3r.py
CHANGED
|
@@ -356,8 +356,10 @@ def main():
|
|
| 356 |
fused = fused_config()
|
| 357 |
print(f"# port path: {'fused ' + str(fused.summary()) if fused.enabled else 'legacy (TT_FUSED=0)'}")
|
| 358 |
# l1_small_size needed by ttnn.conv2d (sliding window state buffer).
|
| 359 |
-
|
| 360 |
-
|
|
|
|
|
|
|
| 361 |
if hasattr(device, "enable_program_cache"):
|
| 362 |
device.enable_program_cache()
|
| 363 |
try:
|
|
|
|
| 356 |
fused = fused_config()
|
| 357 |
print(f"# port path: {'fused ' + str(fused.summary()) if fused.enabled else 'legacy (TT_FUSED=0)'}")
|
| 358 |
# l1_small_size needed by ttnn.conv2d (sliding window state buffer).
|
| 359 |
+
# Same opener as the Python API / server: ETH dispatch + 1 CQ + 12x10 by default (MAST3R_DISPATCH).
|
| 360 |
+
from mast3r_p150.device import open_device as _open_port_device
|
| 361 |
+
device, _dinfo = _open_port_device(args.device_id, None, fused=fused, l1_small_size=32 * 1024)
|
| 362 |
+
print(f"# device: dispatch {_dinfo['dispatch']}, grid {_dinfo['grid']}, {_dinfo['num_cqs']} CQ")
|
| 363 |
if hasattr(device, "enable_program_cache"):
|
| 364 |
device.enable_program_cache()
|
| 365 |
try:
|
code/tools_prof/eth_validate.sh
ADDED
|
@@ -0,0 +1,25 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
#!/bin/bash
|
| 2 |
+
# p150 ETH-dispatch compliance validation (2026-10-05): every run in the p150-equivalent configuration
|
| 3 |
+
# (ETH dispatch, 1 CQ, 12x10) unless the name says w2cq (= the previous served default, Tensix dispatch, 2 CQs,
|
| 4 |
+
# kept here only as the bit-identity reference). Sequential, one device workload at a time.
|
| 5 |
+
# usage: eth_validate.sh CHIP OUTDIR
|
| 6 |
+
ROOT=/home/ttuser/experiments/tt-models; M=$ROOT/models/mast3r-p150
|
| 7 |
+
CHIP=${1:?chip}; L=${2:?outdir}; mkdir -p $L
|
| 8 |
+
cd $M/code || exit 1
|
| 9 |
+
source $ROOT/tools/chipenv.sh $CHIP $M/.venv >/dev/null || { echo "chipenv failed for chip $CHIP"; exit 1; }
|
| 10 |
+
export TT_WEIGHTS_REVISION=61c57447d7b0adc8a1a30b2b0adec7a8935aa2a3
|
| 11 |
+
unset MAST3R_CQS MAST3R_DISPATCH
|
| 12 |
+
echo "chip $CHIP commit $(git -C $M rev-parse --short HEAD) logs $L"
|
| 13 |
+
run() { local name=$1; shift; echo "== $name $(date -u +%H:%M:%S)"; "$@" > $L/$name.log 2>&1; echo "rc=$?"; }
|
| 14 |
+
STEPS=${STEPS:-eth12 eth12_u8 w2cq w2cq_u8 real_eth12 real_eth12_u8 test_e2e sym_check}
|
| 15 |
+
for s in $STEPS; do case $s in
|
| 16 |
+
eth12) run eth12 timeout -s INT 900 python bench_breakdown.py --dump $L/eth12_f.pt ;;
|
| 17 |
+
eth12_u8) run eth12_u8 timeout -s INT 900 python bench_breakdown.py --u8 --dump $L/eth12_u8.pt ;;
|
| 18 |
+
w2cq) run w2cq env MAST3R_CQS=2 timeout -s INT 900 python bench_breakdown.py --mode worker --cmp $L/eth12_f.pt ;;
|
| 19 |
+
w2cq_u8) run w2cq_u8 env MAST3R_CQS=2 timeout -s INT 900 python bench_breakdown.py --mode worker --u8 --cmp $L/eth12_u8.pt ;;
|
| 20 |
+
real_eth12) run real_eth12 timeout -s INT 900 python tools_prof/real_pair_acc.py ;;
|
| 21 |
+
real_eth12_u8) run real_eth12_u8 env MAST3R_U8=1 timeout -s INT 900 python tools_prof/real_pair_acc.py ;;
|
| 22 |
+
test_e2e) run test_e2e timeout -s INT 900 python test_mast3r.py --layer end_to_end --runs 25 ;;
|
| 23 |
+
sym_check) run sym_check timeout -s INT 900 python tools_prof/sym_check.py ;;
|
| 24 |
+
esac; done
|
| 25 |
+
echo "== done $(date -u +%H:%M:%S)"
|
code/tools_prof/real_pair_acc.py
CHANGED
|
@@ -1,7 +1,7 @@
|
|
| 1 |
"""Real-image accuracy gate: media/source_1.png + source_2.png through the traced device
|
| 2 |
graph vs the fp32 torch reference. Reports raw PCC per head, activated-pts3d PCC per head,
|
| 3 |
-
conf PCC, and depth rel-err. Opens the device
|
| 4 |
-
|
| 5 |
import os, sys, time
|
| 6 |
_CODE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
|
| 7 |
sys.path.insert(0, _CODE); sys.path.insert(0, "/home/ttuser/experiments/tt-models/tools")
|
|
@@ -33,7 +33,7 @@ with torch.no_grad():
|
|
| 33 |
ref_hp = load_dust3r(state)(img1, img2)
|
| 34 |
_td.F.interpolate = _orig
|
| 35 |
cfg = fused_config(); kw = cfg.open_device_kwargs(l1_small_size=32 * 1024)
|
| 36 |
-
if os.environ.get("MAST3R_ETH") == "1":
|
| 37 |
from open_device_12x10 import open_device; dev = open_device(grid="12x10", **kw)
|
| 38 |
else:
|
| 39 |
dev = ttnn.open_device(device_id=0, **kw)
|
|
|
|
| 1 |
"""Real-image accuracy gate: media/source_1.png + source_2.png through the traced device
|
| 2 |
graph vs the fp32 torch reference. Reports raw PCC per head, activated-pts3d PCC per head,
|
| 3 |
+
conf PCC, and depth rel-err. Opens the device with ETH dispatch, 12x10 (the p150 configuration,
|
| 4 |
+
default since 2026-10-05) unless MAST3R_ETH=0 (Tensix dispatch, e.g. with MAST3R_CQS=2)."""
|
| 5 |
import os, sys, time
|
| 6 |
_CODE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
|
| 7 |
sys.path.insert(0, _CODE); sys.path.insert(0, "/home/ttuser/experiments/tt-models/tools")
|
|
|
|
| 33 |
ref_hp = load_dust3r(state)(img1, img2)
|
| 34 |
_td.F.interpolate = _orig
|
| 35 |
cfg = fused_config(); kw = cfg.open_device_kwargs(l1_small_size=32 * 1024)
|
| 36 |
+
if os.environ.get("MAST3R_ETH", "1") == "1":
|
| 37 |
from open_device_12x10 import open_device; dev = open_device(grid="12x10", **kw)
|
| 38 |
else:
|
| 39 |
dev = ttnn.open_device(device_id=0, **kw)
|
code/tools_prof/sym_check.py
CHANGED
|
@@ -1,6 +1,6 @@
|
|
| 1 |
"""return_pose B-mode check: forward_symmetric(img1, img2) vs the two-pass path (model(img1, img2), model(img2, img1)):
|
| 2 |
bit-equality of out_ii / out_ji / out_jj and e2e timing (median/min of N). uint8 views (server path) unless FLOAT=1.
|
| 3 |
-
Device opened like test_mast3r / the server (
|
| 4 |
import os, sys, time, statistics, torch, ttnn
|
| 5 |
_CODE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
|
| 6 |
sys.path.insert(0, _CODE); sys.path.insert(0, "/home/ttuser/experiments/tt-models/tools")
|
|
@@ -8,7 +8,8 @@ from models.demos.mast3r.reference.torch_dust3r import load_checkpoint
|
|
| 8 |
from models.demos.mast3r.postprocess import load_image_for_dust3r
|
| 9 |
from models.demos.mast3r.tt.ttnn_dust3r import fused_config, get_model, release_device_caches
|
| 10 |
cfg = fused_config()
|
| 11 |
-
|
|
|
|
| 12 |
N = int(os.environ.get("N", "15"))
|
| 13 |
try:
|
| 14 |
media = os.path.join(os.path.dirname(_CODE), "media")
|
|
|
|
| 1 |
"""return_pose B-mode check: forward_symmetric(img1, img2) vs the two-pass path (model(img1, img2), model(img2, img1)):
|
| 2 |
bit-equality of out_ii / out_ji / out_jj and e2e timing (median/min of N). uint8 views (server path) unless FLOAT=1.
|
| 3 |
+
Device opened like test_mast3r / the server (mast3r_p150.device.open_device: ETH dispatch + 1 CQ by default; MAST3R_CQS=2 -> Tensix)."""
|
| 4 |
import os, sys, time, statistics, torch, ttnn
|
| 5 |
_CODE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
|
| 6 |
sys.path.insert(0, _CODE); sys.path.insert(0, "/home/ttuser/experiments/tt-models/tools")
|
|
|
|
| 8 |
from models.demos.mast3r.postprocess import load_image_for_dust3r
|
| 9 |
from models.demos.mast3r.tt.ttnn_dust3r import fused_config, get_model, release_device_caches
|
| 10 |
cfg = fused_config()
|
| 11 |
+
from mast3r_p150.device import open_device as _open_port_device # ETH + 1 CQ + 12x10 by default
|
| 12 |
+
dev, _dinfo = _open_port_device(0, None, fused=cfg, l1_small_size=32 * 1024); print("# device", _dinfo)
|
| 13 |
N = int(os.environ.get("N", "15"))
|
| 14 |
try:
|
| 15 |
media = os.path.join(os.path.dirname(_CODE), "media")
|