changh95 commited on
Commit
2c79427
·
verified ·
1 Parent(s): b8d0ee3

p150 ETH-dispatch compliance (2026-10-05): default ETH dispatch, 1 CQ, 12x10 in Python API and server; numbers re-measured

Browse files

The Python API and the HTTP server now open the device with dispatch on ETH cores, 1 command queue and a 12x10 compute grid by default (the p150 target). Worker dispatch stays only as an explicit opt-in (11x10 on a p150). Card numbers re-measured in this configuration and independently verified; outputs unchanged. tt-model.yaml and SERVING.md are unchanged; they still describe the container image.

GPU_COMPARISON.md CHANGED
@@ -251,3 +251,28 @@ The 2026-10-04 build runs the encoder and decoder MLPs at LoFi (`MAST3R_OPT` kno
251
  - With `return_pose`, the `sym` graph now takes 36.5 ms (`sym_check.py`). Two GPU forwards with bf16 weights and SDPA take 70.6 ms.
252
  - The newer `code/` writes a stored npz. The npz encode decreases from about 220 ms to about 17 ms. The served `/predict` total decreases from about 285 ms to about 79 ms. The GPU served-like total still uses the deflated npz, so this section does not compare served totals.
253
  - The Python API (`mast3r_p150`) measures 26.8 ms median for the device call (ETH dispatch, 1 CQ).
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
251
  - With `return_pose`, the `sym` graph now takes 36.5 ms (`sym_check.py`). Two GPU forwards with bf16 weights and SDPA take 70.6 ms.
252
  - The newer `code/` writes a stored npz. The npz encode decreases from about 220 ms to about 17 ms. The served `/predict` total decreases from about 285 ms to about 79 ms. The GPU served-like total still uses the deflated npz, so this section does not compare served totals.
253
  - The Python API (`mast3r_p150`) measures 26.8 ms median for the device call (ETH dispatch, 1 CQ).
254
+
255
+ ## Update 2026-10-05: p150 configuration (ETH dispatch, 1 CQ, 12x10)
256
+
257
+ The GPU numbers did not change. The Blackhole numbers in the 2026-10-03 and 2026-10-04 sections used the served path config with worker (Tensix) dispatch and 2 CQ. On a single p150, worker dispatch gives an 11x10 grid, so those numbers do not represent a p150. The newer `code/` uses dispatch on ETH cores, 1 CQ and the 12x10 grid by default. An independent verifier measured the Blackhole numbers again in this configuration (`bench_breakdown.py`, uint8 input, 30 iterations, 2 runs). The outputs are bit-identical to the 2-CQ outputs.
258
+
259
+ - Model call: 27.7 ms (run medians 27.83 / 27.62 ms). The previous value was 26.5 ms (2 CQ).
260
+ - Device trace: 23.8 ms (run medians 23.94 / 23.74 ms). The previous value was 23.7 ms (2 CQ).
261
+ - With 1 CQ, the head-1 readback does not overlap the head-2 compute. Thus the model call is about 1.2 ms slower. The device trace is the same within noise.
262
+
263
+ | GPU variant / precision | GPU incl_h2d | ratio vs 27.7 ms | GPU excl_h2d | ratio vs 23.8 ms |
264
+ |---|---:|---:|---:|---:|
265
+ | ref, fp32 strict | 103.409 | 0.27 (Blackhole 3.73x faster) | 102.278 | 0.23 (Blackhole 4.30x faster) |
266
+ | ref, tf32 | 70.753 | 0.39 (Blackhole 2.55x faster) | 69.983 | 0.34 (Blackhole 2.94x faster) |
267
+ | ref, bf16 autocast | 63.753 | 0.43 (Blackhole 2.30x faster) | 62.613 | 0.38 (Blackhole 2.63x faster) |
268
+ | ref, fp16 autocast | 56.552 | 0.49 (Blackhole 2.04x faster) | 55.505 | 0.43 (Blackhole 2.33x faster) |
269
+ | bf16_weights, eager | 42.329 | 0.65 (Blackhole 1.53x faster) | 41.277 | 0.58 (Blackhole 1.73x faster) |
270
+ | fp16_weights, eager | 39.459 | 0.70 (Blackhole 1.42x faster) | 38.308 | 0.62 (Blackhole 1.61x faster) |
271
+ | SDPA, bf16_weights, eager | 35.327 | 0.78 (Blackhole 1.28x faster) | 34.285 | 0.69 (Blackhole 1.44x faster) |
272
+ | fp16 autocast + compile reduce-overhead | 23.214 | 1.19 (GPU 1.19x faster) | 22.149 | 1.07 (GPU 1.07x faster) |
273
+ | bf16_weights + compile reduce-overhead | 21.106 | 1.31 (GPU 1.31x faster) | 20.109 | 1.18 (GPU 1.18x faster) |
274
+
275
+ - `torch.compile` with CUDA graphs is still faster on the GPU: 1.19-1.31x on the model call and 1.07-1.18x on the device forward.
276
+ - With `return_pose`, the `sym` graph takes 38.0-38.9 ms (`sym_check.py`). Two GPU forwards with bf16 weights and SDPA take 70.6 ms.
277
+ - The served `/predict` npz total is 80-83 ms (stored npz). The GPU served-like total still uses the deflated npz, so this section does not compare served totals.
278
+ - The Python API (`mast3r_p150`) measures 26.9-27.0 ms median for the device call.
OPT_REPORT.md CHANGED
@@ -1,6 +1,10 @@
1
  # mast3r-p150 — 12x10 optimisation report (Galaxy BH chips 9, 3, 18)
2
 
3
- **Status (2026-10-04): audit integration done (section "Audit integration 2026-10-04"): served npz request
 
 
 
 
4
  -72 % server time / -60 % client wall (stored npz), png -34 %, new binary route /predict_npz (-70 % client wall),
5
  and LoFi-RNE MLP linears (served e2e -7.2 %, all gates pass, two ungated metrics slightly worse, disclosed).**
6
 
@@ -21,9 +25,92 @@ Baseline and its profile: `OPT_BASELINE.md`. All knobs are `MAST3R_OPT` entries
21
  Default set now: out, mm, dpt, l1, dptf, gelut, lnfold, phase, hostcol, phase0, decb, resl1, ropes, sym, tsplit,
22
  sdpa2, hrope, hcat, mln, mc2d, mc2dr, resl1d, addln, cdbl, pil, til, mm32, dmm, dadd, addln2, rchain, dlin, dfront, dshard,
23
  ups1, tailf, pemm, pcat, p1hs, hf3, gpoly, qkx, sdbl, pilhs, tups, sgat, tapm, obs, obsd, resmm, resmmd, esplit, demb, hrqk, lofie, lofid
24
- (gelua and lofi are off; lofie / lofid since the 2026-10-04 audit integration, below). The serve config
25
- also pins `MAST3R_CQS=2`: the trace is split into segments and readbacks overlap on CQ1. That needs a 2-CQ device, which on
26
- the Galaxy means Tensix dispatch, still 12x10. With `hostcol` the server uploads the views as uint8.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
27
 
28
  ## Audit integration 2026-10-04 (chip 1, start = ad39150, final = this commit)
29
 
 
1
  # mast3r-p150 — 12x10 optimisation report (Galaxy BH chips 9, 3, 18)
2
 
3
+ **Status (2026-10-05): p150 ETH-dispatch compliance done (section "p150 ETH-dispatch compliance 2026-10-05"): the
4
+ server, the harnesses and the Python API default to ETH dispatch + 1 CQ + 12x10; outputs bit-identical to the old
5
+ Tensix 2-CQ served default; served forward 25.6 -> 26.8 ms (lost CQ1 readback overlap), all gates pass.**
6
+
7
+ Previous status (2026-10-04): audit integration done (section "Audit integration 2026-10-04"): served npz request
8
  -72 % server time / -60 % client wall (stored npz), png -34 %, new binary route /predict_npz (-70 % client wall),
9
  and LoFi-RNE MLP linears (served e2e -7.2 %, all gates pass, two ungated metrics slightly worse, disclosed).**
10
 
 
25
  Default set now: out, mm, dpt, l1, dptf, gelut, lnfold, phase, hostcol, phase0, decb, resl1, ropes, sym, tsplit,
26
  sdpa2, hrope, hcat, mln, mc2d, mc2dr, resl1d, addln, cdbl, pil, til, mm32, dmm, dadd, addln2, rchain, dlin, dfront, dshard,
27
  ups1, tailf, pemm, pcat, p1hs, hf3, gpoly, qkx, sdbl, pilhs, tups, sgat, tapm, obs, obsd, resmm, resmmd, esplit, demb, hrqk, lofie, lofid
28
+ (gelua and lofi are off; lofie / lofid since the 2026-10-04 audit integration, below). Since 2026-10-05 the serve
29
+ config pins `MAST3R_DISPATCH=auto` + `MAST3R_CQS=1` (ETH dispatch, 1 CQ, 12x10 = the p150 configuration; section
30
+ "p150 ETH-dispatch compliance 2026-10-05"). Before that it pinned `MAST3R_CQS=2` (split trace, head-1 readback on CQ1),
31
+ which needs Tensix dispatch: 12x10 on the Galaxy, 11x10 on a p150. `MAST3R_CQS=2` stays as a Galaxy-only opt-in.
32
+ With `hostcol` the server uploads the views as uint8.
33
+
34
+ ## p150 ETH-dispatch compliance 2026-10-05 (chip 16, start = d02a5c2, final = this commit)
35
+
36
+ Requirement: on a single p150 the model must use the 12x10 grid, so dispatch must run on ETH cores (stock Tensix
37
+ dispatch leaves 11x10). ETH dispatch in the patched tt-metal has 1 CQ. On this Galaxy, Tensix dispatch still gives
38
+ 12x10 (13 Tensix columns), so numbers measured with it do not represent a p150. p150-equivalent here = ETH dispatch,
39
+ `num_command_queues=1`, grid capped at 12x10. (The first attempt was cut short by host crash #5; its WIP commit
40
+ 2cf38a7 was reviewed, its host test fixed and every number below measured again after the reboot.)
41
+
42
+ ### Audit (default of every device-open path)
43
+
44
+ | path | before (d02a5c2) | after |
45
+ |---|---|---|
46
+ | Python API `Mast3rP150.from_pretrained` (`mast3r_p150/device.py`) | ETH, 1 CQ, 12x10 (`dispatch="auto"`; `MAST3R_CQS` unset = 1) | unchanged; `MAST3R_CQS=2` now warns (Galaxy-only, 11x10 on p150) |
47
+ | HTTP server `models/server/app.py` | `ttnn.open_device(**open_kwargs)` = Tensix dispatch; CQs from `MAST3R_CQS` (code default 1) | `mast3r_p150.device.open_device`: ETH, 1 CQ, 12x10; `/info` -> `device_config` |
48
+ | `tt-model.yaml` `serve.env` (+ SERVING.md) | `MAST3R_CQS: "2"` -> **Tensix dispatch, 2 CQ**, 12x10 on the Galaxy / 11x10 on a p150 | `MAST3R_DISPATCH: "auto"`, `MAST3R_CQS: "1"` -> ETH, 1 CQ, 12x10 |
49
+ | `test_mast3r.py` (card PCC + latency row) | Tensix; validation ran it with `MAST3R_CQS=2` | shared opener: ETH, 1 CQ, 12x10 |
50
+ | `eval_mast3r.py`, `make_demo.py --local`, `eval_eth3d.py` | Tensix, 1 CQ (eval_eth3d: plain open) | shared opener: ETH, 1 CQ, 12x10 |
51
+ | `bench_breakdown.py` (OPT_REPORT trace / e2e rows) | default `--mode worker11` (Tensix, 11x10); eth12 rows used `--mode eth12`; served rows `MAST3R_CQS=2 --mode worker` | default `--mode eth12` (ETH, 1 CQ, 12x10); worker modes stay as explicit A/B |
52
+ | `tools_prof/real_pair_acc.py` (real-pair gate) | Tensix unless `MAST3R_ETH=1` | ETH unless `MAST3R_ETH=0` |
53
+ | `tools_prof/sym_check.py` (pose path) | Tensix; validation ran it with `MAST3R_CQS=2` | shared opener: ETH, 1 CQ |
54
+ | `test_api_device.py`, `test_warmup_device.py`, `first_call_bench.py` | via `from_pretrained`: ETH, 1 CQ, 12x10 | unchanged |
55
+ | `profile_eager.py` | default `--mode eth12` | unchanged |
56
+
57
+ 1-CQ path: the model already had one (`MAST3R_CQS=1`, one trace on CQ0, both heads read on CQ0 after it); only the
58
+ served default used CQ1. No new device code. Fallback: without the patch (marker in `tt_metal/impl/dispatch/topology.cpp`)
59
+ `auto` uses Tensix dispatch with a `RuntimeWarning` + log line naming the 11x10 grid; an ETH open that raises also
60
+ falls back with a warning (unless `MAST3R_DISPATCH=eth`). `MAST3R_DISPATCH=eth` with `MAST3R_CQS=2` is an error.
61
+
62
+ ### Bit identity (ETH 1 CQ vs the previous served default, Tensix 2 CQ; same chip, same session)
63
+
64
+ - `bench_breakdown.py --dump` (ETH) / `--cmp` (`MAST3R_CQS=2 --mode worker`): float randn pair and uint8 pair, both
65
+ heads bit-identical (n_diff = 0).
66
+ - HTTP: two ETH servers vs one Tensix 2-CQ server: npz arrays (`/predict` and `/predict_npz`) and PNG bytes identical.
67
+ - `test_api_device.py` (ETH): `model(...)` == the server pipeline bit for bit, with and without pose.
68
+
69
+ ### Accuracy (`code/tools_prof/eth_validate.sh 16`; all in ETH, 1 CQ, 12x10; gate code unchanged; all gates pass)
70
+
71
+ | metric | published (audit integration 2026-10-04) | now (ETH, 1 CQ) | gate |
72
+ |---|---|---|---|
73
+ | randn head1 / head2 (all), float | 0.99862 / 0.99869 (0.99891) | 0.99862 / 0.99869 (0.99891) | 0.998 |
74
+ | `test_mast3r.py --layer end_to_end` | 0.9985 PASS (2 CQ) | 0.9985 PASS (head1 0.9985, head2 0.9982) | 0.99 |
75
+ | real pair float pts3d h1 / h2, conf h1 / h2 | 0.99016 / 0.99088, 0.99269 / 0.99486 | same digits | 0.989 / 0.99 |
76
+ | served-path u8 real pair pts3d h1 / h2, conf h1 / h2 | 0.99015 / 0.99097, 0.99246 / 0.99519 (2 CQ) | same digits | 0.989 / 0.99 |
77
+ | sym out_ii / out_ji / out_jj vs two-pass | bit-identical | bit-identical | identical |
78
+ | synthetic u8 head1 / head2 (not gated) | 0.99593 / 0.99927 | 0.99593 / 0.99927 | |
79
+
80
+ ### Performance before / after (chip 16; medians, min in brackets)
81
+
82
+ | metric (method) | published | p150 config now (ETH, 1 CQ, 12x10) | Tensix 2 CQ, same session (Galaxy-only) |
83
+ |---|---|---|---|
84
+ | trace, float (`bench_breakdown`, 30 it) | 23.32 ms eth12 | 23.42 (23.07) | 23.67 (23.15) |
85
+ | e2e `model(img1, img2)`, float | 28.53 eth12 | 28.38 (27.86) | 27.55 (26.89) |
86
+ | e2e, uint8 views (served input path) | 26.27 / 26.36 (worker 2 CQ) | 27.83 (27.37) | 26.64 (26.36) |
87
+ | `test_mast3r.py --layer end_to_end --runs 25` | 27.51 ms (2 CQ, round 9) | 27.77 ms | |
88
+ | pose: single pair / `sym` / two-pass e2e (`sym_check`) | 26.40 / 36.98 / 52.48 (2 CQ) | 27.22 / 38.89 / 54.90 | |
89
+ | server `timing_ms.forward`, npz (2 x 30 req) | 25.5 / 25.57 (2 CQ) | 26.55-26.98 (2 servers x 2 runs) | 25.63 |
90
+ | `/predict` npz total / client wall | 81.5-83.3 / 130-140 | 80.3-82.5 / 123-136 | 79.4-80.0 / 127-130 |
91
+ | png total (2 x 15 req) | 124.5-139.6 | 131.8-143.7 | 123.9-124.7 |
92
+ | `/predict_npz` total / client wall | 72-76 / 96.8-102.0 | 72.7-74.0 / 97.5-104.5 | 71.8-72.0 / 96.8-100.4 |
93
+ | Python API device call `timing_ms['forward']` (30 calls) | 26.7 (26.4) | 26.99 (26.54) | |
94
+ | `model(PIL, PIL)` / `model(path, path)` whole call | 41.9 / 57.1 | 46.75 / 58.65 | |
95
+ | `predict_pairs` per pair | 27.9 | 28.09 | |
96
+ | startup (`from_pretrained`, warm cache) | ~9 s | 7.5 s (device warm-up 5.8 s, host 1.1 s) | |
97
+ | first call / warm same input (`test_warmup_device`): pair, pose, batch | 0.92x, 1.08x, 0.99x | 0.97x, 1.09x, 1.00x (gate 10 % + 5 ms: pass) | |
98
+ | smoke `--require-pose --npz-route` | PASS | PASS x2 | PASS |
99
+
100
+ Reading: on the device the configurations are the same (trace 23.4 vs 23.7 ms, within noise). The p150 config loses
101
+ the CQ1 overlap of the head-1 readback: the served forward is about 1.1 ms (4 %) slower than the Galaxy-only 2-CQ
102
+ mode, and the uint8 e2e about 1.2 ms. Request totals move by 1-3 ms, inside the host noise. The 2-CQ numbers are not
103
+ reachable on a p150 at 12x10, so the ETH column is the honest p150 number. Logs:
104
+ `<scratchpad>/mast3r-p150-eth-reval/` (`val/*.log`, `http/ab_*.txt`, `api.json`, `warmup.json`).
105
+
106
+ ### Incident during this task
107
+
108
+ Before the chip was claimed, a host-test command (`pytest models/tests`) also collected the two device tests
109
+ and ran them for about 2 minutes (04:53-04:56 UTC) without `chipenv.sh`, so with every Galaxy chip visible
110
+ (device 0 = UMD chip 0, not in the pool). The run reached its summary (62 passed, 2 failed) and was interrupted
111
+ (SIGINT) while it closed the device. A chip-2 hang reported
112
+ by another agent at 04:55 overlaps this window; a link was not proven. Host tests are now run with
113
+ `TT_VISIBLE_DEVICES=none` and an explicit file list.
114
 
115
  ## Audit integration 2026-10-04 (chip 1, start = ad39150, final = this commit)
116
 
PYTHON.md CHANGED
@@ -77,6 +77,9 @@ Startup takes about 67 s with an empty kernel cache and about 9 s with a warm ca
77
  Of the 9 s, the default warm-up takes about 7 s (device graphs about 6 s, host about 1 s).
78
  The device-graph knobs (`TT_FUSED`, `MAST3R_OPT`, `MAST3R_CQS`, ...) come from the environment, as in the server.
79
  Their defaults are the published configuration.
 
 
 
80
  You can open only one model at a time in a process. `model.info` shows the dispatch mode, the grid and the weights path.
81
 
82
  ## Calling the model
@@ -256,5 +259,10 @@ The object is thread-safe because a lock serialises the device calls.
256
  | `predict_pairs`, per pair | 27.9 |
257
  | `model(..., return_pose=True)` whole call | 206 / 200 |
258
 
 
 
 
 
 
259
  The host time outside the device call is mostly the PIL bicubic resize of the 779x520 photos (about 6 ms per view).
260
  The resize is the same as the server's, so the outputs stay bit-identical.
 
77
  Of the 9 s, the default warm-up takes about 7 s (device graphs about 6 s, host about 1 s).
78
  The device-graph knobs (`TT_FUSED`, `MAST3R_OPT`, `MAST3R_CQS`, ...) come from the environment, as in the server.
79
  Their defaults are the published configuration.
80
+ The default is the p150 configuration: ETH dispatch, 1 command queue (`MAST3R_CQS=1`), 12x10 grid.
81
+ `MAST3R_CQS=2` is a Galaxy-only opt-in. It needs Tensix dispatch (on a single p150 that gives 11x10), and `from_pretrained` shows a warning.
82
+ The server and the harnesses (`test_mast3r.py`, `eval_mast3r.py`, `make_demo.py`) open the chip with the same function (`mast3r_p150.device.open_device`).
83
  You can open only one model at a time in a process. `model.info` shows the dispatch mode, the grid and the weights path.
84
 
85
  ## Calling the model
 
259
  | `predict_pairs`, per pair | 27.9 |
260
  | `model(..., return_pose=True)` whole call | 206 / 200 |
261
 
262
+ Re-checked on 2026-10-05 (ETH dispatch, 1 CQ, 12x10, `test_api_device.py`, another chip): device call 27.0 / 26.5 ms,
263
+ `model(PIL, PIL)` 46.8 / 36.5 ms, `model(path, path)` 58.7 / 56.9 ms, `predict_pairs` 28.1 ms per pair, pose 210 / 198 ms.
264
+ The outputs are bit-identical to the server pipeline. Startup 7.5 s (warm-up: device 5.8 s, host 1.1 s).
265
+ `test_warmup_device.py`: the first call is 0.97x (`pair`), 1.09x (`pose`) and 1.00x (`batch`) of the warm calls on the same input.
266
+
267
  The host time outside the device call is mostly the PIL bicubic resize of the 779x520 photos (about 6 ms per view).
268
  The resize is the same as the server's, so the outputs stay bit-identical.
README.md CHANGED
@@ -56,7 +56,7 @@ Run the snippet from the repo root, or give absolute paths.
56
 
57
  The first call of `from_pretrained` downloads the NAVER weights to your HF cache. It also compiles the kernels and captures the metal trace (about 65 s). The package does not install ttnn. You do not need the HTTP server.
58
 
59
- `Mast3rP150.from_pretrained()` prepares the plain call, the pose call (`return_pose=True`) and `predict_pairs` before it returns. Thus the first call is as fast as the later calls (within 10 %). Start-up takes about 9 s with a warm kernel cache. Use `warmup_variants=["pair"]` for a faster start-up, or `model.warmup("crop")` to prepare more variants later.
60
 
61
  | | name | type / default | meaning |
62
  |---|---|---|---|
@@ -103,42 +103,45 @@ Kitchen frames 00 / 03 (VGGT example scene) → both predicted pointmaps in the
103
 
104
  ## Demo & Performances
105
 
106
- Warm, batch 1, one 512×512 pair, fused traced graph (`MAST3R_OPT=all`), `bench_breakdown.py` with 30 iterations per run (median / min). The served path config is worker dispatch, 2 CQ and uint8 input (the `tt-model.yaml` serve env).
107
 
108
  | Metric | Performance |
109
  |---|---:|
110
- | Model call `model(img1, img2)` (H2D + trace + D2H of both maps), served path config | **26.5 ms median** (26.47–26.52, 3 runs) · 25.86 ms min |
111
- | Device trace (ViT-L encoder ×2, decoder, two DPT heads), served path config | **23.6–23.8 ms median** · 23.17 ms min |
112
- | Model call, ETH dispatch, 1 CQ, float input | 28.3–28.4 ms median |
113
- | Device trace, ETH dispatch, 1 CQ | **23.2–23.3 ms median** |
114
- | Device span (profiler cycles at 1.35 GHz, ETH dispatch, 1 CQ, 3 replays), measured before the LoFi MLP change | 23.47–23.49 ms · 688 programs |
115
- | Host prep · H2D, served path config (uint8 pixels) | 0.52 ms · 0.60 ms |
116
- | D2H of both maps, ETH dispatch, 1 CQ | 2.95 ms |
117
- | `return_pose` forward on the real pair (`sym_check.py`): symmetric graph (`sym`) · two-pass | **36.5 ms median** · 52.1 ms median |
118
- | Python model() call (`mast3r_p150`, `timing_ms['forward']`, ETH dispatch, 1 CQ, 30 warm calls) | **26.8 ms median** · 26.5 ms min |
 
119
 
120
- The measurement hardware is a Blackhole chip with a 12×10 compute grid and dispatch on ETH cores. The "served path config" rows use the same 12×10 compute grid with worker dispatch and 2 CQ. The 2026-10-04 build runs the encoder and decoder MLPs at LoFi (`MAST3R_OPT` knobs `lofie`, `lofid`). An independent verifier measured the rows above again on this build. The device span row is from the previous build. Accuracy against the torch fp32 reference on the real kitchen pair (served path config): pts3d PCC 0.99015 / 0.99097 (gate 0.989), conf PCC 0.99246 / 0.99519 (gate 0.99); synthetic pair `test_mast3r.py` end_to_end PCC 0.9985. Details: [`VERIFICATION_2026-10-03.md`](VERIFICATION_2026-10-03.md).
121
 
122
- The Python row uses the published ETH dispatch config. A whole `model(path, path)` call takes 56.9 ms median, because it includes PNG decode and resize on the host. `predict_pairs` takes 27.6 ms per pair. The Python API gives arrays that are bit-identical to the server `/predict` and `/predict_npz` outputs, including the pose.
123
 
124
- RTX 5090 reference measurements (2026-09-14) are unchanged. They use the port's torch reference in PyTorch 2.11, batch 1, on the kitchen demo pair. The "incl. H2D/D2H" column compares with our model call (26.5 ms). The "forward only" column compares with our device trace (23.7 ms). Full table: [`GPU_COMPARISON.md`](GPU_COMPARISON.md).
125
 
126
- | RTX 5090 precision | GPU incl. H2D/D2H | vs current build (26.5 ms) | GPU forward only vs ours (23.7 ms) |
 
 
127
  |---|---:|---:|---:|
128
- | fp32 strict, eager | 103.4 ms | **Blackhole 3.90× faster** | 102.3 ms: Blackhole 4.32× faster |
129
- | tf32, eager | 70.8 ms | **Blackhole 2.67× faster** | 70.0 ms: Blackhole 2.95× faster |
130
- | bf16 autocast, eager | 63.8 ms | **Blackhole 2.41× faster** | 62.6 ms: Blackhole 2.64× faster |
131
- | fp16 autocast, eager | 56.6 ms | **Blackhole 2.13× faster** | 55.5 ms: Blackhole 2.34× faster |
132
- | bf16 weights, eager | 42.3 ms | **Blackhole 1.60× faster** | 41.3 ms: Blackhole 1.74× faster |
133
- | bf16 weights + SDPA, eager | 35.3 ms | **Blackhole 1.33× faster** | 34.3 ms: Blackhole 1.45× faster |
134
- | bf16 weights + `torch.compile` (CUDA graphs) | 21.1 ms | GPU 1.26× faster | 20.1 ms: GPU 1.18× faster |
135
 
136
- The Blackhole lead comes from one metal trace that runs fused, block-sharded bf16 kernels on all 120 compute cores. The eager GPU rows pay for kernel launches and explicit softmax attention. The autocast rows also cast the fp32 weights again on each forward. `torch.compile` with CUDA graphs removes these costs, and then the GPU is about 1.2× faster. The newer `code/` writes a stored npz. This change decreases the npz encode from about 220 ms to about 17 ms and the served `/predict` total from about 285 ms to about 79 ms.
137
 
138
  ## Caveats
139
 
140
  - Does not scale to multiple p150a in a mesh configuration. The current build uses a 12x10 compute grid of Tensix cores. To get this grid, the dispatch functions move from 10 Tensix cores to ETH cores (`patches/tt-metal-eth-dispatch.patch`). Thus, this build assumes that you do not need chip-to-chip ethernet communication.
141
- - The serve config sets `MAST3R_CQS=2`. Two CQs need worker (Tensix) dispatch. With ETH dispatch, set `MAST3R_CQS=1`. The outputs are bit-identical.
142
  - Fixed 512×512 input, exactly one image pair per request (batch 1). Non-square images are gray-padded to square (`MAST3R_PREPROC=crop` centre-crops instead). The estimated focal can be 5-8 % high, so pass `intrinsics` when you know them.
143
  - Numbers published before 2026-09-14 describe an older network with wrong decoder taps. This card does not repeat them.
144
  - DUSt3R backbone only: no MASt3R matcher / descriptor head and no N-view global alignment. The pose is single-pair PairViewer (estimated-focal PnP-RANSAC). With `return_pose`, one symmetric graph (`sym`) computes the pose maps. Its outputs are bit-identical to the two-pass path.
@@ -156,7 +159,7 @@ The Blackhole lead comes from one metal trace that runs fused, block-sharded bf1
156
 
157
  ## Provenance
158
 
159
- These are the exact sources the container image was built from. **`code/` has since been updated (2026-10-04 build with the LoFi MLP, the faster npz encode and the `mast3r_p150` Python package; see [`OPT_REPORT.md`](OPT_REPORT.md) and [`PYTHON.md`](PYTHON.md)) and is newer than the image.** `tt-model serve` runs the image's code until the image is rebuilt. `tt-model.yaml` and `SERVING.md` still describe the image:
160
 
161
  | component | built from |
162
  | --- | --- |
 
56
 
57
  The first call of `from_pretrained` downloads the NAVER weights to your HF cache. It also compiles the kernels and captures the metal trace (about 65 s). The package does not install ttnn. You do not need the HTTP server.
58
 
59
+ `Mast3rP150.from_pretrained()` prepares the plain call, the pose call (`return_pose=True`) and `predict_pairs` before it returns. Thus the first call is as fast as the later calls (within 10 %). Start-up takes about 8 s with a warm kernel cache. Use `warmup_variants=["pair"]` for a faster start-up, or `model.warmup("crop")` to prepare more variants later.
60
 
61
  | | name | type / default | meaning |
62
  |---|---|---|---|
 
103
 
104
  ## Demo & Performances
105
 
106
+ Warm, batch 1, one 512×512 pair, fused traced graph (`MAST3R_OPT=all`), `bench_breakdown.py` with 30 iterations per run (median / min). All rows use the p150 configuration: dispatch on ETH cores, 1 command queue (CQ), 12×10 compute grid. The served path config is this configuration with uint8 input (the serve env of the newer `code/`).
107
 
108
  | Metric | Performance |
109
  |---|---:|
110
+ | Model call `model(img1, img2)` (H2D + trace + D2H of both maps), served path config | **27.6–27.8 ms median** (2 runs) · 26.9 ms min |
111
+ | Device trace (ViT-L encoder ×2, decoder, two DPT heads), served path config | **23.7–23.9 ms median** · 23.1 ms min |
112
+ | Model call, float input | 28.4 ms median (2 runs) · 27.9 ms min |
113
+ | Device trace, float input | **23.3–23.4 ms median** · 23.0 ms min |
114
+ | Device span (profiler cycles at 1.35 GHz, 3 replays), measured before the LoFi MLP change | 23.47–23.49 ms · 688 programs |
115
+ | Host prep · H2D, served path config (uint8 pixels) | 0.43 ms · 0.54 ms |
116
+ | D2H of both maps, served path config · float input | 2.75 ms · 3.1–3.4 ms |
117
+ | `return_pose` forward on the real pair (`sym_check.py`): symmetric graph (`sym`) · two-pass | **38.0–38.9 ms median** · 53.9–54.9 ms median |
118
+ | Python model() call (`mast3r_p150`, `timing_ms['forward']`, 30 warm calls) | **26.9–27.0 ms median** · 26.5 ms min |
119
+ | HTTP server, newer `code/` (`timing_ms.forward` · `/predict` npz total · `/predict_npz` total, 30 requests) | **26.6–27.0 ms median** · 80–83 ms · 73–75 ms |
120
 
121
+ The measurement hardware is a Blackhole chip with a 12×10 compute grid and dispatch on ETH cores (`patches/tt-metal-eth-dispatch.patch`), with 1 command queue. This is the configuration of a single p150. The 2026-10-04 build runs the encoder and decoder MLPs at LoFi (`MAST3R_OPT` knobs `lofie`, `lofid`). An independent verifier measured the rows above again on 2026-10-05 in this configuration. The device span row is from the previous build. Accuracy against the torch fp32 reference on the real kitchen pair (served path config): pts3d PCC 0.99015 / 0.99097 (gate 0.989), conf PCC 0.99246 / 0.99519 (gate 0.99); synthetic pair `test_mast3r.py` end_to_end PCC 0.9985. Details: [`VERIFICATION_2026-10-03.md`](VERIFICATION_2026-10-03.md).
122
 
123
+ Before 2026-10-05, this card showed a served model call of 26.5 ms. That number used worker (Tensix) dispatch with 2 CQ, which does not give a 12×10 grid on a p150. With 1 CQ, the head-1 readback does not overlap the head-2 compute, so the model call is about 1.2 ms slower. The outputs are bit-identical.
124
 
125
+ A whole `model(path, path)` call takes 50–59 ms median, because it includes PNG decode and resize on the host. `predict_pairs` takes 27.3–28.1 ms per pair. The Python API gives arrays that are bit-identical to the server `/predict` and `/predict_npz` outputs, including the pose.
126
 
127
+ RTX 5090 reference measurements (2026-09-14) are unchanged. They use the port's torch reference in PyTorch 2.11, batch 1, on the kitchen demo pair. The "incl. H2D/D2H" column compares with our model call (27.7 ms). The "forward only" column compares with our device trace (23.8 ms). Full table: [`GPU_COMPARISON.md`](GPU_COMPARISON.md).
128
+
129
+ | RTX 5090 precision | GPU incl. H2D/D2H | vs current build (27.7 ms) | GPU forward only vs ours (23.8 ms) |
130
  |---|---:|---:|---:|
131
+ | fp32 strict, eager | 103.4 ms | **Blackhole 3.73× faster** | 102.3 ms: Blackhole 4.30× faster |
132
+ | tf32, eager | 70.8 ms | **Blackhole 2.55× faster** | 70.0 ms: Blackhole 2.94× faster |
133
+ | bf16 autocast, eager | 63.8 ms | **Blackhole 2.30× faster** | 62.6 ms: Blackhole 2.63× faster |
134
+ | fp16 autocast, eager | 56.6 ms | **Blackhole 2.04× faster** | 55.5 ms: Blackhole 2.33× faster |
135
+ | bf16 weights, eager | 42.3 ms | **Blackhole 1.53× faster** | 41.3 ms: Blackhole 1.73× faster |
136
+ | bf16 weights + SDPA, eager | 35.3 ms | **Blackhole 1.28× faster** | 34.3 ms: Blackhole 1.44× faster |
137
+ | bf16 weights + `torch.compile` (CUDA graphs) | 21.1 ms | GPU 1.31× faster | 20.1 ms: GPU 1.18× faster |
138
 
139
+ The Blackhole lead comes from one metal trace that runs fused, block-sharded bf16 kernels on all 120 compute cores. The eager GPU rows pay for kernel launches and explicit softmax attention. The autocast rows also cast the fp32 weights again on each forward. `torch.compile` with CUDA graphs removes these costs, and then the GPU is about 1.2-1.3× faster. The newer `code/` writes a stored npz. This change decreases the npz encode from about 220 ms to about 17 ms and the served `/predict` total from about 285 ms to about 80 ms.
140
 
141
  ## Caveats
142
 
143
  - Does not scale to multiple p150a in a mesh configuration. The current build uses a 12x10 compute grid of Tensix cores. To get this grid, the dispatch functions move from 10 Tensix cores to ETH cores (`patches/tt-metal-eth-dispatch.patch`). Thus, this build assumes that you do not need chip-to-chip ethernet communication.
144
+ - The newer `code/` uses ETH dispatch and 1 CQ by default (`MAST3R_DISPATCH=auto`, `MAST3R_CQS=1`). `MAST3R_CQS=2` is an opt-in for Galaxy only. It needs worker (Tensix) dispatch, which gives an 11×10 grid on a p150. The container image and its `tt-model.yaml` still set `MAST3R_CQS=2` until the image is rebuilt. Without the ETH-dispatch patch, `auto` uses Tensix dispatch (11×10 on a p150) and logs a warning.
145
  - Fixed 512×512 input, exactly one image pair per request (batch 1). Non-square images are gray-padded to square (`MAST3R_PREPROC=crop` centre-crops instead). The estimated focal can be 5-8 % high, so pass `intrinsics` when you know them.
146
  - Numbers published before 2026-09-14 describe an older network with wrong decoder taps. This card does not repeat them.
147
  - DUSt3R backbone only: no MASt3R matcher / descriptor head and no N-view global alignment. The pose is single-pair PairViewer (estimated-focal PnP-RANSAC). With `return_pose`, one symmetric graph (`sym`) computes the pose maps. Its outputs are bit-identical to the two-pass path.
 
159
 
160
  ## Provenance
161
 
162
+ These are the exact sources the container image was built from. **`code/` has since been updated (2026-10-04 build with the LoFi MLP, the faster npz encode and the `mast3r_p150` Python package; 2026-10-05: ETH dispatch and 1 CQ are the default for the server and the harnesses; see [`OPT_REPORT.md`](OPT_REPORT.md) and [`PYTHON.md`](PYTHON.md)) and is newer than the image.** `tt-model serve` runs the image's code until the image is rebuilt. `tt-model.yaml` and `SERVING.md` still describe the image:
163
 
164
  | component | built from |
165
  | --- | --- |
VERIFICATION_2026-10-03.md CHANGED
@@ -120,3 +120,77 @@ Python API (`09c7d11`, ETH dispatch, 12×10 grid, 1 CQ, 30 warm calls on the dem
120
  - `predict_pairs` gives 30 results. Each result is equal to one `model()` call.
121
  - `timing_ms['forward']`: 26.82 ms median, 26.52 ms min. A whole `model(path, path)` call: 56.9 ms median. `predict_pairs`: 27.57 ms per pair.
122
  - Startup with a warm cache: 7.4 s. Host tests (`test_api_host.py`, `test_server_host.py`): 21 passed.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
120
  - `predict_pairs` gives 30 results. Each result is equal to one `model()` call.
121
  - `timing_ms['forward']`: 26.82 ms median, 26.52 ms min. A whole `model(path, path)` call: 56.9 ms median. `predict_pairs`: 27.57 ms per pair.
122
  - Startup with a warm cache: 7.4 s. Host tests (`test_api_host.py`, `test_server_host.py`): 21 passed.
123
+
124
+ ## p150 ETH-dispatch compliance (2026-10-05)
125
+
126
+ On a single p150, the model must use the 12×10 compute grid. Thus dispatch must run on ETH cores, and the patched tt-metal then has 1 command queue (CQ). Worker (Tensix) dispatch gives an 11×10 grid on a p150. The measurement Galaxy chips give 12×10 also with worker dispatch, so the earlier "served path config" rows (worker dispatch, 2 CQ) do not represent a p150. Commit `c71e46d` makes ETH dispatch, 1 CQ and 12×10 the default for every device-open path. An independent verifier measured `c71e46d` against `d02a5c2` on one chip.
127
+
128
+ ### Audit (default of each device-open path)
129
+
130
+ | path | before (`d02a5c2`) | after (`c71e46d`) |
131
+ |---|---|---|
132
+ | Python API `Mast3rP150.from_pretrained` | ETH, 1 CQ, 12×10 | unchanged; `MAST3R_CQS=2` gives a warning (Galaxy only) |
133
+ | HTTP server `models/server/app.py` | `ttnn.open_device`: worker dispatch; CQs from `MAST3R_CQS` | `mast3r_p150.device.open_device`: ETH, 1 CQ, 12×10; `/info` has `device_config` |
134
+ | `tt-model.yaml` serve env | `MAST3R_CQS=2`: worker dispatch, 2 CQ | `MAST3R_DISPATCH=auto`, `MAST3R_CQS=1`: ETH, 1 CQ |
135
+ | `test_mast3r.py` | worker dispatch (validation ran with `MAST3R_CQS=2`) | ETH, 1 CQ, 12×10 |
136
+ | `eval_mast3r.py`, `make_demo.py --local`, `eval_eth3d.py` | worker dispatch, 1 CQ | ETH, 1 CQ, 12×10 |
137
+ | `bench_breakdown.py` | default `--mode worker11` | default `--mode eth12`; worker modes stay as explicit A/B options |
138
+ | `tools_prof/real_pair_acc.py` | worker unless `MAST3R_ETH=1` | ETH unless `MAST3R_ETH=0` |
139
+ | `tools_prof/sym_check.py` | worker dispatch (run with `MAST3R_CQS=2`) | ETH, 1 CQ |
140
+ | `test_api_device.py`, `test_warmup_device.py`, `first_call_bench.py`, `profile_eager.py` | ETH, 1 CQ, 12×10 | unchanged |
141
+
142
+ `MAST3R_DISPATCH=eth` together with `MAST3R_CQS=2` raises an error. Without the ETH-dispatch patch, `auto` uses worker dispatch and logs a warning that names the 11×10 grid.
143
+
144
+ ### Configuration check
145
+
146
+ - Python API with defaults: `ttnn.open_device` received `dispatch_core_config=ETH` and no `num_command_queues` (default 1). The device reports a 12×10 grid. A command on CQ 1 fails, so the device has 1 CQ.
147
+ - HTTP server with the serve env of `c71e46d`: the log shows `Device open: dispatch eth, grid 12x10, 1 CQ`. `/info` gives the same `device_config`.
148
+ - The `d02a5c2` server with its own serve env opened with worker dispatch and 2 CQ. This confirms the audit.
149
+
150
+ ### Measured (ETH dispatch, 1 CQ, 12×10; median, min in brackets)
151
+
152
+ | measurement | claim | verifier | previous published value |
153
+ |---|---:|---:|---:|
154
+ | device trace, float (`bench_breakdown`, 30 iterations) | 23.42 (23.07) | 23.29 (23.00) | 23.32 (ETH, 2026-10-04) |
155
+ | model call, float | 28.38 (27.86) | 28.36 (28.09) | 28.53 (ETH, 2026-10-04) |
156
+ | device trace, uint8 (served input) | 23.94 (23.24) | 23.74 (23.12) | 23.57-23.84 (worker, 2 CQ) |
157
+ | model call, uint8 (served input) | 27.83 (27.37) | 27.62 (26.92) | 26.47-26.52 (worker, 2 CQ) |
158
+ | `test_mast3r.py --layer end_to_end --runs 25` | 27.77 | 26.86 | 27.51 (2 CQ) |
159
+ | `sym_check`: single pair / `sym` / two-pass | 27.22 / 38.89 / 54.90 | 27.00 / 37.99 / 53.93 | 26.40 / 36.98 / 52.48 (2 CQ) |
160
+ | HTTP `timing_ms.forward` (npz, 30 requests per run) | 26.55-26.98 | 26.57-26.85 | 25.50-25.61 (worker, 2 CQ, same session) |
161
+ | HTTP `/predict` npz total / client wall | 80.3-82.5 / 123-136 | 80.3-83.1 / 125-140 | 81.5-83.3 / 130-140 (2 CQ) |
162
+ | HTTP png total (15 requests) | 131.8-143.7 | 132.6-137.0 | 124.5-139.6 (2 CQ) |
163
+ | HTTP `/predict_npz` total / client wall | 72.7-74.0 / 97.5-104.5 | 72.8-75.1 / 98.4-100.8 | 72-76 / 96.8-102 (2 CQ) |
164
+ | Python API `timing_ms['forward']` (30 calls) | 26.99 (26.54) | 26.85 (26.47) | 26.82 (26.52) (ETH, 1 CQ) |
165
+ | Python API `predict_pairs`, per pair | 28.09 | 27.33 | 27.57 (ETH, 1 CQ) |
166
+ | Python API startup, warm cache | 7.5 s | 8.1 s | 7.4 s (ETH, 1 CQ) |
167
+
168
+ - The device trace does not change within noise. With 1 CQ, the head-1 readback does not overlap the head-2 compute. Thus the served forward is about 1.1 ms (4 %) slower than the worker 2-CQ mode, and the uint8 model call is about 1.2 ms slower. Request totals change by 1-3 ms, which is in the host noise.
169
+ - First call after `from_pretrained` (`test_warmup_device.py`, 3 runs): `pair` 0.98-1.04x, `pose` 0.95-1.06x, `batch` 0.89-1.13x of the warm calls. All runs pass the gate (10 % + 5 ms). The `batch` call takes about 74 ms and is host-bound, so its spread is host noise.
170
+ - The server smoke test (`--require-pose --npz-route`) passes.
171
+
172
+ ### Accuracy (ETH dispatch, 1 CQ, 12×10; gate code did not change; all gates pass)
173
+
174
+ | metric | gate | `c71e46d` |
175
+ |---|---:|---:|
176
+ | synthetic randn pair, float, PCC head 1 / head 2 / all | ≥ 0.998 | 0.99862 / 0.99869 / 0.99891 |
177
+ | `test_mast3r.py --layer end_to_end` PCC | ≥ 0.998 | 0.9985 (PASS; head 1 0.9985, head 2 0.9982) |
178
+ | real pair, float: pts3d / conf PCC | ≥ 0.989 / ≥ 0.99 | 0.99016 / 0.99088 · 0.99269 / 0.99486 |
179
+ | real pair, uint8: pts3d / conf PCC | ≥ 0.989 / ≥ 0.99 | 0.99015 / 0.99097 · 0.99246 / 0.99519 |
180
+ | synthetic uint8 pair (not gated) | | 0.99593 / 0.99927 |
181
+
182
+ Every value is equal to the 2026-10-04 value to the printed digit.
183
+
184
+ Bit identity against the previous served default (worker dispatch, 2 CQ), same chip and session:
185
+
186
+ - `bench_breakdown.py`, float and uint8 pairs: head 1 and head 2 are bit-identical (n_diff = 0).
187
+ - HTTP: the npz arrays from `/predict` and `/predict_npz` are identical (n_diff = 0 for `pts3d1`, `conf1`, `pts3d2`, `conf2`). The PNG bytes are identical. The pose rotation is the same (43.3 deg).
188
+ - `sym_check`: `out_ii`, `out_ji` and `out_jj` of the `sym` graph are bit-identical to the two-pass path.
189
+ - `test_api_device.py`: `model()` is bit-identical to the server pipeline, with and without pose.
190
+ - Host tests (`test_fused_host.py`, `test_api_host.py`, `test_server_host.py`): 59 passed.
191
+
192
+ ### Review notes
193
+
194
+ - The container image was built before this change. Its `tt-model.yaml` still sets `MAST3R_CQS=2` (worker dispatch) until the image is rebuilt.
195
+ - `auto` selects ETH dispatch only when the tt-metal source tree has the patch marker. If the tree does not have the patch, the server uses worker dispatch (11×10 on a p150) and logs a warning. It does not stop.
196
+ - `tools_prof/real_pair_acc.py` opens the device through its own helper. To use `MAST3R_CQS=2` with it, also set `MAST3R_ETH=0`. This affects only the tooling.
code/bench_breakdown.py CHANGED
@@ -9,9 +9,9 @@ Phases per call (same work as TtDust3r.__call__ / test_mast3r.py end_to_end):
9
  e2e : the public model call (host_prep + h2d + trace + d2h, non-blocking trace)
10
 
11
  Usage (after chipenv.sh):
12
- python bench_breakdown.py --mode worker11 # p150-equivalent: Tensix dispatch, grid capped 11x10
13
- python bench_breakdown.py --mode worker # Galaxy default Tensix dispatch (12x10, Galaxy-only A/B)
14
- python bench_breakdown.py --mode eth12 # ETH dispatch, capped 12x10
15
  MAST3R_CQS=2 python bench_breakdown.py --mode worker # Tensix dispatch 12x10, 2 CQs (split trace, overlapped head-1 read)
16
  add --u8 for uint8 pixel views (HWC-backed, as the server's PIL preprocess produces them)
17
 
@@ -31,7 +31,7 @@ sys.path.insert(0, _CODE_ROOT)
31
  sys.path.insert(0, "/home/ttuser/experiments/tt-models/tools")
32
 
33
  ap = argparse.ArgumentParser()
34
- ap.add_argument("--mode", default="worker11", choices=["worker11", "worker", "eth12"])
35
  ap.add_argument("--iters", type=int, default=30)
36
  ap.add_argument("--no-ref", action="store_true")
37
  ap.add_argument("--u8", action="store_true", help="uint8 pixel views (exact reference input (v/255-0.5)/0.5)")
 
9
  e2e : the public model call (host_prep + h2d + trace + d2h, non-blocking trace)
10
 
11
  Usage (after chipenv.sh):
12
+ python bench_breakdown.py # default --mode eth12: ETH dispatch, 1 CQ, 12x10 (= p150 + ETH dispatch)
13
+ python bench_breakdown.py --mode worker11 # stock p150 Tensix dispatch emulated: grid capped 11x10
14
+ python bench_breakdown.py --mode worker # Galaxy Tensix dispatch (12x10 here, Galaxy-only A/B)
15
  MAST3R_CQS=2 python bench_breakdown.py --mode worker # Tensix dispatch 12x10, 2 CQs (split trace, overlapped head-1 read)
16
  add --u8 for uint8 pixel views (HWC-backed, as the server's PIL preprocess produces them)
17
 
 
31
  sys.path.insert(0, "/home/ttuser/experiments/tt-models/tools")
32
 
33
  ap = argparse.ArgumentParser()
34
+ ap.add_argument("--mode", default="eth12", choices=["worker11", "worker", "eth12"])
35
  ap.add_argument("--iters", type=int, default=30)
36
  ap.add_argument("--no-ref", action="store_true")
37
  ap.add_argument("--u8", action="store_true", help="uint8 pixel views (exact reference input (v/255-0.5)/0.5)")
code/eval_eth3d.py CHANGED
@@ -78,7 +78,10 @@ def main():
78
  pairs = pick_pairs(records, args.pairs, max_rot_deg=args.max_rot_deg)
79
  print(f"# picked {len(pairs)} pairs with GT rot < {args.max_rot_deg}°")
80
 
81
- device = ttnn.open_device(device_id=args.device_id, l1_small_size=32 * 1024)
 
 
 
82
  if hasattr(device, "enable_program_cache"):
83
  device.enable_program_cache()
84
  try:
 
78
  pairs = pick_pairs(records, args.pairs, max_rot_deg=args.max_rot_deg)
79
  print(f"# picked {len(pairs)} pairs with GT rot < {args.max_rot_deg}°")
80
 
81
+ # Same opener as the Python API / server: ETH dispatch + 1 CQ + 12x10 by default (MAST3R_DISPATCH).
82
+ from mast3r_p150.device import open_device as _open_port_device
83
+ device, _dinfo = _open_port_device(args.device_id, None, l1_small_size=32 * 1024)
84
+ print(f"# device: dispatch {_dinfo['dispatch']}, grid {_dinfo['grid']}, {_dinfo['num_cqs']} CQ")
85
  if hasattr(device, "enable_program_cache"):
86
  device.enable_program_cache()
87
  try:
code/eval_mast3r.py CHANGED
@@ -297,8 +297,10 @@ def main():
297
  # of this port runs in); legacy / MAST3R_TRACE=0 keep the plain open.
298
  fused = fused_config()
299
  print(f"# port path: {'fused ' + str(fused.summary()) if fused.enabled else 'legacy (TT_FUSED=0)'}")
300
- device = ttnn.open_device(device_id=args.device_id,
301
- **fused.open_device_kwargs(l1_small_size=32 * 1024))
 
 
302
  if hasattr(device, "enable_program_cache"):
303
  device.enable_program_cache()
304
  try:
 
297
  # of this port runs in); legacy / MAST3R_TRACE=0 keep the plain open.
298
  fused = fused_config()
299
  print(f"# port path: {'fused ' + str(fused.summary()) if fused.enabled else 'legacy (TT_FUSED=0)'}")
300
+ # Same opener as the Python API / server: ETH dispatch + 1 CQ + 12x10 by default (MAST3R_DISPATCH).
301
+ from mast3r_p150.device import open_device as _open_port_device
302
+ device, _dinfo = _open_port_device(args.device_id, None, fused=fused, l1_small_size=32 * 1024)
303
+ print(f"# device: dispatch {_dinfo['dispatch']}, grid {_dinfo['grid']}, {_dinfo['num_cqs']} CQ")
304
  if hasattr(device, "enable_program_cache"):
305
  device.enable_program_cache()
306
  try:
code/make_demo.py CHANGED
@@ -118,8 +118,9 @@ def predict_local(paths, return_pose: bool = False):
118
  mode = preprocess_mode()
119
  imgs, recs = zip(*[preprocess_image(Image.open(p), IMG_SIZE, mode) for p in paths])
120
  fused = port.fused_config()
121
- device = ttnn.open_device(**fused.open_device_kwargs(device_id=int(os.environ.get("TT_DEVICE_ID", "0")),
122
- l1_small_size=32 * 1024))
 
123
  try:
124
  state = load_checkpoint()
125
  with torch.no_grad():
 
118
  mode = preprocess_mode()
119
  imgs, recs = zip(*[preprocess_image(Image.open(p), IMG_SIZE, mode) for p in paths])
120
  fused = port.fused_config()
121
+ from mast3r_p150.device import open_device as _open_port_device
122
+ device, _ = _open_port_device(int(os.environ.get("TT_DEVICE_ID", "0")), None, fused=fused,
123
+ l1_small_size=32 * 1024)
124
  try:
125
  state = load_checkpoint()
126
  with torch.no_grad():
code/mast3r_p150/device.py CHANGED
@@ -10,6 +10,11 @@
10
  * otherwise -> tt-metal's default (Tensix) dispatch, with a warning.
11
 
12
  ``dispatch="eth"`` / ``"worker"`` force one mode (``$MAST3R_DISPATCH`` sets the default).
 
 
 
 
 
13
  """
14
  from __future__ import annotations
15
 
@@ -66,14 +71,18 @@ def resolve_dispatch(dispatch: Optional[str], num_cqs: int) -> Tuple[str, Option
66
  if d == "worker":
67
  return "worker", None
68
  if num_cqs != 1:
69
- return "worker", None # MAST3R_CQS=2 is the served (Tensix-dispatch) configuration
 
 
 
70
  patched = eth_dispatch_patch_present()
71
  if patched:
72
  return "eth", None
73
  why = ("the tt-metal tree has no ETH-dispatch patch" if patched is False
74
  else "the tt-metal source tree was not found, so ETH-dispatch support is unknown")
75
- return "worker", (f"mast3r_p150: {why}; using tt-metal's default (Tensix) dispatch. The published "
76
- "numbers use ETH dispatch with a 12x10 grid; expect a different grid and timing.")
 
77
 
78
 
79
  def open_device(device_id: int = 0, dispatch: Optional[str] = None, *, fused=None, **extra):
@@ -82,7 +91,7 @@ def open_device(device_id: int = 0, dispatch: Optional[str] = None, *, fused=Non
82
  (adds ``trace_region_size`` / ``num_command_queues``)."""
83
  import ttnn
84
 
85
- base = dict(device_id=device_id, l1_small_size=L1_SMALL_SIZE, **extra)
86
  kw = fused.open_device_kwargs(**base) if fused is not None else base
87
  mode, warn = resolve_dispatch(dispatch, int(kw.get("num_command_queues", 1)))
88
  if warn:
@@ -112,4 +121,5 @@ def open_device(device_id: int = 0, dispatch: Optional[str] = None, *, fused=Non
112
  "before the first device open in this process)")
113
  warnings.warn(msg, RuntimeWarning, stacklevel=3)
114
  log.warning(msg)
115
- return device, {"dispatch": mode, "grid": f"{grid[0]}x{grid[1]}", "open_kwargs": {k: v for k, v in kw.items()}}
 
 
10
  * otherwise -> tt-metal's default (Tensix) dispatch, with a warning.
11
 
12
  ``dispatch="eth"`` / ``"worker"`` force one mode (``$MAST3R_DISPATCH`` sets the default).
13
+ ``MAST3R_CQS=2`` (2 command queues) is an explicit opt-in that needs Tensix dispatch: on the
14
+ Galaxy that still gives 12x10, on a single p150 only 11x10; a warning says so.
15
+
16
+ The Python API (``Mast3rP150.from_pretrained``), the HTTP server (``models.server.app``) and the
17
+ harnesses (``test_mast3r.py``, ``eval_mast3r.py``, ``make_demo.py``) all open the chip here.
18
  """
19
  from __future__ import annotations
20
 
 
71
  if d == "worker":
72
  return "worker", None
73
  if num_cqs != 1:
74
+ # MAST3R_CQS=2 is an explicit opt-in: ETH dispatch has 1 CQ, so it needs Tensix dispatch.
75
+ return "worker", ("mast3r_p150: MAST3R_CQS=2 needs Tensix (worker) dispatch. This is a Galaxy-only "
76
+ "opt-in: on a single p150 Tensix dispatch leaves an 11x10 grid, not the tuned 12x10. "
77
+ "Unset MAST3R_CQS (or set it to 1) for the p150 configuration (ETH dispatch, 1 CQ, 12x10).")
78
  patched = eth_dispatch_patch_present()
79
  if patched:
80
  return "eth", None
81
  why = ("the tt-metal tree has no ETH-dispatch patch" if patched is False
82
  else "the tt-metal source tree was not found, so ETH-dispatch support is unknown")
83
+ return "worker", (f"mast3r_p150: {why}; using tt-metal's default (Tensix) dispatch. One Tensix column then "
84
+ "runs dispatch, so a single p150 has an 11x10 compute grid, not the tuned 12x10. The published "
85
+ "numbers use ETH dispatch, 1 command queue and a 12x10 grid; expect lower throughput.")
86
 
87
 
88
  def open_device(device_id: int = 0, dispatch: Optional[str] = None, *, fused=None, **extra):
 
91
  (adds ``trace_region_size`` / ``num_command_queues``)."""
92
  import ttnn
93
 
94
+ base = {"device_id": device_id, "l1_small_size": L1_SMALL_SIZE, **extra}
95
  kw = fused.open_device_kwargs(**base) if fused is not None else base
96
  mode, warn = resolve_dispatch(dispatch, int(kw.get("num_command_queues", 1)))
97
  if warn:
 
121
  "before the first device open in this process)")
122
  warnings.warn(msg, RuntimeWarning, stacklevel=3)
123
  log.warning(msg)
124
+ return device, {"dispatch": mode, "grid": f"{grid[0]}x{grid[1]}",
125
+ "num_cqs": int(kw.get("num_command_queues", 1)), "open_kwargs": {k: v for k, v in kw.items()}}
code/models/server/app.py CHANGED
@@ -289,7 +289,13 @@ async def lifespan(_app: FastAPI):
289
  log.info("Port path: %s", "fused %s" % fused.summary() if fused.enabled else "legacy (TT_FUSED=0)")
290
  log.info("Preprocessing: MAST3R_PREPROC=%s (%s to square, bicubic %dx%d)", cfg.preproc,
291
  "gray-pad" if cfg.preproc == "pad" else "centre-crop", IMG_SIZE, IMG_SIZE)
292
- device = ttnn.open_device(**open_kwargs)
 
 
 
 
 
 
293
  if hasattr(device, "enable_program_cache"): # already on by default; harmless
294
  device.enable_program_cache()
295
  STATE["device"] = device
@@ -534,6 +540,7 @@ def info() -> dict:
534
  "estimated- or known-focal PnP-RANSAC",
535
  },
536
  "limits": {"max_images": 2, "batch": 1, "fixed_input": f"{IMG_SIZE}x{IMG_SIZE}"},
 
537
  "fused": STATE.get("fused"), # TT_FUSED configuration the port was built with (None before startup)
538
  "reference_numbers": {"latency_ms_per_pair_b1": 73, "pcc_vs_torch_reference": 0.9970,
539
  "served_forward_ms_median": 73.6, "legacy_latency_ms_per_pair_b1": 234,
 
289
  log.info("Port path: %s", "fused %s" % fused.summary() if fused.enabled else "legacy (TT_FUSED=0)")
290
  log.info("Preprocessing: MAST3R_PREPROC=%s (%s to square, bicubic %dx%d)", cfg.preproc,
291
  "gray-pad" if cfg.preproc == "pad" else "centre-crop", IMG_SIZE, IMG_SIZE)
292
+ # Dispatch (2026-10-05): the same opener as the Python API. MAST3R_DISPATCH=auto (default): ETH
293
+ # dispatch + 1 CQ + 12x10 grid when tt-metal has the ETH-dispatch patch (the p150 configuration),
294
+ # else Tensix dispatch with a warning. MAST3R_CQS=2 is a Galaxy-only opt-in (Tensix dispatch).
295
+ from mast3r_p150.device import open_device as _open_port_device
296
+ device, dinfo = _open_port_device(cfg.device_id, None, fused=fused, l1_small_size=L1_SMALL_SIZE)
297
+ STATE["dispatch"] = {"dispatch": dinfo["dispatch"], "grid": dinfo["grid"], "num_cqs": dinfo["num_cqs"]}
298
+ log.info("Device open: dispatch %s, grid %s, %d CQ", dinfo["dispatch"], dinfo["grid"], dinfo["num_cqs"])
299
  if hasattr(device, "enable_program_cache"): # already on by default; harmless
300
  device.enable_program_cache()
301
  STATE["device"] = device
 
540
  "estimated- or known-focal PnP-RANSAC",
541
  },
542
  "limits": {"max_images": 2, "batch": 1, "fixed_input": f"{IMG_SIZE}x{IMG_SIZE}"},
543
+ "device_config": STATE.get("dispatch"), # dispatch (eth | worker), compute grid, command queues
544
  "fused": STATE.get("fused"), # TT_FUSED configuration the port was built with (None before startup)
545
  "reference_numbers": {"latency_ms_per_pair_b1": 73, "pcc_vs_torch_reference": 0.9970,
546
  "served_forward_ms_median": 73.6, "legacy_latency_ms_per_pair_b1": 234,
code/models/tests/test_fused_host.py CHANGED
@@ -532,7 +532,8 @@ def test_open_device_kwargs_select_execution_mode():
532
  def test_harness_and_eval_open_the_device_in_the_reported_mode():
533
  """``test_mast3r.py`` exposes the on-device DPT head as its own layer (the ``dpt_head``
534
  layer is the torch bf16 host reference and runs no ttnn op) and, like ``eval_mast3r.py``
535
- and the server, opens the device via ``open_device_kwargs`` and releases the port's
 
536
  device caches before ``close_device``. Source-level checks: the scripts import ttnn only
537
  inside ``main`` and are never executed here."""
538
  import importlib.util
@@ -545,7 +546,8 @@ def test_harness_and_eval_open_the_device_in_the_reported_mode():
545
 
546
  for script in ("test_mast3r.py", "eval_mast3r.py", os.path.join("models", "server", "app.py")):
547
  src = open(os.path.join(_CODE_ROOT, script)).read()
548
- assert "open_device_kwargs(" in src, script
 
549
  assert "release_device_caches()" in src and "close_device(" in src, script
550
  assert src.index("release_device_caches()") < src.rindex("close_device("), script
551
  harness_src = open(os.path.join(_CODE_ROOT, "test_mast3r.py")).read()
 
532
  def test_harness_and_eval_open_the_device_in_the_reported_mode():
533
  """``test_mast3r.py`` exposes the on-device DPT head as its own layer (the ``dpt_head``
534
  layer is the torch bf16 host reference and runs no ttnn op) and, like ``eval_mast3r.py``
535
+ and the server, opens the device via ``mast3r_p150.device.open_device`` (which applies
536
+ ``open_device_kwargs`` and the ETH-dispatch default) and releases the port's
537
  device caches before ``close_device``. Source-level checks: the scripts import ttnn only
538
  inside ``main`` and are never executed here."""
539
  import importlib.util
 
546
 
547
  for script in ("test_mast3r.py", "eval_mast3r.py", os.path.join("models", "server", "app.py")):
548
  src = open(os.path.join(_CODE_ROOT, script)).read()
549
+ assert "_open_port_device(" in src and "fused=fused" in src, script
550
+ assert "ttnn.open_device(" not in src, script
551
  assert "release_device_caches()" in src and "close_device(" in src, script
552
  assert src.index("release_device_caches()") < src.rindex("close_device("), script
553
  harness_src = open(os.path.join(_CODE_ROOT, "test_mast3r.py")).read()
code/test_mast3r.py CHANGED
@@ -356,8 +356,10 @@ def main():
356
  fused = fused_config()
357
  print(f"# port path: {'fused ' + str(fused.summary()) if fused.enabled else 'legacy (TT_FUSED=0)'}")
358
  # l1_small_size needed by ttnn.conv2d (sliding window state buffer).
359
- device = ttnn.open_device(device_id=args.device_id,
360
- **fused.open_device_kwargs(l1_small_size=32 * 1024))
 
 
361
  if hasattr(device, "enable_program_cache"):
362
  device.enable_program_cache()
363
  try:
 
356
  fused = fused_config()
357
  print(f"# port path: {'fused ' + str(fused.summary()) if fused.enabled else 'legacy (TT_FUSED=0)'}")
358
  # l1_small_size needed by ttnn.conv2d (sliding window state buffer).
359
+ # Same opener as the Python API / server: ETH dispatch + 1 CQ + 12x10 by default (MAST3R_DISPATCH).
360
+ from mast3r_p150.device import open_device as _open_port_device
361
+ device, _dinfo = _open_port_device(args.device_id, None, fused=fused, l1_small_size=32 * 1024)
362
+ print(f"# device: dispatch {_dinfo['dispatch']}, grid {_dinfo['grid']}, {_dinfo['num_cqs']} CQ")
363
  if hasattr(device, "enable_program_cache"):
364
  device.enable_program_cache()
365
  try:
code/tools_prof/eth_validate.sh ADDED
@@ -0,0 +1,25 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ #!/bin/bash
2
+ # p150 ETH-dispatch compliance validation (2026-10-05): every run in the p150-equivalent configuration
3
+ # (ETH dispatch, 1 CQ, 12x10) unless the name says w2cq (= the previous served default, Tensix dispatch, 2 CQs,
4
+ # kept here only as the bit-identity reference). Sequential, one device workload at a time.
5
+ # usage: eth_validate.sh CHIP OUTDIR
6
+ ROOT=/home/ttuser/experiments/tt-models; M=$ROOT/models/mast3r-p150
7
+ CHIP=${1:?chip}; L=${2:?outdir}; mkdir -p $L
8
+ cd $M/code || exit 1
9
+ source $ROOT/tools/chipenv.sh $CHIP $M/.venv >/dev/null || { echo "chipenv failed for chip $CHIP"; exit 1; }
10
+ export TT_WEIGHTS_REVISION=61c57447d7b0adc8a1a30b2b0adec7a8935aa2a3
11
+ unset MAST3R_CQS MAST3R_DISPATCH
12
+ echo "chip $CHIP commit $(git -C $M rev-parse --short HEAD) logs $L"
13
+ run() { local name=$1; shift; echo "== $name $(date -u +%H:%M:%S)"; "$@" > $L/$name.log 2>&1; echo "rc=$?"; }
14
+ STEPS=${STEPS:-eth12 eth12_u8 w2cq w2cq_u8 real_eth12 real_eth12_u8 test_e2e sym_check}
15
+ for s in $STEPS; do case $s in
16
+ eth12) run eth12 timeout -s INT 900 python bench_breakdown.py --dump $L/eth12_f.pt ;;
17
+ eth12_u8) run eth12_u8 timeout -s INT 900 python bench_breakdown.py --u8 --dump $L/eth12_u8.pt ;;
18
+ w2cq) run w2cq env MAST3R_CQS=2 timeout -s INT 900 python bench_breakdown.py --mode worker --cmp $L/eth12_f.pt ;;
19
+ w2cq_u8) run w2cq_u8 env MAST3R_CQS=2 timeout -s INT 900 python bench_breakdown.py --mode worker --u8 --cmp $L/eth12_u8.pt ;;
20
+ real_eth12) run real_eth12 timeout -s INT 900 python tools_prof/real_pair_acc.py ;;
21
+ real_eth12_u8) run real_eth12_u8 env MAST3R_U8=1 timeout -s INT 900 python tools_prof/real_pair_acc.py ;;
22
+ test_e2e) run test_e2e timeout -s INT 900 python test_mast3r.py --layer end_to_end --runs 25 ;;
23
+ sym_check) run sym_check timeout -s INT 900 python tools_prof/sym_check.py ;;
24
+ esac; done
25
+ echo "== done $(date -u +%H:%M:%S)"
code/tools_prof/real_pair_acc.py CHANGED
@@ -1,7 +1,7 @@
1
  """Real-image accuracy gate: media/source_1.png + source_2.png through the traced device
2
  graph vs the fp32 torch reference. Reports raw PCC per head, activated-pts3d PCC per head,
3
- conf PCC, and depth rel-err. Opens the device the same way as test_mast3r.py (Tensix
4
- dispatch) unless MAST3R_ETH=1 (ETH dispatch, 12x10)."""
5
  import os, sys, time
6
  _CODE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
7
  sys.path.insert(0, _CODE); sys.path.insert(0, "/home/ttuser/experiments/tt-models/tools")
@@ -33,7 +33,7 @@ with torch.no_grad():
33
  ref_hp = load_dust3r(state)(img1, img2)
34
  _td.F.interpolate = _orig
35
  cfg = fused_config(); kw = cfg.open_device_kwargs(l1_small_size=32 * 1024)
36
- if os.environ.get("MAST3R_ETH") == "1":
37
  from open_device_12x10 import open_device; dev = open_device(grid="12x10", **kw)
38
  else:
39
  dev = ttnn.open_device(device_id=0, **kw)
 
1
  """Real-image accuracy gate: media/source_1.png + source_2.png through the traced device
2
  graph vs the fp32 torch reference. Reports raw PCC per head, activated-pts3d PCC per head,
3
+ conf PCC, and depth rel-err. Opens the device with ETH dispatch, 12x10 (the p150 configuration,
4
+ default since 2026-10-05) unless MAST3R_ETH=0 (Tensix dispatch, e.g. with MAST3R_CQS=2)."""
5
  import os, sys, time
6
  _CODE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
7
  sys.path.insert(0, _CODE); sys.path.insert(0, "/home/ttuser/experiments/tt-models/tools")
 
33
  ref_hp = load_dust3r(state)(img1, img2)
34
  _td.F.interpolate = _orig
35
  cfg = fused_config(); kw = cfg.open_device_kwargs(l1_small_size=32 * 1024)
36
+ if os.environ.get("MAST3R_ETH", "1") == "1":
37
  from open_device_12x10 import open_device; dev = open_device(grid="12x10", **kw)
38
  else:
39
  dev = ttnn.open_device(device_id=0, **kw)
code/tools_prof/sym_check.py CHANGED
@@ -1,6 +1,6 @@
1
  """return_pose B-mode check: forward_symmetric(img1, img2) vs the two-pass path (model(img1, img2), model(img2, img1)):
2
  bit-equality of out_ii / out_ji / out_jj and e2e timing (median/min of N). uint8 views (server path) unless FLOAT=1.
3
- Device opened like test_mast3r / the server (Tensix dispatch; MAST3R_CQS from env)."""
4
  import os, sys, time, statistics, torch, ttnn
5
  _CODE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
6
  sys.path.insert(0, _CODE); sys.path.insert(0, "/home/ttuser/experiments/tt-models/tools")
@@ -8,7 +8,8 @@ from models.demos.mast3r.reference.torch_dust3r import load_checkpoint
8
  from models.demos.mast3r.postprocess import load_image_for_dust3r
9
  from models.demos.mast3r.tt.ttnn_dust3r import fused_config, get_model, release_device_caches
10
  cfg = fused_config()
11
- dev = ttnn.open_device(device_id=0, **cfg.open_device_kwargs(l1_small_size=32 * 1024))
 
12
  N = int(os.environ.get("N", "15"))
13
  try:
14
  media = os.path.join(os.path.dirname(_CODE), "media")
 
1
  """return_pose B-mode check: forward_symmetric(img1, img2) vs the two-pass path (model(img1, img2), model(img2, img1)):
2
  bit-equality of out_ii / out_ji / out_jj and e2e timing (median/min of N). uint8 views (server path) unless FLOAT=1.
3
+ Device opened like test_mast3r / the server (mast3r_p150.device.open_device: ETH dispatch + 1 CQ by default; MAST3R_CQS=2 -> Tensix)."""
4
  import os, sys, time, statistics, torch, ttnn
5
  _CODE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
6
  sys.path.insert(0, _CODE); sys.path.insert(0, "/home/ttuser/experiments/tt-models/tools")
 
8
  from models.demos.mast3r.postprocess import load_image_for_dust3r
9
  from models.demos.mast3r.tt.ttnn_dust3r import fused_config, get_model, release_device_caches
10
  cfg = fused_config()
11
+ from mast3r_p150.device import open_device as _open_port_device # ETH + 1 CQ + 12x10 by default
12
+ dev, _dinfo = _open_port_device(0, None, fused=cfg, l1_small_size=32 * 1024); print("# device", _dinfo)
13
  N = int(os.environ.get("N", "15"))
14
  try:
15
  media = os.path.join(os.path.dirname(_CODE), "media")