mast3r-p150 / OPT_REPORT.md
changh95's picture
p150 ETH-dispatch compliance (2026-10-05): default ETH dispatch, 1 CQ, 12x10 in Python API and server; numbers re-measured
2c79427 verified
|
Raw History Blame Contribute Delete
116 kB

mast3r-p150 — 12x10 optimisation report (Galaxy BH chips 9, 3, 18)

Status (2026-10-05): p150 ETH-dispatch compliance done (section "p150 ETH-dispatch compliance 2026-10-05"): the server, the harnesses and the Python API default to ETH dispatch + 1 CQ + 12x10; outputs bit-identical to the old Tensix 2-CQ served default; served forward 25.6 -> 26.8 ms (lost CQ1 readback overlap), all gates pass.

Previous status (2026-10-04): audit integration done (section "Audit integration 2026-10-04"): served npz request -72 % server time / -60 % client wall (stored npz), png -34 %, new binary route /predict_npz (-70 % client wall), and LoFi-RNE MLP linears (served e2e -7.2 %, all gates pass, two ungated metrics slightly worse, disclosed).**

Previous status: stopped: round 9 finished (this workflow's optimisation round 2); the backlog below is still open.** Round 9 kept three small bit-identical plumbing steps (esplit, demb, hrqk): traced device span 23.531 -> 23.487 ms (-0.044 ms, -0.19 %, chip 9, same session). That is below the host-wall noise; served u8 e2e is unchanged within noise (27.7-28.4 ms). The two planned big items were measured first and turned out to be capped: the 1.3-2.2 ms "host / dispatch gap" is the chip's AICLK throttling under load (not host or dispatch), and a custom SDPA compute kernel alone can save at most ~10 us per encoder call. All gates pass with identical numbers (every kept step is bit-identical on the eth12 float and the served u8 paths); the server smoke test passed 4 times out of 4. No chip faulted.

Model: DUSt3R ViT-L/16 + dual-branch decoder + 2 DPT heads (naver/DUSt3R_ViTLarge_BaseDecoder_512_dpt @ 61c57447), one 512x512 image pair per request, batch 1. tt-metal 8b98410e730 + ETH-dispatch patch (unmodified; every new kernel of round 3 is model-local and runs through ttnn.generic_op). Grid: 12x10 (TT_METAL_CORE_GRID_OVERRIDE_TODEPRECATE=11,9), ETH dispatch, 1 CQ (= p150 + ETH dispatch). Baseline and its profile: OPT_BASELINE.md. All knobs are MAST3R_OPT entries in code/models/demos/mast3r/tt/fused.py (all = default and pinned in tt-model.yaml; none = the pre-optimisation fused graph). Default set now: out, mm, dpt, l1, dptf, gelut, lnfold, phase, hostcol, phase0, decb, resl1, ropes, sym, tsplit, sdpa2, hrope, hcat, mln, mc2d, mc2dr, resl1d, addln, cdbl, pil, til, mm32, dmm, dadd, addln2, rchain, dlin, dfront, dshard, ups1, tailf, pemm, pcat, p1hs, hf3, gpoly, qkx, sdbl, pilhs, tups, sgat, tapm, obs, obsd, resmm, resmmd, esplit, demb, hrqk, lofie, lofid (gelua and lofi are off; lofie / lofid since the 2026-10-04 audit integration, below). Since 2026-10-05 the serve config pins MAST3R_DISPATCH=auto + MAST3R_CQS=1 (ETH dispatch, 1 CQ, 12x10 = the p150 configuration; section "p150 ETH-dispatch compliance 2026-10-05"). Before that it pinned MAST3R_CQS=2 (split trace, head-1 readback on CQ1), which needs Tensix dispatch: 12x10 on the Galaxy, 11x10 on a p150. MAST3R_CQS=2 stays as a Galaxy-only opt-in. With hostcol the server uploads the views as uint8.

p150 ETH-dispatch compliance 2026-10-05 (chip 16, start = d02a5c2, final = this commit)

Requirement: on a single p150 the model must use the 12x10 grid, so dispatch must run on ETH cores (stock Tensix dispatch leaves 11x10). ETH dispatch in the patched tt-metal has 1 CQ. On this Galaxy, Tensix dispatch still gives 12x10 (13 Tensix columns), so numbers measured with it do not represent a p150. p150-equivalent here = ETH dispatch, num_command_queues=1, grid capped at 12x10. (The first attempt was cut short by host crash #5; its WIP commit 2cf38a7 was reviewed, its host test fixed and every number below measured again after the reboot.)

Audit (default of every device-open path)

path before (d02a5c2) after
Python API Mast3rP150.from_pretrained (mast3r_p150/device.py) ETH, 1 CQ, 12x10 (dispatch="auto"; MAST3R_CQS unset = 1) unchanged; MAST3R_CQS=2 now warns (Galaxy-only, 11x10 on p150)
HTTP server models/server/app.py ttnn.open_device(**open_kwargs) = Tensix dispatch; CQs from MAST3R_CQS (code default 1) mast3r_p150.device.open_device: ETH, 1 CQ, 12x10; /info -> device_config
tt-model.yaml serve.env (+ SERVING.md) MAST3R_CQS: "2" -> Tensix dispatch, 2 CQ, 12x10 on the Galaxy / 11x10 on a p150 MAST3R_DISPATCH: "auto", MAST3R_CQS: "1" -> ETH, 1 CQ, 12x10
test_mast3r.py (card PCC + latency row) Tensix; validation ran it with MAST3R_CQS=2 shared opener: ETH, 1 CQ, 12x10
eval_mast3r.py, make_demo.py --local, eval_eth3d.py Tensix, 1 CQ (eval_eth3d: plain open) shared opener: ETH, 1 CQ, 12x10
bench_breakdown.py (OPT_REPORT trace / e2e rows) default --mode worker11 (Tensix, 11x10); eth12 rows used --mode eth12; served rows MAST3R_CQS=2 --mode worker default --mode eth12 (ETH, 1 CQ, 12x10); worker modes stay as explicit A/B
tools_prof/real_pair_acc.py (real-pair gate) Tensix unless MAST3R_ETH=1 ETH unless MAST3R_ETH=0
tools_prof/sym_check.py (pose path) Tensix; validation ran it with MAST3R_CQS=2 shared opener: ETH, 1 CQ
test_api_device.py, test_warmup_device.py, first_call_bench.py via from_pretrained: ETH, 1 CQ, 12x10 unchanged
profile_eager.py default --mode eth12 unchanged

1-CQ path: the model already had one (MAST3R_CQS=1, one trace on CQ0, both heads read on CQ0 after it); only the served default used CQ1. No new device code. Fallback: without the patch (marker in tt_metal/impl/dispatch/topology.cpp) auto uses Tensix dispatch with a RuntimeWarning + log line naming the 11x10 grid; an ETH open that raises also falls back with a warning (unless MAST3R_DISPATCH=eth). MAST3R_DISPATCH=eth with MAST3R_CQS=2 is an error.

Bit identity (ETH 1 CQ vs the previous served default, Tensix 2 CQ; same chip, same session)

  • bench_breakdown.py --dump (ETH) / --cmp (MAST3R_CQS=2 --mode worker): float randn pair and uint8 pair, both heads bit-identical (n_diff = 0).
  • HTTP: two ETH servers vs one Tensix 2-CQ server: npz arrays (/predict and /predict_npz) and PNG bytes identical.
  • test_api_device.py (ETH): model(...) == the server pipeline bit for bit, with and without pose.

Accuracy (code/tools_prof/eth_validate.sh 16; all in ETH, 1 CQ, 12x10; gate code unchanged; all gates pass)

metric published (audit integration 2026-10-04) now (ETH, 1 CQ) gate
randn head1 / head2 (all), float 0.99862 / 0.99869 (0.99891) 0.99862 / 0.99869 (0.99891) 0.998
test_mast3r.py --layer end_to_end 0.9985 PASS (2 CQ) 0.9985 PASS (head1 0.9985, head2 0.9982) 0.99
real pair float pts3d h1 / h2, conf h1 / h2 0.99016 / 0.99088, 0.99269 / 0.99486 same digits 0.989 / 0.99
served-path u8 real pair pts3d h1 / h2, conf h1 / h2 0.99015 / 0.99097, 0.99246 / 0.99519 (2 CQ) same digits 0.989 / 0.99
sym out_ii / out_ji / out_jj vs two-pass bit-identical bit-identical identical
synthetic u8 head1 / head2 (not gated) 0.99593 / 0.99927 0.99593 / 0.99927

Performance before / after (chip 16; medians, min in brackets)

metric (method) published p150 config now (ETH, 1 CQ, 12x10) Tensix 2 CQ, same session (Galaxy-only)
trace, float (bench_breakdown, 30 it) 23.32 ms eth12 23.42 (23.07) 23.67 (23.15)
e2e model(img1, img2), float 28.53 eth12 28.38 (27.86) 27.55 (26.89)
e2e, uint8 views (served input path) 26.27 / 26.36 (worker 2 CQ) 27.83 (27.37) 26.64 (26.36)
test_mast3r.py --layer end_to_end --runs 25 27.51 ms (2 CQ, round 9) 27.77 ms
pose: single pair / sym / two-pass e2e (sym_check) 26.40 / 36.98 / 52.48 (2 CQ) 27.22 / 38.89 / 54.90
server timing_ms.forward, npz (2 x 30 req) 25.5 / 25.57 (2 CQ) 26.55-26.98 (2 servers x 2 runs) 25.63
/predict npz total / client wall 81.5-83.3 / 130-140 80.3-82.5 / 123-136 79.4-80.0 / 127-130
png total (2 x 15 req) 124.5-139.6 131.8-143.7 123.9-124.7
/predict_npz total / client wall 72-76 / 96.8-102.0 72.7-74.0 / 97.5-104.5 71.8-72.0 / 96.8-100.4
Python API device call timing_ms['forward'] (30 calls) 26.7 (26.4) 26.99 (26.54)
model(PIL, PIL) / model(path, path) whole call 41.9 / 57.1 46.75 / 58.65
predict_pairs per pair 27.9 28.09
startup (from_pretrained, warm cache) ~9 s 7.5 s (device warm-up 5.8 s, host 1.1 s)
first call / warm same input (test_warmup_device): pair, pose, batch 0.92x, 1.08x, 0.99x 0.97x, 1.09x, 1.00x (gate 10 % + 5 ms: pass)
smoke --require-pose --npz-route PASS PASS x2 PASS

Reading: on the device the configurations are the same (trace 23.4 vs 23.7 ms, within noise). The p150 config loses the CQ1 overlap of the head-1 readback: the served forward is about 1.1 ms (4 %) slower than the Galaxy-only 2-CQ mode, and the uint8 e2e about 1.2 ms. Request totals move by 1-3 ms, inside the host noise. The 2-CQ numbers are not reachable on a p150 at 12x10, so the ETH column is the honest p150 number. Logs: <scratchpad>/mast3r-p150-eth-reval/ (val/*.log, http/ab_*.txt, api.json, warmup.json).

Incident during this task

Before the chip was claimed, a host-test command (pytest models/tests) also collected the two device tests and ran them for about 2 minutes (04:53-04:56 UTC) without chipenv.sh, so with every Galaxy chip visible (device 0 = UMD chip 0, not in the pool). The run reached its summary (62 passed, 2 failed) and was interrupted (SIGINT) while it closed the device. A chip-2 hang reported by another agent at 04:55 overlaps this window; a link was not proven. Host tests are now run with TT_VISIBLE_DEVICES=none and an explicit file list.

Audit integration 2026-10-04 (chip 1, start = ad39150, final = this commit)

Source: audit/mast3r-p150/AUDIT_REPORT.md (verdict REOPEN) and the audit branch audit/mast3r-p150 (9bfbb82, d2e7aee). The user decided to integrate every demonstrated item of kind A (host / serving, outputs and /predict unchanged), B (additive binary route) and C (precision change that passes every gate, default on, env switch back). Each item is re-implemented on main and measured again on chip 1 (Galaxy BH, PCIe x1, 12x10 grid). Logs and scripts: audit/mast3r-p150/integrate/ (lofi_ab.sh, http_ab.sh, http_seq.sh, r3val_new.log, ab_*.txt, smoke_*.txt, server_*.log).

Integrated

commit item change measured on chip 1 (A/B against the pre-item state)
4c9303a C: lofie, lofid Encoder and decoder MLP fc1 / fc2 at LoFi; the same knobs upload those weights rounded to nearest-even at 4 explicit mantissa bits (LoFi reads the weights as hidden bit + 4 MSBs, so the device multiplies the rounded weights exactly; truncation alone fails the gate, 0.99364). Ported from the audit env knobs MAST3R_LOFI_MLP=both MAST3R_LOFI_RNE=1 as two MAST3R_OPT entries in all: the weight caches key on cfg, so a knob change re-uploads the weights, and MAST3R_OPT=none is unchanged (none_eth12 PCC identical to ad39150). Switch back: MAST3R_OPT=all,-lofie,-lofid (served outputs then bit-identical to ad39150, checked over HTTP) Served path (worker, 2 CQ, uint8, 60 iterations, A/B/A/B): trace 25.45 / 25.89 -> 23.62 / 23.56 ms (-7.7 %); e2e 28.29 / 28.45 -> 26.27 / 26.36 ms (-2.06 ms, -7.2 %). eth12 (ETH dispatch, 1 CQ, float, 30 iterations): trace 25.39 / 25.38 -> 23.32 / 23.24 ms (-8.3 %); e2e 30.01 / 30.98 -> 28.53 / 27.98 ms (-7.3 %). Server forward 27.0 -> 25.5 ms. Pose (sym_check, same run): two-pass 52.48, symmetric 36.98, single pair 26.40 ms
1138587 A: stored npz /predict writes the npz with np.savez (stored) instead of np.savez_compressed. Request field compress_npz (optional, default from env MAST3R_NPZ_COMPRESS, default 0) gives the deflated file back 30 npz requests x 4 runs per side, base = ad39150 server (2 runs, interleaved) vs new: server timing_ms.total 294.5-297.1 -> 81.5-83.3 ms (-72 %); encode 228-230 -> 18.1-18.6 ms; client wall 330.9-333.1 -> 130.0-139.7 ms (-60 %). Response 6.77 -> 9.79 MB (base64 JSON). Same server with compress_npz: true: total 294-304 ms (the old cost; isolates the item). Arrays identical after np.load (n_diff = 0 on all four arrays)
d32f521 A: parallel PNG output_format: "png": the 4 PNGs are encoded in a 4-thread pool (MAST3R_PNG_THREADS, default 4; 1 = serial). The pool is created once under a lock (sync handlers run concurrently in FastAPI's threadpool; host test with 8 racing threads) and shut down with the app 15 png requests x 2 runs per server: new code with MAST3R_PNG_THREADS=1 total 192.1 / 192.9 ms vs default 124.5-139.6 ms (-62 ms, -34 %); vs the ad39150 server: total 195.2-200.8 -> 124.5-139.6 ms, client wall 213.9-220.1 -> 142.1-160.0 ms (-31 %). PNG bytes identical (serial vs parallel, same build)
1eb8d6a B: POST /predict_npz New route, same JSON request; the response body is the npz itself (application/x-npz), the other fields are compact ASCII JSON in the X-Mast3r-Meta header. /predict unchanged (same keys in the same order, host test). smoke_test.py --npz-route; host tests models/tests/test_server_host.py (9 tests) 30 requests x 4 runs: total 72-76 ms, client wall 96.8-102.0 ms (-70 % vs the ad39150 /predict npz 331-333 ms, -35 ms vs stored /predict). Body byte-identical to /predict's decoded npz_b64; arrays n_diff = 0

Accuracy (code/tools_prof/r3_validate.sh 1 at 1eb8d6a; gate code unchanged; all gates pass)

metric ad39150 (audit base run) now (lofie + lofid) gate
randn head1 / head2 (all), eth12 float 0.99848 / 0.99855 (0.99875) 0.99862 / 0.99869 (0.99891) 0.998
test_e2e (2 CQ) 0.9985 PASS 0.9985 PASS (head1 0.9985, head2 0.9982) 0.99
real pair float pts3d h1 / h2 0.99035 / 0.99081 0.99016 / 0.99088 0.989
real pair float conf h1 / h2 0.99262 / 0.99511 0.99269 / 0.99486 0.99
served u8 real pair pts3d h1 / h2 0.99021 / 0.99086 0.99015 / 0.99097 0.989
served u8 real pair conf h1 / h2 0.99232 / 0.99497 0.99246 / 0.99519 0.99
sym out_ii / out_ji / out_jj vs two-pass bit-identical bit-identical identical
MAST3R_OPT=none randn 0.99836 / 0.99895 0.99836 / 0.99895 (unchanged)
not gated: synthetic u8 pair head1 / head2 0.99733 / 0.99941 0.99593 / 0.99927
not gated: real pair float median |dz|/z h1 / h2 vs fp32 ref 1.17 / 1.10 % 1.24 / 1.17 %
not gated: same vs the half-pixel reference 0.37 / 0.41 % 0.40 / 0.48 %
not gated: served u8 median |dz|/z h1 / h2 1.18 / 1.09 % 1.24 / 1.16 %

Disclosure: the LoFi item is not bit-identical. Every gate passes with the same or better margin, but the ungated synthetic-u8 head1 PCC drops by 0.0014 (from the encoder group) and the median relative depth error on the real pair grows by about 0.07 percentage points per head. The audit flagged these two as "owner review"; the user decision for this stage was to adopt precision changes that pass all gates, with an env switch back (MAST3R_OPT=all,-lofie,-lofid, verified bit-identical to ad39150 on the served path).

Served smoke tests

smoke_test.py --require-pose PASS on both ad39150 servers; --require-pose --npz-route PASS on all four new-code servers (default, MAST3R_PNG_THREADS=1, old-numerics env, default again). Host tests: test_fused_host.py 32/32, test_server_host.py 9/9.

Skipped (from the audit list)

  • json splice of the base64 payloads into the response body (audit d2e7aee): about -5 ms, inside the +-10 ms noise (audit: inconclusive). /predict_npz removes that cost for clients that opt in.
  • par (decode + preprocess of the two views in a thread pool): measured negative / bimodal in the audit (GIL).
  • fmed (np.partition median in the summary): too small to measure (audit), and the summary is outside timing_ms.
  • Every device item that the audit did not demonstrate (SDPA redesign, LN fold, hrope fold, polyphase border strips, bfloat8_b K/V, HiFi2 DPT phase convs with bias correction, megakernel, concurrent DPT heads on sub-devices): not attempted in the audit, so out of scope here.
  • Serving on an x8 Galaxy chip (audit #17): a placement item, not demonstrated, and not relevant on a real p150 (x16).
  • Nothing integrated needs a grid other than 12x10 or Tensix-only dispatch: lofie / lofid were measured on both eth12 (ETH dispatch, 1 CQ) and the served worker 2-CQ path.

Serving files

tt-model.yaml: serve.env pins MAST3R_NPZ_COMPRESS: "0" and MAST3R_PNG_THREADS: "4", the MAST3R_OPT comment lists lofie / lofid, two new verify lines (knobs in all; /predict_npz route and compress_npz field), and the card quickstart lists compress_npz and /predict_npz. No new runtime package (stdlib json, concurrent.futures only). SERVING.md: env table, compress_npz, the stored-npz change note, and the /predict_npz contract.

Round 9 (2026-10-03, chip 9, round start = 5b918db, final = ad39150)

Verifier notes on round 8, resolved

  • Span extraction script. code/tools_prof/trace_span.py (commit 0c5f9fe) is the device-span tool: per trace replay (grouped by METAL TRACE ID + METAL TRACE REPLAY SESSION ID) it prints max(FW end) - min(FW start) at 1.35 GHz, the program count, the kernel-duration sum and the summed FW gaps; --list prints the per-program timeline. Usage: MAST3R_CQS=1 python -m tracy -r -p -v -o <dir> --op-support-count 6000 tools_prof/prof_trace.py then python3 tools_prof/trace_span.py <dir>.
  • Re-validation of 5b918db on chip 9 (14:26-14:30 UTC). It reproduces the verified numbers: eth12 trace 24.96 / 24.68, e2e 29.99 ms; served u8 trace 25.32, e2e 28.44 ms; all accuracy lines identical to the verified table; device span 23.531 / 23.531 / 23.545 ms (693 programs, kernel sum 22.90 ms).
  • Server smoke test re-run (worker dispatch, 2 CQ, the tt-model.yaml serve env, chip 9, at e9396dd): PASS 4 / 4, ready 6.2 s after start, forward 37-39 ms per pose request, pose rot 43.2 deg, f1 / f2 446 / 440 (same as round 8).
  • Host wall vs device span. Explained below (AICLK throttling). The device span stays the reliable A/B metric; every host-wall A/B in this round is reported with that caveat.
  • Process. Every new kernel / kernel variant of this round ran first alone, on the smallest shape, under TT_METAL_WATCHER=2 and timeout -s INT 60/120, one device process per command. No hang, no fault.

Measured: where the 1.3-2.2 ms between trace wall and device span goes (backlog item 1)

  • Fixed trace launch + completion latency is 0.02 ms (tools_prof/r9_launch_probe.py, ETH dispatch: a 1-program trace, execute + synchronize, median 0.023 ms; event_synchronize 0.019 ms; idle synchronize 0.010 ms). Tiny programs cost 6.3 us each when dispatch-bound (700-program trace of 1-core adds: 4.43 ms), which is not the case in this graph (in-trace FW gaps 0.24 ms in total).
  • The chip clock drops under load. tt_aiclk (sysfs, read-only, chip 9) sampled every 50 ms during a 200-iteration bench: 1350 MHz when idle or lightly loaded, 1231-1343 MHz (mostly 1256-1287) while the forward runs at 100-170 W. The device profiler counts cycles and converts at the nominal 1.35 GHz, so the span understates wall time: tools_prof/r9_clock_probe.py (traces of 20 / 120 stock matmuls, wall slope vs cycle slope) gives an effective clock of 1.296 GHz (144 us matmuls) and 1.314 GHz (53 us matmuls). For this graph: 23.50 ms of cycles at ~1.28 GHz = 24.8 ms, which is the measured eth12 trace wall (24.76-24.96 ms). The gap is DVFS, not host or dispatch overhead; model code cannot remove it (less energy per forward would raise the clock a little).
  • Served-path tail (tools_prof/r9_tail_timeline.py, worker 2 CQ, u8, host event waits as read_outputs): prep + H2D + enqueue 0.55 ms; segment ends without readbacks A 22.51 / B 25.68 / C 26.30 / D 26.92 ms; with the overlapped readbacks the last segment ends ~0.9 ms later, and the e2e is 28.0 ms. The tail after the last segment is only ~0.15 ms.
  • The overlapped CQ1 readbacks stall CQ0 kernels (Galaxy x1 PCIe). tools_prof/r9_cq2_prof.py (served path under the device profiler, overlapped vs synchronize-then-read replays): segment A is unchanged (19.55 ms); segment B (head 2 main rows) goes 2.92 -> 3.44-3.54 ms and segment C 0.54 -> 1.13 ms, each because ONE program (a different one per replay: e.g. the 64^2 dadd GenericOp 24 -> 637 us, a 2-us I2S -> 565 us) stalls while the 2 MB D2H transfer runs. Reproduced with stock ops only (tools_prof/r9_d2h_stall.py, trace of 300 I2S programs + a concurrent 2 MB CQ1 read): one program stalls 645 us; all its cores start on time and finish late (their NoC transactions wait). 8 x 256 KB reads give one 80 us stall per chunk; an L1 source gives many 30-40 us stalls; total stall is the same. The PCIe link of chip 9 is x1 (32 GT/s, current_link_width 1): the device -> host writes back up into the NoC for the whole transfer (2 GB/s). A p150 (x16) drains 16x faster. Effect on the served e2e: about 1.2 ms per pair on this host; the overlap still saves ~1 ms vs reading after the last segment (2 x ~1.1 ms). No model-side fix found.

Measured: SDPA cost structure with the exact production config (backlog item 2)

Method: the stock streaming SDPA compute kernel resolved from a scratch CWD (tt-metal searches the current directory first, tt_metal/impl/kernels/kernel.cpp), with compute-only edits, separate TT_METAL_CACHE per variant, watcher first run each; tools_prof/r7_sdpa_cost.py, traced, host wall per call. Stock reader / writer and program factory unchanged. Diagnostic only (wrong outputs), nothing of this is in the model.

variant (compute kernel edit) encoder B2 H16 q160 k512 decoder B2 H12 q224 k512
stock (copy, unchanged) 94.6 us 67.4 us
no exp 91.0 60.4
no row-sum packs 92.7 65.2
no P @ V matmul 91.1 60.8
no Q @ K^T matmul 91.2 65.1
none of the four 84.2 51.2
  • Even with all matmuls, the exp and the sum packs removed, the encoder SDPA takes 84 us: the stock KV reader / chain (K and V streamed twice per core: 224 q chunks on 120 cores) is a co-bottleneck. A custom compute kernel with the stock reader can save at most ~10 us per encoder call (the planner's threshold was 15 us) and ~16 us per decoder call. Not started.
  • Zone timeline (stock zones enabled, top-level zones only): the math thread is busy for the whole kernel, ~2.4 us per q row and 16-tile K chunk for Q @ K^T + max + sub/exp, ~6 us per K chunk for the P @ V drain; ~300 cycles per score tile.
  • Chunk sweep, stock: encoder q160 k512 94 us is the best of q96-q512 / k256-k1024 (q352 110, q384 114, q256 160 us); decoder q224 k512 67 us is the best (q256 79, q192 118 us); exp_approx 66.5 vs 67.3 us (decoder).
  • What a real gain needs: per-core K / V delivered once (multicast from the hrope writer into the SDPA cores' L1) plus a leaner compute kernel and balanced 8-9 rows per core across head boundaries. Estimate -0.7 to -1.0 ms in total, a multi-kernel redesign (hrope writer + SDPA reader + compute).

Measured: other cost splits

  • hrope (tools_prof/r9_hrope_probe.py on the obs block-sharded qkv; MAST3R_HROPE_DIAG defines in the model-local kernels, default build unchanged and bit-identical): stock 29.1 us (host wall per call), no reads 25.8, no writes 23.3, no compute 25.3, none of them 11.5. No single stage dominates.
  • LayerNorm (tools_prof/r9_ln_probe.py, enc shape on the obs residual): stock 20.2 us; without reads 16.1 us (tools_prof/kern/ln_reader_fast.cpp, D_NOREAD); 1 or 2 blocks per barrier 20.2-20.3, 4 or 8 blocks per barrier 27.2-27.8 us. LN is compute-bound.
  • fc1 output writes (obsc probe below): fc1 with a block-sharded output is 125.7 vs 127.7 us interleaved; the output NoC writes are not a cost.

Kept steps

commit knob change effect
de679b6 esplit The final enc_norm (multi_layernorm, one tile row per core) writes view 1's and view 2's encoder outputs as two tensors (each core's writer gets its view's base address and start tile) instead of one [2, N, 1024] tensor + two DRAM slices (14.6 us each). Bit-identical (eth12 float, sym_check). device span 23.531 -> 23.508 ms (693 -> 691 programs)
e9396dd demb The two decoder embeds (1024 x 1024 x 768 per view, two full-grid linears 19.6 + 20.0 us + two I2S 2.7 + 2.4 us) as ONE dual program on the top / bottom 12x5 halves (dmm machinery) writing decoder step 0's per-half block-sharded residual specs directly (29.7 us). Same in0_block_w and compute config -> bit-identical. 23.508 -> 23.498 ms (688 programs)
b117533 hrqk hrope compute: q and k of a unit in one pass per stage (4 tiles per fp32 dest half instead of 2 x 2 tiles), same per-tile op sequence -> bit-identical. Probe enc 28.9 -> 28.2 us. 23.498 -> 23.487 ms

Bit-identity of the final commit vs the round start: bench_breakdown.py --dump/--cmp, eth12 float n_diff 0 on both heads, served u8 (worker 2 CQ) n_diff 0 on both heads (5b918db worktree dump vs 6d6e155).

Tried and rejected / measured only (round 9)

  • lnx: a fewer-pass LayerNorm compute kernel (row sums as x @ J, sum of squares as the diagonal of sum x x^T, SFPU statistics in fp32, (x - mean) * rstd with a dest-reuse multiply; tools_prof/rejected_lnx/). Correct; the host emulation on the real pair's LN inputs (tools_prof/r9_ln_numerics.py, |mean|/std <= 0.57) predicted equal accuracy. On device it is slower: enc 23.7 vs 20.1 us, decoder pair 20.9 vs 17.0 us; mean |err| vs fp64 1.19e-3 vs 1.17e-3. The single-tile HiFi4 matmuls cost more than the stock reduce passes.
  • obsc: fc1 writes its block-sharded output in its own packing order (column-major tiles in each shard, 7x1 subblocks) and fc2 reads it through the stock interleaved in0 sender with patched strides and a "virtual transposed" TensorAccessor (heads_rope.colmajor_ta_args, linear_descriptor(in0_colmajor=True); tools_prof/r9_obsc_probe.py). Bit-identical on the first run (watcher), but fc1 125.7 + fc2 81.5 us vs 127.7 + 80.3 us: -0.8 us per pair. Not wired in.
  • LN reader with more reads per barrier: slower (above).

Final validation (code/tools_prof/r3_validate.sh 9 at 6d6e155, 15:40-15:44 UTC, 30-iteration median / min)

none baseline (eth12, 1 CQ, float) now: eth12, 1 CQ, float now: worker 2 CQ, float now: served (worker 2 CQ, uint8)
trace (host wall, synchronized) 62.21 / 62.11 24.76 / 24.45 25.08 / 24.53 25.19 / 24.56
host prep / H2D 0.88 / 1.08 1.00 / 1.03 0.76 / 1.06 0.51 / 0.62
D2H, synchronized 28.70 3.31 2.84 3.02 (overlapped in e2e)
e2e model(img1, img2) 88.55 / 84.36 29.83 / 29.62 28.49 / 28.25 27.74 / 27.35
test_mast3r.py --layer end_to_end (2 CQ, float) 27.51 ms, PCC 0.9985, PASS
real pair: single pair / sym pose / two-pass pose e2e 27.81 / 38.35 / 55.83 (min 27.30 / 37.82 / 52.96)
device span, traced (eth12 1 CQ, 3 replays) 23.491 / 23.479 / 23.490 ms (688 programs, kernel sum 22.86 ms)

Same-session A/B (abco.sh, 3 alternating pairs, mean of medians, round start 5b918db -> new): eth12 trace 24.93 -> 25.01 ms, e2e 29.51 -> 30.05 ms (new = 3feb9b3, i.e. esplit + demb); served u8 trace 25.29 -> 25.20 ms, e2e 28.06 -> 28.12 ms (new = 6d6e155). All differences are inside the host-wall noise (run-to-run medians move by 0.3-0.5 ms); the device span (-0.044 ms) is the measurable effect.

Accuracy: unchanged to every printed digit (all kept steps bit-identical): randn head1 / head2 / all 0.99848 / 0.99855 / 0.99875 (gate 0.998); test_e2e 0.9985 PASS; real pair float pts3d 0.99035 / 0.99081, conf 0.99262 / 0.99511; served u8 pts3d 0.99021 / 0.99086, conf 0.99232 / 0.99497; raw PCC vs the half-pixel reference 0.999882 / 0.999814 (float) and 0.999886 / 0.999782 (u8); sym out_ii / out_ji / out_jj bit-identical to the two-pass path. Gate code unchanged.

vs RTX 5090 (GPU_COMPARISON.md; like-for-like = the same network, bf16, batch 1, one pair)

  • vs bf16 + SDPA (35.3 ms incl_h2d, 34.3 ms device-only): TT served u8 e2e 27.7-28.1 ms is 20-21 % faster; TT device span 23.49 ms (cycles at 1.35 GHz; ~24.8 ms at the throttled clock) is 28-32 % faster.
  • vs torch.compile + CUDA graphs (21.1 ms served, 20.1 ms device): still out of reach; the GPU is 1.31-1.33x faster served and 1.17x (cycle span) to 1.23x (throttled wall) faster on device. Of the ~6.8 ms served gap, ~1.2 ms is the D2H stall of this host's x1 PCIe link and ~1.3-1.7 ms is the clock throttling under load (trace wall vs cycle span). The stall is specific to the x1 link; whether a p150 throttles the same way was not measured.
  • Pose (sym, mode-specific, reported separately): 38.35 ms vs 2 x 35.3 = 70.6 ms for two GPU forwards.

Remaining backlog (device span 23.49 ms; ceilings measured this round)

  • SDPA redesign (hrope writes K / V once per head into the SDPA cores' L1 by multicast, balanced 8-9 rows per core, leaner compute): -0.7 to -1.0 ms. A compute-only kernel is capped at ~-10 us (enc) / -16 us (dec) per call.
  • LayerNorm 1.45 ms in 84 programs: compute-bound (reads ~4 of 20 us). A fewer-pass kernel built from single-tile HiFi4 matmuls is slower (lnx); a win needs a kernel that keeps the stock reduce LLKs but drops the two fp32 round trips.
  • fc1 GELU epilogue (~28 us x 24 + ~26 us x 12): the pack-thread SFPU can only overlap the matmul if the final sum is in dest at the last K block, which needs fp32 dest (4-tile halves) and breaks the 7x1 / 7x11 blocking.
  • Served path on this host: the D2H stall (~1.2 ms per pair) is a Galaxy x1-PCIe property; on an x8 chip (5, 13, 21,
    1. it should be ~8x smaller.
  • Small items left: pemm residual straight into the obs spec (-5 us), DPT 128^2 plumbing (L1-capacity bound).

Round 8 (2026-10-03, chips 3 -> 18, round start = 96fe46e, final = this commit)

Re-validation of the round start. HEAD 96fe46e on chip 3 at 12:15 UTC reproduced round 6/7:

  • eth12 trace 25.94 / 24.71 ms; PCC 0.99843 / 0.99872.
  • Served u8: trace 26.53 / 25.37, e2e 29.04 / 28.78 ms.
  • Real pair u8: conf h1 0.99188.

Nothing was broken, so nothing was reverted.

Measurement method (new this round: device span)

  • Device span. tools_prof/prof_trace.py under tracy (MAST3R_CQS=1, 3 trace replays). From ops_perf_results*.csv, take max(FW end) - min(FW start) of one replay's programs at 1.35 GHz. Replays agree to about 0.03 ms. This is the device time without any host latency. The bench's "trace" number (execute + synchronize, host wall) moves by 0.3-0.5 ms between runs on this shared host.
  • Host-side numbers. bench_breakdown.py, 30 or 60 iterations, median / min, as in earlier rounds.
  • A/B. Same chip, same session, alternating runs, worktree vs working tree (tools_prof/abco.sh).
chip 18, same session 96fe46e (round start) 5fbbe2d (round 8) change
device span, traced (3 replays) 23.884 / 23.889 / 23.913 ms (690 programs) 23.525 / 23.542 / 23.550 ms (693 programs) -0.35 ms (-1.5 %)
device kernel sum per replay 23.25 ms 22.91 ms -0.35 ms
eth12 1 CQ float: trace (3 pairs, mean of medians) 26.44 ms (min 25.64-25.88) 26.21 ms (min 25.40-25.55) -0.23 ms
served (worker 2 CQ, u8): trace 26.17 ms 25.85 ms -0.32 ms
served (worker 2 CQ, u8): e2e 28.68 ms (min 28.28-28.58) 28.53 ms (min 28.06-28.34) -0.15 ms

Final validation: code/tools_prof/r3_validate.sh 18 at 5fbbe2d, 13:31-13:36 UTC, 30-iteration median / min.

none baseline (eth12, 1 CQ, float) now: eth12, 1 CQ, float now: worker 2 CQ, float now: served (worker 2 CQ, uint8)
trace (host wall, synchronized) 62.75 / 62.67 26.01 / 25.46 25.20 / 24.74 25.68 / 24.65
host prep / H2D 0.82 / 1.06 0.90 / 1.03 0.89 / 1.09 0.44 / 0.58
D2H, synchronized 29.21 2.81 3.75 3.07 (overlapped in e2e)
e2e model(img1, img2) 92.57 / 84.59 30.48 / 30.21 29.06 / 28.42 28.52 / 28.31
test_mast3r.py --layer end_to_end (2 CQ, float) 27.95 ms, PCC 0.9985, PASS
real pair: single pair / sym pose / two-pass pose e2e 28.57 / 39.41 / 57.00 (min 28.40 / 39.10 / 55.22)

Chip 18's host wall numbers are about 1 ms above chip 9's (round 6: eth12 trace 25.15). The device span is the chip-independent comparison.

Server smoke test: worker dispatch, 2 CQ, MAST3R_OPT=all, chip 18, at 5fbbe2d, with the tt-model.yaml serve env.

  • PASS 4 times out of 4. Ready 7.8 s after start.
  • forward 37-39 ms per pose request (round 6: 38-39).
  • Pose rot 43.2 deg, f1 / f2 446 / 440 (round 6: 43.3 deg, 445 / 440).

Kept steps

commit knob change effect
8a7b17a resmm Encoder residual adds folded into the proj / fc2 matmul epilogues. The model-local mm_gelu_compute.cpp gets MAST3R_RESID_CB: after the FUSE_BIAS bias add, dst += resid (dest-reuse ELWADD). The residual comes from a CB globally allocated on the residual tensor's shard (heads_rope._add_resid). The residual stream stays BLOCK_SHARDED in the obs output spec, which equals the per-core output block (row-major, 1-row subblocks, asserted). Block 0 converts the patch-embed residual once (one I2S, 5.5 us). The two LayerNorms per block read only the sum, and the final enc_norm is a plain affine LN of the sum. Probe, traced, enc: proj + add/LN 51.5 -> 46.9 us, fc2 + add/LN 114.2 -> 110.0 us. Graph, chip 3, 3 A/B pairs: eth12 trace 26.00 -> 25.72 ms (mean of medians)
8fbb2e3 resmmd Same for the decoder: dual proj / cproj / fc2 with per-half residual shards (obsd specs). multi_layernorm now groups reader kernels per input layout, because the two branches are sharded on different grid halves. Step 0 converts both decoder-embed residuals once (2 x I2S, 2.5-2.9 us). The last fc2 sum goes straight into dec_norm. Taps (blocks 5 / 8) copy the sum to DRAM as before. Chip 3, 3 A/B pairs: eth12 trace 26.01 -> 25.81 ms

Accuracy (gate code unchanged; all gates pass). Neither step is bit-identical. The sum is rounded to bf16 once (fp32 dest -> one pack) instead of twice (matmul output, then the add). The residual add happens in fp32 in the matmul's dest. The matmul + bias value passes the dest -> srcA move as tf32: the partials CB of these mm32 linears is fp32, so srcA is in tf32 format.

round start 96fe46e round 8 (5fbbe2d)
randn head1 / head2 (all), gate 0.998 0.99843 / 0.99872 (0.99893) 0.99848 / 0.99855 (0.99875)
test_mast3r end_to_end (PCC h1 / h2) 0.9988 (0.9987 / 0.9986) 0.9985 (0.9986 / 0.9982), PASS
real pair float: pts3d h1 / h2 (gate 0.989) 0.99023 / 0.99074 0.99035 / 0.99081
real pair float: conf h1 / h2 (gate 0.99) 0.99209 / 0.99509 0.99262 / 0.99511
real pair float: raw PCC vs half-pixel fp32 ref h1 / h2 0.999863 / 0.999798 0.999882 / 0.999814
real pair float: median abs(dz)/z h1 / h2 1.26 / 1.12 % 1.17 / 1.10 %
real pair served u8: pts3d h1 / h2 0.99047 / 0.99085 0.99021 / 0.99086
real pair served u8: conf h1 / h2 0.99188 / 0.99483 0.99232 / 0.99497
real pair served u8: raw PCC vs half-pixel ref h1 / h2 0.999878 / 0.999770 0.999886 / 0.999782
synthetic u8 pair (not gated) 0.99690 / 0.99939 0.99733 / 0.99941
sym out_ii / out_ji / out_jj vs two-pass bit-identical bit-identical

Summary of the accuracy changes:

  • Against the fp32 reference, the real-pair metrics improved or stayed equal: raw PCC vs the half-pixel reference went up on both heads, and the median depth error went down.
  • The thin served-u8 conf-h1 margin widened from 0.0019 to 0.0023.
  • With resmm alone (chip 3), u8 conf h1 was 0.99151. resmmd moved it back up. These 1e-4-level moves are bf16 rounding variation on one pair, as in round 5.
  • The synthetic randn head2 PCC (0.99872 -> 0.99855) and test_e2e head2 (0.9986 -> 0.9982) went down slightly. Both stay well above their gates.

Tried and rejected / measured only (round 8)

  • hf2h: HiFi2 for the head.0 / head.2 phase convs (HiFi3 now).
    • It is fast: eth12 trace 24.58 ms, about -1 ms.
    • It is rejected for accuracy. Median depth error vs the half-pixel reference goes from 0.37 / 0.41 % to 0.73 / 0.91 % (both head convs), or 0.51 / 0.63 % (head.0 only).
    • This matches the earlier finding recorded in the code comment (HiFi2 there was 0.26 -> 0.68 %). Not kept.
  • bsln: LayerNorm of the block-sharded residual in its own layout.
    • Design: per-grid-row exchange of the token sums / sums of squares (bf16 hi + lo, so the fp32 sums are exact through identity matmuls) plus a fused SFPU normalise.
    • Correct, and more accurate than the row-per-core LN: enc mean abs error vs fp64 1.135e-3 vs 1.195e-3.
    • Slower: enc 30.7 us (multicast) / 27.9 us (unicast) vs 20.4 us. The exchange (8 KB per peer) and the many small compute passes cost more than the 64 -> 110 core spread saves.
    • The stock sharded ttnn.layer_norm on the same layout is also slower (26.0 us) and less accurate (1.54e-3).
    • Code: tools_prof/rejected_bsln/.
  • gelin: fc1 bias + GELU epilogue interleaved per output subblock in the last K block.
    • Bit-identical after a fix, but no gain: enc fc1 137.4 vs 136.1 us. Per subblock the chain "pack partial (L1 acc) -> unpack + bias -> SFPU GELU" is serial on the pack thread, so the SFPU epilogue still does not overlap the matmul.
    • Adding the bias before the last K block instead (so that the GELU could overlap the last block's products) was 40 % less accurate: bf16 dest accumulation at full magnitude.
    • The first version of this kernel hung chip 3 (see Known hangs).
    • Code: tools_prof/rejected_gelin/.
  • hrf: RoPE in two dest-reuse stages instead of four.
    • More accurate: mean abs error vs fp64 1.75e-3 vs 2.06e-3. This needs the srcA format set to fp32 before the dest-reuse add; otherwise the dest -> srcA move truncates to bf16.
    • But enc hrope only goes from 28.6 to 27.7 us.
    • Batching more reads per barrier is slower (ub 2 / 4 / 8: 35 / 45 / 59 us), and so is a single read in flight (28.2 us).
    • Zones show about 0.9 us of read latency per unit. hrope moves about 24 MB through the NoC per call at about 1 TB/s aggregate: it is NoC-bandwidth bound, not compute bound.
    • Code: tools_prof/rejected_hrf/.

Fresh device profile (eager, after resmm / resmmd; 717 programs, kernel sum 22.94 ms)

  • Matmuls (generic_op + ttnn): enc fc1 + GELU 127.5 us x 24, enc fc2 80.3 x 24, enc qkv 54.1 x 24, dec fc1 + GELU pair 84.1 x 12, dec qkx 56.8 x 12, dec fc2 pair 52.2 x 12.
  • SDPA 3.59 ms: enc 88 us x 24, dec 61 us x 24.
  • DPT convs 4.25 ms + halos 0.77 ms. The biggest are the head.2 phase convs (128 -> 128 at 256^2, HiFi3, 132.5 us x 8) and the 128^2 convs (resconv 256 -> 256 113 us x 8, head.0 phase 84.7 us x 8).
  • LayerNorm (now LN only, no add): enc 18.8 us x 48, dec 15 us x 36.
  • hrope: enc 25 us x 24, dec 21 us x 24.

Remaining backlog (device span 23.54 ms)

  • Custom SDPA compute kernel, about -0.7 to -1.2 ms (round-7 analysis: P @ V at 2x the FPU peak, normalise tail).
  • hrope data movement, about 1.1 ms. It is NoC-bandwidth bound: it reads and writes all of q / k / v. Writing v (and possibly q / k) straight into the SDPA layout from the qkv matmul (a tile-remapping output writer) would remove about 1/3 of that traffic. Estimate -0.2 to -0.35 ms.
  • GELU epilogue, about 0.9 ms of SFPU. It is serial on the pack thread. Only a cheaper GELU at equal accuracy, or real FPU / SFPU overlap (a different blocking with fp32 partials), would help.
  • LayerNorm-only programs, 1.55 ms on 64 cores. A layout-local LN needs a much cheaper statistics exchange than bsln.
  • DPT head convs. HiFi2 is not an option (accuracy). Conv blocking for the 128^2 / 256^2 phase convs was explored in rounds 3-5.
  • Small items:
    • pemm writing the residual straight into the obs shard spec would save the encoder I2S (5.5 us) and let block 0 use the sharded LN.
    • The decoder-embed I2S pair (5.4 us).

vs RTX 5090 (GPU_COMPARISON.md; like-for-like = the same network, bf16, batch 1, one pair)

  • Against the module in bf16 + SDPA (35.3 ms incl_h2d, 34.3 ms device-only):
    • TT served u8 e2e 28.5 ms: 19 % faster.
    • TT device span 23.5 ms vs GPU 34.3 ms device-only: 31 % faster.
  • torch.compile + CUDA graphs (21.1 ms served, 20.1 ms device) is still out of reach: the GPU is 1.35x faster served and 1.17x faster on device time (TT device span 23.5 vs 20.1 ms).
  • Pose (sym, mode-specific, reported separately): 39.4 ms vs 2 x 35.3 = 70.6 ms for two GPU forwards.

Known hangs (round 8)

  • chip 3, 13:10-13:20 UTC: tools_prof/r8_gelin_probe.py enc 20 with the first version of the new model-local fc1 compute kernel tt/kernels/mm_gelu_inline.cpp (bias + GELU epilogue interleaved per subblock in the last K block). Cause (found by review, not by re-running): the kernel waited for the bias CB at its START, but ttnn's in1 sender/writer pushes the bias only after it has pushed every in1 K block; with 8 K blocks the in1 CB fills, the reader blocks, and the compute never consumes -> deadlock. The 2-K-block small shape fit the in1 CB and passed. Fixed by waiting for the bias in the last K block (where the stock kernel waits). The fixed kernel then ran on chip 18 (watcher, 8-K-block small shape and the enc shape, bit-identical, no hang). FAULT written via chip_pool (new chip 18). A follow-on r8_gelin_probe.py dec from the same shell loop opened chip 3 for ~3 s before it was killed. Lesson: never chain several device runs in one shell loop after a new kernel; a hang must stop the loop.

Round 7 (2026-10-03, chip 9, round start = e9c87d6): no kept speedup, chip faulted

Verifier notes on round 6 (all minor), resolved in this report:

  • The independently measured round-6 A/B was -2.5 % eth12 trace (claimed -2.9 %), -2.8 % served u8 trace and -1.4 % served e2e (claimed -1.7 %). Those verified figures replace the round-6 claims. Absolute verified values at e9c87d6: eth12 trace 24.89 / 24.73 ms, e2e 30.00 / 29.35 ms; served u8 2-CQ trace 25.19 / 24.77 ms, e2e 28.13 / 27.45 ms.
  • Baseline reference: the first commit 20bdff0 has no bench_breakdown.py. The baseline numbers use 81ff339, the earliest commit with the bench tooling (trace 62.2 ms). MAST3R_OPT=none at the final commit reproduces it (62.19 / 62.11 ms).
  • chipenv's "expected 32 chips on the bus, found 31" warning is host-level and appeared on every run of every agent.
  • The server smoke test and the tracy stage profile were not re-verified independently. They were not re-run in round 7 either, because the default path did not change.

Nothing new is enabled by default. MAST3R_OPT=all is the same knob set and code path as e9c87d6. The new code below is reachable only through explicit arguments or probes. The round-7 commit changes no default kernel source.

Measured (chip 9, traced, isolated ops)

1. SDPA cost structure (tools_prof/r7_sdpa_cost.py, stock streaming SDPA, 12x10). Varying the K length at a fixed q chunking:

  • decoder B2 H12 q224 / k512: 40.8 / 67.3 / 117.6 / 214.8 us at Sk = 512 / 1024 / 2048 / 4096. That is about 300 cycles per 32x32 score tile in steady state, plus about 16 us fixed per call.
  • encoder B2 H16 q160 / k512: about 317 cycles per score tile, plus about 18 us fixed.
  • k256 instead of k512 costs about 4.6 us per extra K chunk.
  • Other points:
    • exp_approx on/off: 93.4 vs 93.5 us.
    • head dim 32 / 64 / 128: 69 / 94 / 190 us.
    • The HiFi2 FPU floor would be about 128 cycles per score tile (2 + 2 tile products).

2. fattn: stock streaming SDPA compute with a single K chunk (exact softmax, no rescale), plus a model-local reader that keeps each head's K^T / V resident in L1, plus row-balanced work (9 rows per core instead of 10) (tt/fattn.py, tt/kernels/fattn_reader.cpp, tools_prof/fattn_probe.py):

  • Correct: PCC vs fp32 0.999613 (stock 0.999613); mean |err| 0.00344 vs 0.00352.
  • Slower: 120 us vs 94 us (encoder), 103 vs 87 us (decoder shape). Without the K/V loads (diagnostic, wrong output) the compute alone is 80.7 us. So the stock kernel's KV chain / multicast forwarding costs only about 14 us, and loading full K/V per core costs about 40 us.
  • Not kept. A win here needs a new SDPA compute kernel, not a new schedule.

3. Zone profile of the stock streaming SDPA compute (model-local copy tt/kernels/fsdpa/ with the stock zones enabled, FSDPA_PROF=1, tools_prof/zonesum.py). Per 2-row q chunk x 32 K tiles:

  • Phase 1 (Q K^T + row max + partial sub/exp): about 4 us. The Q K^T matmul runs near the FPU peak.
  • Phase 2 (softmax @ V + normalise): about 12 us. The P @ V matmul alone is about 6 us, i.e. about 63 cycles per tile product, 2x the peak. With dh = 64 the output subblock is only 2 tiles wide, so every product needs about one unpack.
  • Taller P @ V subblocks (h = 4) would be about 8 % faster, but the stock streaming kernel produces wrong output with them (it assumes h <= 2).
  • This is the concrete target for a custom attention kernel: P @ V with V-tile reuse across 4+ q rows, and the normalise / rescale tail.

4. L1-interleaved NoC read bandwidth (tools_prof/r7_read_bw.py, 120 cores each reading 128 tiles):

  • Aggregate bandwidth by reads in flight per barrier: 1 -> 707 GB/s, 2 -> 955, 4 -> 879, 8 -> 731, 16 -> 514, unthrottled -> 270 GB/s.
  • The same reads split over two RISCs (reader + writer kernels, both NoCs) reach 1.66 TB/s.
  • ttnn's SDPA reader already throttles to 2. Model-local readers that issue many reads before a barrier should be checked against this.

5. Fused add + LayerNorm (tools_prof/r7_ln_probe.py, encoder shape, 64 rows on 64 cores; stock-config kernel 24.7 us traced, about 20 us kernel time):

  • Phase marks show the add phase is bound by reading the two operands (about 9 us). The remaining five compute passes take about 10 us.
  • lnf (tt/kernels/ln_fast_compute.cpp, fewer passes: ones-matmul row sums, SFPU square with packer L1 accumulation, dest-reuse normalise): correct (mean |err| vs fp64 1.62e-3 vs 1.56e-3 stock) but slower: 28.0 vs 24.7 us. The packer L1-accumulate packs cost about 350 cycles each.
  • lnsplit (operand b read by the writer on the other NoC: ln_reader_split.cpp + ln_add_writer_rb.cpp): bit-identical, slower (26.0 vs 24.7 us), because the writer serialises the b reads with the residual writes.
  • Not kept, and neither is wired into the model.

Chip fault (round 7)

tools_prof/r7_ln_probe.py enc with a variant of ln_add_writer_rb.cpp that read ALL of operand b before writing any residual block deadlocked:

  • The compute blocked on the full 8-tile CB 17 while the writer blocked on the full 8-tile cb_inb.
  • The process ignored SIGINT and was SIGKILLed.
  • A follow-on loop iteration opened the device for about 2 s before it was killed too.
  • FAULT file: chipstate/chip9/FAULT.

The committed ln_add_writer_rb.cpp is the interleaved (deadlock-free, previously run) variant, with a comment warning against the hung one. Lesson for the next round: any writer that consumes a compute output must drain it inside the same loop that feeds the compute, unless the CB holds the whole row.

Round-7 backlog (unchanged priorities, with the new measurements)

  • Custom SDPA compute kernel, about -0.7 to -1.2 ms.
    • Phase 2 (P @ V at 2x peak, plus normalise) is about 75 % of the per-chunk time.
    • Design: P @ V with 4-row output subblocks (V-tile reuse) and the row sum folded into the P @ V matmul (a ones column).
    • Keep ttnn's KV chain / multicast reader. It is only about 14 us of overhead.
  • Add + LN, -0.3 to -0.5 ms. The add phase is read-bound (about 9 of 20 us). Two-NoC operand reads with large CBs (cb_inb and CB 17 sized to the whole row, so that no deadlock is possible) should overlap the two operand streams. A pass-fusion compute kernel is not a win with packer L1-accumulation.
  • GELU epilogue overlap in fc1, about -0.5 to -0.9 ms. Not started.

Round 6 summary (2026-10-03, chip 9, HEAD 8d1c6b0)

Round start = 242bfe7. Its independent verification: eth12 trace 25.88 / 25.63, e2e 30.08 / 29.80; served u8 trace 26.18 / 25.80, e2e 28.86 / 28.50. The verifier listed no blocking problems. Its one note, the thin served-u8 conf-h1 margin (0.99188 vs the 0.99 gate), is unchanged because every round-6 step is bit-identical.

Same-session A/B, round start vs now. 3 pairs each, alternating runs. A = a git worktree at 242bfe7, B = 8d1c6b0 (tools_prof/abco.sh). Each run is bench_breakdown.py, 30 iterations, median / min.

242bfe7 (3 runs) 8d1c6b0 (3 runs) change (mean of medians)
eth12 1 CQ float: trace (device) 25.77 / 25.70 / 26.01 (min 25.50-25.69) 25.00 / 25.29 / 24.98 (min 24.77-24.84) 25.83 -> 25.09 ms, -2.9 %
eth12 1 CQ float: e2e 30.86 / 30.88 / 30.95 29.99 / 30.01 / 30.19 30.90 -> 30.06 ms, -2.7 %
served (worker 2 CQ, u8): trace 25.95 / 26.29 / 26.18 25.05 / 25.44 / 25.64 26.14 -> 25.38 ms, -2.9 %
served (worker 2 CQ, u8): e2e 28.76 / 28.88 / 28.86 (min 28.20-28.55) 28.22 / 28.38 / 28.46 (min 27.48-27.94) 28.83 -> 28.35 ms, -1.7 %

Final validation: one code/tools_prof/r3_validate.sh 9 session at 8d1c6b0, 10:24-10:28 UTC, 30-iteration median / min.

none baseline (eth12, 1 CQ, float) now: eth12, 1 CQ, float now: worker 2 CQ, float now: served (worker 2 CQ, uint8)
device forward (trace, synchronized) 62.22 / 62.12 25.15 / 24.68 25.10 / 24.88 25.36 / 25.09
host prep / H2D 0.91 / 1.07 0.85 / 1.02 0.89 / 1.05 0.41 / 0.58
D2H, synchronized 28.14 2.89 3.51 3.18 (overlapped in e2e)
e2e model(img1, img2) 91.66 / 84.02 29.91 / 29.48 28.91 / 27.98 28.32 / 27.46
test_mast3r.py --layer end_to_end (2 CQ, float) 28.12 ms, PCC 0.9988, PASS
real pair: single pair / sym pose / two-pass pose e2e 28.34 / 39.33 / 56.84 (min 27.91 / 38.80 / 54.96)
eager profile: programs / kernel sum 1252 / 60.97 714 / 23.29 (round start 714 / 24.19)

Server smoke test (worker dispatch, 2 CQ, MAST3R_OPT=all, chip 9, at 8d1c6b0) passed 4 times out of 4. The server was ready 6.4 s after start. forward took 38-39 ms per pose request (round 5: 39-41 ms). Pose rot was 43.3 deg and f1/f2 were 445/440, the same as round 5.

Accuracy: every round-6 step is bit-identical. Each step was compared with bench_breakdown.py --dump/--cmp (n_diff = 0 on both heads), on the eth12 float path and on the served u8 2-CQ path. The validation run reproduces the round-5 numbers to every printed digit:

  • randn head1 / head2 / all: 0.99843 / 0.99872 / 0.99893 (gate 0.998).
  • test_mast3r end_to_end: PCC 0.9988, PASS.
  • Real pair, float: pts3d 0.99023 / 0.99074, conf 0.99209 / 0.99509, median dz/z 1.26 / 1.12 %. Raw PCC vs the half-pixel reference is 0.999863 / 0.999798.
  • Real pair, served u8: pts3d 0.99047 / 0.99085, conf 0.99188 / 0.99483, median dz/z 1.17 / 1.09 %.
  • Synthetic u8: 0.99690 / 0.99939.
  • sym: out_ii / out_ji / out_jj are bit-identical to the two-pass path.

Matmul efficiency audit (fresh eager device profile, 12x10, HiFi2 + fp32 dest except the fc1 GELU kernels)

linear (per forward) M x K x N round start us (TFLOP/s) now us (TFLOP/s) note
enc qkv (x24) 2048 x 1024 x 3072 68.7 (188) 54.1 (238) obs
enc proj (x24, pcat) 2048 x 1024 x 1024 29.2 (147) 24.2 (178) obs
enc fc1 + GELU (x24, gpoly) 2048 x 1024 x 4096 127.5 (135) 127.4 about 100 matmul + about 27-30 SFPU epilogue
enc fc2 (x24) 2048 x 4096 x 1024 84.9 (202) 79.1 (217) obs
dec qkx pair (x12, dmm) 2 x 1024 x 768 x 3840 74.6 (81) 56.9 (106) obsd
dec fc1 + GELU pair (x12) 2 x 1024 x 768 x 3072 84.0 84.0 matmul-only 59; GELU about 26
dec fc2 pair (x12) 2 x 1024 x 3072 x 768 54.9 51.3 obsd

tools_prof/r6_mm_sweep.py sweeps the full 2-D mcast config space on 12x10:

  • grid shapes, transposed mcast, per_core_M/N incl. 11-/12-column N splits, in0_block_w 2-16, subblocks, out_block_h/w splits.

It confirmed the shipped configs as the fastest for every encoder and decoder shape. The remaining loss was not in the compute configuration:

  • The per-core compute runs at about 40 cycles per 32x32x32 tile product (HiFi2 peak is 32).
  • With an L1-interleaved output, the in1-sender/writer finished ~15 us after the compute threads in the encoder qkv (its 56 output tiles are NoC-written to 120 L1 banks at the end).
  • A BLOCK_SHARDED output equal to the per-core block lets the matmul pack straight into its own shard: qkv 75.1 -> 60.2, proj 31.0 -> 25.7, fc2 92.0 -> 87.4, decoder qkx 74.2 -> 58.4 us (in isolation, traced). That is what obs / obsd do.

The remaining grid-quantisation loss (64 tile rows on 10 grid rows -> 7 rows per core vs 6.4 ideal; fc2 N = 32 tiles on 11 of 12 columns) was not recoverable with uniform 2-D blocks. A 60 + 4 tile-row split into two programs estimates at only -5 us per fc1, because the 4-row tail job is inefficient. It was not done.

In-trace dispatch gaps are small. A traced device profile (tools_prof/prof_trace.py, 3 iterations) gives a span of 23.90 ms for the 690 programs of the main segment, with a kernel sum of 23.27 ms. So in-trace gaps total about 0.63 ms (about 0.9 us per program), not the ~1.9 ms estimated in round 5 from (bench trace - kernel sum). The rest of the bench's "trace" time is host enqueue / completion latency. Program-count reduction is therefore worth about 1 us per removed program.

Round-6 steps (A/B in one session; trace = eth12 1-CQ median)

commit knob change effect
667e43e obs Encoder qkv / proj / fc2 write BLOCK_SHARDED outputs whose shard is the 2-D mcast per-core block (_bs_mc), so there are no output NoC writes. The last shard row may be partial: 64 tile rows = 9 x 7 + 1. Consumers: hrope reads the qkv shards (reader kind = sharded layout); the fused add + LN reads proj / fc2 through its stock TensorAccessor (compile-time args of the sharded tensor) 26.12 / 25.95 -> 25.43 / 25.45 ms (2 pairs); bit-identical
5e87349 obsd Decoder dual (top / bottom 12x5) qkx, proj, cproj and fc2 write per-half BLOCK_SHARDED outputs. hrope takes two sharded layouts (self: a1 top / a2 bottom; cross: the other branch's qkx half). The fused add + LN builds one reader kernel per residual layout, each on the cores of its tensor's rows 25.38 / 25.72 -> 24.92 / 25.20 ms (2 pairs); served u8 trace 26.13 -> 25.31 (obs + obsd); bit-identical (eth12 float, served u8)
20ce5de (helper) _patch_in0_ta: a block-sharded in0 read by the stock interleaved in0 sender through a patched TensorAccessor (descriptor built on a shape-equal placeholder). Bit-identical (obsf_probe.py); used only when an in0 is sharded
8d1c6b0 (serve) tt-model.yaml knob comment (obs, obsd; no new kernel files, the verify list is unchanged)

Round-6 tried and rejected / measured only

  • obsf (encoder / decoder fc1 output block-sharded -> fc2 in0 via _patch_in0_ta): a sharded matmul output needs 1-row subblocks (a 7x1 subblock writes column-major blocks: obsf_probe.py n_diff 8.0M). The 1x1-subblock GELU fc1 runs 135 -> 151 us. Not kept.
  • hrope with direct block-shard addressing (NoC table, no TensorAccessor) + decoder cq block-sharded: no gain in 2 A/B pairs (25.06 / 24.95 vs 24.94 / 25.17 ms). Not kept.
  • fc1 with fp32 dest + the GELU epilogue (1x4 subblocks): 144.5 vs 135 us. fc1 out_block_w splits (overlap the epilogue of one block with the next): 151-162 us. Decoder fc1 config sweep (in0_block_w 2-24, subblocks): none faster than the shipped 84.8 us.
  • LayerNorm: a 2-way width split needs 128 cores for the 64 encoder (and 2 x 32 decoder) tile rows, which is more than 120. Switching the fp32 intermediate CBs to bf16 (r6_ln_probe.py) does not change the time (24.6-24.8 us), so the stock LN compute is not CB-traffic bound. Not pursued without a new LN compute kernel.

Remaining backlog (device ~25.1 ms, eager kernel sum 23.29 ms, in-trace gaps ~0.6 ms)

  • GELU epilogue, ~1.0 ms (enc fc1 ~28 us x 24, dec fc1 pairs ~26 us x 12). It is SFPU-bound on the math thread after the last K block. Only overlapping it with the FPU (a new low-level compute kernel) or a cheaper approximation at equal accuracy would cut it.
  • Add + LayerNorm, ~1.85 ms (48 x 24 us enc on 64 cores, 36 x 19.5 us dec). Compute-bound in the stock multi-pass kernel. A new LN compute kernel (fewer passes; not bit-identical) is estimated at -0.5 to -0.8 ms, with the served-u8 conf-h1 margin of 0.0019 as the accuracy risk.
  • SDPA, 3.6 ms (enc 88 us x 24 at ~97 TFLOP/s, dec 61 us x 24). Stock-kernel floor at q160/k512 and q224/k512; exp-approx, LoFi and other chunkings were measured earlier with no gain.
  • DPT ring strips: ~0.46 ms of kernels in ~46 small programs per forward. The halos on 6-row strips take 23-25 us each. An exact border correction of the polyphase convs (1-row correction terms instead of re-running upsample + conv on strips) is estimated at -0.4 ms. It is a redesign of the phase path with rounding changes.
  • Grid quantisation of the encoder matmuls (7 vs 6.4 tile rows per core): ~0.2-0.3 ms, needs non-uniform per-core blocks (a custom matmul) to recover.

vs RTX 5090 (GPU_COMPARISON.md; like-for-like = the same network, bf16, batch 1, one pair)

  • Module in bf16 + SDPA, 35.3 ms incl_h2d (34.3 ms device-only):
    • TT served 28.3 ms: 20 % faster.
    • TT eth12 1-CQ 29.9-30.1 ms: 15 % faster.
    • Device only: TT 25.1 vs 34.3 ms, 27 % faster.
  • torch.compile + CUDA graphs, 21.1 ms served (20.1 ms device): still out of reach. GPU is 1.34x faster served and 1.25x device-only. Closing the gap needs ~19 ms of device time, the practical floor estimated by the planner, at the gated precision.
  • Pose (sym, mode-specific, reported separately): 39.3 ms vs 2 x 35.3 = 70.6 ms for two GPU forwards.

Round 5 summary (2026-10-03, chip 9, HEAD 2a5f1ed)

Round start = e845c4c (verified: eth12 trace 28.76, e2e 33.71; served u8 trace 28.98, e2e 31.91). Re-measured at the start of this session (07:44 UTC): eth12 trace 28.72 / 28.60, e2e 33.23 / 32.88; served u8 trace 28.97 / 28.78, e2e 32.01 / 31.41. Final numbers: one code/tools_prof/r3_validate.sh 9 session at 2a5f1ed, 09:31-09:35 UTC, 30-iteration median / min.

none baseline (eth12, 1 CQ, float) round start (07:44) now: eth12, 1 CQ, float now: worker 2 CQ, float now: served (worker 2 CQ, uint8)
device forward (trace, synchronized) 62.21 / 62.08 28.72 / 28.60 (served 28.97 / 28.78) 25.85 / 25.52 26.14 / 25.67 26.23 / 25.72
host prep / H2D 0.79 / 1.07 0.80 / 1.01 0.99 / 1.03 0.87 / 1.06 0.39 / 0.56
D2H, synchronized 28.30 3.04 3.04 2.83 2.68 (overlapped in e2e)
e2e model(img1, img2) 91.73 / 91.39 33.23 / 32.88 (served 32.01 / 31.41) 30.96 / 30.78 29.34 / 28.94 28.88 / 28.16
test_mast3r.py --layer end_to_end (2 CQ, float) 28.84 ms, PCC 0.9988, PASS
real pair: single pair / sym pose / two-pass pose e2e 31.53 / 43.80 / 63.12 (verifier) 28.56 / 39.41 / 56.98
eager profile: programs / kernel sum 1252 / 60.97 754 / 26.86 ~710 / ~23.9 (724 / 24.20 at 740fb2f, before the 16^2 tups)

Device -10.0 % this round (28.72 -> 25.85 ms eth12; 28.97 -> 26.23 ms served), -58 % overall vs none (62.2 ms). Served e2e 32.0 -> 28.9 ms (-9.8 %); eth12 1-CQ e2e 33.2 -> 30.7-31.0 ms (-7 to -8 %, host phases noisy). Server smoke test (worker dispatch, 2 CQ, MAST3R_OPT=all, chip 9, at 2a5f1ed): PASS x4, ready 7.8 s after start, forward 39-41 ms per pose request (round 4: 43-45 ms), pose rot 43.3 deg, f1/f2 445/440 (round 4: 443/438; the focal estimate moves by 2 px with the hf3 / gpoly / tups numerics).

Accuracy (all gates pass; gate code unchanged). Three steps change numerics (hf3, gpoly, tups); the other seven are bit-identical (bench_breakdown.py --dump/--cmp, n_diff = 0 on eth12 float and on served u8 2-CQ).

round 4 / start (float ; u8) now, eth12 float now, served u8 none baseline
randn head1 / head2 (all), gate 0.998 0.99845 / 0.99868 (0.99887) 0.99843 / 0.99872 (0.99893) 0.99836 / 0.99895 (0.99909)
test_mast3r end_to_end 0.99844 PASS 0.9988 (h1 0.9987 / h2 0.9986) PASS 0.9984
real pair pts3d h1 / h2, gate 0.989 0.99065 / 0.99056 ; 0.99046 / 0.99080 0.99023 / 0.99074 0.99047 / 0.99085 0.98988 / 0.99076
real pair conf h1 / h2, gate 0.99 0.99149 / 0.99433 ; 0.99232 / 0.99478 0.99209 / 0.99509 0.99188 / 0.99483 0.99237 / 0.99439
real pair median abs(dz)/z h1 / h2 1.26 / 1.14 % ; 1.29 / 1.18 % 1.26 / 1.12 % 1.17 / 1.09 % 1.3 %
raw PCC vs the half-pixel fp32 reference h1 / h2 0.999828 / 0.999682 ; 0.999873 / 0.999748 0.999863 / 0.999798 0.999878 / 0.999770
synthetic u8 pair (not gated) 0.99589 / 0.99929 0.99690 / 0.99939 (0.99851) 0.99752 / 0.99934
sym out_ii / out_ji / out_jj vs two-pass bit-identical bit-identical

The thin float conf-h1 margin of rounds 3-4 (0.99149) widened to 0.99209. The u8 conf h1 moved the other way (0.99232 -> 0.99188, still 0.0019 above the gate). pts3d h1 float moved 0.99065 -> 0.99023 (gate 0.989), while the median depth errors improved (float 1.26 / 1.14 -> 1.26 / 1.12 %, u8 1.29 / 1.18 -> 1.17 / 1.09 %). For scale: switching the round-4 graph to ttnn's exact erf GELU (all,-gelut, slower) gives randn 0.99855 / 0.99844, real-pair float conf h1 0.99233 and pts3d 0.99039 / 0.99074. That is the same spread, so these 1e-4-level moves look like bf16 rounding variation on one pair, not a trend.

Round-5 steps (A/B in one session; trace = eth12 1-CQ median unless noted)

commit knob change effect
e4005d2 (tests) host mock tests: fake_ttnn gets uint8 / BufferType / to_torch(cq_id), the mock graph runs MAST3R_OPT=none (the optimised graph runs model-local generic_op kernels a shape-only fake cannot execute; it is device-gated) test_fused_host.py 27/32 -> 32/32 pass
463160f hf3 HiFi3 (fp32 acc) instead of HiFi4 for the two output-resolution head convs: head.0 / head.2 phase convs and their ring-strip convs. HiFi3 only drops the lo x lo mantissa partial product. Not bit-identical (n_diff 6-7 %, PCC vs previous 0.99998); real-pair metrics equal to 4-5 digits (conf h1 0.99149 -> 0.99150) 28.72 -> 28.32; served trace 28.92 -> 28.48, e2e 31.61 -> 31.04
feb7859 gpoly fc1 GELU epilogue as a model-local erf-GELU polynomial. An identity epilogue through the same pack-side SFPU path costs nothing (enc fc1 100.5 us vs 198 us with ttnn's gelu_tanh), so the epilogue IS bound by the SFPU instruction count. The round-3 conclusion was wrong: its exp-based variant was simply not shorter. New: gelu(x) = 0.5 x + abs(x) s q(s^2), s = min(abs(x), 4.25), q a 9-coefficient weighted-LP minimax fit of (Phi(s) - 0.5)/s with s q(s^2) = 0.5 at s = 4.25 (tools_prof/gelu_fit.py). It goes in the GELU_TANH slot of a model-local copy of ttnn's matmul compute kernel (tt/kernels/mm_gelu_compute.cpp, mm_gelu_activation.hpp, mast3r_gelu_poly.h); the program descriptor is otherwise unchanged. It approximates the reference's exact erf GELU with max abs error 6.2e-5; the tanh form gelut is 4.7e-4 away. After bf16 rounding, mean abs error vs erf GELU at sigma 1 / 2 / 4: this kernel 5.655e-4 / 1.126e-3 / 2.252e-3; gelut 5.804e-4 / 1.174e-3 / 2.297e-3; correctly rounded erf 5.648e-4 / 1.122e-3 / 2.246e-3. In the graph: encoder fc1 188 -> 143 us, decoder fc1 pairs 128 -> ~96 us 28.35 -> 26.98; served trace 27.16, e2e 29.70
81a44aa (gpoly) two dst rows per SFPU step: each coefficient is loaded once (SFPLOADI pair) for both Horner chains and the two MAD chains interleave 26.98 -> 26.56, bit-identical
e70a7ad qkx decoder: n_b = LN(x_b) feeds branch b's qkv AND the other branch's cross-attention k/v (norm_y(x_b) == n_b with the affine folded), so each branch runs ONE linear with concatenated weights [qkv_b, ckv_other] (N 2304 + 1536, same in0_block_w 4 -> bit-identical); the self-attention RoPE reads at row stride 5*Ht, the cross RoPE reads k/v from the other branch's output. One dual program per step instead of two 26.6 -> 26.45 (3 A/B pairs), bit-identical
36c2b6a (served path) host-side event_synchronize before each CQ1 readback instead of a device-side wait_for_event on CQ1. tools_prof/read_slow_probe2.py: a pending CQ1 wait alone slows the CQ0 trace by 0.3-0.5 ms (enqueue -> done: none 27.47, wait0 27.81, wait2 28.00 ms); device-side wait + 3 reads 28.92 vs host-side 28.52 ms. MAST3R_HOSTWAIT=0 restores the old path served e2e 29.66 -> 29.48 ms (5 of 6 A/B pairs, 60 iterations each), bit-identical
add4ceb sdbl ring-strip convs (head.0 / head.2 on the 6- / 8-row strips) with act + weight double buffering, height-sharded (strip_conv_probe.py: 68 -> 60 and 74 -> 69 us per strip conv) served e2e -0.15 to -0.33 ms (3 A/B pairs), bit-identical
fd9b448 pilhs the phase0 interleave kernel writes the head.2 phase convs' HEIGHT_SHARDED input spec directly: one NoC write per 32-row unit into its shard (tt/kernels/phase_il_rm_writer_hs.cpp). Removes the ~21 us interleaved-to-sharded copy per head -0.03 to -0.21 ms (3 A/B pairs), bit-identical
e8f518c tups DPT refinenet bilinear x2 upsample (32^2 -> 64^2, 64^2 -> 128^2) TILE -> TILE in one model-local kernel (tt/kernels/ups2_*.cpp). Each output tile is a sum of <= 4 tile matmuls with 12 constant interpolation tiles (half-pixel, clamped == ttnn.upsample; all entries exact in bf16). HiFi4 + fp32 dest, so products and sums are exact and there is one bf16 rounding. It replaces untilize + I2S + halo + upsample + S2I + tilize (tups_probe.py: 64^2 x 256 136 -> 43 us, 32^2 108 -> 20 us). Not bit-identical: the stock upsample rounds in a bf16 dest, so 23 % of its outputs differ from the correctly rounded value; for this kernel it is 3.5 %. All gates equal or slightly better (above) -0.34 to -0.57 ms (2 A/B pairs, 60 iterations); served 26.28 / e2e 28.76 ms
740fb2f tapm the 4 DPT tap projections (DRAM-latency-bound 1x1 linears, 19-22 us each) as ONE program on disjoint column blocks (6 / 3 / 2 / 1 grid columns for ap3 / ap2 / ap1 / ap0; ttnn's 2-D mcast descriptors with allowed_worker_cores, merged; dlin's in0_block_w) 79 -> 44 us per head (device profile), bit-identical
2a5f1ed (tups) the 16^2 -> 32^2 upsample too (W = 16: two image rows per input tile row, ups2h_* kernels with 4 constant tiles): replaces 7 small programs min -0.12 ms (2 A/B pairs); gates unchanged
a4a2051 sgat ring-strip inputs gathered straight from refinenet1's height shards by a model-local data-movement kernel (tt/kernels/strip_gather.cpp, 32-byte face-row reads, lr already transposed). The gathered strips are the carry for the strip segment. This replaces the 8 MB S2I to DRAM + untilize + 4 slices + 2 concats + permute eth12 -0.06 to -0.10 ms (2 pairs), served trace -0.15 to -0.40 ms (3 pairs), bit-identical
e70a7ad..2a5f1ed (serve) tt-model.yaml: knob list + verify list for the new model-local kernels (mm_gelu_compute.cpp, mm_gelu_activation.hpp, mast3r_gelu_poly.h, phase_il_rm_writer_hs, ups2_*, ups2h_*, strip_gather) verify command passes on the host

Round-5 tried and rejected / measured only

  • LayerNorm math fidelity HiFi3 / HiFi2 (the 84 add+LN programs, ~1.95 ms on 64 cores): no time change (28.35 / 28.31 / 28.21 ms within noise) and the outputs change; the LN is not bound by its multiplies. Not kept.
  • hrope2 (q and k of a unit batched per RoPE stage, bit-identical; tools_prof/rejected_hrope2/): within noise in 3 A/B pairs.
  • Phase-conv configs (tools_prof/phase_cfg_probe.py, HiFi3): act_block_h 32-128 and act / weight double buffering are bit-identical but within ~4 us per conv. Auto act_block_h for the head.0 phase convs: 0.0-0.1 ms in the graph (noise). Merging the head.0 phase convs into 2 x 256 or 1 x 512 output channels (p0_merge_probe.py): 16 % faster at 32^2, but at 128^2 every merged variant throws the L1 / CB clash.
  • cdbl for the layer_rn convs (in_c != out_c, MAST3R_FDBL A/B): bit-identical, no gain in 3 pairs.
  • gpoly variants (rejected_gelu/gelu_fast_probe.py MAST3R_GELU_EXP): 7 / 8 / 9 coefficients give enc fc1 145 / 148 / 155 us (one row per step). 9 is kept for accuracy; 8 has max error 1.1e-4 and would be about -0.07 ms. Hoisting 4 coefficients into LREGs makes the SFPU compiler run out of registers. The full-graph 8-coefficient variant had randn 0.99835 / 0.99846; the 9-coefficient one 0.99838 / 0.99868.
  • fc1 program config sweep with the gpoly epilogue (tools_prof/fc1_gpoly_sweep.py): the current configs stay best.
  • Served tail: with the readbacks overlapped, each 2 MB CQ1 read still costs the CQ0 trace ~0.5-0.7 ms (read_slow_probe.py), and a read after the trace costs ~1.36 ms. Only the device-side event wait was avoidable (above).

Remaining backlog (device 25.9 ms, kernel sum ~23.9 ms, ~710 programs)

  • Add + LayerNorm, ~1.95 ms (48 x 23 us encoder, 36 x 20 us decoder) on 64 cores. There are only 64 tile rows, every row must stay on one core for a bit-identical reduction, and 16-row face splits give 128 units for 120 cores. A width split needs a cross-core (semaphore) reduction and changes the summation order. Estimated -0.4 to -0.6 ms, with accuracy and hang risk.
  • The 128^2 add3s reads the upsampled x from DRAM; writing x into the add3s shard spec would save about 0.05 ms but keeps 8 MB in L1 across the resconv convs (CB risk).
  • head.2 tail: the 4 per-phase 1x1 linears write 87 % padding (4 valid of 32 columns each), and tailf re-reads 16 MB. A fused phase-lin + tail kernel would save about 0.1 ms but needs the 4 phase-conv outputs live together in L1.
  • Dispatch gaps: ~1.9 ms (trace 25.9 vs kernel sum ~23.9 ms).
  • Served tail: ~2.5 ms (prep 0.6, H2D 0.6, overlapped readback slowdown ~1.0, last read). It is bounded by the x1 link and the CQ0 slowdown during CQ1 reads.

vs RTX 5090 (GPU_COMPARISON.md; like-for-like = the same network, bf16, batch 1, one pair)

  • Module in bf16 + SDPA, 35.3 ms incl_h2d (34.3 ms device-only):
    • TT served 28.9 ms: 18 % faster.
    • TT eth12 1-CQ 31.0 ms: 12 % faster.
    • Device only: TT 25.9 vs GPU 34.3 ms, 25 % faster.
  • bf16 module 42.3 ms: TT served is 32 % faster.
  • torch.compile + CUDA graphs, 21.1 ms (20.1 ms device): still out of reach. GPU is 1.37x faster (served) / 1.29x (device only).
  • Pose (sym), reported separately as mode-specific: 39.4 ms vs 2 x 35.3 = 70.6 ms for two GPU forwards.
  • gpoly evaluates the same erf GELU the GPU reference runs, more accurately than the earlier tanh form. tups computes the same half-pixel bilinear upsample the port always used, with exact products. Neither is an algorithmic shortcut, so both stay in the like-for-like comparison.

Round 4 summary (2026-10-03, chip 9, HEAD 8af42e2)

Round start = c9679bd (verified: eth12 trace 30.45 / 30.23, e2e 35.71 / 35.49; served trace 30.55, e2e 33.16 / 32.81). Final numbers: one code/tools_prof/r3_validate.sh 9 session, 07:29-07:33 UTC, 30-iteration median / min. The script now takes the chip and the checkout as arguments: r3_validate.sh [CHIP] [CHECKOUT_DIR].

none baseline (eth12, 1 CQ, float) round start c9679bd (verified) now: eth12, 1 CQ, float now: worker 2 CQ, float now: served (worker 2 CQ, uint8)
device forward (trace, synchronized) 62.20 / 62.12 30.45 / 30.23 (served 30.55 / 30.34) 28.81 / 28.62 28.91 / 28.72 28.90 / 28.75
host prep / H2D 0.61 / 1.04 0.93 / 1.02 0.87 / 1.04 0.90 / 1.07 0.57 / 0.57
D2H, synchronized 20.89 3.59 3.20 3.24 3.12 (overlapped in e2e)
e2e model(img1, img2) 92.72 / 84.50 35.71 / 35.49 (served 33.16 / 32.81) 33.98 / 33.09 32.05 / 31.83 31.99 / 31.48
test_mast3r.py --layer end_to_end (2 CQ, float) 33.67 31.07 ms, PCC 0.99844 PASS
real pair: single pair / sym pose / two-pass pose e2e 32.89 / 46.22 / 65.99 31.49 / 43.46 / 62.66
eager profile: programs / kernel sum 1252 / 60.97 905 / 28.48 803 / 27.20 (at 20b64b5; pcat/p1hs remove ~50 more)

Device -5.4 % this round (30.45 -> 28.81 ms eth12; 30.55 -> 28.90 ms served), -54 % overall vs none. The served e2e moves 33.2 -> 31.4-32.0 ms between same-day runs (07:26 A/B run 31.44 / 31.07; validation run 31.99 / 31.48). That is -3.5 % to -5 %. The eth12 1-CQ e2e moves 35.7 -> 34.0 ms (-4.8 %). Server smoke test (worker dispatch, 2 CQ, MAST3R_OPT=all, chip 9, at 8af42e2): PASS x4. Pose rot is 43.3 deg and f1/f2 are 443/438, the same as round 3. forward is 43-45 ms per pose request (round 3: 46-48 ms).

Accuracy: every round-4 step is bit-identical. Each step was A/B'd with bench_breakdown.py --dump/--cmp on the eth12 float path and on the served uint8 2-CQ path; the n_diff count was 0 for both heads. The gated numbers equal round 3's to every printed digit:

  • randn head1 / head2 / all: 0.99845 / 0.99868 / 0.99887 (gate 0.998).
  • test_mast3r end_to_end: PCC 0.99844, PASS.
  • Real pair, float: pts3d 0.99065 / 0.99056, conf 0.99149 / 0.99433, median dz/z 1.26 / 1.14 %, raw PCC vs the half-pixel reference 0.999828 / 0.999682.
  • Real pair, served u8: pts3d 0.99046 / 0.99080, conf 0.99232 / 0.99478, median dz/z 1.29 / 1.18 %.
  • Synthetic u8 (not gated): 0.99589 / 0.99929.
  • sym: out_ii, out_ji and out_jj bit-identical to the two-pass path.

The thin conf h1 float margin (0.99149 vs the 0.99 gate) is unchanged. This round did not touch numerics.

Round-4 steps (A/B in one session; trace = eth12 1-CQ median unless noted)

commit knob change effect
12205a0 (tooling) r3_validate.sh [CHIP] [CHECKOUT] fixes the verifier's "hardcoded chip 9 / main checkout" note
27f59fc rchain DPT resConfUnit conv1 -> conv2 on ttnn's L1 conv path. conv1's sharded output feeds conv2 directly, which removes an S2I to DRAM + I2S per resconv and the DRAM-path NHWC reshape copy at 16^2. At >= 64^2 the relu input is height-sharded into the conv's own spec and freed after halo (that is the only way 128^2 fits the L1 CB budget). Below 64^2 it lives in L1 interleaved. Per-resconv probe (rchain_probe.py): 128^2 376 -> 322 us, 64^2 148 -> 128, 16^2 76 -> 47 30.46 -> 30.05 ms
6850a3b dlin refinenet out_conv (1x1) with its input in L1, and output in L1 below 128^2. Uses an explicit full-grid config with the auto config's in0_block_w and compute config, so every output tile sees the same K sequence (dlin_probe.py: 16384x256x256 96 -> 30 us) 30.06 -> 29.86
bd8aaef dfront DPT tap projections write L1, so the ConvTranspose / ap3_down / layer_rn convs use the L1 conv path (removes 16^2 reshape copies). The ap1 ConvTranspose output goes to DRAM: l2_rn on its sharded output picks another config and is not bit-identical 29.96 -> 29.82
ab3928d dshard >= 64^2: the refinenet adds read conv2's height shards in place and write relu(s) straight into the next conv1's shard spec. Model-local add3s / add2 kernels, same FPU adds. The leading relu writes the shard spec directly (add3s_probe.py at 128^2: S2I + add3 + I2S 156 -> 72 us, S2I + add 77 -> 29 us) 29.76 -> 29.34; served e2e 31.98
cb12ce3 ups1 upsample output staged in L1 before the tilize 29.36 -> 29.27 (2 A/B pairs)
108afa4 tailf head tail after the 4 per-phase 1x1s as ONE model-local program: phase sum + bias + WH transpose (compute) and the 16-bit NCHW interleave (writer). Replaces 3 adds + transpose + bias add + slice + untilize + tail_il (tailf_probe.py: 113 -> 55 us per head) 29.27 -> 29.18 (3 A/B pairs)
20b64b5 pemm patch-embed linear with an explicit config (auto in0_block_w) writing the L1 residual directly (drops the 4 MB copy) 29.17 -> 29.12
ba4971d pcat attention output projections (encoder proj, decoder self / cross proj) read the SDPA output directly. Built from ttnn's 2-D mcast matmul descriptor with the in0 sender reader swapped for a model-local copy that remaps tile ids (tt/kernels/mm_in0_heads_reader.cpp), so 48 concat-heads programs go away (pcat_probe.py: enc 39.9 -> 33.3 us, dec pair 31.1 -> 26.7 us) 29.16 -> 28.90
194c1a0 p1hs refinenet1's out_conv writes head.0's height-shard spec directly; the DRAM copy for the strips is one S2I. dlin uses the auto 1-D layout for the N=96 tap projection 28.91 -> 28.73
8af42e2 serve tt-model.yaml: knob list + verify list for the new kernels (add3s/add2/tailf/mm_in0_heads)

Round-4 tried and rejected

  • thalf (served path): head 2's main rows in two trace segments (rows 2i | 2i+1, tailf in half mode, exact), so the first 1 MB readback overlaps the second half. In 3 A/B pairs the e2e did not move (31.60-31.98 vs 31.68-32.03 ms). The e2e_tail_probe.py / e2e_timeline2.py timelines show why. With no reads, the device is done at 30.3 ms (including prep + H2D). With concurrent CQ1 readbacks the last trace ends about 1 ms later, so overlapped D2H on the x1 link slows the device instead of hiding behind it.
  • strip1: both heads' ring strips in one last segment with one readback. e2e 31.66 -> 32.17 ms (worse).
  • r1 out_conv output in L1 (instead of DRAM): CB clash in the 2-CQ served path (p1 is carried across trace segments).
  • Device im2col for the uint8 upload: the host copy is only ~0.2 ms either way (measured), so it does not pay.
  • Decoder 4-way cq+ckv (or qkv+ckv) merge: not done. These pairs already run at 120-160 TFLOP/s on their halves, so a merge saves only one 2 us dispatch per step (25 us per forward).

Remaining backlog (device 28.8 ms, kernel sum ~27 ms, ~750 programs)

  • Encoder / decoder fc1 GELU epilogue: about 100 us x 24 + 70 us x 12, about 3.1 ms. This is the largest item. It is SFPU-bound on the pack thread. Only an accuracy-gated activation change could cut it (round-3 rejection notes).
  • LayerNorm programs (add + LN) run on 64 cores, because there are only 64 tile rows. That is 48 x 23 us (enc) + 36 x 20 us (dec), about 1.8 ms. A width-split LN with a cross-core reduction could reach ~120 cores, but it is not bit-identical. Estimated -0.5 ms.
  • DPT ring strips: about 0.25 ms per head across ~25 programs. Halos at 23-25 us come from 6-row, 256-wide strips on 96 cores. Possible fixes: a transposed or width-sharded layout, or one batched program. Estimated -0.2 ms.
  • DPT 128^2 plumbing: tilize 35 us, relu 23 us, add3s 70 us (DRAM-bound) per head. Keeping the upsampled x in L1 would need the refinenet split moved (do u1's convs before the upsample). Estimated -0.1 ms.
  • Dispatch gaps: about 2.0 ms (2.5 us x ~750 programs).
  • Host mock tests (code/models/tests/test_fused_host.py): 5 tests already fail at c9679bd and still fail. fake_ttnn lacks uint8 and the newer generic_op paths. These tests are not device gates, but they are stale.

vs RTX 5090 (GPU_COMPARISON.md)

  • bf16 + SDPA module 35.3 ms (34.3 device-only): TT served 31.4-32.0 ms (9-11 % faster); eth12 1-CQ 34.0 ms (4 % faster, previously a tie); device-only TT 28.8 ms vs 34.3 ms (16 % faster).
  • bf16 module 42.3 ms: TT served 24-26 % faster.
  • torch.compile + CUDA graphs 21.1 ms (20.1 device): still out of reach. GPU is 1.5x faster.
  • Pose (sym): 43.5 ms vs 2 x 35.3 = 70.6 ms for two GPU forwards (mode-specific; reported separately).

Round 3 summary (2026-10-03, chip 9, HEAD aae6921)

Round start = b1f4930, the verified round-2 result. All numbers below are 30-iteration median / min from one session (code/tools_prof/r3_validate.sh, 2026-10-03 06:21-06:28 UTC), except the round-start column, which was measured at the start of this session (04:28 UTC served path; the eth12 round-start trace/e2e is the all,-sdpa2 A/B run of 04:4x, i.e. the b1f4930 graph).

baseline MAST3R_OPT=none (eth12, 1 CQ, float) round start b1f4930 now: eth12, 1 CQ, float now: worker 12x10, 2 CQ, float now: served path (worker 12x10, 2 CQ, uint8)
device forward (trace, synchronized) 62.18 / 62.08 39.73 / 39.57 (eth12); 39.93 / 39.69 (served) 30.44 / 30.21 30.49 / 30.35 30.59 / 30.41
host prep 0.81 0.88 0.80 0.62
H2D 1.05 1.04 1.05 0.59
D2H, synchronized, incl. host assembly 20.63 3.59 2.96 3.76 (overlapped in e2e)
e2e model(img1, img2) 92.67 / 83.87 45.16 / 44.56 (eth12); 42.43 / 42.09 (served) 35.60 / 35.09 (previous run 35.53 / 34.78) 33.77 / 33.41 33.42 / 33.13 (previous run 33.39 / 32.91)
real pair (kitchen, sym_check.py) single-pair e2e 33.01 / 32.57
return_pose (sym B-mode), real pair e2e 60.41 / 60.07 (round 2) 46.01 / 45.62 (two-pass: 65.77 / 65.07)
test_mast3r.py --layer end_to_end (2 CQ, float) 42.69 (round-2 verification) 33.58 ms
eager profile: programs / kernel sum 1252 / 60.97 1232 / 37.64 925 / 28.68 ms (at 3627397; dadd/addln2 remove ~22 more)

Device 39.7 -> 30.4 ms (-23 %) this round, 62.2 -> 30.4 ms (-51 %) overall. e2e (served path) 42.4 -> 33.4 ms (-21 %); eth12 1-CQ like-for-like 45.2 -> 35.6 ms (-21 %). Host-side phases (prep, D2H, e2e) move by up to +-0.5 ms between runs on the shared host.

Accuracy (all gates pass; gate code in test_mast3r.py unchanged since the baseline commit)

round 2 (b1f4930) float / u8 now, eth12 float now, served u8 none baseline
randn pair head1 / head2 (all) — gate 0.998 0.99846 / 0.99870 (0.99890) 0.99845 / 0.99868 (0.99887) n/a 0.99836 / 0.99895 (0.99909)
test_mast3r end_to_end PCC (h1 / h2) 0.9984 (0.9984 / 0.9980) 0.9984 (0.9985 / 0.9982), PASS 0.9984
real pair pts3d PCC h1 / h2 — gate 0.989 0.99046 / 0.99081 ; 0.99043 / 0.99091 0.99065 / 0.99056 0.99046 / 0.99080 0.98988 / 0.99076
real pair conf PCC h1 / h2 — gate 0.99 0.99182 / 0.99459 ; 0.99220 / 0.99481 0.99149 / 0.99433 0.99232 / 0.99478 0.99237 / 0.99439
real pair median abs(dz)/z h1 / h2 1.13 / 1.18 % ; 1.20 / 1.21 % 1.26 / 1.14 % 1.29 / 1.18 % 1.3 %
real pair raw PCC vs half-pixel fp32 ref h1 / h2 0.999837 / 0.999849 ; 0.999879 / 0.999884 0.999828 / 0.999682 0.999873 / 0.999748
synthetic smooth uint8 pair head1 / head2 (all) (not a gate) ; 0.99711 / 0.99929 (0.99855) 0.99589 / 0.99929 (0.99740) 0.99752 / 0.99934 (0.99900)
sym B-mode out_ii / out_ji / out_jj vs two-pass bit-identical bit-identical

Honest notes on accuracy:

  • Most round-3 steps are bit-identical (hrope, hcat, mln, addln, resl1d, cdbl, pil, til, dmm, the common-args change; each A/B'd with bench_breakdown.py --dump/--cmp). The numerics changed only through sdpa2 (larger SDPA chunks: fewer online-softmax rescales), mc2d/mc2dr (2-D mcast linears instead of minimal_matmul / the fused dit residual op) and mm32 (fp32 dest accumulation in those linears).
  • mc2d/mc2dr alone (bf16 dest, larger in0_block_w) worsened the served-path real pair: conf PCC h1 0.99071, median depth error 1.43 %. Probing showed that a larger in0_block_w accumulates more K tiles in the 16-bit dest (tools_prof/resid_acc_probe.py: mean |err| vs fp64 9.1e-3 / 13.1e-3 at in0_block_w 8 / 16 vs 7.8e-3 for the old dit op). mm32 (fp32 dest, out subblocks <= 4) gives 6.0e-3 — more accurate than the round-2 matmuls — at the same speed, and it restored the real-pair gates (conf h1 0.99232). It is part of the default set.
  • Remaining drift vs round 2: the real-pair median relative depth error is 1.26 / 1.14 % (float) and 1.29 / 1.18 % (u8), vs 1.13 / 1.18 % and 1.20 / 1.21 % in round 2, and vs 1.3 % for the none baseline. Head-2's raw PCC vs the half-pixel reference fell from 0.99985-0.99988 to 0.99968-0.99975. These metrics moved by a similar amount (±0.1 % median, ±0.0002 PCC) under every individual numeric knob in A/B runs (e.g. SDPA k-chunk 256 vs 512, or removing sdpa2), so they look like bf16-rounding variation on one image pair rather than a systematic loss, but they are reported as measured.
  • The synthetic smooth uint8 pair (not a gate) moved 0.99711 -> 0.99589 on head 1. Its value jumps non-monotonically with the knob set (0.99604 without the mc2d family, 0.99680 without mm32, 0.99675 without sdpa2, 0.99711 with only the round-2 knobs); none gives 0.99752.

vs RTX 5090 (GPU_COMPARISON.md; "incl_h2d" = upload of both views + forward + readback of both maps)

GPU variant GPU incl_h2d (excl_h2d) ms TT served path (2 CQ, uint8) TT eth12 1 CQ float (like-for-like p150 config)
fp32 strict 103.4 33.4 (TT 3.1x faster) 35.5 (2.9x)
bf16 autocast 63.8 33.4 (1.9x) 35.5 (1.8x)
module converted to bf16 42.3 (41.3) 33.4 (TT 21 % faster) 35.6 (TT 16 % faster)
module bf16 + SDPA 35.3 (34.3) 33.4 (TT 5 % faster; min 33.1) 35.5-35.6 / min 34.8-35.1 (median 0.6-0.8 % slower, min 0.6-1.5 % faster: a tie)
bf16 module + torch.compile + CUDA graphs 21.1 (20.1) 33.4 (GPU 1.6x faster) 35.5
device only (TT trace vs GPU excl_h2d) 34.3 (bf16 + SDPA) 30.6 (TT 11 % faster) 30.4 (TT 11 % faster)

Caveats: the GPU rows upload fp32 views (the TT served path uploads uint8, 1.5 MB); with uint8 input the GPU would save part of its ~1 ms of transfers, so the bf16 + SDPA comparison is a near tie in transfer-neutral terms (device-only: TT 30.4-30.6 vs GPU 34.3 ms). The compiled-graph GPU row (21.1 ms) stays out of reach at the gated precision. Pose requests: 46.0 ms with the sym B-mode vs 2 x 35.3 = 70.6 ms for two GPU forwards (mode-specific saving; reported separately). Served-like totals (PNG decode + npz encode, ~180 ms of host work) are host-bound on both accelerators.

Server smoke test (worker dispatch, 2 CQ, MAST3R_OPT=all, chip 9, at 3627397 and again at aae6921): PASS x4 each, ready 6.5 s after start (warm-up captures both traces), forward=46-48 ms per pose request (the sym graph), pose rot 43.3 deg, f1/f2 443/438 in both runs.

Round-3 steps (each kept step committed; A/B in one session; bit-identical unless noted)

commit knob change effect (trace, eth12 1 CQ)
5f2e4cc sdpa2 per-shape SDPA chunks balanced for 120 cores (flat BHq-chunk scheduling): encoder q160/k512 (7 q-chunks per head -> 224 chunks <= 2 per core), decoder q224/k512 (5 per head -> 120 chunks, 1 per core); the kernel pads the partial last chunk. Probe (sdpa_probe3.py): 133 -> 94 us, 93 -> 67 us 39.73 -> 38.69; numerics change (randn 0.99851 / 0.99867)
8668258 hrope model-local generic_op kernel (tt/kernels/heads_rope_*.cpp): split heads + RoPE(q, k) in one program on 120 cores, reading the per-branch qkv / q / kv matmul outputs directly (also removes the decoder's batch concats); the compute kernel runs rotary_embedding_llama's exact op sequence. Probe: 72.4 -> 26.5 us (enc), 74.3 -> 24.2 / 75.3 -> 24.0 us (dec self / cross). Also: the encoder's L1 ones vector became a per-pass copy, so no L1 buffer sits under the DPT conv CBs 38.33 -> 35.81
2f6e591 hcat model-local concat-heads kernel (pure data movement, 120 cores); the decoder writes the two branch outputs directly (no batch slices). 13.5 -> 9.6 us enc, 19.4 -> 8.3 us dec 35.81 -> 35.50
3852783 mln the decoder's branch LayerNorm pairs as ONE program: ttnn's own LN kernels with the exact compile-time config of its factory, via generic_op, each core's runtime args pointing at its branch's tensor (64 cores instead of 2 x 32). 30.9 -> 21.3 us per pair 35.50 -> 35.10
caad9ac mc2d, mc2dr plain and residual linears as full-grid 2-D mcast ttnn.linear with per-shape in0_block_w (8 / 16 for K = 768 / 3072 / 4096) instead of minimal_matmul / the fused dit op (residual as a separate add): dec 768x768 25 -> 12.4 us, fc2 56 -> 33 us, enc fc2 115 -> 85 us 35.10 -> 34.33 (mc2d) -> 32.97; numerics change, see mm32
795536d resl1d decoder residual streams in L1 (block-5/8 taps copied to DRAM, dec_norm out in DRAM) 32.97 -> 32.52
eae74ab addln every residual add that feeds a LayerNorm runs inside the LN program: ttnn's FUSE_PRE_ADD reader + a model-local copy of layernorm.cpp that also packs the bf16 sum (+ a model-local writer); cb_x kept bf16 so the LN sees the same bf16 sum ttnn.add produces (bit-identical). Dec pair 41.9 -> 24.6 us, enc 54.1 -> 31.2 us 32.52 -> 32.06
d1fb6ac cdbl DPT 3x3 convs <= 64^2: activation + weight double buffering (64^2 height-sharded, 32^2 / 16^2 block-sharded): 96 -> 76, 51 -> 40, 45 -> 39 us 32.06 -> 31.83
644b5c5 pil phase0 interleave by a model-local data-movement kernel on ROW_MAJOR pixel rows (the head.0 phase convs emit ROW_MAJOR) instead of concat + transpose + 0/1 permutation matmul + transpose: 30 us vs about 200 us per head 31.83 -> 31.53
ed25302 til head-tail phase interleave by a model-local kernel (16-bit element interleave straight into the NCHW rows) instead of RM reshape (65 us) + permute + reshape + tilize + 0/1 matmul + untilize: 35 us vs about 130 us per head 31.53 -> 31.29
784b1d2 mm32 mc2d / mc2dr linears with fp32 dest accumulation (out subblocks <= 4) same speed; restores accuracy (above)
2c106b6 dmm the decoder's same-shape branch linears (qkv, proj, cq, ckv, cproj, fc1+GELU, fc2) as ONE program each: ttnn's own 2-D mcast matmul descriptors (MatmulMultiCoreReuseMcast2DProgramFactory.create_descriptor) placed on the top / bottom 12x5 halves of the grid with allowed_worker_cores, merged with ttnn.merge_program_descriptors, run through generic_op. Pairs: qkv 60 -> 51, 768x768 30 -> 25, ckv 44 -> 37, fc1 154 -> 132, fc2 72 -> 63 us 31.44 -> 30.64
3627397 (hrope) shared addresses as common runtime args no change
9082ce4 dadd DPT refinenet with skip: x + (skip + c2) and the next unit's leading relu as one model-local program (FPU adds in ttnn.add's order + SFPU relu): 128^2 165 -> 104, 64^2 56 -> 36, 32^2 21 -> 11 us 30.65 -> 30.49
aae6921 addln2 encoder / decoder last add + enc_norm / dec_norm (affine) as the fused add + LN program with gamma/beta (stock FUSE_GAMMA/BETA path); decoder block-5/8 taps copied from the next step's fused sums 30.49 -> 30.43

Round-3 tried and rejected

  • Faster GELU epilogue (x / (1 + exp(-2u)), exp_21f or fp32-accurate exp + Newton reciprocal) in a model-local copy of the fc1 matmul compute kernel, built from ttnn's own matmul descriptor (tools_prof/rejected_gelu/): 191-207 us vs 194-210 us stock; the GELU epilogue (about 96 of 198 us in the encoder fc1) is not bound by the SFPU instruction count. The exp_21f variant was also less accurate (up to 1.8 half-ulp on x in [-5, 1)).
  • Separate GELU op after a plain fc1: 102 + 110 us vs 198 us fused.
  • fc1 out-subblock / in0_block_w / grid sweep (fc1_sub_probe.py): best = current.
  • Merging the 4 head-tail phase convs into one 128 -> 512 conv: fits L1 only at 64^2 (129 -> 95 us there); at 256^2 the height-sharded variants throw a CB/L1 clash and the block-sharded ones are 5-8x slower.
  • Double buffering for the phase / phase0 convs and for the 128^2 refinenet convs: no gain.
  • TILE-layout phase0 interleave reading 32-byte face rows over the NoC: 144 us (vs 30 us for the ROW_MAJOR kernel).
  • SDPA q64 / q32 / q96 / q192 / q256 / q512 and k1024 variants: slower (sdpa_probe3.py).
  • minimal_matmul grid / block sweep for the 768-wide decoder linears (mm_grid_sweep.py): at most 2-3 us; superseded by mc2d.
  • Half-grid decoder linears split by columns (6x10 halves): only 2-4 us per pair; the row split (12x5) is what dmm uses.

Megakernel / parameter taxonomy status (round 3)

  • No parameter classification changed: A (launch-fixed) = resolution 512, batch 1, weights, MAST3R_CQS, uint8 input, the knob set; B = return_pose (precompiled sym graph, captured at warm-up); C = pixel contents.
  • Fusion: 1232 -> about 903 programs per forward. Model-local generic_op kernels: heads_rope, heads_concat, ln_add (copy of layernorm.cpp + writer), phase_il_rm, tail_il, add3; plus ttnn's own LN and matmul kernels re-launched through generic_op with per-core tensor bindings / merged disjoint-core programs (mln, addln, dmm). Still one trace per forward segment, one H2D, readbacks overlapped on CQ1.
  • The per-program dispatch gap is now about 2.1 us (trace 30.65 vs kernel sum 28.68 ms over 925 programs): fewer programs remain the main lever for the gaps.

Remaining backlog (round-3 view; device 30.7 ms, kernel sum 28.7 ms)

  • Encoder fc1 + GELU_TANH: 188 us x 24 = 4.5 ms; the GELU epilogue is ~96 us of it and is not reducible with fewer SFPU instructions (above). Only a different, accuracy-gated activation form could cut it.
  • DPT: the remaining resconv add (sum + c2' before out_conv) could fold into the out_conv input path; conv in/out reshards (I2S / S2I / halo, 1.8 ms total) are bound by the stock conv L1 CB sizes.
  • Head-tail post-conv chain (4 x 128->4 linears, 3 adds, transpose, bias add, slice, untilize: about 125 us per head): a fused kernel could save about 0.2 ms.
  • e2e tail (served path): e2e - trace = 2.6 ms, of which about 1 ms is host prep + H2D and the rest is mostly the 2 MB head-2 main-row readback that the remaining strip work (about 0.9 ms) cannot fully hide on the x1 link.
  • Compiled-GPU parity (21.1 ms) is not reachable at the gated precision.

Round 2 summary (2026-10-03, chip 9, HEAD f6cbac0)

Round start = aff19e9, the verified round-1 result. Same-session numbers, all 30-iteration median / min, measured 2026-10-03 04:13-04:16 UTC.

baseline MAST3R_OPT=none (eth12, 1 CQ, float) round start aff19e9 (eth12, 1 CQ, float) now: eth12, 1 CQ, float now: worker 12x10, 2 CQ, float now: served path (worker 12x10, 2 CQ, uint8)
device forward (sum of the trace segments, synchronized) 62.22 / 62.10 40.01 / 39.81 39.64 / 39.37 39.78 / 39.44 39.78 / 39.51
host prep 0.86 0.94 0.98 0.88 0.46
H2D 1.07 1.04 1.03 1.07 0.60 (1.5 MB uint8)
D2H, synchronized, both maps incl. host assembly 29.00 3.34 3.09 3.09 3.01 (hidden in e2e)
e2e model(img1, img2) 92.15 / 91.89 45.06 / 44.44 44.57 / 44.06 43.22 / 42.72 42.01 / 41.84
real pair (kitchen, tools_prof/sym_check.py) e2e 42.03 / 41.54
return_pose (two maps + swapped pass), real pair e2e 84.14 / 83.73 (two-pass, today's graph) 60.41 / 60.07 (sym B-mode)
test_mast3r end_to_end (2 CQ, float) 44.79 42.82
eager profile: programs / kernel sum 1252 / 60.97 1230 / 38.07 1232 / 37.64

Accuracy, all gates pass:

round start / float path now uint8 served path now
synthetic randn pair head1 / head2 (all) 0.99846 / 0.99870 (0.99890); float outputs are bit-identical to the round start n/a (randn is not a uint8 image)
synthetic smooth uint8 pair head1 / head2 (all) float path on the same pixels 0.99667 / 0.99937 (0.99856); MAST3R_OPT=none 0.99752 / 0.99934 (0.99900) 0.99711 / 0.99929 (0.99855)
real pair pts3d PCC h1 / h2 0.99046 / 0.99081 0.99043 / 0.99091
real pair conf PCC h1 / h2 0.99182 / 0.99459 0.99220 / 0.99481
real pair median abs(dz)/z h1 / h2 1.13% / 1.18% 1.20% / 1.21%
real pair raw PCC vs half-pixel fp32 reference h1 / h2 0.999837 / 0.999849 0.999879 / 0.999884
test_mast3r end_to_end 0.9984 (h1 0.9984 / h2 0.9980) PASS

Honest notes:

  • The uint8 path changes the arithmetic. The input is exact (v - 127.5), and 1/127.5 is folded into the patch-embed weight in fp64 with a single bf16 rounding. Against the half-pixel fp32 reference it is closer than the float path. Against the align-corners reference, its median relative depth error is 1.20 / 1.21% (float path 1.13 / 1.18%); the PCCs are equal or higher.
  • On the smooth synthetic uint8 pair, every variant is below 0.998 on head 1, including the none baseline at 0.99752. The 0.998 gate is defined on the randn pair, where the float path is unchanged.
  • The served-path e2e uses Tensix dispatch with 2 CQs. On the Galaxy that is 12x10, like ETH dispatch. On a real p150 that mode has 11x10 unless ETH dispatch supports 2 CQs there; the Galaxy ETH path has only 2 idle ETH cores, so it is 1 CQ. The like-for-like eth12 1-CQ float number is 44.57 / 44.06 ms.
  • The host is shared by about 11 agents, and e2e medians moved by up to ±0.5 ms between runs in this round. The same-session A/B pairs are in the step table.

vs RTX 5090 (GPU_COMPARISON.md; the GPU "incl_h2d" row = upload of both views + forward + readback of both maps)

GPU variant GPU ms TT served path now (worker 2 CQ, uint8) TT eth12 1 CQ float now
fp32 strict 103.4 42.0 (TT 2.5x faster) 44.6
tf32 70.8 42.0 (1.7x) 44.6
bf16 autocast 63.8 42.0 (1.5x) 44.6
fp16 autocast 56.6 42.0 (1.35x) 44.6
module converted to bf16 42.3 42.0 (TT 0.7% faster; real pair 42.03 / 41.54 min) 44.6 (GPU 5% faster)
module bf16 + SDPA 35.3 42.0 (GPU 1.19x faster)
bf16 module + torch.compile + CUDA graphs 21.1 42.0 (GPU 2x faster)
device-only (GPU excl_h2d bf16 module 41.3) 41.3 device 39.8 (TT faster) 39.6

The served-path TT number uses the server's uint8 views (1.5 MB upload) and readbacks that overlap compute on a second CQ. The GPU row uploads fp32 views; a GPU could also take uint8, which would save it about 0.5 ms of its 1.1 ms of transfers. So the bf16-module win is narrow (0.3 ms median), and it does not hold for the 1-CQ ETH configuration. Pose requests (return_pose) take 60.4 ms with the sym B-mode, against 2 x 42.3 = 84.6 ms for the GPU's two forwards. This is a mode-specific algorithmic saving (the encoder runs once and an unused head is skipped), so it is reported separately.

Round-2 steps (each kept step committed; A/B in one session where host noise matters)

commit knob change effect
72b75eb u8 (with hostcol) uint8 pixel views. Host im2col is uint8 (1.5 MB instead of 3 MB). The device does typecast (exact) and v - 127.5 (exact in bf16). W/127.5 is folded into the patch-embed weight in fp64. The server preprocess returns uint8 HWC views (preprocess_image(..., uint8=True)), and warm-up uses a uint8 dummy H2D 1.04 -> 0.58, host prep 0.94 -> 0.76 ms
72b75eb MAST3R_CQS=2 forward captured as 2 trace segments (A: encoder, decoder, head 1; B: head 2). An event after A, then a CQ1 read of head 1 while CQ0 runs B e2e 45.06 -> 43.95 (float) / 43.05 (u8); outputs bit-identical
167a332 u8 (py, px, c) im2col row order for uint8 (48-byte runs from the HWC array; weight rows permuted to match) host_input 0.40 -> 0.18 ms, e2e 43.05 -> 42.68
3d850ff phase 0/1 permutation matmuls at HiFi3. This is exact: the data operand sits in srcB and is fully covered by HiFi3, as probed with perm_mm_probe.py. Phase0 strip inputs are untilized into L1 bit-identical, trace 39.96 -> 39.88
2fa1785 ropes RoPE trans_mat HEIGHT_SHARDED, one tile per core on all 120 cores (per-pass L1 copy), which selects the prefill-sharded rotary_embedding_llama factory. Probe: 26.7 -> 21.3 us (enc), 22.5 -> 18.0 us (dec) bit-identical, trace 39.88 -> 39.49 ms (RoPE total 1.96 -> 1.61 ms)
440121c sym (B-mode) return_pose: one symmetric graph. The encoder runs once (it is per-view, so pass-2 inputs are the pass-1 encoder outputs), the decoder runs for both orders, then DPT head 1 + head 2 of (1,2) and head 1 of (2,1). Head 2 of (2,1) is skipped because PairViewer never reads it. Captured at server warm-up out_ii / out_ji / out_jj bit-identical to the two-pass path; pose e2e 84.1 -> 60.4 ms; served forward 86 -> 60 ms (smoke test PASS, same pose / focals)
d087a52, f6cbac0 tsplit (needs 2 CQ) 4 segments: enc+dec+head-1 main rows, then head-2 main rows, then head-2 ring strips, then head-1 ring strips. Each 2 MB main-row readback overlaps later compute; only head 1's 128 KB strips are read after the last trace. Host assembly is incremental bit-identical; e2e 42.49 / 42.58 -> 42.08 / 42.03 ms (same-session A/B)
f6a7dc5 serve MAST3R_CQS=2 pinned in tt-model.yaml and SERVING.md server smoke test PASS (forward 42-43 ms, pose 60 ms)

Round-2 tried and rejected

  • RoPE math fidelity HiFi2 / HiFi3: no trace gain (RoPE is data-movement bound) and the outputs change. Reverted.
  • cos/sin HEIGHT_SHARDED (with or without trans_mat), and RoPE with heads folded into batch: slower (rope_probe.py).
  • Decoder fc1 with transposed-mcast 11x10 (110 cores) and encoder fc1 transposed: slower (99 vs 80 us, 214 vs 198 us; fc1_probe.py).
  • MinimalMatmul K_block 2 for the decoder M=1024 linears and encoder proj (mm_sweep.py: 2-4 us faster per op in isolation): no trace gain in the model (39.58 vs 39.49 ms), synthetic head2 PCC 0.99862 vs 0.99870. Reverted.
  • DPT 1x1 linears with explicit 1-D/2-D configs (lin1x1_probe.py): at most -28 us (128^2) and the outputs change; not taken.
  • 128^2 refinenet relu/add in L1: conv circular buffers clash with the L1 tensors (TT_THROW, no hang). Reverted.
  • Preallocated host buffers for D2H (copy_device_to_host_tensor): it does not write through to the torch buffer; net -0.03 ms. Not taken.
  • glibc malloc tunables against page faults in the host assembly: no change.

Method (identical for every number below)

ROOT=/home/ttuser/experiments/tt-models; M=$ROOT/models/mast3r-p150; cd $M/code
source $ROOT/tools/chipenv.sh 9 $M/.venv >/dev/null
export TT_WEIGHTS_REVISION=61c57447d7b0adc8a1a30b2b0adec7a8935aa2a3
timeout -s INT 900 python bench_breakdown.py --mode eth12          # synthetic pair: PCC + per-phase timing (30 iters)
MAST3R_ETH=1 timeout -s INT 900 python tools_prof/real_pair_acc.py   # real pair accuracy gate
timeout -s INT 900 python test_mast3r.py --layer end_to_end --runs 10   # authors' gate (Tensix dispatch, 12x10)
MAST3R_TRACE=0 python -m tracy -r -p -v -o $TT_METAL_PROFILER_DIR/x --op-support-count 6000 tools_prof/prof_eager.py  # op profile

trace = execute_trace + synchronize_device of the whole forward (pure device time, one trace per forward); e2e = the public model(img1, img2) call (host prep + H2D + trace + D2H incl. host output assembly). The host is shared by ~11 agents, so host-side phases (host_prep, D2H, e2e) are noisy (D2H 2.4-4.1 ms run to run); the device trace is stable to about +-0.2 ms. Median and min are reported.

Accuracy gates (from OPT_BASELINE.md, kept): synthetic e2e PCC >= 0.998; real pair pts3d PCC >= 0.989 per head; conf PCC >= 0.99.

Round 1 results (chip 9, eth12, synthetic pair)

Final numbers measured 2026-10-03 02:45-02:52 UTC on HEAD f6b97f8 (two default runs, one MAST3R_OPT=none run in the same session).

baseline (MAST3R_OPT=none) round start (8dc072d) now (f6b97f8)
trace (device forward) median / min 62.20 / 62.10 ms 46.47 / 46.23 ms 40.08 / 39.73 ms (2nd run 40.08 / 39.82)
host prep 0.83 0.84 0.88-0.98
H2D 1.06 1.03 1.03
D2H (incl. host output assembly) 28.00 (min 20.49) 2.49 3.13-3.49 (min 2.62)
e2e median / min 87.26 / 84.75 ms 50.48 / 50.13 ms 44.97-45.64 / 44.55 ms
test_mast3r end_to_end best-of-10 (Tensix dispatch 12x10) 84.49 ms — 44.87 ms
p150-equivalent (--mode worker11, Tensix dispatch, 11x10) trace / e2e 62.62 / 89.16 — 41.41 / 47.07 ms
device programs per forward (eager profile) 1252 1212 1230
kernel-time sum (eager profile) 60.97 ms 44.91 ms 38.02 ms
synthetic PCC head1 / head2 (all) 0.99836 / 0.99895 (0.99909) 0.99845 / 0.99871 0.99846 / 0.99870 (0.99890)
test_mast3r end_to_end PCC (4 digits) 0.9984 0.9984 0.9984 (head1 0.9984 / head2 0.9980)
real pair pts3d PCC head1 / head2 0.98988 / 0.99076 0.99044 / 0.99079 0.99046 / 0.99081
real pair conf PCC head1 / head2 0.99237 / 0.99439 0.99178 / 0.99457 0.99182 / 0.99459
real pair raw PCC vs half-pixel-upsample fp32 ref — 0.999839 / 0.999847 0.999837 / 0.999849
real pair median abs(dz)/z vs fp32 ref (vs half-pixel ref) 1.3% 1.05% / 1.11% 1.13% / 1.18% (0.37% / 0.33%)

All accuracy gates pass. Honest notes: the synthetic head2 PCC is 0.99870 vs 0.99895 at the baseline (head1 is higher, 0.99846 vs 0.99836); that shift came with the earlier sessions' gelut/lnfold steps, not this round. The real-pair median relative depth error moved 1.05/1.11% -> 1.13/1.18% within this round (phase/phase0 change the head arithmetic: folded kernels, no bf16 rounding of the upsampled activations); the PCCs are equal or higher.

Device: 62.2 -> 40.1 ms (-36%); this round 46.5 -> 40.1 ms (-14%). e2e: ~85-87 -> ~45 ms (-48%). Device time by stage now (eager profile, kernel sums): patch-embed + encoder 16.4 ms, decoder (incl. embed) 11.9 ms, DPT x2 9.8 ms; dispatch gaps ~2.0 ms (trace - kernel sum, ~1.6 us per program).

vs RTX 5090 (GPU_COMPARISON.md, same pair size, forward incl. H2D + readback of both maps)

GPU variant GPU ms TT now (e2e incl. H2D/D2H) ratio TT/GPU
reference as shipped, fp32 strict 103.4 45.0 0.44 (TT faster)
tf32 70.8 45.0 0.64 (TT faster)
bf16 autocast 63.8 45.0 0.71 (TT faster)
fp16 autocast 56.6 45.0 0.80 (TT faster)
module converted to bf16 42.3 45.0 1.06 (GPU faster)
module bf16 + SDPA 35.3 45.0 1.27 (GPU faster)
bf16 module + torch.compile + CUDA graphs 21.1 45.0 2.13 (GPU faster)

The TT device forward alone (40.1 ms) is below the bf16-module GPU number but the x1-PCIe Galaxy transfers (H2D 1.0 + D2H ~3 ms incl. host assembly) put e2e just above it. The compiled GPU graph (21 ms) is out of reach without lower precision: LoFi matmuls were tried again this round per group (below) and fail the accuracy gates. Served-like e2e (PNG decode + preprocess + forward + npz encode, ~180 ms of identical host work on both sides) was not re-measured; it is host-bound on either accelerator (GPU 241-301 ms).

Kept steps (cumulative, each measured with the method above)

Earlier sessions (before this round; numbers from their commits):

commit knob change trace
0e84074 out compact channel-first head output, D2H 29 -> 3 ms (e2e 89.7 -> 67.2)
c64edc6 mm per-shape 12x10 matmul configs, exact GELU fused into fc1 62.5 -> 57.4
3fe76db dpt refinenet out_conv before the upsample, RM conv in/out 57.4 -> 55.9
4401161 l1 per-block transient activations in L1 55.9 -> 51.8
772e075 dpt height-sharded refinenet convs >= 64x64 51.8 -> 50.3
13363fc/30c038f dptf HiFi2 (fp32 acc) DPT convs except the two head convs 50.3 -> 49.3
f7587b7 gelut tanh-form GELU in the fc1 epilogue 49.3 -> 47.9
316339b lnfold LayerNorm affine folded into the following linear 47.9 -> 47.1
8dc072d l1 per-pass L1 copies of the RoPE tables 47.1 -> 46.65

This round (re-validated HEAD 8dc072d first: trace 46.47, PCC 0.99845/0.99871, real pair gates pass):

commit knob change trace notes
206768d phase Polyphase DPT head tail. head.1-2 (bilinear x2, half-pixel, clamped == ttnn.upsample) + 3x3 conv == 4 phase 3x3 convs on the 256^2 grid with kernels K_ab = sum A[a] A[b] K (fp64 fold on host). Exact except the 2-px border ring, which is recomputed with the original upsample+conv on 2-row/2-col strips. The 512^2 x 128 upsample, its reshards and the 6-way DRAM-sliced conv are gone; head.4 (1x1, 128->4) is 4 tiny height-sharded linears summed. 46.47 -> 42.67 tail probe 3.21 -> 1.29 ms per head. D2H went up (host phase interleave)
ee66e6b phase phase interleave on device (RM permute + one exact 0/1 permutation matmul) -> NCHW rows; host only writes the ring 42.67 -> 43.20 D2H 4.5 -> ~2.7 ms, e2e -0.4..-2 ms
683d128 hostcol host cast writes the patch-embed im2col layout [2, N, 768] bf16 directly (one strided copy, 0.14 ms, replaces cat+cast 0.97 ms); device drops reshape/permute im2col (0.7 ms) 43.20 -> 42.46 same 3 MB H2D
82bbe3e phase full-grid 2-D mcast config for the interleave matmul (was 16 cores, 123 us) 42.42 -> 42.22
c141100 lnfold decoder: norm1(x)==norm_y(x) once the affine is folded, so each decoder step normalises x1 and x2 once for both branch blocks (-24 LN) 42.22 -> 41.93 bit-identical
07250e9 phase0 head.0 polyphase too: refinenet1's upsample folded into 4 phase convs at 128^2 (256->128); phases interleaved with an exact batched 0/1 matmul; final 6-px ring recomputed exactly by a thin-strip run of the original op sequence 41.93 -> 41.51 removes the 128->256 upsample + head.0's 3-way DRAM-sliced conv
ed59a76 phase permutation matmuls at HiFi4 without fp32 acc (bit-exact; LoFi/HiFi2 are not, tools_prof/perm_mm_probe.py) 41.51 -> 41.45
8c74544 decb decoder: the two branch blocks of a step share one B=2 split-heads / 2xRoPE / SDPA / concat-heads per attention (these ops cost ~the same at B=2 as at B=1: 28/21/12 vs 26/19/11 us); per-branch linears stay separate (a per-branch-weight bmm is 3-7x slower), outputs concatenated / sliced on batch 41.45 -> 40.25 same PCC
f6b97f8 resl1 encoder residual stream in L1 (patch-embed output copied to L1; residual-matmul outputs and their ones vector in L1; enc_norm output in DRAM) 40.25 -> 40.02 encoder LN 25.5 -> 20.0 us; same PCC

Exactness of the polyphase fold: host check (fp64) interior max abs diff 2e-8; device output vs fp32 torch: PCC 0.999993 for both the old and the new tail (tools_prof/phase_probe.py); real-pair raw PCC vs the half-pixel reference unchanged.

Tried and rejected (round 1)

  • Row-sliced in-L1 tail (upsample + conv + 1x1 per row slice): correct but slower (3.5 vs 3.2 ms per head; the upsample itself was the cost).
  • DPT resconv chain kept height-sharded in L1 (dpts): conv CBs of the 128^2x256 height-sharded convs (~1.2 MB, full weights per core) clash with any extra L1 tensor. Reverted.
  • DPT relu/add in L1 interleaved for <= 64^2: -0.3 ms but changed PCC (conv picked a different config); not worth the accuracy churn. Reverted.
  • Feeding the sharded upsample output straight into head.0's conv: CBs exceed L1 (2.4 MB). Superseded by phase0.
  • Strip convs on fewer cores (16/32): no gain or CB clash.
  • SDPA sweep on 12x10 (tools_prof/sdpa_probe2.py): q128/k256 HiFi2 (current) is the fastest; LoFi -2%, exp-approx 0%, fp32-acc +16%.
  • LoFi for selected transformer matmul groups (experiment, not committed): decoder MLP only -0.25 ms but synthetic head1 PCC 0.99679; encoder MLP -1.0 ms, head1 0.99327; both MLPs -1.1 ms, head1 0.98269. All fail the 0.998 gate.
  • Explicit minimal_matmul patch embed with L1 output (instead of linear + copy): bf16 packer accumulation over 6 K blocks changed outputs (head2 0.99852); with fp32 acc the PCC rose (0.99853/0.99889) but the real-pair depth error moved (1.29%/1.22%). Kept the bit-identical linear + copy.
  • reallocate_halo_output=False for the DPT convs (drop the Move ops): no gain.
  • GELU: the fused GELU_TANH epilogue costs ~100 us (enc fc1) / ~40 us (dec fc1) per call (3.4 ms per forward); no cheaper accurate SFPU variant exists in this tt-metal (the approximate GELU was rejected earlier for accuracy, PCC 0.993).

Megakernel / parameter taxonomy status

Round 2 changes:

  • A (launch-fixed): MAST3R_CQS = 2 (trace segmentation + CQ1 readbacks), and the input dtype (uint8 in the server) is fixed at launch through hostcol.
  • B (return_pose): now a precompiled second variant, the sym graph with 3-4 segments, captured at warm-up and selected per request. The other per-request fields remain C (pixel contents) or D (host only).
  • Per forward: one H2D (1.5 MB uint8); trace segments replay back to back; each output is read on CQ1 as soon as its segment's event fires.

Round 1 text:

Per param_taxonomy/mast3r-p150.md: no request parameter changes device shapes (A: resolution 512, 1 pair, weights; C: pixel contents; B: return_pose = run the same trace twice). The whole forward is one metal trace with one H2D (3 MB im2col input) and two D2H (2 x 2.2 MB: NCHW maps + ring strips) per pair. Default knob set frozen as MAST3R_OPT=all (pinned in tt-model.yaml serve.env, documented in SERVING.md).

Remaining backlog (estimated device gain)

Round-2 view. Device time is now 39.6-39.8 ms; kernel sum 37.6 ms over 1232 programs.

  • Phase0 ring strips: about 0.48 ms per head (two strip runs of about 25 small ops each). A cheaper exact border formulation, or asymmetric-padding convs that compute only the needed rows, could save about 0.3 ms in total.
  • Tail phase-interleave chain: about 0.21 ms per head, incl. a 65 us RM reshape with a page-size change. An alternative layout for the 4 phase linears could remove the reshape and permute, about -0.2 ms in total.
  • act_postprocess 1x1 + ConvTranspose fold (ap0 / ap1): about 190 us per head now. The fold is exact, but it needs a depth-to-space data movement (RM page-size change), so the net is uncertain, perhaps -0.2 ms.
  • Decoder per-branch plumbing (concat 31 us/step, batch slices 14 us/step): blocked by per-branch weights. The bmm with per-branch weights is 3-7x slower.
  • GELU_TANH epilogue (about 2.7 ms), SDPA (5.2 ms) and create-heads (1.4 ms) are at the stock-kernel floor at the gated precision.
  • sym B-mode: branch-2 step 11 and dec_norm of the swapped pass are still computed (dead work, about 0.5 ms per pose request).

Round 1 list:

  • Phase0 thin-strip pipeline: ~0.4 ms per head of small ops (halo/conv on tiny inputs); could be cut with a cheaper exact ring formulation.
  • DPT refinenet plumbing at 64^2/128^2 (S2I/I2S per conv, DRAM relu/add): ~0.5 ms per head, blocked by conv L1 CB size.
  • Dispatch gaps ~2 ms over ~1230 programs: fewer/fused small ops (DPT conv plumbing is ~4-5 programs per conv).
  • GELU epilogue (3.4 ms) and SDPA (5.2 ms) are at the floor of the stock kernels at the gated precision; further device gains there need custom kernels (not done: tt-metal is frozen for this workflow).
  • Decoder residual stream in L1: probe says ~1 us per LN, not worth it.
  • models/tests/test_fused_host.py: 5 mock-graph tests fail since the earlier OPT commits (fake_ttnn lacks to_memory_config etc.); pre-existing, unchanged by this round.