File size: 116,272 Bytes
00e0473
 
2c79427
 
 
 
 
b15da4f
 
 
 
00e0473
 
 
 
 
 
 
 
 
 
 
 
 
 
 
b15da4f
2c79427
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
00e0473
b15da4f
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
00e0473
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
522
523
524
525
526
527
528
529
530
531
532
533
534
535
536
537
538
539
540
541
542
543
544
545
546
547
548
549
550
551
552
553
554
555
556
557
558
559
560
561
562
563
564
565
566
567
568
569
570
571
572
573
574
575
576
577
578
579
580
581
582
583
584
585
586
587
588
589
590
591
592
593
594
595
596
597
598
599
600
601
602
603
604
605
606
607
608
609
610
611
612
613
614
615
616
617
618
619
620
621
622
623
624
625
626
627
628
629
630
631
632
633
634
635
636
637
638
639
640
641
642
643
644
645
646
647
648
649
650
651
652
653
654
655
656
657
658
659
660
661
662
663
664
665
666
667
668
669
670
671
672
673
674
675
676
677
678
679
680
681
682
683
684
685
686
687
688
689
690
691
692
693
694
695
696
697
698
699
700
701
702
703
704
705
706
707
708
709
710
711
712
713
714
715
716
717
718
719
720
721
722
723
724
725
726
727
728
729
730
731
732
733
734
735
736
737
738
739
740
741
742
743
744
745
746
747
748
749
750
751
752
753
754
755
756
757
758
759
760
761
762
763
764
765
766
767
768
769
770
771
772
773
774
775
776
777
778
779
780
781
782
783
784
785
786
787
788
789
790
791
792
793
794
795
796
797
798
799
800
801
802
803
804
805
806
807
808
809
810
811
812
813
814
815
816
817
818
819
820
821
822
823
824
825
826
827
828
829
830
831
832
833
834
835
836
837
838
839
840
841
842
843
844
845
846
847
848
849
850
851
852
853
854
855
856
857
858
859
860
861
862
863
864
865
866
867
868
869
870
871
872
873
874
875
876
877
878
879
880
881
882
883
884
885
886
887
888
889
890
891
892
893
894
895
896
897
898
899
900
901
902
903
904
905
906
907
908
909
910
911
912
913
914
915
916
917
918
919
920
921
922
923
924
925
926
927
928
929
930
931
932
933
934
935
936
937
938
939
940
941
942
943
944
945
946
947
948
949
950
951
952
953
954
955
956
957
958
959
960
961
962
963
964
965
966
967
968
969
970
971
972
973
974
975
976
977
978
979
980
981
982
983
984
985
986
987
988
989
990
991
992
993
994
995
996
997
998
999
1000
1001
1002
1003
1004
1005
1006
1007
1008
1009
1010
1011
1012
1013
1014
1015
1016
1017
1018
1019
1020
1021
1022
1023
1024
1025
1026
1027
1028
1029
1030
1031
1032
1033
1034
1035
1036
1037
1038
1039
1040
1041
1042
1043
1044
1045
1046
1047
1048
1049
1050
1051
1052
1053
1054
1055
1056
1057
1058
1059
1060
1061
1062
1063
1064
1065
1066
1067
1068
1069
1070
1071
1072
1073
1074
1075
1076
1077
1078
1079
1080
1081
1082
1083
1084
1085
1086
1087
1088
1089
1090
1091
1092
1093
1094
1095
1096
1097
1098
1099
1100
1101
1102
1103
1104
1105
1106
1107
1108
1109
1110
1111
1112
1113
1114
1115
1116
1117
1118
1119
1120
1121
1122
1123
1124
1125
1126
1127
1128
1129
1130
1131
1132
1133
1134
1135
1136
1137
1138
1139
1140
1141
1142
1143
1144
1145
1146
1147
1148
1149
1150
1151
1152
1153
1154
1155
1156
1157
1158
1159
1160
1161
1162
1163
1164
1165
1166
1167
1168
1169
1170
1171
1172
1173
1174
1175
1176
1177
1178
1179
1180
1181
1182
1183
1184
1185
1186
1187
1188
1189
1190
1191
1192
1193
1194
1195
1196
1197
1198
1199
1200
1201
1202
1203
1204
1205
1206
1207
1208
1209
1210
1211
1212
1213
1214
1215
1216
1217
1218
1219
1220
1221
1222
1223
1224
1225
1226
1227
1228
1229
1230
1231
1232
1233
1234
1235
1236
1237
1238
1239
1240
1241
1242
1243
1244
1245
1246
1247
1248
1249
1250
1251
1252
1253
1254
1255
1256
1257
# mast3r-p150 — 12x10 optimisation report (Galaxy BH chips 9, 3, 18)

**Status (2026-10-05): p150 ETH-dispatch compliance done (section "p150 ETH-dispatch compliance 2026-10-05"): the
server, the harnesses and the Python API default to ETH dispatch + 1 CQ + 12x10; outputs bit-identical to the old
Tensix 2-CQ served default; served forward 25.6 -> 26.8 ms (lost CQ1 readback overlap), all gates pass.**

Previous status (2026-10-04): audit integration done (section "Audit integration 2026-10-04"): served npz request
-72 % server time / -60 % client wall (stored npz), png -34 %, new binary route /predict_npz (-70 % client wall),
and LoFi-RNE MLP linears (served e2e -7.2 %, all gates pass, two ungated metrics slightly worse, disclosed).**

Previous status: stopped: round 9 finished (this workflow's optimisation round 2); the backlog below is still open.**
Round 9 kept three small bit-identical plumbing steps (esplit, demb, hrqk): traced device span 23.531 -> 23.487 ms
(-0.044 ms, -0.19 %, chip 9, same session). That is below the host-wall noise; served u8 e2e is unchanged within noise
(27.7-28.4 ms). The two planned big items were measured first and turned out to be capped: the 1.3-2.2 ms "host /
dispatch gap" is the chip's AICLK throttling under load (not host or dispatch), and a custom SDPA compute kernel alone
can save at most ~10 us per encoder call. All gates pass with identical numbers (every kept step is bit-identical on the
eth12 float and the served u8 paths); the server smoke test passed 4 times out of 4. No chip faulted.

Model: DUSt3R ViT-L/16 + dual-branch decoder + 2 DPT heads (`naver/DUSt3R_ViTLarge_BaseDecoder_512_dpt` @ `61c57447`),
one 512x512 image pair per request, batch 1. tt-metal `8b98410e730` + ETH-dispatch patch (unmodified; every new kernel
of round 3 is model-local and runs through `ttnn.generic_op`).
Grid: 12x10 (`TT_METAL_CORE_GRID_OVERRIDE_TODEPRECATE=11,9`), ETH dispatch, 1 CQ (= p150 + ETH dispatch).
Baseline and its profile: `OPT_BASELINE.md`. All knobs are `MAST3R_OPT` entries in `code/models/demos/mast3r/tt/fused.py`
(`all` = default and pinned in `tt-model.yaml`; `none` = the pre-optimisation fused graph).
Default set now: out, mm, dpt, l1, dptf, gelut, lnfold, phase, hostcol, phase0, decb, resl1, ropes, sym, tsplit,
sdpa2, hrope, hcat, mln, mc2d, mc2dr, resl1d, addln, cdbl, pil, til, mm32, dmm, dadd, addln2, rchain, dlin, dfront, dshard,
ups1, tailf, pemm, pcat, p1hs, hf3, gpoly, qkx, sdbl, pilhs, tups, sgat, tapm, obs, obsd, resmm, resmmd, esplit, demb, hrqk, lofie, lofid
(gelua and lofi are off; lofie / lofid since the 2026-10-04 audit integration, below). Since 2026-10-05 the serve
config pins `MAST3R_DISPATCH=auto` + `MAST3R_CQS=1` (ETH dispatch, 1 CQ, 12x10 = the p150 configuration; section
"p150 ETH-dispatch compliance 2026-10-05"). Before that it pinned `MAST3R_CQS=2` (split trace, head-1 readback on CQ1),
which needs Tensix dispatch: 12x10 on the Galaxy, 11x10 on a p150. `MAST3R_CQS=2` stays as a Galaxy-only opt-in.
With `hostcol` the server uploads the views as uint8.

## p150 ETH-dispatch compliance 2026-10-05 (chip 16, start = d02a5c2, final = this commit)

Requirement: on a single p150 the model must use the 12x10 grid, so dispatch must run on ETH cores (stock Tensix
dispatch leaves 11x10). ETH dispatch in the patched tt-metal has 1 CQ. On this Galaxy, Tensix dispatch still gives
12x10 (13 Tensix columns), so numbers measured with it do not represent a p150. p150-equivalent here = ETH dispatch,
`num_command_queues=1`, grid capped at 12x10. (The first attempt was cut short by host crash #5; its WIP commit
2cf38a7 was reviewed, its host test fixed and every number below measured again after the reboot.)

### Audit (default of every device-open path)

| path | before (d02a5c2) | after |
|---|---|---|
| Python API `Mast3rP150.from_pretrained` (`mast3r_p150/device.py`) | ETH, 1 CQ, 12x10 (`dispatch="auto"`; `MAST3R_CQS` unset = 1) | unchanged; `MAST3R_CQS=2` now warns (Galaxy-only, 11x10 on p150) |
| HTTP server `models/server/app.py` | `ttnn.open_device(**open_kwargs)` = Tensix dispatch; CQs from `MAST3R_CQS` (code default 1) | `mast3r_p150.device.open_device`: ETH, 1 CQ, 12x10; `/info` -> `device_config` |
| `tt-model.yaml` `serve.env` (+ SERVING.md) | `MAST3R_CQS: "2"` -> **Tensix dispatch, 2 CQ**, 12x10 on the Galaxy / 11x10 on a p150 | `MAST3R_DISPATCH: "auto"`, `MAST3R_CQS: "1"` -> ETH, 1 CQ, 12x10 |
| `test_mast3r.py` (card PCC + latency row) | Tensix; validation ran it with `MAST3R_CQS=2` | shared opener: ETH, 1 CQ, 12x10 |
| `eval_mast3r.py`, `make_demo.py --local`, `eval_eth3d.py` | Tensix, 1 CQ (eval_eth3d: plain open) | shared opener: ETH, 1 CQ, 12x10 |
| `bench_breakdown.py` (OPT_REPORT trace / e2e rows) | default `--mode worker11` (Tensix, 11x10); eth12 rows used `--mode eth12`; served rows `MAST3R_CQS=2 --mode worker` | default `--mode eth12` (ETH, 1 CQ, 12x10); worker modes stay as explicit A/B |
| `tools_prof/real_pair_acc.py` (real-pair gate) | Tensix unless `MAST3R_ETH=1` | ETH unless `MAST3R_ETH=0` |
| `tools_prof/sym_check.py` (pose path) | Tensix; validation ran it with `MAST3R_CQS=2` | shared opener: ETH, 1 CQ |
| `test_api_device.py`, `test_warmup_device.py`, `first_call_bench.py` | via `from_pretrained`: ETH, 1 CQ, 12x10 | unchanged |
| `profile_eager.py` | default `--mode eth12` | unchanged |

1-CQ path: the model already had one (`MAST3R_CQS=1`, one trace on CQ0, both heads read on CQ0 after it); only the
served default used CQ1. No new device code. Fallback: without the patch (marker in `tt_metal/impl/dispatch/topology.cpp`)
`auto` uses Tensix dispatch with a `RuntimeWarning` + log line naming the 11x10 grid; an ETH open that raises also
falls back with a warning (unless `MAST3R_DISPATCH=eth`). `MAST3R_DISPATCH=eth` with `MAST3R_CQS=2` is an error.

### Bit identity (ETH 1 CQ vs the previous served default, Tensix 2 CQ; same chip, same session)

- `bench_breakdown.py --dump` (ETH) / `--cmp` (`MAST3R_CQS=2 --mode worker`): float randn pair and uint8 pair, both
  heads bit-identical (n_diff = 0).
- HTTP: two ETH servers vs one Tensix 2-CQ server: npz arrays (`/predict` and `/predict_npz`) and PNG bytes identical.
- `test_api_device.py` (ETH): `model(...)` == the server pipeline bit for bit, with and without pose.

### Accuracy (`code/tools_prof/eth_validate.sh 16`; all in ETH, 1 CQ, 12x10; gate code unchanged; all gates pass)

| metric | published (audit integration 2026-10-04) | now (ETH, 1 CQ) | gate |
|---|---|---|---|
| randn head1 / head2 (all), float | 0.99862 / 0.99869 (0.99891) | 0.99862 / 0.99869 (0.99891) | 0.998 |
| `test_mast3r.py --layer end_to_end` | 0.9985 PASS (2 CQ) | 0.9985 PASS (head1 0.9985, head2 0.9982) | 0.99 |
| real pair float pts3d h1 / h2, conf h1 / h2 | 0.99016 / 0.99088, 0.99269 / 0.99486 | same digits | 0.989 / 0.99 |
| served-path u8 real pair pts3d h1 / h2, conf h1 / h2 | 0.99015 / 0.99097, 0.99246 / 0.99519 (2 CQ) | same digits | 0.989 / 0.99 |
| sym out_ii / out_ji / out_jj vs two-pass | bit-identical | bit-identical | identical |
| synthetic u8 head1 / head2 (not gated) | 0.99593 / 0.99927 | 0.99593 / 0.99927 | |

### Performance before / after (chip 16; medians, min in brackets)

| metric (method) | published | p150 config now (ETH, 1 CQ, 12x10) | Tensix 2 CQ, same session (Galaxy-only) |
|---|---|---|---|
| trace, float (`bench_breakdown`, 30 it) | 23.32 ms eth12 | 23.42 (23.07) | 23.67 (23.15) |
| e2e `model(img1, img2)`, float | 28.53 eth12 | 28.38 (27.86) | 27.55 (26.89) |
| e2e, uint8 views (served input path) | 26.27 / 26.36 (worker 2 CQ) | 27.83 (27.37) | 26.64 (26.36) |
| `test_mast3r.py --layer end_to_end --runs 25` | 27.51 ms (2 CQ, round 9) | 27.77 ms | |
| pose: single pair / `sym` / two-pass e2e (`sym_check`) | 26.40 / 36.98 / 52.48 (2 CQ) | 27.22 / 38.89 / 54.90 | |
| server `timing_ms.forward`, npz (2 x 30 req) | 25.5 / 25.57 (2 CQ) | 26.55-26.98 (2 servers x 2 runs) | 25.63 |
| `/predict` npz total / client wall | 81.5-83.3 / 130-140 | 80.3-82.5 / 123-136 | 79.4-80.0 / 127-130 |
| png total (2 x 15 req) | 124.5-139.6 | 131.8-143.7 | 123.9-124.7 |
| `/predict_npz` total / client wall | 72-76 / 96.8-102.0 | 72.7-74.0 / 97.5-104.5 | 71.8-72.0 / 96.8-100.4 |
| Python API device call `timing_ms['forward']` (30 calls) | 26.7 (26.4) | 26.99 (26.54) | |
| `model(PIL, PIL)` / `model(path, path)` whole call | 41.9 / 57.1 | 46.75 / 58.65 | |
| `predict_pairs` per pair | 27.9 | 28.09 | |
| startup (`from_pretrained`, warm cache) | ~9 s | 7.5 s (device warm-up 5.8 s, host 1.1 s) | |
| first call / warm same input (`test_warmup_device`): pair, pose, batch | 0.92x, 1.08x, 0.99x | 0.97x, 1.09x, 1.00x (gate 10 % + 5 ms: pass) | |
| smoke `--require-pose --npz-route` | PASS | PASS x2 | PASS |

Reading: on the device the configurations are the same (trace 23.4 vs 23.7 ms, within noise). The p150 config loses
the CQ1 overlap of the head-1 readback: the served forward is about 1.1 ms (4 %) slower than the Galaxy-only 2-CQ
mode, and the uint8 e2e about 1.2 ms. Request totals move by 1-3 ms, inside the host noise. The 2-CQ numbers are not
reachable on a p150 at 12x10, so the ETH column is the honest p150 number. Logs:
`<scratchpad>/mast3r-p150-eth-reval/` (`val/*.log`, `http/ab_*.txt`, `api.json`, `warmup.json`).

### Incident during this task

Before the chip was claimed, a host-test command (`pytest models/tests`) also collected the two device tests
and ran them for about 2 minutes (04:53-04:56 UTC) without `chipenv.sh`, so with every Galaxy chip visible
(device 0 = UMD chip 0, not in the pool). The run reached its summary (62 passed, 2 failed) and was interrupted
(SIGINT) while it closed the device. A chip-2 hang reported
by another agent at 04:55 overlaps this window; a link was not proven. Host tests are now run with
`TT_VISIBLE_DEVICES=none` and an explicit file list.

## Audit integration 2026-10-04 (chip 1, start = ad39150, final = this commit)

Source: `audit/mast3r-p150/AUDIT_REPORT.md` (verdict REOPEN) and the audit branch `audit/mast3r-p150`
(9bfbb82, d2e7aee). The user decided to integrate every demonstrated item of kind A (host / serving, outputs and
/predict unchanged), B (additive binary route) and C (precision change that passes every gate, default on, env
switch back). Each item is re-implemented on main and measured again on chip 1 (Galaxy BH, PCIe x1, 12x10 grid).
Logs and scripts: `audit/mast3r-p150/integrate/` (`lofi_ab.sh`, `http_ab.sh`, `http_seq.sh`, `r3val_new.log`,
`ab_*.txt`, `smoke_*.txt`, `server_*.log`).

### Integrated

| commit | item | change | measured on chip 1 (A/B against the pre-item state) |
|---|---|---|---|
| 4c9303a | C: `lofie`, `lofid` | Encoder and decoder MLP fc1 / fc2 at LoFi; the same knobs upload those weights rounded to nearest-even at 4 explicit mantissa bits (LoFi reads the weights as hidden bit + 4 MSBs, so the device multiplies the rounded weights exactly; truncation alone fails the gate, 0.99364). Ported from the audit env knobs `MAST3R_LOFI_MLP=both MAST3R_LOFI_RNE=1` as two `MAST3R_OPT` entries in `all`: the weight caches key on cfg, so a knob change re-uploads the weights, and `MAST3R_OPT=none` is unchanged (none_eth12 PCC identical to ad39150). Switch back: `MAST3R_OPT=all,-lofie,-lofid` (served outputs then bit-identical to ad39150, checked over HTTP) | Served path (worker, 2 CQ, uint8, 60 iterations, A/B/A/B): trace 25.45 / 25.89 -> 23.62 / 23.56 ms (-7.7 %); e2e 28.29 / 28.45 -> 26.27 / 26.36 ms (**-2.06 ms, -7.2 %**). eth12 (ETH dispatch, 1 CQ, float, 30 iterations): trace 25.39 / 25.38 -> 23.32 / 23.24 ms (-8.3 %); e2e 30.01 / 30.98 -> 28.53 / 27.98 ms (-7.3 %). Server `forward` 27.0 -> 25.5 ms. Pose (sym_check, same run): two-pass 52.48, symmetric 36.98, single pair 26.40 ms |
| 1138587 | A: stored npz | `/predict` writes the npz with `np.savez` (stored) instead of `np.savez_compressed`. Request field `compress_npz` (optional, default from env `MAST3R_NPZ_COMPRESS`, default 0) gives the deflated file back | 30 npz requests x 4 runs per side, base = ad39150 server (2 runs, interleaved) vs new: server `timing_ms.total` 294.5-297.1 -> 81.5-83.3 ms (**-72 %**); encode 228-230 -> 18.1-18.6 ms; client wall 330.9-333.1 -> 130.0-139.7 ms (**-60 %**). Response 6.77 -> 9.79 MB (base64 JSON). Same server with `compress_npz: true`: total 294-304 ms (the old cost; isolates the item). Arrays identical after `np.load` (n_diff = 0 on all four arrays) |
| d32f521 | A: parallel PNG | `output_format: "png"`: the 4 PNGs are encoded in a 4-thread pool (`MAST3R_PNG_THREADS`, default 4; 1 = serial). The pool is created once under a lock (sync handlers run concurrently in FastAPI's threadpool; host test with 8 racing threads) and shut down with the app | 15 png requests x 2 runs per server: new code with `MAST3R_PNG_THREADS=1` total 192.1 / 192.9 ms vs default 124.5-139.6 ms (-62 ms, **-34 %**); vs the ad39150 server: total 195.2-200.8 -> 124.5-139.6 ms, client wall 213.9-220.1 -> 142.1-160.0 ms (-31 %). PNG bytes identical (serial vs parallel, same build) |
| 1eb8d6a | B: `POST /predict_npz` | New route, same JSON request; the response body is the npz itself (`application/x-npz`), the other fields are compact ASCII JSON in the `X-Mast3r-Meta` header. `/predict` unchanged (same keys in the same order, host test). `smoke_test.py --npz-route`; host tests `models/tests/test_server_host.py` (9 tests) | 30 requests x 4 runs: total 72-76 ms, client wall 96.8-102.0 ms (**-70 %** vs the ad39150 `/predict` npz 331-333 ms, -35 ms vs stored `/predict`). Body byte-identical to `/predict`'s decoded `npz_b64`; arrays n_diff = 0 |

### Accuracy (`code/tools_prof/r3_validate.sh 1` at 1eb8d6a; gate code unchanged; all gates pass)

| metric | ad39150 (audit base run) | now (lofie + lofid) | gate |
|---|---|---|---|
| randn head1 / head2 (all), eth12 float | 0.99848 / 0.99855 (0.99875) | 0.99862 / 0.99869 (0.99891) | 0.998 |
| test_e2e (2 CQ) | 0.9985 PASS | 0.9985 PASS (head1 0.9985, head2 0.9982) | 0.99 |
| real pair float pts3d h1 / h2 | 0.99035 / 0.99081 | 0.99016 / 0.99088 | 0.989 |
| real pair float conf h1 / h2 | 0.99262 / 0.99511 | 0.99269 / 0.99486 | 0.99 |
| served u8 real pair pts3d h1 / h2 | 0.99021 / 0.99086 | 0.99015 / 0.99097 | 0.989 |
| served u8 real pair conf h1 / h2 | 0.99232 / 0.99497 | 0.99246 / 0.99519 | 0.99 |
| sym out_ii / out_ji / out_jj vs two-pass | bit-identical | bit-identical | identical |
| `MAST3R_OPT=none` randn | 0.99836 / 0.99895 | 0.99836 / 0.99895 (unchanged) | |
| **not gated**: synthetic u8 pair head1 / head2 | 0.99733 / 0.99941 | **0.99593** / 0.99927 | |
| **not gated**: real pair float median \|dz\|/z h1 / h2 vs fp32 ref | 1.17 / 1.10 % | **1.24 / 1.17 %** | |
| **not gated**: same vs the half-pixel reference | 0.37 / 0.41 % | **0.40 / 0.48 %** | |
| **not gated**: served u8 median \|dz\|/z h1 / h2 | 1.18 / 1.09 % | 1.24 / 1.16 % | |

Disclosure: the LoFi item is not bit-identical. Every gate passes with the same or better margin, but the
ungated synthetic-u8 head1 PCC drops by 0.0014 (from the encoder group) and the median relative depth error on the
real pair grows by about 0.07 percentage points per head. The audit flagged these two as "owner review"; the user
decision for this stage was to adopt precision changes that pass all gates, with an env switch back
(`MAST3R_OPT=all,-lofie,-lofid`, verified bit-identical to ad39150 on the served path).

### Served smoke tests

`smoke_test.py --require-pose` PASS on both ad39150 servers; `--require-pose --npz-route` PASS on all four new-code
servers (default, `MAST3R_PNG_THREADS=1`, old-numerics env, default again). Host tests: `test_fused_host.py` 32/32,
`test_server_host.py` 9/9.

### Skipped (from the audit list)

- `json` splice of the base64 payloads into the response body (audit d2e7aee): about -5 ms, inside the +-10 ms noise
  (audit: inconclusive). `/predict_npz` removes that cost for clients that opt in.
- `par` (decode + preprocess of the two views in a thread pool): measured negative / bimodal in the audit (GIL).
- `fmed` (`np.partition` median in the summary): too small to measure (audit), and the summary is outside `timing_ms`.
- Every device item that the audit did not demonstrate (SDPA redesign, LN fold, hrope fold, polyphase border strips,
  bfloat8_b K/V, HiFi2 DPT phase convs with bias correction, megakernel, concurrent DPT heads on sub-devices): not
  attempted in the audit, so out of scope here.
- Serving on an x8 Galaxy chip (audit #17): a placement item, not demonstrated, and not relevant on a real p150 (x16).
- Nothing integrated needs a grid other than 12x10 or Tensix-only dispatch: lofie / lofid were measured on both
  eth12 (ETH dispatch, 1 CQ) and the served worker 2-CQ path.

### Serving files

`tt-model.yaml`: `serve.env` pins `MAST3R_NPZ_COMPRESS: "0"` and `MAST3R_PNG_THREADS: "4"`, the `MAST3R_OPT` comment
lists lofie / lofid, two new `verify` lines (knobs in `all`; `/predict_npz` route and `compress_npz` field), and
the card quickstart lists `compress_npz` and `/predict_npz`. No new runtime package (stdlib `json`,
`concurrent.futures` only). `SERVING.md`: env table, `compress_npz`, the stored-npz change note, and the
`/predict_npz` contract.

## Round 9 (2026-10-03, chip 9, round start = 5b918db, final = ad39150)

### Verifier notes on round 8, resolved
- **Span extraction script.** `code/tools_prof/trace_span.py` (commit 0c5f9fe) is the device-span tool: per trace replay
  (grouped by `METAL TRACE ID` + `METAL TRACE REPLAY SESSION ID`) it prints max(FW end) - min(FW start) at 1.35 GHz, the
  program count, the kernel-duration sum and the summed FW gaps; `--list` prints the per-program timeline. Usage:
  `MAST3R_CQS=1 python -m tracy -r -p -v -o <dir> --op-support-count 6000 tools_prof/prof_trace.py` then
  `python3 tools_prof/trace_span.py <dir>`.
- **Re-validation of 5b918db on chip 9 (14:26-14:30 UTC).** It reproduces the verified numbers: eth12 trace 24.96 / 24.68,
  e2e 29.99 ms; served u8 trace 25.32, e2e 28.44 ms; all accuracy lines identical to the verified table; device span
  23.531 / 23.531 / 23.545 ms (693 programs, kernel sum 22.90 ms).
- **Server smoke test re-run** (worker dispatch, 2 CQ, the tt-model.yaml serve env, chip 9, at e9396dd): PASS 4 / 4,
  ready 6.2 s after start, `forward` 37-39 ms per pose request, pose rot 43.2 deg, f1 / f2 446 / 440 (same as round 8).
- **Host wall vs device span.** Explained below (AICLK throttling). The device span stays the reliable A/B metric; every
  host-wall A/B in this round is reported with that caveat.
- **Process.** Every new kernel / kernel variant of this round ran first alone, on the smallest shape, under
  `TT_METAL_WATCHER=2` and `timeout -s INT 60/120`, one device process per command. No hang, no fault.

### Measured: where the 1.3-2.2 ms between trace wall and device span goes (backlog item 1)
- **Fixed trace launch + completion latency is 0.02 ms** (`tools_prof/r9_launch_probe.py`, ETH dispatch: a 1-program
  trace, execute + synchronize, median 0.023 ms; event_synchronize 0.019 ms; idle synchronize 0.010 ms). Tiny programs
  cost 6.3 us each when dispatch-bound (700-program trace of 1-core adds: 4.43 ms), which is not the case in this graph
  (in-trace FW gaps 0.24 ms in total).
- **The chip clock drops under load.** `tt_aiclk` (sysfs, read-only, chip 9) sampled every 50 ms during a 200-iteration
  bench: 1350 MHz when idle or lightly loaded, 1231-1343 MHz (mostly 1256-1287) while the forward runs at 100-170 W.
  The device profiler counts cycles and converts at the nominal 1.35 GHz, so the span understates wall time:
  `tools_prof/r9_clock_probe.py` (traces of 20 / 120 stock matmuls, wall slope vs cycle slope) gives an effective
  clock of 1.296 GHz (144 us matmuls) and 1.314 GHz (53 us matmuls). For this graph: 23.50 ms of cycles at ~1.28 GHz =
  24.8 ms, which is the measured eth12 trace wall (24.76-24.96 ms). **The gap is DVFS, not host or dispatch overhead;
  model code cannot remove it** (less energy per forward would raise the clock a little).
- **Served-path tail** (`tools_prof/r9_tail_timeline.py`, worker 2 CQ, u8, host event waits as `read_outputs`):
  prep + H2D + enqueue 0.55 ms; segment ends without readbacks A 22.51 / B 25.68 / C 26.30 / D 26.92 ms; with the
  overlapped readbacks the last segment ends ~0.9 ms later, and the e2e is 28.0 ms. The tail after the last segment is
  only ~0.15 ms.
- **The overlapped CQ1 readbacks stall CQ0 kernels (Galaxy x1 PCIe).** `tools_prof/r9_cq2_prof.py` (served path under
  the device profiler, overlapped vs synchronize-then-read replays): segment A is unchanged (19.55 ms); segment B
  (head 2 main rows) goes 2.92 -> 3.44-3.54 ms and segment C 0.54 -> 1.13 ms, each because ONE program (a different one
  per replay: e.g. the 64^2 dadd GenericOp 24 -> 637 us, a 2-us I2S -> 565 us) stalls while the 2 MB D2H transfer runs.
  Reproduced with stock ops only (`tools_prof/r9_d2h_stall.py`, trace of 300 I2S programs + a concurrent 2 MB CQ1
  read): one program stalls 645 us; all its cores start on time and finish late (their NoC transactions wait).
  8 x 256 KB reads give one ~80 us stall per chunk; an L1 source gives many 30-40 us stalls; total stall is the same.
  The PCIe link of chip 9 is x1 (32 GT/s, `current_link_width` 1): the device -> host writes back up into the NoC for
  the whole transfer (~2 GB/s). A p150 (x16) drains 16x faster. **Effect on the served e2e: about 1.2 ms per pair on
  this host; the overlap still saves ~1 ms vs reading after the last segment (2 x ~1.1 ms).** No model-side fix found.

### Measured: SDPA cost structure with the exact production config (backlog item 2)
Method: the stock streaming SDPA compute kernel resolved from a scratch CWD (tt-metal searches the current directory
first, `tt_metal/impl/kernels/kernel.cpp`), with compute-only edits, separate `TT_METAL_CACHE` per variant, watcher
first run each; `tools_prof/r7_sdpa_cost.py`, traced, host wall per call. Stock reader / writer and program factory
unchanged. Diagnostic only (wrong outputs), nothing of this is in the model.

| variant (compute kernel edit) | encoder B2 H16 q160 k512 | decoder B2 H12 q224 k512 |
|---|---|---|
| stock (copy, unchanged) | 94.6 us | 67.4 us |
| no exp | 91.0 | 60.4 |
| no row-sum packs | 92.7 | 65.2 |
| no P @ V matmul | 91.1 | 60.8 |
| no Q @ K^T matmul | 91.2 | 65.1 |
| none of the four | **84.2** | **51.2** |

- Even with all matmuls, the exp and the sum packs removed, the encoder SDPA takes 84 us: the stock KV reader / chain
  (K and V streamed twice per core: 224 q chunks on 120 cores) is a co-bottleneck. **A custom compute kernel with the
  stock reader can save at most ~10 us per encoder call (the planner's threshold was 15 us) and ~16 us per decoder
  call.** Not started.
- Zone timeline (stock zones enabled, top-level zones only): the math thread is busy for the whole kernel, ~2.4 us per
  q row and 16-tile K chunk for Q @ K^T + max + sub/exp, ~6 us per K chunk for the P @ V drain; ~300 cycles per score
  tile.
- Chunk sweep, stock: encoder q160 k512 94 us is the best of q96-q512 / k256-k1024 (q352 110, q384 114, q256 160 us);
  decoder q224 k512 67 us is the best (q256 79, q192 118 us); exp_approx 66.5 vs 67.3 us (decoder).
- What a real gain needs: per-core K / V delivered once (multicast from the hrope writer into the SDPA cores' L1) plus a
  leaner compute kernel and balanced 8-9 rows per core across head boundaries. Estimate -0.7 to -1.0 ms in total, a
  multi-kernel redesign (hrope writer + SDPA reader + compute).

### Measured: other cost splits
- **hrope** (`tools_prof/r9_hrope_probe.py` on the obs block-sharded qkv; `MAST3R_HROPE_DIAG` defines in the model-local
  kernels, default build unchanged and bit-identical): stock 29.1 us (host wall per call), no reads 25.8, no writes 23.3,
  no compute 25.3, none of them 11.5. No single stage dominates.
- **LayerNorm** (`tools_prof/r9_ln_probe.py`, enc shape on the obs residual): stock 20.2 us; without reads 16.1 us
  (`tools_prof/kern/ln_reader_fast.cpp`, D_NOREAD); 1 or 2 blocks per barrier 20.2-20.3, 4 or 8 blocks per barrier
  27.2-27.8 us. LN is compute-bound.
- **fc1 output writes** (obsc probe below): fc1 with a block-sharded output is 125.7 vs 127.7 us interleaved; the output
  NoC writes are not a cost.

### Kept steps

| commit | knob | change | effect |
|---|---|---|---|
| de679b6 | esplit | The final enc_norm (multi_layernorm, one tile row per core) writes view 1's and view 2's encoder outputs as two tensors (each core's writer gets its view's base address and start tile) instead of one [2, N, 1024] tensor + two DRAM slices (14.6 us each). Bit-identical (eth12 float, sym_check). | device span 23.531 -> 23.508 ms (693 -> 691 programs) |
| e9396dd | demb | The two decoder embeds (1024 x 1024 x 768 per view, two full-grid linears 19.6 + 20.0 us + two I2S 2.7 + 2.4 us) as ONE dual program on the top / bottom 12x5 halves (dmm machinery) writing decoder step 0's per-half block-sharded residual specs directly (29.7 us). Same in0_block_w and compute config -> bit-identical. | 23.508 -> 23.498 ms (688 programs) |
| b117533 | hrqk | hrope compute: q and k of a unit in one pass per stage (4 tiles per fp32 dest half instead of 2 x 2 tiles), same per-tile op sequence -> bit-identical. Probe enc 28.9 -> 28.2 us. | 23.498 -> 23.487 ms |

Bit-identity of the final commit vs the round start: `bench_breakdown.py --dump/--cmp`, eth12 float n_diff 0 on both
heads, served u8 (worker 2 CQ) n_diff 0 on both heads (5b918db worktree dump vs 6d6e155).

### Tried and rejected / measured only (round 9)
- **lnx: a fewer-pass LayerNorm compute kernel** (row sums as x @ J, sum of squares as the diagonal of sum x x^T, SFPU
  statistics in fp32, (x - mean) * rstd with a dest-reuse multiply; `tools_prof/rejected_lnx/`). Correct; the host
  emulation on the real pair's LN inputs (`tools_prof/r9_ln_numerics.py`, |mean|/std <= 0.57) predicted equal accuracy.
  On device it is slower: enc 23.7 vs 20.1 us, decoder pair 20.9 vs 17.0 us; mean |err| vs fp64 1.19e-3 vs 1.17e-3.
  The single-tile HiFi4 matmuls cost more than the stock reduce passes.
- **obsc: fc1 writes its block-sharded output in its own packing order (column-major tiles in each shard, 7x1
  subblocks) and fc2 reads it through the stock interleaved in0 sender with patched strides and a "virtual transposed"
  TensorAccessor** (`heads_rope.colmajor_ta_args`, `linear_descriptor(in0_colmajor=True)`; `tools_prof/r9_obsc_probe.py`).
  Bit-identical on the first run (watcher), but fc1 125.7 + fc2 81.5 us vs 127.7 + 80.3 us: -0.8 us per pair. Not wired in.
- **LN reader with more reads per barrier**: slower (above).

### Final validation (`code/tools_prof/r3_validate.sh 9` at 6d6e155, 15:40-15:44 UTC, 30-iteration median / min)

| | `none` baseline (eth12, 1 CQ, float) | **now: eth12, 1 CQ, float** | now: worker 2 CQ, float | **now: served (worker 2 CQ, uint8)** |
|---|---|---|---|---|
| trace (host wall, synchronized) | 62.21 / 62.11 | **24.76 / 24.45** | 25.08 / 24.53 | 25.19 / 24.56 |
| host prep / H2D | 0.88 / 1.08 | 1.00 / 1.03 | 0.76 / 1.06 | 0.51 / 0.62 |
| D2H, synchronized | 28.70 | 3.31 | 2.84 | 3.02 (overlapped in e2e) |
| **e2e** `model(img1, img2)` | 88.55 / 84.36 | **29.83 / 29.62** | 28.49 / 28.25 | **27.74 / 27.35** |
| `test_mast3r.py --layer end_to_end` (2 CQ, float) | | | 27.51 ms, PCC 0.9985, PASS | |
| real pair: single pair / `sym` pose / two-pass pose e2e | | | | **27.81 / 38.35 / 55.83** (min 27.30 / 37.82 / 52.96) |
| device span, traced (eth12 1 CQ, 3 replays) | | **23.491 / 23.479 / 23.490 ms** (688 programs, kernel sum 22.86 ms) | | |

Same-session A/B (`abco.sh`, 3 alternating pairs, mean of medians, round start 5b918db -> new): eth12 trace 24.93 ->
25.01 ms, e2e 29.51 -> 30.05 ms (new = 3feb9b3, i.e. esplit + demb); served u8 trace 25.29 -> 25.20 ms, e2e 28.06 ->
28.12 ms (new = 6d6e155). All differences are inside the host-wall noise (run-to-run medians move by 0.3-0.5 ms); the
device span (-0.044 ms) is the measurable effect.

**Accuracy: unchanged to every printed digit** (all kept steps bit-identical): randn head1 / head2 / all 0.99848 /
0.99855 / 0.99875 (gate 0.998); test_e2e 0.9985 PASS; real pair float pts3d 0.99035 / 0.99081, conf 0.99262 / 0.99511;
served u8 pts3d 0.99021 / 0.99086, conf 0.99232 / 0.99497; raw PCC vs the half-pixel reference 0.999882 / 0.999814
(float) and 0.999886 / 0.999782 (u8); sym out_ii / out_ji / out_jj bit-identical to the two-pass path. Gate code
unchanged.

### vs RTX 5090 (GPU_COMPARISON.md; like-for-like = the same network, bf16, batch 1, one pair)
- vs bf16 + SDPA (35.3 ms incl_h2d, 34.3 ms device-only): TT served u8 e2e 27.7-28.1 ms is 20-21 % faster; TT device
  span 23.49 ms (cycles at 1.35 GHz; ~24.8 ms at the throttled clock) is 28-32 % faster.
- vs torch.compile + CUDA graphs (21.1 ms served, 20.1 ms device): still out of reach; the GPU is 1.31-1.33x faster served
  and 1.17x (cycle span) to 1.23x (throttled wall) faster on device. Of the ~6.8 ms served gap, ~1.2 ms is the D2H stall
  of this host's x1 PCIe link and ~1.3-1.7 ms is the clock throttling under load (trace wall vs cycle span). The stall is
  specific to the x1 link; whether a p150 throttles the same way was not measured.
- Pose (`sym`, mode-specific, reported separately): 38.35 ms vs 2 x 35.3 = 70.6 ms for two GPU forwards.

### Remaining backlog (device span 23.49 ms; ceilings measured this round)
- **SDPA redesign** (hrope writes K / V once per head into the SDPA cores' L1 by multicast, balanced 8-9 rows per core,
  leaner compute): -0.7 to -1.0 ms. A compute-only kernel is capped at ~-10 us (enc) / -16 us (dec) per call.
- **LayerNorm** 1.45 ms in 84 programs: compute-bound (reads ~4 of 20 us). A fewer-pass kernel built from single-tile
  HiFi4 matmuls is slower (lnx); a win needs a kernel that keeps the stock reduce LLKs but drops the two fp32 round trips.
- **fc1 GELU epilogue** (~28 us x 24 + ~26 us x 12): the pack-thread SFPU can only overlap the matmul if the final sum
  is in dest at the last K block, which needs fp32 dest (4-tile halves) and breaks the 7x1 / 7x11 blocking.
- **Served path on this host:** the D2H stall (~1.2 ms per pair) is a Galaxy x1-PCIe property; on an x8 chip (5, 13, 21,
  29) it should be ~8x smaller.
- Small items left: pemm residual straight into the obs spec (-5 us), DPT 128^2 plumbing (L1-capacity bound).

## Round 8 (2026-10-03, chips 3 -> 18, round start = 96fe46e, final = this commit)

**Re-validation of the round start.** HEAD 96fe46e on chip 3 at 12:15 UTC reproduced round 6/7:
- eth12 trace 25.94 / 24.71 ms; PCC 0.99843 / 0.99872.
- Served u8: trace 26.53 / 25.37, e2e 29.04 / 28.78 ms.
- Real pair u8: conf h1 0.99188.

Nothing was broken, so nothing was reverted.

### Measurement method (new this round: device span)
- **Device span.** `tools_prof/prof_trace.py` under tracy (`MAST3R_CQS=1`, 3 trace replays). From
  `ops_perf_results*.csv`, take max(FW end) - min(FW start) of one replay's programs at 1.35 GHz. Replays agree to
  about 0.03 ms. This is the device time without any host latency. The bench's "trace" number (execute + synchronize,
  host wall) moves by 0.3-0.5 ms between runs on this shared host.
- **Host-side numbers.** `bench_breakdown.py`, 30 or 60 iterations, median / min, as in earlier rounds.
- **A/B.** Same chip, same session, alternating runs, worktree vs working tree (`tools_prof/abco.sh`).

| chip 18, same session | 96fe46e (round start) | **5fbbe2d (round 8)** | change |
|---|---|---|---|
| device span, traced (3 replays) | 23.884 / 23.889 / 23.913 ms (690 programs) | **23.525 / 23.542 / 23.550 ms** (693 programs) | **-0.35 ms (-1.5 %)** |
| device kernel sum per replay | 23.25 ms | 22.91 ms | -0.35 ms |
| eth12 1 CQ float: trace (3 pairs, mean of medians) | 26.44 ms (min 25.64-25.88) | **26.21 ms** (min 25.40-25.55) | -0.23 ms |
| served (worker 2 CQ, u8): trace | 26.17 ms | 25.85 ms | -0.32 ms |
| served (worker 2 CQ, u8): e2e | 28.68 ms (min 28.28-28.58) | **28.53 ms** (min 28.06-28.34) | -0.15 ms |

Final validation: `code/tools_prof/r3_validate.sh 18` at 5fbbe2d, 13:31-13:36 UTC, 30-iteration median / min.

| | `none` baseline (eth12, 1 CQ, float) | **now: eth12, 1 CQ, float** | now: worker 2 CQ, float | **now: served (worker 2 CQ, uint8)** |
|---|---|---|---|---|
| trace (host wall, synchronized) | 62.75 / 62.67 | **26.01 / 25.46** | 25.20 / 24.74 | 25.68 / 24.65 |
| host prep / H2D | 0.82 / 1.06 | 0.90 / 1.03 | 0.89 / 1.09 | 0.44 / 0.58 |
| D2H, synchronized | 29.21 | 2.81 | 3.75 | 3.07 (overlapped in e2e) |
| **e2e** `model(img1, img2)` | 92.57 / 84.59 | **30.48 / 30.21** | 29.06 / 28.42 | **28.52 / 28.31** |
| `test_mast3r.py --layer end_to_end` (2 CQ, float) | | | 27.95 ms, PCC 0.9985, PASS | |
| real pair: single pair / `sym` pose / two-pass pose e2e | | | | **28.57 / 39.41 / 57.00** (min 28.40 / 39.10 / 55.22) |

Chip 18's host wall numbers are about 1 ms above chip 9's (round 6: eth12 trace 25.15). The device span is the
chip-independent comparison.

Server smoke test: worker dispatch, 2 CQ, `MAST3R_OPT=all`, chip 18, at 5fbbe2d, with the tt-model.yaml serve env.
- PASS 4 times out of 4. Ready 7.8 s after start.
- `forward` 37-39 ms per pose request (round 6: 38-39).
- Pose rot 43.2 deg, f1 / f2 446 / 440 (round 6: 43.3 deg, 445 / 440).

### Kept steps

| commit | knob | change | effect |
|---|---|---|---|
| 8a7b17a | resmm | **Encoder residual adds folded into the proj / fc2 matmul epilogues.** The model-local `mm_gelu_compute.cpp` gets `MAST3R_RESID_CB`: after the FUSE_BIAS bias add, `dst += resid` (dest-reuse ELWADD). The residual comes from a CB globally allocated on the residual tensor's shard (`heads_rope._add_resid`). The residual stream stays BLOCK_SHARDED in the obs output spec, which equals the per-core output block (row-major, 1-row subblocks, asserted). Block 0 converts the patch-embed residual once (one I2S, 5.5 us). The two LayerNorms per block read only the sum, and the final enc_norm is a plain affine LN of the sum. | Probe, traced, enc: proj + add/LN 51.5 -> 46.9 us, fc2 + add/LN 114.2 -> 110.0 us. Graph, chip 3, 3 A/B pairs: eth12 trace 26.00 -> 25.72 ms (mean of medians) |
| 8fbb2e3 | resmmd | Same for the decoder: dual proj / cproj / fc2 with per-half residual shards (obsd specs). `multi_layernorm` now groups reader kernels per input layout, because the two branches are sharded on different grid halves. Step 0 converts both decoder-embed residuals once (2 x I2S, 2.5-2.9 us). The last fc2 sum goes straight into dec_norm. Taps (blocks 5 / 8) copy the sum to DRAM as before. | Chip 3, 3 A/B pairs: eth12 trace 26.01 -> 25.81 ms |

**Accuracy (gate code unchanged; all gates pass).** Neither step is bit-identical. The sum is rounded to bf16 once
(fp32 dest -> one pack) instead of twice (matmul output, then the add). The residual add happens in fp32 in the
matmul's dest. The matmul + bias value passes the dest -> srcA move as tf32: the partials CB of these mm32 linears is
fp32, so srcA is in tf32 format.

| | round start 96fe46e | **round 8 (5fbbe2d)** |
|---|---|---|
| randn head1 / head2 (all), gate 0.998 | 0.99843 / 0.99872 (0.99893) | 0.99848 / 0.99855 (0.99875) |
| test_mast3r end_to_end (PCC h1 / h2) | 0.9988 (0.9987 / 0.9986) | 0.9985 (0.9986 / 0.9982), PASS |
| real pair float: pts3d h1 / h2 (gate 0.989) | 0.99023 / 0.99074 | 0.99035 / 0.99081 |
| real pair float: conf h1 / h2 (gate 0.99) | 0.99209 / 0.99509 | **0.99262** / 0.99511 |
| real pair float: raw PCC vs half-pixel fp32 ref h1 / h2 | 0.999863 / 0.999798 | 0.999882 / 0.999814 |
| real pair float: median abs(dz)/z h1 / h2 | 1.26 / 1.12 % | 1.17 / 1.10 % |
| real pair served u8: pts3d h1 / h2 | 0.99047 / 0.99085 | 0.99021 / 0.99086 |
| real pair served u8: conf h1 / h2 | 0.99188 / 0.99483 | **0.99232** / 0.99497 |
| real pair served u8: raw PCC vs half-pixel ref h1 / h2 | 0.999878 / 0.999770 | 0.999886 / 0.999782 |
| synthetic u8 pair (not gated) | 0.99690 / 0.99939 | 0.99733 / 0.99941 |
| `sym` out_ii / out_ji / out_jj vs two-pass | bit-identical | bit-identical |

Summary of the accuracy changes:
- Against the fp32 reference, the real-pair metrics improved or stayed equal: raw PCC vs the half-pixel reference went
  up on both heads, and the median depth error went down.
- The thin served-u8 conf-h1 margin widened from 0.0019 to 0.0023.
- With resmm alone (chip 3), u8 conf h1 was 0.99151. resmmd moved it back up. These 1e-4-level moves are bf16 rounding
  variation on one pair, as in round 5.
- The synthetic randn head2 PCC (0.99872 -> 0.99855) and test_e2e head2 (0.9986 -> 0.9982) went down slightly. Both
  stay well above their gates.

### Tried and rejected / measured only (round 8)
- **hf2h: HiFi2 for the head.0 / head.2 phase convs (HiFi3 now).**
  - It is fast: eth12 trace 24.58 ms, about -1 ms.
  - It is rejected for accuracy. Median depth error vs the half-pixel reference goes from 0.37 / 0.41 % to
    0.73 / 0.91 % (both head convs), or 0.51 / 0.63 % (head.0 only).
  - This matches the earlier finding recorded in the code comment (HiFi2 there was 0.26 -> 0.68 %). Not kept.
- **bsln: LayerNorm of the block-sharded residual in its own layout.**
  - Design: per-grid-row exchange of the token sums / sums of squares (bf16 hi + lo, so the fp32 sums are exact through
    identity matmuls) plus a fused SFPU normalise.
  - Correct, and more accurate than the row-per-core LN: enc mean abs error vs fp64 1.135e-3 vs 1.195e-3.
  - Slower: enc 30.7 us (multicast) / 27.9 us (unicast) vs 20.4 us. The exchange (8 KB per peer) and the many small
    compute passes cost more than the 64 -> 110 core spread saves.
  - The stock sharded `ttnn.layer_norm` on the same layout is also slower (26.0 us) and less accurate (1.54e-3).
  - Code: `tools_prof/rejected_bsln/`.
- **gelin: fc1 bias + GELU epilogue interleaved per output subblock in the last K block.**
  - Bit-identical after a fix, but no gain: enc fc1 137.4 vs 136.1 us. Per subblock the chain "pack partial (L1 acc) ->
    unpack + bias -> SFPU GELU" is serial on the pack thread, so the SFPU epilogue still does not overlap the matmul.
  - Adding the bias before the last K block instead (so that the GELU could overlap the last block's products) was
    40 % less accurate: bf16 dest accumulation at full magnitude.
  - The first version of this kernel hung chip 3 (see Known hangs).
  - Code: `tools_prof/rejected_gelin/`.
- **hrf: RoPE in two dest-reuse stages instead of four.**
  - More accurate: mean abs error vs fp64 1.75e-3 vs 2.06e-3. This needs the srcA format set to fp32 before the
    dest-reuse add; otherwise the dest -> srcA move truncates to bf16.
  - But enc hrope only goes from 28.6 to 27.7 us.
  - Batching more reads per barrier is slower (ub 2 / 4 / 8: 35 / 45 / 59 us), and so is a single read in flight
    (28.2 us).
  - Zones show about 0.9 us of read latency per unit. hrope moves about 24 MB through the NoC per call at about 1 TB/s
    aggregate: it is NoC-bandwidth bound, not compute bound.
  - Code: `tools_prof/rejected_hrf/`.

### Fresh device profile (eager, after resmm / resmmd; 717 programs, kernel sum 22.94 ms)
- Matmuls (generic_op + ttnn): enc fc1 + GELU 127.5 us x 24, enc fc2 80.3 x 24, enc qkv 54.1 x 24, dec fc1 + GELU pair
  84.1 x 12, dec qkx 56.8 x 12, dec fc2 pair 52.2 x 12.
- SDPA 3.59 ms: enc 88 us x 24, dec 61 us x 24.
- DPT convs 4.25 ms + halos 0.77 ms. The biggest are the head.2 phase convs (128 -> 128 at 256^2, HiFi3, 132.5 us x 8)
  and the 128^2 convs (resconv 256 -> 256 113 us x 8, head.0 phase 84.7 us x 8).
- LayerNorm (now LN only, no add): enc 18.8 us x 48, dec 15 us x 36.
- hrope: enc 25 us x 24, dec 21 us x 24.

### Remaining backlog (device span 23.54 ms)
- **Custom SDPA compute kernel**, about -0.7 to -1.2 ms (round-7 analysis: P @ V at 2x the FPU peak, normalise tail).
- **hrope data movement**, about 1.1 ms. It is NoC-bandwidth bound: it reads and writes all of q / k / v. Writing v
  (and possibly q / k) straight into the SDPA layout from the qkv matmul (a tile-remapping output writer) would remove
  about 1/3 of that traffic. Estimate -0.2 to -0.35 ms.
- **GELU epilogue**, about 0.9 ms of SFPU. It is serial on the pack thread. Only a cheaper GELU at equal accuracy, or
  real FPU / SFPU overlap (a different blocking with fp32 partials), would help.
- **LayerNorm-only programs**, 1.55 ms on 64 cores. A layout-local LN needs a much cheaper statistics exchange than
  bsln.
- **DPT head convs.** HiFi2 is not an option (accuracy). Conv blocking for the 128^2 / 256^2 phase convs was explored in
  rounds 3-5.
- Small items:
  - pemm writing the residual straight into the obs shard spec would save the encoder I2S (5.5 us) and let block 0 use
    the sharded LN.
  - The decoder-embed I2S pair (5.4 us).

### vs RTX 5090 (GPU_COMPARISON.md; like-for-like = the same network, bf16, batch 1, one pair)
- Against the module in bf16 + SDPA (35.3 ms incl_h2d, 34.3 ms device-only):
  - TT served u8 e2e 28.5 ms: 19 % faster.
  - TT device span 23.5 ms vs GPU 34.3 ms device-only: 31 % faster.
- torch.compile + CUDA graphs (21.1 ms served, 20.1 ms device) is still out of reach: the GPU is 1.35x faster served
  and 1.17x faster on device time (TT device span 23.5 vs 20.1 ms).
- Pose (`sym`, mode-specific, reported separately): 39.4 ms vs 2 x 35.3 = 70.6 ms for two GPU forwards.

### Known hangs (round 8)
- **chip 3, 13:10-13:20 UTC**: `tools_prof/r8_gelin_probe.py enc 20` with the first version of the new model-local fc1
  compute kernel `tt/kernels/mm_gelu_inline.cpp` (bias + GELU epilogue interleaved per subblock in the last K block).
  Cause (found by review, not by re-running): the kernel waited for the bias CB at its START, but ttnn's in1
  sender/writer pushes the bias only after it has pushed every in1 K block; with 8 K blocks the in1 CB fills, the reader
  blocks, and the compute never consumes -> deadlock. The 2-K-block small shape fit the in1 CB and passed. Fixed by
  waiting for the bias in the last K block (where the stock kernel waits). The fixed kernel then ran on chip 18 (watcher,
  8-K-block small shape and the enc shape, bit-identical, no hang). FAULT written via chip_pool (new chip 18).
  A follow-on `r8_gelin_probe.py dec` from the same shell loop opened chip 3 for ~3 s before it was killed.
  Lesson: never chain several device runs in one shell loop after a new kernel; a hang must stop the loop.

## Round 7 (2026-10-03, chip 9, round start = e9c87d6): no kept speedup, chip faulted

**Verifier notes on round 6 (all minor), resolved in this report:**
- The independently measured round-6 A/B was -2.5 % eth12 trace (claimed -2.9 %), -2.8 % served u8 trace and -1.4 %
  served e2e (claimed -1.7 %). Those verified figures replace the round-6 claims. Absolute verified values at e9c87d6:
  eth12 trace 24.89 / 24.73 ms, e2e 30.00 / 29.35 ms; served u8 2-CQ trace 25.19 / 24.77 ms, e2e 28.13 / 27.45 ms.
- Baseline reference: the first commit 20bdff0 has no `bench_breakdown.py`. The baseline numbers use 81ff339, the
  earliest commit with the bench tooling (trace 62.2 ms). `MAST3R_OPT=none` at the final commit reproduces it
  (62.19 / 62.11 ms).
- chipenv's "expected 32 chips on the bus, found 31" warning is host-level and appeared on every run of every agent.
- The server smoke test and the tracy stage profile were not re-verified independently. They were not re-run in round 7
  either, because the default path did not change.

**Nothing new is enabled by default.** `MAST3R_OPT=all` is the same knob set and code path as e9c87d6. The new code
below is reachable only through explicit arguments or probes. The round-7 commit changes no default kernel source.

### Measured (chip 9, traced, isolated ops)

**1. SDPA cost structure** (`tools_prof/r7_sdpa_cost.py`, stock streaming SDPA, 12x10). Varying the K length at a fixed
q chunking:
- decoder B2 H12 q224 / k512: 40.8 / 67.3 / 117.6 / 214.8 us at Sk = 512 / 1024 / 2048 / 4096. That is about
  300 cycles per 32x32 score tile in steady state, plus about 16 us fixed per call.
- encoder B2 H16 q160 / k512: about 317 cycles per score tile, plus about 18 us fixed.
- k256 instead of k512 costs about 4.6 us per extra K chunk.
- Other points:
  - exp_approx on/off: 93.4 vs 93.5 us.
  - head dim 32 / 64 / 128: 69 / 94 / 190 us.
  - The HiFi2 FPU floor would be about 128 cycles per score tile (2 + 2 tile products).

**2. `fattn`: stock streaming SDPA compute with a single K chunk (exact softmax, no rescale), plus a model-local reader
that keeps each head's K^T / V resident in L1, plus row-balanced work (9 rows per core instead of 10)**
(`tt/fattn.py`, `tt/kernels/fattn_reader.cpp`, `tools_prof/fattn_probe.py`):
- Correct: PCC vs fp32 0.999613 (stock 0.999613); mean |err| 0.00344 vs 0.00352.
- **Slower: 120 us vs 94 us (encoder), 103 vs 87 us (decoder shape).** Without the K/V loads (diagnostic, wrong
  output) the compute alone is 80.7 us. So the stock kernel's KV chain / multicast forwarding costs only about 14 us, and
  loading full K/V per core costs about 40 us.
- Not kept. A win here needs a new SDPA compute kernel, not a new schedule.

**3. Zone profile of the stock streaming SDPA compute** (model-local copy `tt/kernels/fsdpa/` with the stock zones
enabled, `FSDPA_PROF=1`, `tools_prof/zonesum.py`). Per 2-row q chunk x 32 K tiles:
- Phase 1 (Q K^T + row max + partial sub/exp): about 4 us. The Q K^T matmul runs near the FPU peak.
- Phase 2 (softmax @ V + normalise): about 12 us. The P @ V matmul alone is about 6 us, i.e. about 63 cycles per tile
  product, 2x the peak. With dh = 64 the output subblock is only 2 tiles wide, so every product needs about one
  unpack.
- Taller P @ V subblocks (h = 4) would be about 8 % faster, but the stock streaming kernel produces wrong output with
  them (it assumes h <= 2).
- This is the concrete target for a custom attention kernel: P @ V with V-tile reuse across 4+ q rows, and the
  normalise / rescale tail.

**4. L1-interleaved NoC read bandwidth** (`tools_prof/r7_read_bw.py`, 120 cores each reading 128 tiles):
- Aggregate bandwidth by reads in flight per barrier: 1 -> 707 GB/s, 2 -> 955, 4 -> 879, 8 -> 731, 16 -> 514,
  unthrottled -> 270 GB/s.
- The same reads split over two RISCs (reader + writer kernels, both NoCs) reach 1.66 TB/s.
- ttnn's SDPA reader already throttles to 2. Model-local readers that issue many reads before a barrier should be
  checked against this.

**5. Fused add + LayerNorm** (`tools_prof/r7_ln_probe.py`, encoder shape, 64 rows on 64 cores; stock-config kernel 24.7 us
traced, about 20 us kernel time):
- Phase marks show the add phase is bound by reading the two operands (about 9 us). The remaining five compute passes
  take about 10 us.
- **`lnf`** (`tt/kernels/ln_fast_compute.cpp`, fewer passes: ones-matmul row sums, SFPU square with packer L1
  accumulation, dest-reuse normalise): correct (mean |err| vs fp64 1.62e-3 vs 1.56e-3 stock) but **slower**:
  28.0 vs 24.7 us. The packer L1-accumulate packs cost about 350 cycles each.
- **`lnsplit`** (operand b read by the writer on the other NoC: `ln_reader_split.cpp` + `ln_add_writer_rb.cpp`):
  bit-identical, **slower** (26.0 vs 24.7 us), because the writer serialises the b reads with the residual writes.
- Not kept, and neither is wired into the model.

### Chip fault (round 7)
`tools_prof/r7_ln_probe.py enc` with a variant of `ln_add_writer_rb.cpp` that read ALL of operand b before writing any
residual block deadlocked:
- The compute blocked on the full 8-tile CB 17 while the writer blocked on the full 8-tile cb_inb.
- The process ignored SIGINT and was SIGKILLed.
- A follow-on loop iteration opened the device for about 2 s before it was killed too.
- FAULT file: `chipstate/chip9/FAULT`.

The committed `ln_add_writer_rb.cpp` is the interleaved (deadlock-free, previously run) variant, with a comment
warning against the hung one. **Lesson for the next round:** any writer that consumes a compute output must drain it
inside the same loop that feeds the compute, unless the CB holds the whole row.

### Round-7 backlog (unchanged priorities, with the new measurements)
- **Custom SDPA compute kernel, about -0.7 to -1.2 ms.**
  - Phase 2 (P @ V at 2x peak, plus normalise) is about 75 % of the per-chunk time.
  - Design: P @ V with 4-row output subblocks (V-tile reuse) and the row sum folded into the P @ V matmul (a ones column).
  - Keep ttnn's KV chain / multicast reader. It is only about 14 us of overhead.
- **Add + LN, -0.3 to -0.5 ms.** The add phase is read-bound (about 9 of 20 us). Two-NoC operand reads with large CBs
  (cb_inb and CB 17 sized to the whole row, so that no deadlock is possible) should overlap the two operand streams.
  A pass-fusion compute kernel is not a win with packer L1-accumulation.
- **GELU epilogue overlap in fc1, about -0.5 to -0.9 ms.** Not started.

## Round 6 summary (2026-10-03, chip 9, HEAD 8d1c6b0)

Round start = 242bfe7. Its independent verification: eth12 trace 25.88 / 25.63, e2e 30.08 / 29.80; served u8 trace
26.18 / 25.80, e2e 28.86 / 28.50. The verifier listed no blocking problems. Its one note, the thin served-u8 conf-h1 margin
(0.99188 vs the 0.99 gate), is unchanged because every round-6 step is bit-identical.

**Same-session A/B, round start vs now.** 3 pairs each, alternating runs. A = a git worktree at 242bfe7, B = 8d1c6b0
(`tools_prof/abco.sh`). Each run is `bench_breakdown.py`, 30 iterations, median / min.

| | 242bfe7 (3 runs) | **8d1c6b0 (3 runs)** | change (mean of medians) |
|---|---|---|---|
| eth12 1 CQ float: trace (device) | 25.77 / 25.70 / 26.01 (min 25.50-25.69) | **25.00 / 25.29 / 24.98** (min 24.77-24.84) | 25.83 -> 25.09 ms, **-2.9 %** |
| eth12 1 CQ float: e2e | 30.86 / 30.88 / 30.95 | 29.99 / 30.01 / 30.19 | 30.90 -> 30.06 ms, -2.7 % |
| served (worker 2 CQ, u8): trace | 25.95 / 26.29 / 26.18 | 25.05 / 25.44 / 25.64 | 26.14 -> 25.38 ms, -2.9 % |
| served (worker 2 CQ, u8): e2e | 28.76 / 28.88 / 28.86 (min 28.20-28.55) | **28.22 / 28.38 / 28.46** (min 27.48-27.94) | 28.83 -> 28.35 ms, **-1.7 %** |

Final validation: one `code/tools_prof/r3_validate.sh 9` session at 8d1c6b0, 10:24-10:28 UTC, 30-iteration median / min.

| | `none` baseline (eth12, 1 CQ, float) | **now: eth12, 1 CQ, float** | now: worker 2 CQ, float | **now: served (worker 2 CQ, uint8)** |
|---|---|---|---|---|
| device forward (trace, synchronized) | 62.22 / 62.12 | **25.15 / 24.68** | 25.10 / 24.88 | 25.36 / 25.09 |
| host prep / H2D | 0.91 / 1.07 | 0.85 / 1.02 | 0.89 / 1.05 | 0.41 / 0.58 |
| D2H, synchronized | 28.14 | 2.89 | 3.51 | 3.18 (overlapped in e2e) |
| **e2e** `model(img1, img2)` | 91.66 / 84.02 | **29.91 / 29.48** | 28.91 / 27.98 | **28.32 / 27.46** |
| `test_mast3r.py --layer end_to_end` (2 CQ, float) | | | 28.12 ms, PCC 0.9988, PASS | |
| real pair: single pair / `sym` pose / two-pass pose e2e | | | | **28.34 / 39.33 / 56.84** (min 27.91 / 38.80 / 54.96) |
| eager profile: programs / kernel sum | 1252 / 60.97 | 714 / 23.29 (round start 714 / 24.19) | | |

Server smoke test (worker dispatch, 2 CQ, `MAST3R_OPT=all`, chip 9, at 8d1c6b0) passed 4 times out of 4. The server was
ready 6.4 s after start. `forward` took 38-39 ms per pose request (round 5: 39-41 ms). Pose rot was 43.3 deg and f1/f2
were 445/440, the same as round 5.

**Accuracy: every round-6 step is bit-identical.** Each step was compared with `bench_breakdown.py --dump/--cmp` (n_diff = 0
on both heads), on the eth12 float path and on the served u8 2-CQ path. The validation run reproduces the round-5 numbers
to every printed digit:
- randn head1 / head2 / all: 0.99843 / 0.99872 / 0.99893 (gate 0.998).
- test_mast3r end_to_end: PCC 0.9988, PASS.
- Real pair, float: pts3d 0.99023 / 0.99074, conf 0.99209 / 0.99509, median dz/z 1.26 / 1.12 %. Raw PCC vs the half-pixel
  reference is 0.999863 / 0.999798.
- Real pair, served u8: pts3d 0.99047 / 0.99085, conf 0.99188 / 0.99483, median dz/z 1.17 / 1.09 %.
- Synthetic u8: 0.99690 / 0.99939.
- `sym`: out_ii / out_ji / out_jj are bit-identical to the two-pass path.

### Matmul efficiency audit (fresh eager device profile, 12x10, HiFi2 + fp32 dest except the fc1 GELU kernels)

| linear (per forward) | M x K x N | round start us (TFLOP/s) | **now us (TFLOP/s)** | note |
|---|---|---|---|---|
| enc qkv (x24) | 2048 x 1024 x 3072 | 68.7 (188) | **54.1 (238)** | obs |
| enc proj (x24, pcat) | 2048 x 1024 x 1024 | 29.2 (147) | **24.2 (178)** | obs |
| enc fc1 + GELU (x24, gpoly) | 2048 x 1024 x 4096 | 127.5 (135) | 127.4 | about 100 matmul + about 27-30 SFPU epilogue |
| enc fc2 (x24) | 2048 x 4096 x 1024 | 84.9 (202) | **79.1 (217)** | obs |
| dec qkx pair (x12, dmm) | 2 x 1024 x 768 x 3840 | 74.6 (81) | **56.9 (106)** | obsd |
| dec fc1 + GELU pair (x12) | 2 x 1024 x 768 x 3072 | 84.0 | 84.0 | matmul-only 59; GELU about 26 |
| dec fc2 pair (x12) | 2 x 1024 x 3072 x 768 | 54.9 | **51.3** | obsd |

`tools_prof/r6_mm_sweep.py` sweeps the full 2-D mcast config space on 12x10:
- grid shapes, transposed mcast, per_core_M/N incl. 11-/12-column N splits, in0_block_w 2-16, subblocks,
  out_block_h/w splits.

It confirmed the shipped configs as the fastest for every encoder and decoder shape. The remaining loss was not in the
compute configuration:
- The per-core compute runs at about 40 cycles per 32x32x32 tile product (HiFi2 peak is 32).
- With an L1-interleaved output, the in1-sender/writer finished ~15 us after the compute threads in the encoder qkv
  (its 56 output tiles are NoC-written to 120 L1 banks at the end).
- A BLOCK_SHARDED output equal to the per-core block lets the matmul pack straight into its own shard: qkv 75.1 -> 60.2,
  proj 31.0 -> 25.7, fc2 92.0 -> 87.4, decoder qkx 74.2 -> 58.4 us (in isolation, traced). That is what obs / obsd do.

The remaining grid-quantisation loss (64 tile rows on 10 grid rows -> 7 rows per core vs 6.4 ideal; fc2 N = 32 tiles on 11
of 12 columns) was not recoverable with uniform 2-D blocks. A 60 + 4 tile-row split into two programs estimates at only
-5 us per fc1, because the 4-row tail job is inefficient. It was not done.

**In-trace dispatch gaps are small.** A traced device profile (`tools_prof/prof_trace.py`, 3 iterations) gives a span of
23.90 ms for the 690 programs of the main segment, with a kernel sum of 23.27 ms. So in-trace gaps total about 0.63 ms (about
0.9 us per program), not the ~1.9 ms estimated in round 5 from (bench trace - kernel sum). The rest of the bench's
"trace" time is host enqueue / completion latency. Program-count reduction is therefore worth about 1 us per removed
program.

### Round-6 steps (A/B in one session; trace = eth12 1-CQ median)

| commit | knob | change | effect |
|---|---|---|---|
| 667e43e | obs | **Encoder qkv / proj / fc2 write BLOCK_SHARDED outputs** whose shard is the 2-D mcast per-core block (`_bs_mc`), so there are no output NoC writes. The last shard row may be partial: 64 tile rows = 9 x 7 + 1. Consumers: hrope reads the qkv shards (reader `kind` = sharded layout); the fused add + LN reads proj / fc2 through its stock TensorAccessor (compile-time args of the sharded tensor) | 26.12 / 25.95 -> 25.43 / 25.45 ms (2 pairs); bit-identical |
| 5e87349 | obsd | **Decoder dual (top / bottom 12x5) qkx, proj, cproj and fc2 write per-half BLOCK_SHARDED outputs.** hrope takes two sharded layouts (self: a1 top / a2 bottom; cross: the other branch's qkx half). The fused add + LN builds one reader kernel per residual layout, each on the cores of its tensor's rows | 25.38 / 25.72 -> 24.92 / 25.20 ms (2 pairs); served u8 trace 26.13 -> 25.31 (obs + obsd); bit-identical (eth12 float, served u8) |
| 20ce5de | (helper) | `_patch_in0_ta`: a block-sharded in0 read by the stock interleaved in0 sender through a patched TensorAccessor (descriptor built on a shape-equal placeholder). Bit-identical (`obsf_probe.py`); used only when an in0 is sharded | |
| 8d1c6b0 | (serve) | tt-model.yaml knob comment (obs, obsd; no new kernel files, the verify list is unchanged) | |

### Round-6 tried and rejected / measured only
- **obsf** (encoder / decoder fc1 output block-sharded -> fc2 in0 via `_patch_in0_ta`): a sharded matmul output needs
  1-row subblocks (a 7x1 subblock writes column-major blocks: `obsf_probe.py` n_diff 8.0M). The 1x1-subblock GELU fc1 runs
  135 -> 151 us. Not kept.
- hrope with direct block-shard addressing (NoC table, no TensorAccessor) + decoder cq block-sharded: no gain in 2 A/B
  pairs (25.06 / 24.95 vs 24.94 / 25.17 ms). Not kept.
- fc1 with fp32 dest + the GELU epilogue (1x4 subblocks): 144.5 vs 135 us. fc1 out_block_w splits (overlap the epilogue of one
  block with the next): 151-162 us. Decoder fc1 config sweep (in0_block_w 2-24, subblocks): none faster than the
  shipped 84.8 us.
- LayerNorm: a 2-way width split needs 128 cores for the 64 encoder (and 2 x 32 decoder) tile rows, which is more than 120.
  Switching the fp32 intermediate CBs to bf16 (`r6_ln_probe.py`) does not change the time (24.6-24.8 us), so the stock
  LN compute is not CB-traffic bound. Not pursued without a new LN compute kernel.

### Remaining backlog (device ~25.1 ms, eager kernel sum 23.29 ms, in-trace gaps ~0.6 ms)
- **GELU epilogue, ~1.0 ms** (enc fc1 ~28 us x 24, dec fc1 pairs ~26 us x 12). It is SFPU-bound on the math thread after the
  last K block. Only overlapping it with the FPU (a new low-level compute kernel) or a cheaper approximation at equal
  accuracy would cut it.
- **Add + LayerNorm, ~1.85 ms** (48 x 24 us enc on 64 cores, 36 x 19.5 us dec). Compute-bound in the stock multi-pass
  kernel. A new LN compute kernel (fewer passes; not bit-identical) is estimated at -0.5 to -0.8 ms, with the served-u8
  conf-h1 margin of 0.0019 as the accuracy risk.
- **SDPA, 3.6 ms** (enc 88 us x 24 at ~97 TFLOP/s, dec 61 us x 24). Stock-kernel floor at q160/k512 and q224/k512; exp-approx,
  LoFi and other chunkings were measured earlier with no gain.
- **DPT ring strips**: ~0.46 ms of kernels in ~46 small programs per forward. The halos on 6-row strips take 23-25 us each.
  An exact border correction of the polyphase convs (1-row correction terms instead of re-running upsample + conv on strips)
  is estimated at -0.4 ms. It is a redesign of the phase path with rounding changes.
- Grid quantisation of the encoder matmuls (7 vs 6.4 tile rows per core): ~0.2-0.3 ms, needs non-uniform per-core blocks
  (a custom matmul) to recover.

### vs RTX 5090 (GPU_COMPARISON.md; like-for-like = the same network, bf16, batch 1, one pair)
- Module in bf16 + SDPA, 35.3 ms incl_h2d (34.3 ms device-only):
  - TT served 28.3 ms: 20 % faster.
  - TT eth12 1-CQ 29.9-30.1 ms: 15 % faster.
  - Device only: TT 25.1 vs 34.3 ms, 27 % faster.
- torch.compile + CUDA graphs, 21.1 ms served (20.1 ms device): still out of reach. GPU is 1.34x faster served and 1.25x
  device-only. Closing the gap needs ~19 ms of device time, the practical floor estimated by the planner, at the gated
  precision.
- Pose (`sym`, mode-specific, reported separately): 39.3 ms vs 2 x 35.3 = 70.6 ms for two GPU forwards.

## Round 5 summary (2026-10-03, chip 9, HEAD 2a5f1ed)

Round start = e845c4c (verified: eth12 trace 28.76, e2e 33.71; served u8 trace 28.98, e2e 31.91). Re-measured at the start
of this session (07:44 UTC): eth12 trace 28.72 / 28.60, e2e 33.23 / 32.88; served u8 trace 28.97 / 28.78, e2e 32.01 / 31.41.
Final numbers: one `code/tools_prof/r3_validate.sh 9` session at 2a5f1ed, 09:31-09:35 UTC, 30-iteration median / min.

| | `none` baseline (eth12, 1 CQ, float) | round start (07:44) | **now: eth12, 1 CQ, float** | now: worker 2 CQ, float | **now: served (worker 2 CQ, uint8)** |
|---|---|---|---|---|---|
| device forward (trace, synchronized) | 62.21 / 62.08 | 28.72 / 28.60 (served 28.97 / 28.78) | **25.85 / 25.52** | 26.14 / 25.67 | 26.23 / 25.72 |
| host prep / H2D | 0.79 / 1.07 | 0.80 / 1.01 | 0.99 / 1.03 | 0.87 / 1.06 | 0.39 / 0.56 |
| D2H, synchronized | 28.30 | 3.04 | 3.04 | 2.83 | 2.68 (overlapped in e2e) |
| **e2e** `model(img1, img2)` | 91.73 / 91.39 | 33.23 / 32.88 (served 32.01 / 31.41) | **30.96 / 30.78** | 29.34 / 28.94 | **28.88 / 28.16** |
| `test_mast3r.py --layer end_to_end` (2 CQ, float) | | | | 28.84 ms, PCC 0.9988, PASS | |
| real pair: single pair / `sym` pose / two-pass pose e2e | | 31.53 / 43.80 / 63.12 (verifier) | | | **28.56 / 39.41 / 56.98** |
| eager profile: programs / kernel sum | 1252 / 60.97 | 754 / 26.86 | ~710 / ~23.9 (724 / 24.20 at 740fb2f, before the 16^2 tups) | | |

Device -10.0 % this round (28.72 -> 25.85 ms eth12; 28.97 -> 26.23 ms served), -58 % overall vs `none` (62.2 ms). Served e2e
32.0 -> 28.9 ms (-9.8 %); eth12 1-CQ e2e 33.2 -> 30.7-31.0 ms (-7 to -8 %, host phases noisy). Server smoke test (worker
dispatch, 2 CQ, `MAST3R_OPT=all`, chip 9, at 2a5f1ed): PASS x4, ready 7.8 s after start, `forward` 39-41 ms per pose request (round 4:
43-45 ms), pose rot 43.3 deg, f1/f2 445/440 (round 4: 443/438; the focal estimate moves by 2 px with the hf3 / gpoly / tups
numerics).

**Accuracy (all gates pass; gate code unchanged).** Three steps change numerics (hf3, gpoly, tups); the other seven are
bit-identical (`bench_breakdown.py --dump/--cmp`, n_diff = 0 on eth12 float and on served u8 2-CQ).

| | round 4 / start (float ; u8) | **now, eth12 float** | **now, served u8** | `none` baseline |
|---|---|---|---|---|
| randn head1 / head2 (all), gate 0.998 | 0.99845 / 0.99868 (0.99887) | 0.99843 / 0.99872 (0.99893) | | 0.99836 / 0.99895 (0.99909) |
| test_mast3r end_to_end | 0.99844 PASS | 0.9988 (h1 0.9987 / h2 0.9986) PASS | | 0.9984 |
| real pair pts3d h1 / h2, gate 0.989 | 0.99065 / 0.99056 ; 0.99046 / 0.99080 | 0.99023 / 0.99074 | 0.99047 / 0.99085 | 0.98988 / 0.99076 |
| real pair conf h1 / h2, gate 0.99 | **0.99149** / 0.99433 ; 0.99232 / 0.99478 | **0.99209** / 0.99509 | 0.99188 / 0.99483 | 0.99237 / 0.99439 |
| real pair median abs(dz)/z h1 / h2 | 1.26 / 1.14 % ; 1.29 / 1.18 % | 1.26 / 1.12 % | 1.17 / 1.09 % | 1.3 % |
| raw PCC vs the half-pixel fp32 reference h1 / h2 | 0.999828 / 0.999682 ; 0.999873 / 0.999748 | 0.999863 / 0.999798 | 0.999878 / 0.999770 | |
| synthetic u8 pair (not gated) | 0.99589 / 0.99929 | | 0.99690 / 0.99939 (0.99851) | 0.99752 / 0.99934 |
| `sym` out_ii / out_ji / out_jj vs two-pass | bit-identical | | bit-identical | |

The thin float conf-h1 margin of rounds 3-4 (0.99149) widened to 0.99209. The u8 conf h1 moved the other way (0.99232 ->
0.99188, still 0.0019 above the gate). pts3d h1 float moved 0.99065 -> 0.99023 (gate 0.989), while the median depth errors
improved (float 1.26 / 1.14 -> 1.26 / 1.12 %, u8 1.29 / 1.18 -> 1.17 / 1.09 %). For scale: switching the round-4 graph to
ttnn's exact erf GELU (`all,-gelut`, slower) gives randn 0.99855 / 0.99844, real-pair float conf h1 0.99233 and pts3d
0.99039 / 0.99074. That is the same spread, so these 1e-4-level moves look like bf16 rounding variation on one pair, not a trend.

### Round-5 steps (A/B in one session; trace = eth12 1-CQ median unless noted)

| commit | knob | change | effect |
|---|---|---|---|
| e4005d2 | (tests) | host mock tests: `fake_ttnn` gets uint8 / BufferType / `to_torch(cq_id)`, the mock graph runs `MAST3R_OPT=none` (the optimised graph runs model-local generic_op kernels a shape-only fake cannot execute; it is device-gated) | `test_fused_host.py` 27/32 -> 32/32 pass |
| 463160f | hf3 | HiFi3 (fp32 acc) instead of HiFi4 for the two output-resolution head convs: head.0 / head.2 phase convs and their ring-strip convs. HiFi3 only drops the lo x lo mantissa partial product. Not bit-identical (n_diff 6-7 %, PCC vs previous 0.99998); real-pair metrics equal to 4-5 digits (conf h1 0.99149 -> 0.99150) | 28.72 -> 28.32; served trace 28.92 -> 28.48, e2e 31.61 -> 31.04 |
| feb7859 | gpoly | **fc1 GELU epilogue as a model-local erf-GELU polynomial.** An identity epilogue through the same pack-side SFPU path costs nothing (enc fc1 100.5 us vs 198 us with ttnn's gelu_tanh), so the epilogue IS bound by the SFPU instruction count. The round-3 conclusion was wrong: its exp-based variant was simply not shorter. New: gelu(x) = 0.5 x + abs(x) s q(s^2), s = min(abs(x), 4.25), q a 9-coefficient weighted-LP minimax fit of (Phi(s) - 0.5)/s with s q(s^2) = 0.5 at s = 4.25 (`tools_prof/gelu_fit.py`). It goes in the GELU_TANH slot of a model-local copy of ttnn's matmul compute kernel (`tt/kernels/mm_gelu_compute.cpp`, `mm_gelu_activation.hpp`, `mast3r_gelu_poly.h`); the program descriptor is otherwise unchanged. It approximates the reference's **exact erf GELU** with max abs error 6.2e-5; the tanh form gelut is 4.7e-4 away. After bf16 rounding, mean abs error vs erf GELU at sigma 1 / 2 / 4: this kernel 5.655e-4 / 1.126e-3 / 2.252e-3; gelut 5.804e-4 / 1.174e-3 / 2.297e-3; correctly rounded erf 5.648e-4 / 1.122e-3 / 2.246e-3. In the graph: encoder fc1 188 -> 143 us, decoder fc1 pairs 128 -> ~96 us | 28.35 -> 26.98; served trace 27.16, e2e 29.70 |
| 81a44aa | (gpoly) | two dst rows per SFPU step: each coefficient is loaded once (SFPLOADI pair) for both Horner chains and the two MAD chains interleave | 26.98 -> 26.56, bit-identical |
| e70a7ad | qkx | decoder: n_b = LN(x_b) feeds branch b's qkv AND the other branch's cross-attention k/v (norm_y(x_b) == n_b with the affine folded), so each branch runs ONE linear with concatenated weights [qkv_b, ckv_other] (N 2304 + 1536, same in0_block_w 4 -> bit-identical); the self-attention RoPE reads at row stride 5*Ht, the cross RoPE reads k/v from the other branch's output. One dual program per step instead of two | 26.6 -> 26.45 (3 A/B pairs), bit-identical |
| 36c2b6a | (served path) | host-side `event_synchronize` before each CQ1 readback instead of a device-side `wait_for_event` on CQ1. `tools_prof/read_slow_probe2.py`: a pending CQ1 wait alone slows the CQ0 trace by 0.3-0.5 ms (enqueue -> done: none 27.47, wait0 27.81, wait2 28.00 ms); device-side wait + 3 reads 28.92 vs host-side 28.52 ms. `MAST3R_HOSTWAIT=0` restores the old path | served e2e 29.66 -> 29.48 ms (5 of 6 A/B pairs, 60 iterations each), bit-identical |
| add4ceb | sdbl | ring-strip convs (head.0 / head.2 on the 6- / 8-row strips) with act + weight double buffering, height-sharded (`strip_conv_probe.py`: 68 -> 60 and 74 -> 69 us per strip conv) | served e2e -0.15 to -0.33 ms (3 A/B pairs), bit-identical |
| fd9b448 | pilhs | the phase0 interleave kernel writes the head.2 phase convs' HEIGHT_SHARDED input spec directly: one NoC write per 32-row unit into its shard (`tt/kernels/phase_il_rm_writer_hs.cpp`). Removes the ~21 us interleaved-to-sharded copy per head | -0.03 to -0.21 ms (3 A/B pairs), bit-identical |
| e8f518c | tups | **DPT refinenet bilinear x2 upsample (32^2 -> 64^2, 64^2 -> 128^2) TILE -> TILE in one model-local kernel** (`tt/kernels/ups2_*.cpp`). Each output tile is a sum of <= 4 tile matmuls with 12 constant interpolation tiles (half-pixel, clamped == ttnn.upsample; all entries exact in bf16). HiFi4 + fp32 dest, so products and sums are exact and there is one bf16 rounding. It replaces untilize + I2S + halo + upsample + S2I + tilize (`tups_probe.py`: 64^2 x 256 136 -> 43 us, 32^2 108 -> 20 us). Not bit-identical: the stock upsample rounds in a bf16 dest, so 23 % of its outputs differ from the correctly rounded value; for this kernel it is 3.5 %. All gates equal or slightly better (above) | -0.34 to -0.57 ms (2 A/B pairs, 60 iterations); served 26.28 / e2e 28.76 ms |
| 740fb2f | tapm | the 4 DPT tap projections (DRAM-latency-bound 1x1 linears, 19-22 us each) as ONE program on disjoint column blocks (6 / 3 / 2 / 1 grid columns for ap3 / ap2 / ap1 / ap0; ttnn's 2-D mcast descriptors with allowed_worker_cores, merged; dlin's in0_block_w) | 79 -> 44 us per head (device profile), bit-identical |
| 2a5f1ed | (tups) | the 16^2 -> 32^2 upsample too (W = 16: two image rows per input tile row, `ups2h_*` kernels with 4 constant tiles): replaces 7 small programs | min -0.12 ms (2 A/B pairs); gates unchanged |
| a4a2051 | sgat | ring-strip inputs gathered straight from refinenet1's height shards by a model-local data-movement kernel (`tt/kernels/strip_gather.cpp`, 32-byte face-row reads, lr already transposed). The gathered strips are the carry for the strip segment. This replaces the 8 MB S2I to DRAM + untilize + 4 slices + 2 concats + permute | eth12 -0.06 to -0.10 ms (2 pairs), served trace -0.15 to -0.40 ms (3 pairs), bit-identical |
| e70a7ad..2a5f1ed | (serve) | `tt-model.yaml`: knob list + verify list for the new model-local kernels (mm_gelu_compute.cpp, mm_gelu_activation.hpp, mast3r_gelu_poly.h, phase_il_rm_writer_hs, ups2_*, ups2h_*, strip_gather) | verify command passes on the host |

### Round-5 tried and rejected / measured only
- LayerNorm math fidelity HiFi3 / HiFi2 (the 84 add+LN programs, ~1.95 ms on 64 cores): no time change (28.35 / 28.31 /
  28.21 ms within noise) and the outputs change; the LN is not bound by its multiplies. Not kept.
- hrope2 (q and k of a unit batched per RoPE stage, bit-identical; `tools_prof/rejected_hrope2/`): within noise in 3 A/B pairs.
- Phase-conv configs (`tools_prof/phase_cfg_probe.py`, HiFi3): act_block_h 32-128 and act / weight double buffering are
  bit-identical but within ~4 us per conv. Auto act_block_h for the head.0 phase convs: 0.0-0.1 ms in the graph (noise).
  Merging the head.0 phase convs into 2 x 256 or 1 x 512 output channels (`p0_merge_probe.py`): 16 % faster at 32^2, but at
  128^2 every merged variant throws the L1 / CB clash.
- cdbl for the layer_rn convs (in_c != out_c, `MAST3R_FDBL` A/B): bit-identical, no gain in 3 pairs.
- gpoly variants (`rejected_gelu/gelu_fast_probe.py` MAST3R_GELU_EXP): 7 / 8 / 9 coefficients give enc fc1 145 / 148 / 155 us
  (one row per step). 9 is kept for accuracy; 8 has max error 1.1e-4 and would be about -0.07 ms. Hoisting 4 coefficients
  into LREGs makes the SFPU compiler run out of registers. The full-graph 8-coefficient variant had randn 0.99835 / 0.99846;
  the 9-coefficient one 0.99838 / 0.99868.
- fc1 program config sweep with the gpoly epilogue (`tools_prof/fc1_gpoly_sweep.py`): the current configs stay best.
- Served tail: with the readbacks overlapped, each 2 MB CQ1 read still costs the CQ0 trace ~0.5-0.7 ms
  (`read_slow_probe.py`), and a read after the trace costs ~1.36 ms. Only the device-side event wait was avoidable (above).

### Remaining backlog (device 25.9 ms, kernel sum ~23.9 ms, ~710 programs)
- Add + LayerNorm, ~1.95 ms (48 x 23 us encoder, 36 x 20 us decoder) on 64 cores. There are only 64 tile rows, every row must
  stay on one core for a bit-identical reduction, and 16-row face splits give 128 units for 120 cores. A width split needs a
  cross-core (semaphore) reduction and changes the summation order. Estimated -0.4 to -0.6 ms, with accuracy and hang risk.
- The 128^2 add3s reads the upsampled x from DRAM; writing x into the add3s shard spec would save about 0.05 ms but keeps
  8 MB in L1 across the resconv convs (CB risk).
- head.2 tail: the 4 per-phase 1x1 linears write 87 % padding (4 valid of 32 columns each), and tailf re-reads 16 MB.
  A fused phase-lin + tail kernel would save about 0.1 ms but needs the 4 phase-conv outputs live together in L1.
- Dispatch gaps: ~1.9 ms (trace 25.9 vs kernel sum ~23.9 ms).
- Served tail: ~2.5 ms (prep 0.6, H2D 0.6, overlapped readback slowdown ~1.0, last read). It is bounded by the x1 link and
  the CQ0 slowdown during CQ1 reads.

### vs RTX 5090 (GPU_COMPARISON.md; like-for-like = the same network, bf16, batch 1, one pair)
- Module in bf16 + SDPA, 35.3 ms incl_h2d (34.3 ms device-only):
  - TT served 28.9 ms: 18 % faster.
  - TT eth12 1-CQ 31.0 ms: 12 % faster.
  - Device only: TT 25.9 vs GPU 34.3 ms, 25 % faster.
- bf16 module 42.3 ms: TT served is 32 % faster.
- torch.compile + CUDA graphs, 21.1 ms (20.1 ms device): still out of reach. GPU is 1.37x faster (served) / 1.29x
  (device only).
- Pose (`sym`), reported separately as mode-specific: 39.4 ms vs 2 x 35.3 = 70.6 ms for two GPU forwards.
- gpoly evaluates the same erf GELU the GPU reference runs, more accurately than the earlier tanh form. tups computes the
  same half-pixel bilinear upsample the port always used, with exact products. Neither is an algorithmic shortcut, so both
  stay in the like-for-like comparison.

## Round 4 summary (2026-10-03, chip 9, HEAD 8af42e2)

Round start = c9679bd (verified: eth12 trace 30.45 / 30.23, e2e 35.71 / 35.49; served trace 30.55, e2e 33.16 / 32.81).
Final numbers: one `code/tools_prof/r3_validate.sh 9` session, 07:29-07:33 UTC, 30-iteration median / min. The script now
takes the chip and the checkout as arguments: `r3_validate.sh [CHIP] [CHECKOUT_DIR]`.

| | `none` baseline (eth12, 1 CQ, float) | round start c9679bd (verified) | **now: eth12, 1 CQ, float** | now: worker 2 CQ, float | **now: served (worker 2 CQ, uint8)** |
|---|---|---|---|---|---|
| device forward (trace, synchronized) | 62.20 / 62.12 | 30.45 / 30.23 (served 30.55 / 30.34) | **28.81 / 28.62** | 28.91 / 28.72 | 28.90 / 28.75 |
| host prep / H2D | 0.61 / 1.04 | 0.93 / 1.02 | 0.87 / 1.04 | 0.90 / 1.07 | 0.57 / 0.57 |
| D2H, synchronized | 20.89 | 3.59 | 3.20 | 3.24 | 3.12 (overlapped in e2e) |
| **e2e** `model(img1, img2)` | 92.72 / 84.50 | 35.71 / 35.49 (served 33.16 / 32.81) | **33.98 / 33.09** | 32.05 / 31.83 | **31.99 / 31.48** |
| `test_mast3r.py --layer end_to_end` (2 CQ, float) | | 33.67 | | 31.07 ms, PCC 0.99844 PASS | |
| real pair: single pair / `sym` pose / two-pass pose e2e | | 32.89 / 46.22 / 65.99 | | | **31.49 / 43.46 / 62.66** |
| eager profile: programs / kernel sum | 1252 / 60.97 | 905 / 28.48 | 803 / 27.20 (at 20b64b5; pcat/p1hs remove ~50 more) | | |

Device -5.4 % this round (30.45 -> 28.81 ms eth12; 30.55 -> 28.90 ms served), -54 % overall vs `none`. The served e2e moves
33.2 -> 31.4-32.0 ms between same-day runs (07:26 A/B run 31.44 / 31.07; validation run 31.99 / 31.48). That is -3.5 % to
-5 %. The eth12 1-CQ e2e moves 35.7 -> 34.0 ms (-4.8 %). Server smoke test (worker dispatch, 2 CQ, `MAST3R_OPT=all`,
chip 9, at 8af42e2): PASS x4. Pose rot is 43.3 deg and f1/f2 are 443/438, the same as round 3. `forward` is 43-45 ms per pose
request (round 3: 46-48 ms).

**Accuracy: every round-4 step is bit-identical.** Each step was A/B'd with `bench_breakdown.py --dump/--cmp` on the eth12
float path and on the served uint8 2-CQ path; the `n_diff` count was 0 for both heads. The gated numbers equal round 3's to
every printed digit:
- randn head1 / head2 / all: 0.99845 / 0.99868 / 0.99887 (gate 0.998).
- test_mast3r end_to_end: PCC 0.99844, PASS.
- Real pair, float: pts3d 0.99065 / 0.99056, conf 0.99149 / 0.99433, median dz/z 1.26 / 1.14 %, raw PCC vs the
  half-pixel reference 0.999828 / 0.999682.
- Real pair, served u8: pts3d 0.99046 / 0.99080, conf 0.99232 / 0.99478, median dz/z 1.29 / 1.18 %.
- Synthetic u8 (not gated): 0.99589 / 0.99929.
- `sym`: out_ii, out_ji and out_jj bit-identical to the two-pass path.

The thin conf h1 float margin (0.99149 vs the 0.99 gate) is unchanged. This round did not touch numerics.

### Round-4 steps (A/B in one session; trace = eth12 1-CQ median unless noted)

| commit | knob | change | effect |
|---|---|---|---|
| 12205a0 | (tooling) | `r3_validate.sh [CHIP] [CHECKOUT]` | fixes the verifier's "hardcoded chip 9 / main checkout" note |
| 27f59fc | rchain | DPT resConfUnit conv1 -> conv2 on ttnn's **L1 conv path**. conv1's sharded output feeds conv2 directly, which removes an S2I to DRAM + I2S per resconv and the DRAM-path NHWC reshape copy at 16^2. At >= 64^2 the relu input is height-sharded into the conv's own spec and freed after halo (that is the only way 128^2 fits the L1 CB budget). Below 64^2 it lives in L1 interleaved. Per-resconv probe (`rchain_probe.py`): 128^2 376 -> 322 us, 64^2 148 -> 128, 16^2 76 -> 47 | 30.46 -> 30.05 ms |
| 6850a3b | dlin | refinenet out_conv (1x1) with its input in L1, and output in L1 below 128^2. Uses an explicit full-grid config with the auto config's in0_block_w and compute config, so every output tile sees the same K sequence (`dlin_probe.py`: 16384x256x256 96 -> 30 us) | 30.06 -> 29.86 |
| bd8aaef | dfront | DPT tap projections write L1, so the ConvTranspose / ap3_down / layer_rn convs use the L1 conv path (removes 16^2 reshape copies). The ap1 ConvTranspose output goes to DRAM: l2_rn on its sharded output picks another config and is not bit-identical | 29.96 -> 29.82 |
| ab3928d | dshard | >= 64^2: the refinenet adds read conv2's **height shards in place** and write relu(s) straight into the next conv1's shard spec. Model-local `add3s` / `add2` kernels, same FPU adds. The leading relu writes the shard spec directly (`add3s_probe.py` at 128^2: S2I + add3 + I2S 156 -> 72 us, S2I + add 77 -> 29 us) | 29.76 -> 29.34; served e2e 31.98 |
| cb12ce3 | ups1 | upsample output staged in L1 before the tilize | 29.36 -> 29.27 (2 A/B pairs) |
| 108afa4 | tailf | head tail after the 4 per-phase 1x1s as ONE model-local program: phase sum + bias + WH transpose (compute) and the 16-bit NCHW interleave (writer). Replaces 3 adds + transpose + bias add + slice + untilize + tail_il (`tailf_probe.py`: 113 -> 55 us per head) | 29.27 -> 29.18 (3 A/B pairs) |
| 20b64b5 | pemm | patch-embed linear with an explicit config (auto in0_block_w) writing the L1 residual directly (drops the 4 MB copy) | 29.17 -> 29.12 |
| ba4971d | pcat | attention output projections (encoder proj, decoder self / cross proj) **read the SDPA output directly**. Built from ttnn's 2-D mcast matmul descriptor with the in0 sender reader swapped for a model-local copy that remaps tile ids (`tt/kernels/mm_in0_heads_reader.cpp`), so 48 concat-heads programs go away (`pcat_probe.py`: enc 39.9 -> 33.3 us, dec pair 31.1 -> 26.7 us) | 29.16 -> 28.90 |
| 194c1a0 | p1hs | refinenet1's out_conv writes head.0's height-shard spec directly; the DRAM copy for the strips is one S2I. dlin uses the auto 1-D layout for the N=96 tap projection | 28.91 -> 28.73 |
| 8af42e2 | serve | tt-model.yaml: knob list + verify list for the new kernels (add3s/add2/tailf/mm_in0_heads) | |

### Round-4 tried and rejected
- **thalf** (served path): head 2's main rows in two trace segments (rows 2i | 2i+1, tailf in half mode, exact), so the
  first 1 MB readback overlaps the second half. In 3 A/B pairs the e2e did not move (31.60-31.98 vs 31.68-32.03 ms).
  The `e2e_tail_probe.py` / `e2e_timeline2.py` timelines show why. With no reads, the device is done at 30.3 ms
  (including prep + H2D). With concurrent CQ1 readbacks the last trace ends about 1 ms later, so overlapped D2H on
  the x1 link slows the device instead of hiding behind it.
- **strip1**: both heads' ring strips in one last segment with one readback. e2e 31.66 -> 32.17 ms (worse).
- r1 out_conv output in L1 (instead of DRAM): CB clash in the 2-CQ served path (p1 is carried across trace segments).
- Device im2col for the uint8 upload: the host copy is only ~0.2 ms either way (measured), so it does not pay.
- Decoder 4-way cq+ckv (or qkv+ckv) merge: not done. These pairs already run at 120-160 TFLOP/s on their halves, so a merge
  saves only one ~2 us dispatch per step (~25 us per forward).

### Remaining backlog (device 28.8 ms, kernel sum ~27 ms, ~750 programs)
- Encoder / decoder fc1 GELU epilogue: about 100 us x 24 + 70 us x 12, about 3.1 ms. This is the largest item.
  It is SFPU-bound on the pack thread. Only an accuracy-gated activation change could cut it (round-3 rejection notes).
- LayerNorm programs (add + LN) run on 64 cores, because there are only 64 tile rows. That is 48 x 23 us (enc) +
  36 x 20 us (dec), about 1.8 ms. A width-split LN with a cross-core reduction could reach ~120 cores, but it is not
  bit-identical. Estimated -0.5 ms.
- DPT ring strips: about 0.25 ms per head across ~25 programs. Halos at 23-25 us come from 6-row, 256-wide strips on
  96 cores. Possible fixes: a transposed or width-sharded layout, or one batched program. Estimated -0.2 ms.
- DPT 128^2 plumbing: tilize 35 us, relu 23 us, add3s 70 us (DRAM-bound) per head. Keeping the upsampled x in L1 would
  need the refinenet split moved (do u1's convs before the upsample). Estimated -0.1 ms.
- Dispatch gaps: about 2.0 ms (2.5 us x ~750 programs).
- Host mock tests (`code/models/tests/test_fused_host.py`): 5 tests already fail at c9679bd and still fail. `fake_ttnn`
  lacks `uint8` and the newer generic_op paths. These tests are not device gates, but they are stale.

### vs RTX 5090 (GPU_COMPARISON.md)
- bf16 + SDPA module 35.3 ms (34.3 device-only): TT served 31.4-32.0 ms (9-11 % faster); eth12 1-CQ 34.0 ms (4 % faster,
  previously a tie); device-only TT 28.8 ms vs 34.3 ms (16 % faster).
- bf16 module 42.3 ms: TT served 24-26 % faster.
- torch.compile + CUDA graphs 21.1 ms (20.1 device): still out of reach. GPU is 1.5x faster.
- Pose (`sym`): 43.5 ms vs 2 x 35.3 = 70.6 ms for two GPU forwards (mode-specific; reported separately).

## Round 3 summary (2026-10-03, chip 9, HEAD aae6921)

Round start = b1f4930, the verified round-2 result. All numbers below are 30-iteration median / min from one session
(`code/tools_prof/r3_validate.sh`, 2026-10-03 06:21-06:28 UTC), except the round-start column, which was measured at
the start of this session (04:28 UTC served path; the eth12 round-start trace/e2e is the `all,-sdpa2` A/B run of 04:4x,
i.e. the b1f4930 graph).

| | baseline `MAST3R_OPT=none` (eth12, 1 CQ, float) | round start b1f4930 | **now: eth12, 1 CQ, float** | now: worker 12x10, 2 CQ, float | **now: served path (worker 12x10, 2 CQ, uint8)** |
|---|---|---|---|---|---|
| device forward (trace, synchronized) | 62.18 / 62.08 | 39.73 / 39.57 (eth12); 39.93 / 39.69 (served) | **30.44 / 30.21** | 30.49 / 30.35 | 30.59 / 30.41 |
| host prep | 0.81 | | 0.88 | 0.80 | 0.62 |
| H2D | 1.05 | | 1.04 | 1.05 | 0.59 |
| D2H, synchronized, incl. host assembly | 20.63 | | 3.59 | 2.96 | 3.76 (overlapped in e2e) |
| **e2e** `model(img1, img2)` | 92.67 / 83.87 | 45.16 / 44.56 (eth12); 42.43 / 42.09 (served) | **35.60 / 35.09** (previous run 35.53 / 34.78) | 33.77 / 33.41 | **33.42 / 33.13** (previous run 33.39 / 32.91) |
| real pair (kitchen, `sym_check.py`) single-pair e2e | | | | | 33.01 / 32.57 |
| return_pose (`sym` B-mode), real pair e2e | | 60.41 / 60.07 (round 2) | | | **46.01 / 45.62** (two-pass: 65.77 / 65.07) |
| `test_mast3r.py --layer end_to_end` (2 CQ, float) | | 42.69 (round-2 verification) | | 33.58 ms | |
| eager profile: programs / kernel sum | 1252 / 60.97 | 1232 / 37.64 | 925 / 28.68 ms (at 3627397; dadd/addln2 remove ~22 more) | | |

Device 39.7 -> 30.4 ms (-23 %) this round, 62.2 -> 30.4 ms (-51 %) overall. e2e (served path) 42.4 -> 33.4 ms (-21 %); eth12 1-CQ like-for-like 45.2 -> 35.6 ms (-21 %). Host-side phases (prep, D2H, e2e) move by up to +-0.5 ms between runs on the shared host.

### Accuracy (all gates pass; gate code in `test_mast3r.py` unchanged since the baseline commit)

| | round 2 (b1f4930) float / u8 | **now, eth12 float** | **now, served u8** | `none` baseline |
|---|---|---|---|---|
| randn pair head1 / head2 (all) — gate 0.998 | 0.99846 / 0.99870 (0.99890) | 0.99845 / 0.99868 (0.99887) | n/a | 0.99836 / 0.99895 (0.99909) |
| test_mast3r end_to_end PCC (h1 / h2) | 0.9984 (0.9984 / 0.9980) | 0.9984 (0.9985 / 0.9982), PASS | | 0.9984 |
| real pair pts3d PCC h1 / h2 — gate 0.989 | 0.99046 / 0.99081 ; 0.99043 / 0.99091 | 0.99065 / 0.99056 | 0.99046 / 0.99080 | 0.98988 / 0.99076 |
| real pair conf PCC h1 / h2 — gate 0.99 | 0.99182 / 0.99459 ; 0.99220 / 0.99481 | 0.99149 / 0.99433 | 0.99232 / 0.99478 | 0.99237 / 0.99439 |
| real pair median abs(dz)/z h1 / h2 | 1.13 / 1.18 % ; 1.20 / 1.21 % | 1.26 / 1.14 % | 1.29 / 1.18 % | 1.3 % |
| real pair raw PCC vs half-pixel fp32 ref h1 / h2 | 0.999837 / 0.999849 ; 0.999879 / 0.999884 | 0.999828 / 0.999682 | 0.999873 / 0.999748 | |
| synthetic smooth uint8 pair head1 / head2 (all) (not a gate) | ; 0.99711 / 0.99929 (0.99855) | | 0.99589 / 0.99929 (0.99740) | 0.99752 / 0.99934 (0.99900) |
| sym B-mode out_ii / out_ji / out_jj vs two-pass | bit-identical | | bit-identical | |

Honest notes on accuracy:
- Most round-3 steps are bit-identical (hrope, hcat, mln, addln, resl1d, cdbl, pil, til, dmm, the common-args change; each
  A/B'd with `bench_breakdown.py --dump/--cmp`). The numerics changed only through `sdpa2` (larger SDPA chunks: fewer
  online-softmax rescales), `mc2d`/`mc2dr` (2-D mcast linears instead of minimal_matmul / the fused dit residual op) and
  `mm32` (fp32 dest accumulation in those linears).
- `mc2d`/`mc2dr` alone (bf16 dest, larger in0_block_w) worsened the served-path real pair: conf PCC h1 0.99071, median
  depth error 1.43 %. Probing showed that a larger in0_block_w accumulates more K tiles in the 16-bit dest
  (`tools_prof/resid_acc_probe.py`: mean |err| vs fp64 9.1e-3 / 13.1e-3 at in0_block_w 8 / 16 vs 7.8e-3 for the old dit op).
  `mm32` (fp32 dest, out subblocks <= 4) gives 6.0e-3 — more accurate than the round-2 matmuls — at the same speed, and it
  restored the real-pair gates (conf h1 0.99232). It is part of the default set.
- Remaining drift vs round 2: the real-pair median relative depth error is 1.26 / 1.14 % (float) and 1.29 / 1.18 % (u8),
  vs 1.13 / 1.18 % and 1.20 / 1.21 % in round 2, and vs 1.3 % for the `none` baseline. Head-2's raw PCC vs the half-pixel
  reference fell from 0.99985-0.99988 to 0.99968-0.99975. These metrics moved by a similar amount (±0.1 % median, ±0.0002 PCC)
  under every individual numeric knob in A/B runs (e.g. SDPA k-chunk 256 vs 512, or removing sdpa2), so they look like
  bf16-rounding variation on one image pair rather than a systematic loss, but they are reported as measured.
- The synthetic smooth uint8 pair (not a gate) moved 0.99711 -> 0.99589 on head 1. Its value jumps non-monotonically
  with the knob set (0.99604 without the mc2d family, 0.99680 without mm32, 0.99675 without sdpa2, 0.99711 with only the
  round-2 knobs); `none` gives 0.99752.

### vs RTX 5090 (GPU_COMPARISON.md; "incl_h2d" = upload of both views + forward + readback of both maps)

| GPU variant | GPU incl_h2d (excl_h2d) ms | TT served path (2 CQ, uint8) | TT eth12 1 CQ float (like-for-like p150 config) |
|---|---|---|---|
| fp32 strict | 103.4 | 33.4 (TT 3.1x faster) | 35.5 (2.9x) |
| bf16 autocast | 63.8 | 33.4 (1.9x) | 35.5 (1.8x) |
| module converted to bf16 | 42.3 (41.3) | **33.4 (TT 21 % faster)** | **35.6 (TT 16 % faster)** |
| module bf16 + SDPA | 35.3 (34.3) | **33.4 (TT 5 % faster; min 33.1)** | 35.5-35.6 / min 34.8-35.1 (median 0.6-0.8 % slower, min 0.6-1.5 % faster: a tie) |
| bf16 module + torch.compile + CUDA graphs | 21.1 (20.1) | 33.4 (GPU 1.6x faster) | 35.5 |
| device only (TT trace vs GPU excl_h2d) | 34.3 (bf16 + SDPA) | 30.6 (TT 11 % faster) | 30.4 (TT 11 % faster) |

Caveats: the GPU rows upload fp32 views (the TT served path uploads uint8, 1.5 MB); with uint8 input the GPU would save
part of its ~1 ms of transfers, so the bf16 + SDPA comparison is a near tie in transfer-neutral terms (device-only:
TT 30.4-30.6 vs GPU 34.3 ms). The compiled-graph GPU row (21.1 ms) stays out of reach at the gated precision.
Pose requests: 46.0 ms with the `sym` B-mode vs 2 x 35.3 = 70.6 ms for two GPU forwards (mode-specific saving; reported
separately). Served-like totals (PNG decode + npz encode, ~180 ms of host work) are host-bound on both accelerators.

Server smoke test (worker dispatch, 2 CQ, `MAST3R_OPT=all`, chip 9, at 3627397 and again at aae6921): PASS x4 each,
ready 6.5 s after start (warm-up captures both traces), `forward=46-48 ms` per pose request (the `sym` graph), pose rot
43.3 deg, f1/f2 443/438 in both runs.

### Round-3 steps (each kept step committed; A/B in one session; bit-identical unless noted)

| commit | knob | change | effect (trace, eth12 1 CQ) |
|---|---|---|---|
| 5f2e4cc | sdpa2 | per-shape SDPA chunks balanced for 120 cores (flat B*H*q-chunk scheduling): encoder q160/k512 (7 q-chunks per head -> 224 chunks <= 2 per core), decoder q224/k512 (5 per head -> 120 chunks, 1 per core); the kernel pads the partial last chunk. Probe (`sdpa_probe3.py`): 133 -> 94 us, 93 -> 67 us | 39.73 -> 38.69; numerics change (randn 0.99851 / 0.99867) |
| 8668258 | hrope | **model-local generic_op kernel** (`tt/kernels/heads_rope_*.cpp`): split heads + RoPE(q, k) in one program on 120 cores, reading the per-branch qkv / q / kv matmul outputs directly (also removes the decoder's batch concats); the compute kernel runs rotary_embedding_llama's exact op sequence. Probe: 72.4 -> 26.5 us (enc), 74.3 -> 24.2 / 75.3 -> 24.0 us (dec self / cross). Also: the encoder's L1 ones vector became a per-pass copy, so no L1 buffer sits under the DPT conv CBs | 38.33 -> 35.81 |
| 2f6e591 | hcat | model-local concat-heads kernel (pure data movement, 120 cores); the decoder writes the two branch outputs directly (no batch slices). 13.5 -> 9.6 us enc, 19.4 -> 8.3 us dec | 35.81 -> 35.50 |
| 3852783 | mln | the decoder's branch LayerNorm pairs as ONE program: ttnn's own LN kernels with the exact compile-time config of its factory, via generic_op, each core's runtime args pointing at its branch's tensor (64 cores instead of 2 x 32). 30.9 -> 21.3 us per pair | 35.50 -> 35.10 |
| caad9ac | mc2d, mc2dr | plain and residual linears as full-grid 2-D mcast `ttnn.linear` with per-shape in0_block_w (8 / 16 for K = 768 / 3072 / 4096) instead of minimal_matmul / the fused dit op (residual as a separate add): dec 768x768 25 -> 12.4 us, fc2 56 -> 33 us, enc fc2 115 -> 85 us | 35.10 -> 34.33 (mc2d) -> 32.97; numerics change, see mm32 |
| 795536d | resl1d | decoder residual streams in L1 (block-5/8 taps copied to DRAM, dec_norm out in DRAM) | 32.97 -> 32.52 |
| eae74ab | addln | every residual add that feeds a LayerNorm runs inside the LN program: ttnn's FUSE_PRE_ADD reader + a model-local copy of layernorm.cpp that also packs the bf16 sum (+ a model-local writer); cb_x kept bf16 so the LN sees the same bf16 sum ttnn.add produces (bit-identical). Dec pair 41.9 -> 24.6 us, enc 54.1 -> 31.2 us | 32.52 -> 32.06 |
| d1fb6ac | cdbl | DPT 3x3 convs <= 64^2: activation + weight double buffering (64^2 height-sharded, 32^2 / 16^2 block-sharded): 96 -> 76, 51 -> 40, 45 -> 39 us | 32.06 -> 31.83 |
| 644b5c5 | pil | phase0 interleave by a model-local data-movement kernel on ROW_MAJOR pixel rows (the head.0 phase convs emit ROW_MAJOR) instead of concat + transpose + 0/1 permutation matmul + transpose: 30 us vs about 200 us per head | 31.83 -> 31.53 |
| ed25302 | til | head-tail phase interleave by a model-local kernel (16-bit element interleave straight into the NCHW rows) instead of RM reshape (65 us) + permute + reshape + tilize + 0/1 matmul + untilize: 35 us vs about 130 us per head | 31.53 -> 31.29 |
| 784b1d2 | mm32 | mc2d / mc2dr linears with fp32 dest accumulation (out subblocks <= 4) | same speed; restores accuracy (above) |
| 2c106b6 | dmm | the decoder's same-shape branch linears (qkv, proj, cq, ckv, cproj, fc1+GELU, fc2) as ONE program each: ttnn's own 2-D mcast matmul descriptors (`MatmulMultiCoreReuseMcast2DProgramFactory.create_descriptor`) placed on the top / bottom 12x5 halves of the grid with `allowed_worker_cores`, merged with `ttnn.merge_program_descriptors`, run through generic_op. Pairs: qkv 60 -> 51, 768x768 30 -> 25, ckv 44 -> 37, fc1 154 -> 132, fc2 72 -> 63 us | 31.44 -> 30.64 |
| 3627397 | (hrope) | shared addresses as common runtime args | no change |
| 9082ce4 | dadd | DPT refinenet with skip: x + (skip + c2) and the next unit's leading relu as one model-local program (FPU adds in ttnn.add's order + SFPU relu): 128^2 165 -> 104, 64^2 56 -> 36, 32^2 21 -> 11 us | 30.65 -> 30.49 |
| aae6921 | addln2 | encoder / decoder last add + enc_norm / dec_norm (affine) as the fused add + LN program with gamma/beta (stock FUSE_GAMMA/BETA path); decoder block-5/8 taps copied from the next step's fused sums | 30.49 -> 30.43 |

### Round-3 tried and rejected

- Faster GELU epilogue (x / (1 + exp(-2u)), exp_21f or fp32-accurate exp + Newton reciprocal) in a model-local copy of the
  fc1 matmul compute kernel, built from ttnn's own matmul descriptor (`tools_prof/rejected_gelu/`): 191-207 us vs 194-210 us
  stock; the GELU epilogue (about 96 of 198 us in the encoder fc1) is not bound by the SFPU instruction count. The exp_21f
  variant was also less accurate (up to 1.8 half-ulp on x in [-5, 1)).
- Separate GELU op after a plain fc1: 102 + 110 us vs 198 us fused.
- fc1 out-subblock / in0_block_w / grid sweep (`fc1_sub_probe.py`): best = current.
- Merging the 4 head-tail phase convs into one 128 -> 512 conv: fits L1 only at 64^2 (129 -> 95 us there); at 256^2 the
  height-sharded variants throw a CB/L1 clash and the block-sharded ones are 5-8x slower.
- Double buffering for the phase / phase0 convs and for the 128^2 refinenet convs: no gain.
- TILE-layout phase0 interleave reading 32-byte face rows over the NoC: 144 us (vs 30 us for the ROW_MAJOR kernel).
- SDPA q64 / q32 / q96 / q192 / q256 / q512 and k1024 variants: slower (`sdpa_probe3.py`).
- minimal_matmul grid / block sweep for the 768-wide decoder linears (`mm_grid_sweep.py`): at most 2-3 us; superseded by mc2d.
- Half-grid decoder linears split by columns (6x10 halves): only 2-4 us per pair; the row split (12x5) is what `dmm` uses.

### Megakernel / parameter taxonomy status (round 3)

- No parameter classification changed: A (launch-fixed) = resolution 512, batch 1, weights, `MAST3R_CQS`, uint8 input,
  the knob set; B = `return_pose` (precompiled `sym` graph, captured at warm-up); C = pixel contents.
- Fusion: 1232 -> about 903 programs per forward. Model-local generic_op kernels: heads_rope, heads_concat, ln_add
  (copy of layernorm.cpp + writer), phase_il_rm, tail_il, add3; plus ttnn's own LN and matmul kernels re-launched through
  generic_op with per-core tensor bindings / merged disjoint-core programs (mln, addln, dmm). Still one trace per forward segment, one H2D, readbacks overlapped on CQ1.
- The per-program dispatch gap is now about 2.1 us (trace 30.65 vs kernel sum 28.68 ms over 925 programs): fewer programs
  remain the main lever for the gaps.

### Remaining backlog (round-3 view; device 30.7 ms, kernel sum 28.7 ms)

- Encoder fc1 + GELU_TANH: 188 us x 24 = 4.5 ms; the GELU epilogue is ~96 us of it and is not reducible with fewer SFPU
  instructions (above). Only a different, accuracy-gated activation form could cut it.
- DPT: the remaining resconv add (sum + c2' before out_conv) could fold into the out_conv input path; conv in/out
  reshards (I2S / S2I / halo, 1.8 ms total) are bound by the stock conv L1 CB sizes.
- Head-tail post-conv chain (4 x 128->4 linears, 3 adds, transpose, bias add, slice, untilize: about 125 us per head):
  a fused kernel could save about 0.2 ms.
- e2e tail (served path): e2e - trace = 2.6 ms, of which about 1 ms is host prep + H2D and the rest is mostly the 2 MB
  head-2 main-row readback that the remaining strip work (about 0.9 ms) cannot fully hide on the x1 link.
- Compiled-GPU parity (21.1 ms) is not reachable at the gated precision.

## Round 2 summary (2026-10-03, chip 9, HEAD f6cbac0)

Round start = aff19e9, the verified round-1 result. Same-session numbers, all 30-iteration median / min, measured
2026-10-03 04:13-04:16 UTC.

| | baseline `MAST3R_OPT=none` (eth12, 1 CQ, float) | round start aff19e9 (eth12, 1 CQ, float) | now: eth12, 1 CQ, float | now: worker 12x10, 2 CQ, float | **now: served path (worker 12x10, 2 CQ, uint8)** |
|---|---|---|---|---|---|
| device forward (sum of the trace segments, synchronized) | 62.22 / 62.10 | 40.01 / 39.81 | **39.64 / 39.37** | 39.78 / 39.44 | 39.78 / 39.51 |
| host prep | 0.86 | 0.94 | 0.98 | 0.88 | **0.46** |
| H2D | 1.07 | 1.04 | 1.03 | 1.07 | **0.60** (1.5 MB uint8) |
| D2H, synchronized, both maps incl. host assembly | 29.00 | 3.34 | 3.09 | 3.09 | 3.01 (hidden in e2e) |
| **e2e** `model(img1, img2)` | 92.15 / 91.89 | 45.06 / 44.44 | 44.57 / 44.06 | 43.22 / 42.72 | **42.01 / 41.84** |
| real pair (kitchen, `tools_prof/sym_check.py`) e2e | | | | | **42.03 / 41.54** |
| return_pose (two maps + swapped pass), real pair e2e | | 84.14 / 83.73 (two-pass, today's graph) | | | **60.41 / 60.07** (`sym` B-mode) |
| test_mast3r end_to_end (2 CQ, float) | | 44.79 | | 42.82 | |
| eager profile: programs / kernel sum | 1252 / 60.97 | 1230 / 38.07 | 1232 / 37.64 | | |

Accuracy, all gates pass:

| | round start / float path now | uint8 served path now |
|---|---|---|
| synthetic randn pair head1 / head2 (all) | 0.99846 / 0.99870 (0.99890); float outputs are **bit-identical** to the round start | n/a (randn is not a uint8 image) |
| synthetic smooth uint8 pair head1 / head2 (all) | float path on the same pixels 0.99667 / 0.99937 (0.99856); `MAST3R_OPT=none` 0.99752 / 0.99934 (0.99900) | 0.99711 / 0.99929 (0.99855) |
| real pair pts3d PCC h1 / h2 | 0.99046 / 0.99081 | 0.99043 / 0.99091 |
| real pair conf PCC h1 / h2 | 0.99182 / 0.99459 | 0.99220 / 0.99481 |
| real pair median abs(dz)/z h1 / h2 | 1.13% / 1.18% | 1.20% / 1.21% |
| real pair raw PCC vs half-pixel fp32 reference h1 / h2 | 0.999837 / 0.999849 | **0.999879 / 0.999884** |
| test_mast3r end_to_end | 0.9984 (h1 0.9984 / h2 0.9980) PASS | |

Honest notes:
- The uint8 path changes the arithmetic. The input is exact (v - 127.5), and 1/127.5 is folded into the patch-embed weight in fp64 with a single bf16 rounding.
  Against the half-pixel fp32 reference it is closer than the float path. Against the align-corners reference, its median relative depth error is 1.20 / 1.21% (float path 1.13 / 1.18%); the PCCs are equal or higher.
- On the smooth synthetic uint8 pair, every variant is below 0.998 on head 1, including the `none` baseline at 0.99752. The 0.998 gate is defined on the randn pair, where the float path is unchanged.
- The served-path e2e uses Tensix dispatch with 2 CQs. On the Galaxy that is 12x10, like ETH dispatch. On a real p150 that mode has 11x10 unless ETH dispatch supports 2 CQs there; the Galaxy ETH path has only 2 idle ETH cores, so it is 1 CQ.
  The like-for-like eth12 1-CQ float number is 44.57 / 44.06 ms.
- The host is shared by about 11 agents, and e2e medians moved by up to ±0.5 ms between runs in this round. The same-session A/B pairs are in the step table.

### vs RTX 5090 (GPU_COMPARISON.md; the GPU "incl_h2d" row = upload of both views + forward + readback of both maps)

| GPU variant | GPU ms | TT served path now (worker 2 CQ, uint8) | TT eth12 1 CQ float now |
|---|---|---|---|
| fp32 strict | 103.4 | 42.0 (TT 2.5x faster) | 44.6 |
| tf32 | 70.8 | 42.0 (1.7x) | 44.6 |
| bf16 autocast | 63.8 | 42.0 (1.5x) | 44.6 |
| fp16 autocast | 56.6 | 42.0 (1.35x) | 44.6 |
| **module converted to bf16** | **42.3** | **42.0 (TT 0.7% faster; real pair 42.03 / 41.54 min)** | 44.6 (GPU 5% faster) |
| module bf16 + SDPA | 35.3 | 42.0 (GPU 1.19x faster) | |
| bf16 module + torch.compile + CUDA graphs | 21.1 | 42.0 (GPU 2x faster) | |
| device-only (GPU excl_h2d bf16 module 41.3) | 41.3 | device 39.8 (TT faster) | 39.6 |

The served-path TT number uses the server's uint8 views (1.5 MB upload) and readbacks that overlap compute on a second CQ.
The GPU row uploads fp32 views; a GPU could also take uint8, which would save it about 0.5 ms of its 1.1 ms of transfers.
So the bf16-module win is narrow (0.3 ms median), and it does not hold for the 1-CQ ETH configuration.
Pose requests (return_pose) take 60.4 ms with the `sym` B-mode, against 2 x 42.3 = 84.6 ms for the GPU's two forwards. This is a
mode-specific algorithmic saving (the encoder runs once and an unused head is skipped), so it is reported separately.

### Round-2 steps (each kept step committed; A/B in one session where host noise matters)

| commit | knob | change | effect |
|---|---|---|---|
| 72b75eb | `u8` (with hostcol) | uint8 pixel views. Host im2col is uint8 (1.5 MB instead of 3 MB). The device does typecast (exact) and v - 127.5 (exact in bf16). W/127.5 is folded into the patch-embed weight in fp64. The server preprocess returns uint8 HWC views (`preprocess_image(..., uint8=True)`), and warm-up uses a uint8 dummy | H2D 1.04 -> 0.58, host prep 0.94 -> 0.76 ms |
| 72b75eb | `MAST3R_CQS=2` | forward captured as 2 trace segments (A: encoder, decoder, head 1; B: head 2). An event after A, then a CQ1 read of head 1 while CQ0 runs B | e2e 45.06 -> 43.95 (float) / 43.05 (u8); outputs bit-identical |
| 167a332 | u8 | (py, px, c) im2col row order for uint8 (48-byte runs from the HWC array; weight rows permuted to match) | host_input 0.40 -> 0.18 ms, e2e 43.05 -> 42.68 |
| 3d850ff | phase | 0/1 permutation matmuls at HiFi3. This is exact: the data operand sits in srcB and is fully covered by HiFi3, as probed with `perm_mm_probe.py`. Phase0 strip inputs are untilized into L1 | bit-identical, trace 39.96 -> 39.88 |
| 2fa1785 | ropes | RoPE trans_mat HEIGHT_SHARDED, one tile per core on all 120 cores (per-pass L1 copy), which selects the prefill-sharded rotary_embedding_llama factory. Probe: 26.7 -> 21.3 us (enc), 22.5 -> 18.0 us (dec) | **bit-identical, trace 39.88 -> 39.49 ms** (RoPE total 1.96 -> 1.61 ms) |
| 440121c | sym (B-mode) | return_pose: one symmetric graph. The encoder runs once (it is per-view, so pass-2 inputs are the pass-1 encoder outputs), the decoder runs for both orders, then DPT head 1 + head 2 of (1,2) and head 1 of (2,1). Head 2 of (2,1) is skipped because PairViewer never reads it. Captured at server warm-up | out_ii / out_ji / out_jj bit-identical to the two-pass path; pose e2e 84.1 -> 60.4 ms; served forward 86 -> 60 ms (smoke test PASS, same pose / focals) |
| d087a52, f6cbac0 | tsplit (needs 2 CQ) | 4 segments: enc+dec+head-1 main rows, then head-2 main rows, then head-2 ring strips, then head-1 ring strips. Each 2 MB main-row readback overlaps later compute; only head 1's 128 KB strips are read after the last trace. Host assembly is incremental | bit-identical; e2e 42.49 / 42.58 -> 42.08 / 42.03 ms (same-session A/B) |
| f6a7dc5 | serve | `MAST3R_CQS=2` pinned in tt-model.yaml and SERVING.md | server smoke test PASS (forward 42-43 ms, pose 60 ms) |

### Round-2 tried and rejected

- RoPE math fidelity HiFi2 / HiFi3: no trace gain (RoPE is data-movement bound) and the outputs change. Reverted.
- cos/sin HEIGHT_SHARDED (with or without trans_mat), and RoPE with heads folded into batch: slower (`rope_probe.py`).
- Decoder fc1 with transposed-mcast 11x10 (110 cores) and encoder fc1 transposed: slower (99 vs 80 us, 214 vs 198 us; `fc1_probe.py`).
- MinimalMatmul K_block 2 for the decoder M=1024 linears and encoder proj (`mm_sweep.py`: 2-4 us faster per op in isolation): no
  trace gain in the model (39.58 vs 39.49 ms), synthetic head2 PCC 0.99862 vs 0.99870. Reverted.
- DPT 1x1 linears with explicit 1-D/2-D configs (`lin1x1_probe.py`): at most -28 us (128^2) and the outputs change; not taken.
- 128^2 refinenet relu/add in L1: conv circular buffers clash with the L1 tensors (TT_THROW, no hang). Reverted.
- Preallocated host buffers for D2H (`copy_device_to_host_tensor`): it does not write through to the torch buffer; net -0.03 ms. Not taken.
- glibc malloc tunables against page faults in the host assembly: no change.


## Method (identical for every number below)

```bash
ROOT=/home/ttuser/experiments/tt-models; M=$ROOT/models/mast3r-p150; cd $M/code
source $ROOT/tools/chipenv.sh 9 $M/.venv >/dev/null
export TT_WEIGHTS_REVISION=61c57447d7b0adc8a1a30b2b0adec7a8935aa2a3
timeout -s INT 900 python bench_breakdown.py --mode eth12          # synthetic pair: PCC + per-phase timing (30 iters)
MAST3R_ETH=1 timeout -s INT 900 python tools_prof/real_pair_acc.py   # real pair accuracy gate
timeout -s INT 900 python test_mast3r.py --layer end_to_end --runs 10   # authors' gate (Tensix dispatch, 12x10)
MAST3R_TRACE=0 python -m tracy -r -p -v -o $TT_METAL_PROFILER_DIR/x --op-support-count 6000 tools_prof/prof_eager.py  # op profile
```
`trace` = `execute_trace` + `synchronize_device` of the whole forward (pure device time, one trace per forward);
`e2e` = the public `model(img1, img2)` call (host prep + H2D + trace + D2H incl. host output assembly).
The host is shared by ~11 agents, so host-side phases (host_prep, D2H, e2e) are noisy (D2H 2.4-4.1 ms run to run);
the device `trace` is stable to about +-0.2 ms. Median and min are reported.

Accuracy gates (from OPT_BASELINE.md, kept): synthetic e2e PCC >= 0.998; real pair pts3d PCC >= 0.989 per head; conf PCC >= 0.99.

## Round 1 results (chip 9, eth12, synthetic pair)

Final numbers measured 2026-10-03 02:45-02:52 UTC on HEAD f6b97f8 (two default runs, one `MAST3R_OPT=none` run in the same session).

| | baseline (`MAST3R_OPT=none`) | round start (8dc072d) | now (f6b97f8) |
|---|---|---|---|
| trace (device forward) median / min | 62.20 / 62.10 ms | 46.47 / 46.23 ms | **40.08 / 39.73 ms** (2nd run 40.08 / 39.82) |
| host prep | 0.83 | 0.84 | 0.88-0.98 |
| H2D | 1.06 | 1.03 | 1.03 |
| D2H (incl. host output assembly) | 28.00 (min 20.49) | 2.49 | 3.13-3.49 (min 2.62) |
| e2e median / min | 87.26 / 84.75 ms | 50.48 / 50.13 ms | **44.97-45.64 / 44.55 ms** |
| test_mast3r end_to_end best-of-10 (Tensix dispatch 12x10) | 84.49 ms | — | 44.87 ms |
| p150-equivalent (`--mode worker11`, Tensix dispatch, 11x10) trace / e2e | 62.62 / 89.16 | — | 41.41 / 47.07 ms |
| device programs per forward (eager profile) | 1252 | 1212 | 1230 |
| kernel-time sum (eager profile) | 60.97 ms | 44.91 ms | 38.02 ms |
| synthetic PCC head1 / head2 (all) | 0.99836 / 0.99895 (0.99909) | 0.99845 / 0.99871 | 0.99846 / 0.99870 (0.99890) |
| test_mast3r end_to_end PCC (4 digits) | 0.9984 | 0.9984 | 0.9984 (head1 0.9984 / head2 0.9980) |
| real pair pts3d PCC head1 / head2 | 0.98988 / 0.99076 | 0.99044 / 0.99079 | 0.99046 / 0.99081 |
| real pair conf PCC head1 / head2 | 0.99237 / 0.99439 | 0.99178 / 0.99457 | 0.99182 / 0.99459 |
| real pair raw PCC vs half-pixel-upsample fp32 ref | — | 0.999839 / 0.999847 | 0.999837 / 0.999849 |
| real pair median abs(dz)/z vs fp32 ref (vs half-pixel ref) | 1.3% | 1.05% / 1.11% | 1.13% / 1.18% (0.37% / 0.33%) |

All accuracy gates pass. Honest notes: the synthetic head2 PCC is 0.99870 vs 0.99895 at the baseline (head1 is higher,
0.99846 vs 0.99836); that shift came with the earlier sessions' gelut/lnfold steps, not this round. The real-pair median
relative depth error moved 1.05/1.11% -> 1.13/1.18% within this round (phase/phase0 change the head arithmetic: folded
kernels, no bf16 rounding of the upsampled activations); the PCCs are equal or higher.

Device: 62.2 -> 40.1 ms (-36%); this round 46.5 -> 40.1 ms (-14%). e2e: ~85-87 -> ~45 ms (-48%).
Device time by stage now (eager profile, kernel sums): patch-embed + encoder 16.4 ms, decoder (incl. embed) 11.9 ms,
DPT x2 9.8 ms; dispatch gaps ~2.0 ms (trace - kernel sum, ~1.6 us per program).

### vs RTX 5090 (GPU_COMPARISON.md, same pair size, forward incl. H2D + readback of both maps)

| GPU variant | GPU ms | TT now (e2e incl. H2D/D2H) | ratio TT/GPU |
|---|---|---|---|
| reference as shipped, fp32 strict | 103.4 | 45.0 | 0.44 (TT faster) |
| tf32 | 70.8 | 45.0 | 0.64 (TT faster) |
| bf16 autocast | 63.8 | 45.0 | 0.71 (TT faster) |
| fp16 autocast | 56.6 | 45.0 | 0.80 (TT faster) |
| module converted to bf16 | 42.3 | 45.0 | 1.06 (GPU faster) |
| module bf16 + SDPA | 35.3 | 45.0 | 1.27 (GPU faster) |
| bf16 module + torch.compile + CUDA graphs | 21.1 | 45.0 | 2.13 (GPU faster) |

The TT device forward alone (40.1 ms) is below the bf16-module GPU number but the x1-PCIe Galaxy transfers (H2D 1.0 +
D2H ~3 ms incl. host assembly) put e2e just above it. The compiled GPU graph (21 ms) is out of reach without lower
precision: LoFi matmuls were tried again this round per group (below) and fail the accuracy gates. Served-like e2e
(PNG decode + preprocess + forward + npz encode, ~180 ms of identical host work on both sides) was not re-measured;
it is host-bound on either accelerator (GPU 241-301 ms).

## Kept steps (cumulative, each measured with the method above)

Earlier sessions (before this round; numbers from their commits):

| commit | knob | change | trace |
|---|---|---|---|
| 0e84074 | out | compact channel-first head output, D2H 29 -> 3 ms | (e2e 89.7 -> 67.2) |
| c64edc6 | mm | per-shape 12x10 matmul configs, exact GELU fused into fc1 | 62.5 -> 57.4 |
| 3fe76db | dpt | refinenet out_conv before the upsample, RM conv in/out | 57.4 -> 55.9 |
| 4401161 | l1 | per-block transient activations in L1 | 55.9 -> 51.8 |
| 772e075 | dpt | height-sharded refinenet convs >= 64x64 | 51.8 -> 50.3 |
| 13363fc/30c038f | dptf | HiFi2 (fp32 acc) DPT convs except the two head convs | 50.3 -> 49.3 |
| f7587b7 | gelut | tanh-form GELU in the fc1 epilogue | 49.3 -> 47.9 |
| 316339b | lnfold | LayerNorm affine folded into the following linear | 47.9 -> 47.1 |
| 8dc072d | l1 | per-pass L1 copies of the RoPE tables | 47.1 -> 46.65 |

This round (re-validated HEAD 8dc072d first: trace 46.47, PCC 0.99845/0.99871, real pair gates pass):

| commit | knob | change | trace | notes |
|---|---|---|---|---|
| 206768d | phase | **Polyphase DPT head tail.** head.1-2 (bilinear x2, half-pixel, clamped == `ttnn.upsample`) + 3x3 conv == 4 phase 3x3 convs on the 256^2 grid with kernels K_ab = sum A[a] A[b] K (fp64 fold on host). Exact except the 2-px border ring, which is recomputed with the original upsample+conv on 2-row/2-col strips. The 512^2 x 128 upsample, its reshards and the 6-way DRAM-sliced conv are gone; head.4 (1x1, 128->4) is 4 tiny height-sharded linears summed. | 46.47 -> 42.67 | tail probe 3.21 -> 1.29 ms per head. D2H went up (host phase interleave) |
| ee66e6b | phase | phase interleave on device (RM permute + one exact 0/1 permutation matmul) -> NCHW rows; host only writes the ring | 42.67 -> 43.20 | D2H 4.5 -> ~2.7 ms, e2e -0.4..-2 ms |
| 683d128 | hostcol | host cast writes the patch-embed im2col layout `[2, N, 768]` bf16 directly (one strided copy, 0.14 ms, replaces cat+cast 0.97 ms); device drops reshape/permute im2col (0.7 ms) | 43.20 -> 42.46 | same 3 MB H2D |
| 82bbe3e | phase | full-grid 2-D mcast config for the interleave matmul (was 16 cores, 123 us) | 42.42 -> 42.22 | |
| c141100 | lnfold | decoder: norm1(x)==norm_y(x) once the affine is folded, so each decoder step normalises x1 and x2 once for both branch blocks (-24 LN) | 42.22 -> 41.93 | bit-identical |
| 07250e9 | phase0 | **head.0 polyphase too**: refinenet1's upsample folded into 4 phase convs at 128^2 (256->128); phases interleaved with an exact batched 0/1 matmul; final 6-px ring recomputed exactly by a thin-strip run of the original op sequence | 41.93 -> 41.51 | removes the 128->256 upsample + head.0's 3-way DRAM-sliced conv |
| ed59a76 | phase | permutation matmuls at HiFi4 without fp32 acc (bit-exact; LoFi/HiFi2 are not, `tools_prof/perm_mm_probe.py`) | 41.51 -> 41.45 | |
| 8c74544 | decb | decoder: the two branch blocks of a step share one B=2 split-heads / 2xRoPE / SDPA / concat-heads per attention (these ops cost ~the same at B=2 as at B=1: 28/21/12 vs 26/19/11 us); per-branch linears stay separate (a per-branch-weight bmm is 3-7x slower), outputs concatenated / sliced on batch | 41.45 -> 40.25 | same PCC |
| f6b97f8 | resl1 | encoder residual stream in L1 (patch-embed output copied to L1; residual-matmul outputs and their ones vector in L1; enc_norm output in DRAM) | 40.25 -> 40.02 | encoder LN 25.5 -> 20.0 us; same PCC |

Exactness of the polyphase fold: host check (fp64) interior max abs diff 2e-8; device output vs fp32 torch: PCC 0.999993
for both the old and the new tail (`tools_prof/phase_probe.py`); real-pair raw PCC vs the half-pixel reference unchanged.

## Tried and rejected (round 1)

- Row-sliced in-L1 tail (upsample + conv + 1x1 per row slice): correct but slower (3.5 vs 3.2 ms per head; the upsample itself was the cost).
- DPT resconv chain kept height-sharded in L1 (`dpts`): conv CBs of the 128^2x256 height-sharded convs (~1.2 MB, full weights per core) clash with any extra L1 tensor. Reverted.
- DPT relu/add in L1 interleaved for <= 64^2: -0.3 ms but changed PCC (conv picked a different config); not worth the accuracy churn. Reverted.
- Feeding the sharded upsample output straight into head.0's conv: CBs exceed L1 (2.4 MB). Superseded by phase0.
- Strip convs on fewer cores (16/32): no gain or CB clash.
- SDPA sweep on 12x10 (`tools_prof/sdpa_probe2.py`): q128/k256 HiFi2 (current) is the fastest; LoFi -2%, exp-approx 0%, fp32-acc +16%.
- LoFi for selected transformer matmul groups (experiment, not committed): decoder MLP only -0.25 ms but synthetic head1
  PCC 0.99679; encoder MLP -1.0 ms, head1 0.99327; both MLPs -1.1 ms, head1 0.98269. All fail the 0.998 gate.
- Explicit minimal_matmul patch embed with L1 output (instead of linear + copy): bf16 packer accumulation over 6 K blocks
  changed outputs (head2 0.99852); with fp32 acc the PCC rose (0.99853/0.99889) but the real-pair depth error moved
  (1.29%/1.22%). Kept the bit-identical linear + copy.
- `reallocate_halo_output=False` for the DPT convs (drop the Move ops): no gain.
- GELU: the fused GELU_TANH epilogue costs ~100 us (enc fc1) / ~40 us (dec fc1) per call (3.4 ms per forward); no cheaper
  accurate SFPU variant exists in this tt-metal (the approximate GELU was rejected earlier for accuracy, PCC 0.993).

## Megakernel / parameter taxonomy status

Round 2 changes:
- **A (launch-fixed):** `MAST3R_CQS` = 2 (trace segmentation + CQ1 readbacks), and the input dtype (uint8 in the server) is
  fixed at launch through `hostcol`.
- **B (`return_pose`):** now a precompiled second variant, the `sym` graph with 3-4 segments, captured at warm-up and
  selected per request. The other per-request fields remain C (pixel contents) or D (host only).
- **Per forward:** one H2D (1.5 MB uint8); trace segments replay back to back; each output is read on CQ1 as soon as its
  segment's event fires.

Round 1 text:

Per `param_taxonomy/mast3r-p150.md`: no request parameter changes device shapes (A: resolution 512, 1 pair, weights;
C: pixel contents; B: return_pose = run the same trace twice). The whole forward is one metal trace with one H2D
(3 MB im2col input) and two D2H (2 x 2.2 MB: NCHW maps + ring strips) per pair. Default knob set frozen as
`MAST3R_OPT=all` (pinned in `tt-model.yaml` serve.env, documented in SERVING.md).

## Remaining backlog (estimated device gain)

Round-2 view. Device time is now 39.6-39.8 ms; kernel sum 37.6 ms over 1232 programs.
- Phase0 ring strips: about 0.48 ms per head (two strip runs of about 25 small ops each). A cheaper exact border formulation, or
  asymmetric-padding convs that compute only the needed rows, could save about 0.3 ms in total.
- Tail phase-interleave chain: about 0.21 ms per head, incl. a 65 us RM reshape with a page-size change. An alternative
  layout for the 4 phase linears could remove the reshape and permute, about -0.2 ms in total.
- act_postprocess 1x1 + ConvTranspose fold (ap0 / ap1): about 190 us per head now. The fold is exact, but it needs a
  depth-to-space data movement (RM page-size change), so the net is uncertain, perhaps -0.2 ms.
- Decoder per-branch plumbing (concat 31 us/step, batch slices 14 us/step): blocked by per-branch weights. The bmm with
  per-branch weights is 3-7x slower.
- GELU_TANH epilogue (about 2.7 ms), SDPA (5.2 ms) and create-heads (1.4 ms) are at the stock-kernel floor at the gated precision.
- sym B-mode: branch-2 step 11 and dec_norm of the swapped pass are still computed (dead work, about 0.5 ms per pose request).

Round 1 list:

- Phase0 thin-strip pipeline: ~0.4 ms per head of small ops (halo/conv on tiny inputs); could be cut with a cheaper exact ring formulation.
- DPT refinenet plumbing at 64^2/128^2 (S2I/I2S per conv, DRAM relu/add): ~0.5 ms per head, blocked by conv L1 CB size.
- Dispatch gaps ~2 ms over ~1230 programs: fewer/fused small ops (DPT conv plumbing is ~4-5 programs per conv).
- GELU epilogue (3.4 ms) and SDPA (5.2 ms) are at the floor of the stock kernels at the gated precision; further
  device gains there need custom kernels (not done: tt-metal is frozen for this workflow).
- Decoder residual stream in L1: probe says ~1 us per LN, not worth it.
- `models/tests/test_fused_host.py`: 5 mock-graph tests fail since the earlier OPT commits (fake_ttnn lacks
  to_memory_config etc.); pre-existing, unchanged by this round.