File size: 29,990 Bytes
997f91e 5904b64 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 | # vggt-1b-p150 (VGGT-1B, 518x518, S=1 and S=2) -- Blackhole p150a vs RTX 5090 (same host, same weights, same input)
Date 2026-09-14. Facts only; every GPU number below was measured in this pass, every p150a number is copied (with its source line) from the validation / publish reports. The p150a was NOT touched.
## What was run
| | |
|---|---|
| Model | The port's own torch reference: `models/vggt-1b-p150/code/models/demos/vggt/reference/torch_vggt.py::load_vggt(eval_mode=True) -> upstream vggt.models.vggt.VGGT(enable_track=False) @ 44b3afbd1869d8bde4894dd8ea1e293112dd5eba (checkout /home/deepgadget/experiments/tt-models/logs/gpu-vs-p150/vggt-1b/vggt_upstream, byte-identical to the image's /opt/tt-metal/vggt): Aggregator (DINOv2 ViT-L/14 + 24 frame/global alternating blocks, dim 1024, 2D RoPE, F.scaled_dot_product_attention) + CameraHead + DPTHead x2; forward(images) -> dict`. 1190.6 M parameters (track head dropped, as the port and the server load it). This is the network the p150a port is PCC-gated against (`test_vggt.py`, the real-image set, the served A/B, the point-map study) |
| Weights | `facebook/VGGT-1B` @ `860abec7937da0a4c03c41d3c269c366e82abdf9` (tt-model.yaml `weights.revision` = `serve.env.TT_WEIGHTS_REVISION`), `model.safetensors` (5.03 GB fp32) from the HF cache `/home/deepgadget/.cache/huggingface/hub/models--facebook--VGGT-1B/snapshots/860abec7937da0a4c03c41d3c269c366e82abdf9/model.safetensors`, resolved by the port's `torch_vggt.resolve_weights_path`, `HF_HUB_OFFLINE=1` |
| Input | S=1: media/input.png (518x518 RGB PNG); S=2: media/source_1.png + media/source_2.png (kitchen 00/03, 779x520 RGB PNG) -> models.server.app.preprocess_pad (imported): longer side 518, shorter round(x/14)*14, PIL BICUBIC, /255, white pad to 518x518 -> (1, S, 3, 518, 518) fp32 in [0,1], batch 1 (== server/app.py::predict; the Aggregator normalises with ImageNet mean/std internally) |
| Per-view preprocess records | S1: ['media/input.png'] -> [{'orig_w': 518, 'orig_h': 518, 'resized_w': 518, 'resized_h': 518, 'scale_x': 1.0, 'scale_y': 1.0, 'pad_left': 0, 'pad_top': 0}]; S2: ['media/source_1.png', 'media/source_2.png'] -> [{'orig_w': 779, 'orig_h': 520, 'resized_w': 518, 'resized_h': 350, 'scale_x': 0.6649550706033376, 'scale_y': 0.6730769230769231, 'pad_left': 0, 'pad_top': 84}, {'orig_w': 779, 'orig_h': 520, 'resized_w': 518, 'resized_h': 350, 'scale_x': 0.6649550706033376, 'scale_y': 0.6730769230769231, 'pad_left': 0, 'pad_top': 84}] |
| Output | `pose_enc (1,S,9)`, `depth (1,S,518,518,1)`, `depth_conf (1,S,518,518)`, `world_points (1,S,518,518,3)`, `world_points_conf (1,S,518,518)` -- the five tensors `server/app.py` reads back from the port; readback 6.4 MB at S=1, 12.9 MB at S=2 |
| GPU | NVIDIA GeForce RTX 5090 (sm_120), driver 580.126.18, power limit 600.00 W, 32607 MiB; idle 31.2 W |
| venv | `/home/deepgadget/experiments/tt-models/.venv-gpu/main` -- Python 3.12.13, torch 2.11.0+cu128, CUDA 12.8, cuDNN 9.19.0, torchvision 0.26.0+cu128, numpy 1.26.4, safetensors 0.8.0, huggingface_hub 1.31.0, pillow 12.3.0, einops 0.8.2, triton 3.6.0; fastapi 0.141.1 / pydantic 2.13.5 added in this pass so `models.server.app` imports unchanged |
| Repo state | `models/vggt-1b-p150` @ `9586c93` (branch `tt-model-package`), read-only; `reference/torch_vggt.py` and `server/app.py` (`preprocess_pad`, `_decode_image`, `_b64_npz`) imported unchanged; upstream `vggt` package cloned to `logs/gpu-vs-p150/vggt-1b/vggt_upstream` @ `44b3afbd1869d8bde4894dd8ea1e293112dd5eba` (tt-model.yaml `source.extra_code` ref; `diff -rq` against the image's `/opt/tt-metal/vggt` copy: identical) |
| Script | `logs/gpu-vs-p150/vggt-1b/bench_vggt_gpu.py` (uses `logs/gpu-vs-p150/bench_common.py`); logs `full_run.log` (the run below), `smoke.log` (3-iteration dry run, incl. the CPU reference pass), `smoke2_lowprec.log` (dry run of the converted-module path after the fix below), `full_run_attempt1_bf16weights_dtype_error.log` (first full run: reference rows complete and within 1 ms of the run below, then aborted at the `bf16_weights` variant because upstream `DPTHead._apply_pos_embed` promotes a bf16 activation to fp32 -- fixed by casting that constant to `x.dtype` for the converted copies only, see `bench_vggt_gpu.py`); raw JSON `reports/gpu-vs-p150/vggt-1b.json` (= `logs/gpu-vs-p150/vggt-1b/result.json`, incl. every raw wall-clock sample); CPU reference tensors `cpu_fp32_reference.pt`; `render_report.py` renders this file from the JSON |
| Command | `cd logs/gpu-vs-p150/vggt-1b && HF_HUB_OFFLINE=1 /home/deepgadget/experiments/tt-models/.venv-gpu/main/bin/python bench_vggt_gpu.py --iters 50 --warmup 10 --skip-cpu > full_run.log 2>&1` (one process; exit 0) |
| Loop | per S, precision and variant 10 warm-ups + 50 timed iterations, `torch.cuda.synchronize()` before and after each; wall-clock (`perf_counter`) is the primary number, CUDA-event time recorded alongside |
| p150a source | `reports/gpu-vs-p150/p150_numbers.json` -> `reports/megakernel/POINTMAP_SUMMARY.md:89` (Hub `tt serve` after the demo republish: S=1 x30 forward **420.1 ms** median (414.1 min / 437.7 max), total **528.5**; S=2 x20 forward **887.3** (873.7 / 1021.8), total **1104.6**); `models/vggt-1b-p150/DEVICE_VALIDATION.md:411` (pre-publish served A/B, final default: S=1 411.4 / 407.9 / 435.0 forward, 525.5 total; S=2 880.6), `DEVICE_VALIDATION.md:416` (`encode` host npz 105-125 ms), `VALIDATION_SUMMARY.md:40`, `PUBLISH_SUMMARY.md:20` (419.9 / 881.0) |
Timing definitions (they match the p150a `timing_ms` keys of `server/app.py`):
- **incl_h2d** = `images.to("cuda")` (one `(1,S,3,518,518)` fp32 pageable host tensor, 3.2 MB per view, as the server preprocess produces) + forward + the five outputs `.float().cpu()`. Compare with p150a `timing_ms.forward` = `vggt_forward(images)`: host ImageNet-normalise + DINOv2 patch-embed prelude (fp32 torch), upload of the `(S, 1374, 1024)` token tensor, one whole-graph metal-trace replay (aggregator + camera head + per-frame DPT heads), readback of the raw head outputs, host `activate_head` / `activate_pose` (**420.1 ms** S=1, **887.3** S=2).
- **excl_h2d** = forward only, input already resident, outputs left on the device.
- **served-like** = base64 decode + PNG decode + `preprocess_pad` of every view + `torch.stack` (`preprocess`) + incl_h2d forward (`forward`) + `pose_encoding_to_extri_intri` + float16/float32 casts + `np.savez_compressed` + base64 of the npz (`encode`) = the interval `server/app.py` reports as `timing_ms.total` (HTTP/JSON framing is outside on both sides). Compare with p150a `timing_ms.total` (**528.5 ms** S=1, **1104.6** S=2, of which the `encode` npz stage is 105-125 ms of identical host work on the p150a host).
Precision rows: the upstream `VGGT.forward` wraps the camera head and both DPT heads in `torch.cuda.amp.autocast(enabled=False)`, so in the bf16 / fp16 autocast rows only the aggregator (DINOv2 + 48 alternating-attention blocks) runs in low precision and the heads run fp32 (with TF32 allowed). The `bf16_weights` / `fp16_weights` rows convert the whole module once (heads included, closer to the p150a's bf16-on-device class) and are informational: not the reference as shipped.
## Accuracy check (GPU vs the CPU fp32 reference)
CPU fp32 reference: the same module on the host (16 torch threads, per smoke.log): forward 5540 ms at S=1, 8988 ms at S=2 (`smoke.log`; the p150a point-map study quotes 7.9 s for torch CPU S=2 on the same host under load). PCC is Pearson correlation over the flattened tensor; the port's gate is the minimum over the five output keys.
| S | GPU precision / variant | pose_enc | depth | depth_conf | world_points | world_points_conf | min PCC | max abs diff depth | depth median (ref) |
|---|---|---:|---:|---:|---:|---:|---:|---:|---:|
| 1 | GPU fp32 strict (PCC check) | 1.0000000 | 1.0000000 | 1.0000000 | 1.0000000 | 1.0000000 | **1.0000000** | 4.09e-05 | 0.96509 (0.96509) |
| 1 | fp32 + TF32 ('high') | 1.0000000 | 0.9999991 | 0.9999995 | 0.9999968 | 0.9999997 | **0.9999968** | 1.16e-02 | 0.96508 (0.96509) |
| 1 | bf16 autocast (aggregator; heads fp32/TF32) | 0.9999999 | 0.9992216 | 0.9999176 | 0.9992729 | 0.9999892 | **0.9992216** | 2.54e-01 | 0.9636 (0.96509) |
| 1 | fp16 autocast (aggregator; heads fp32/TF32) | 1.0000000 | 0.9999975 | 0.9999985 | 0.9999957 | 0.9999994 | **0.9999957** | 1.52e-02 | 0.96493 (0.96509) |
| 1 | bf16_weights (module converted) | 0.9999980 | 0.9992063 | 0.9998382 | 0.9991676 | 0.9998973 | **0.9991676** | 2.79e-01 | 0.96484 (0.96509) |
| 1 | fp16_weights (module converted) | 0.9999998 | 0.9999607 | 0.9999485 | 0.9999338 | 0.9999979 | **0.9999338** | 8.07e-02 | 0.96533 (0.96509) |
| 1 | torch.compile fp16_autocast+default | 1.0000000 | 0.9999938 | 0.9999991 | 0.9999885 | 0.9999994 | **0.9999885** | 2.78e-02 | 0.96497 (0.96509) |
| 1 | torch.compile fp16_autocast+reduce-overhead | 1.0000000 | 0.9999938 | 0.9999991 | 0.9999885 | 0.9999994 | **0.9999885** | 2.78e-02 | 0.96497 (0.96509) |
| 1 | torch.compile bf16_weights+reduce-overhead | 0.9999954 | 0.9996907 | 0.9998696 | 0.9995378 | 0.9999276 | **0.9995378** | 1.95e-01 | 0.96094 (0.96509) |
| 2 | GPU fp32 strict (PCC check) | 1.0000000 | 1.0000000 | 1.0000000 | 1.0000000 | 1.0000000 | **1.0000000** | 4.49e-05 | 1.0232 (1.0232) |
| 2 | fp32 + TF32 ('high') | 1.0000000 | 0.9999990 | 0.9999997 | 0.9999963 | 0.9999997 | **0.9999963** | 1.62e-02 | 1.02322 (1.0232) |
| 2 | bf16 autocast (aggregator; heads fp32/TF32) | 0.9999999 | 0.9994821 | 0.9999570 | 0.9993312 | 0.9999758 | **0.9993312** | 2.83e-01 | 1.02244 (1.0232) |
| 2 | fp16 autocast (aggregator; heads fp32/TF32) | 1.0000000 | 0.9999936 | 0.9999981 | 0.9999855 | 0.9999993 | **0.9999855** | 3.77e-02 | 1.02336 (1.0232) |
| 2 | bf16_weights (module converted) | 0.9999892 | 0.9992768 | 0.9996268 | 0.9992485 | 0.9998776 | **0.9992485** | 2.53e-01 | 1.02344 (1.0232) |
| 2 | fp16_weights (module converted) | 0.9999999 | 0.9999718 | 0.9999807 | 0.9999426 | 0.9999974 | **0.9999426** | 8.18e-02 | 1.02344 (1.0232) |
| 2 | torch.compile fp16_autocast+default | 1.0000000 | 0.9999945 | 0.9999986 | 0.9999878 | 0.9999994 | **0.9999878** | 3.48e-02 | 1.02302 (1.0232) |
| 2 | torch.compile fp16_autocast+reduce-overhead | 1.0000000 | 0.9999945 | 0.9999986 | 0.9999878 | 0.9999994 | **0.9999878** | 3.48e-02 | 1.02302 (1.0232) |
| 2 | torch.compile bf16_weights+reduce-overhead | 0.9999959 | 0.9996145 | 0.9997712 | 0.9995052 | 0.9998895 | **0.9995052** | 1.67e-01 | 1.02344 (1.0232) |
PCC gate (GPU fp32 strict vs CPU fp32, > 0.999): **PASS** -- min PCC 1.0000000 (S=1) / 1.0000000 (S=2); max |d| pose_enc 4.8e-07 / 4.9e-07, world_points 6.4e-04 / 3.2e-04. For scale: the p150a served S=2 output vs the same torch fp32 reference is depth 0.99927 / world_points 0.99911 / pose_enc 0.9999994 (POINTMAP_SUMMARY.md:69), synthetic min-PCC 0.9952 (S=1) / 0.9981 (S=2) (VALIDATION_SUMMARY.md:40).
## GPU timing (median / min / p90 of 50 iterations after 10 warm-ups, ms)
Model load (`load_vggt` incl. the 5.0 GB safetensors read from the page cache + `.cuda()`): 5.312 s. First call (fp32 strict, incl. cuDNN/SDPA autotune): 390.3 ms (S=1) / 256.8 ms (S=2).
### Reference as shipped
| S | precision | incl_h2d wall med / min / p90 | excl_h2d wall med / min / p90 | excl_h2d CUDA-event med | first call under this precision | GPU power mean (max) excl loop | peak mem alloc (reserved) MiB |
|---|---|---:|---:|---:|---:|---:|---:|
| 1 | fp32 strict (no TF32) | 134.07 / 133.78 / 140.52 | 133.24 / 132.90 / 139.45 | 133.23 | 132.66 | 524.3 (526.0) | 5420 (6354) |
| 1 | fp32 + TF32 ('high') | 93.04 / 92.90 / 98.83 | 92.22 / 92.07 / 97.49 | 92.20 | 98.49 | 402.8 (407.6) | 5420 (6354) |
| 1 | bf16 autocast (aggregator; heads fp32/TF32) | 65.58 / 65.44 / 69.13 | 64.64 / 64.50 / 69.36 | 64.63 | 64.88 | 378.6 (387.1) | 5420 (6354) |
| 1 | fp16 autocast (aggregator; heads fp32/TF32) | 63.32 / 63.15 / 67.65 | 62.43 / 62.23 / 67.74 | 62.41 | 66.63 | 422.0 (425.7) | 5420 (6354) |
| 2 | fp32 strict (no TF32) | 244.97 / 243.95 / 251.15 | 246.11 / 242.75 / 249.88 | 245.51 | 247.86 | 517.4 (520.4) | 5815 (6682) |
| 2 | fp32 + TF32 ('high') | 166.20 / 165.96 / 172.72 | 170.75 / 164.70 / 179.88 | 170.72 | 170.88 | 453.5 (461.3) | 5815 (6682) |
| 2 | bf16 autocast (aggregator; heads fp32/TF32) | 103.37 / 102.94 / 109.45 | 101.66 / 101.44 / 107.60 | 101.64 | 107.3 | 434.4 (436.1) | 5816 (6682) |
| 2 | fp16 autocast (aggregator; heads fp32/TF32) | 95.72 / 95.55 / 101.50 | 94.21 / 94.01 / 100.41 | 94.20 | 99.73 | 489.9 (496.6) | 5816 (6682) |
Pinned (page-locked) host input, bf16 autocast, incl_h2d median: S=1 65.49 ms, S=2 102.91 ms (vs pageable above; informational).
### Module converted once to bf16 / fp16 (informational; heads included)
| S | variant | incl_h2d wall med / min / p90 | excl_h2d wall med / min / p90 | min PCC vs fp32 ref | power mean W | peak mem MiB |
|---|---|---:|---:|---:|---:|---:|
| 1 | bf16_weights | 54.47 / 54.30 / 55.67 | 53.69 / 53.59 / 56.66 | 0.9991676 | 379.0 | 7497 |
| 2 | bf16_weights | 84.59 / 84.43 / 90.62 | 83.10 / 82.95 / 89.06 | 0.9992485 | 415.0 | 7698 |
| 1 | fp16_weights | 52.32 / 52.16 / 53.15 | 51.46 / 51.24 / 56.89 | 0.9999338 | 409.7 | 9779 |
| 2 | fp16_weights | 76.80 / 76.60 / 82.50 | 75.42 / 75.14 / 81.09 | 0.9999426 | 482.3 | 9981 |
### torch.compile (informational)
torch.compile(dynamic=False) of the as-shipped module; best autocast precision by S=1 excl_h2d median with min PCC >= 0.99 -> fp16_autocast; one compile per (S, mode); budget 300 s per compile (a variant whose compile exceeds it is recorded and the remaining variants are skipped)
| S | compiled variant | compile (first call) s | incl_h2d wall med / min / p90 | excl_h2d wall med / min / p90 | min PCC | power mean W | peak mem MiB |
|---|---|---:|---:|---:|---:|---:|---:|
| 1 | fp16_autocast+default | 26.5 | 83.63 / 83.32 / 88.85 | 82.63 / 82.47 / 88.44 | 0.9999885 | 339.2 | 9735 |
| 2 | fp16_autocast+default | 23.1 | 111.58 / 111.40 / 117.77 | 109.92 / 109.66 / 116.32 | 0.9999878 | 400.4 | 10466 |
| 1 | fp16_autocast+reduce-overhead | 17.9 | 82.58 / 82.15 / 87.70 | 81.61 / 81.15 / 87.49 | 0.9999885 | 341.3 | 9373 |
| 2 | fp16_autocast+reduce-overhead | 18.2 | 111.56 / 110.51 / 117.38 | 109.00 / 108.74 / 115.80 | 0.9999878 | 401.1 | 9636 |
| 1 | bf16_weights+reduce-overhead | 22.1 | 74.23 / 73.29 / 79.15 | 73.12 / 72.50 / 78.68 | 0.9995378 | 317.0 | 9243 |
| 2 | bf16_weights+reduce-overhead | 20.7 | 103.49 / 102.77 / 109.45 | 101.34 / 100.96 / 107.55 | 0.9995052 | 333.0 | 9379 |
### Served-like loop (the `server/app.py::predict` host stages around the GPU forward, npz float16, all four dense keys)
| S | precision | preprocess med | forward (incl_h2d) med | encode med | **total med / min / p90** | power mean W | npz base64 bytes |
|---|---|---:|---:|---:|---:|---:|---:|
| 1 | fp32 strict (no TF32) | 4.31 | 133.59 | 109.68 | **247.65 / 246.71 / 253.79** | 322.6 | 4578776 |
| 1 | fp32 + TF32 ('high') | 4.30 | 93.48 | 110.09 | **208.00 / 206.88 / 213.78** | 233.5 | 4581076 |
| 1 | bf16 autocast (aggregator; heads fp32/TF32) | 4.32 | 65.89 | 109.98 | **180.34 / 179.32 / 182.55** | 192.4 | 4580536 |
| 1 | fp16 autocast (aggregator; heads fp32/TF32) | 4.30 | 63.47 | 110.20 | **178.04 / 177.25 / 180.62** | 197.2 | 4579212 |
| 2 | fp32 strict (no TF32) | 20.94 | 248.68 | 223.44 | **493.07 / 486.43 / 495.52** | 306.3 | 9303936 |
| 2 | fp32 + TF32 ('high') | 21.44 | 166.48 | 226.01 | **414.50 / 412.68 / 420.39** | 234.6 | 9310592 |
| 2 | bf16 autocast (aggregator; heads fp32/TF32) | 20.95 | 103.86 | 224.14 | **349.11 / 348.21 / 355.11** | 190.1 | 9309164 |
| 2 | fp16 autocast (aggregator; heads fp32/TF32) | 20.97 | 96.49 | 224.77 | **342.27 / 341.07 / 349.02** | 201.2 | 9310104 |
`preprocess` here is PNG decode + PIL bicubic resize + pad; `encode` is the camera conversion + `np.savez_compressed` (zlib) of ~6.4 MB (S=1) / 12.9 MB (S=2) of dense arrays + base64 -- pure host work, single-threaded numpy/zlib, identical in kind to the p150a server's `encode` (105-125 ms at S=1 per DEVICE_VALIDATION.md:416).
## Comparison with the p150a (matching definitions)
Ratio = p150a ms / GPU ms (> 1 means the GPU is faster). p150a precision: bf16 weights and matmul inputs on device, fp32 residual stream / scores / softmax, HiFi4 + fp32 accumulation on proj / fc2 / DPT convs, exact fp32 eltwise bilinear interpolation (`VGGT_FUSED_INTERP=exact`), matmul attention (SDPA not default), one whole-graph metal trace per S (`TT_FUSED=1`, tt-metal 0.65.2.dev9100 @ 8b98410e7). GPU precision per row as stated.
| S | row | p150a ms (definition) | GPU variant / precision | GPU ms | ratio p150a/GPU |
|---|---|---:|---|---:|---:|
| 1 | device forward (p150a `timing_ms.forward` vs GPU incl_h2d) | 420.1 (POINTMAP_SUMMARY.md:89 Hub tt serve, DEVICE_VALIDATION.md:411 411.4) | ref, fp32 strict (no TF32) | 134.07 | **3.13** |
| 1 | | | ref, fp32 + TF32 ('high') | 93.04 | **4.52** |
| 1 | | | ref, bf16 autocast (aggregator; heads fp32/TF32) | 65.58 | **6.41** |
| 1 | | | ref, fp16 autocast (aggregator; heads fp32/TF32) | 63.32 | **6.63** |
| 1 | | | bf16_weights (module converted, eager; informational) | 54.47 | 7.71 |
| 1 | | | fp16_weights (module converted, eager; informational) | 52.32 | 8.03 |
| 1 | | | torch.compile fp16_autocast+default (informational) | 83.63 | 5.02 |
| 1 | | | torch.compile fp16_autocast+reduce-overhead (informational) | 82.58 | 5.09 |
| 1 | | | torch.compile bf16_weights+reduce-overhead (informational) | 74.23 | 5.66 |
| 1 | GPU forward only (excl_h2d, no PCIe) vs the same p150a number (which cannot exclude its transfers) | 420.1 | ref, fp32 strict (no TF32) | 133.24 | 3.15 |
| 1 | | | ref, fp32 + TF32 ('high') | 92.22 | 4.56 |
| 1 | | | ref, bf16 autocast (aggregator; heads fp32/TF32) | 64.64 | 6.50 |
| 1 | | | ref, fp16 autocast (aggregator; heads fp32/TF32) | 62.43 | 6.73 |
| 1 | served e2e (p150a `timing_ms.total` vs GPU served-like total, same host stages, npz float16) | 528.5 (POINTMAP_SUMMARY.md:89) | ref, fp32 strict (no TF32) | 247.65 | **2.13** |
| 1 | | | ref, fp32 + TF32 ('high') | 208.00 | **2.54** |
| 1 | | | ref, bf16 autocast (aggregator; heads fp32/TF32) | 180.34 | **2.93** |
| 1 | | | ref, fp16 autocast (aggregator; heads fp32/TF32) | 178.04 | **2.97** |
| 2 | device forward (p150a `timing_ms.forward` vs GPU incl_h2d) | 887.3 (POINTMAP_SUMMARY.md:89 Hub tt serve, DEVICE_VALIDATION.md:411 880.6) | ref, fp32 strict (no TF32) | 244.97 | **3.62** |
| 2 | | | ref, fp32 + TF32 ('high') | 166.20 | **5.34** |
| 2 | | | ref, bf16 autocast (aggregator; heads fp32/TF32) | 103.37 | **8.58** |
| 2 | | | ref, fp16 autocast (aggregator; heads fp32/TF32) | 95.72 | **9.27** |
| 2 | | | bf16_weights (module converted, eager; informational) | 84.59 | 10.49 |
| 2 | | | fp16_weights (module converted, eager; informational) | 76.80 | 11.55 |
| 2 | | | torch.compile fp16_autocast+default (informational) | 111.58 | 7.95 |
| 2 | | | torch.compile fp16_autocast+reduce-overhead (informational) | 111.56 | 7.95 |
| 2 | | | torch.compile bf16_weights+reduce-overhead (informational) | 103.49 | 8.57 |
| 2 | GPU forward only (excl_h2d, no PCIe) vs the same p150a number (which cannot exclude its transfers) | 887.3 | ref, fp32 strict (no TF32) | 246.11 | 3.61 |
| 2 | | | ref, fp32 + TF32 ('high') | 170.75 | 5.20 |
| 2 | | | ref, bf16 autocast (aggregator; heads fp32/TF32) | 101.66 | 8.73 |
| 2 | | | ref, fp16 autocast (aggregator; heads fp32/TF32) | 94.21 | 9.42 |
| 2 | served e2e (p150a `timing_ms.total` vs GPU served-like total, same host stages, npz float16) | 1104.6 (POINTMAP_SUMMARY.md:89) | ref, fp32 strict (no TF32) | 493.07 | **2.24** |
| 2 | | | ref, fp32 + TF32 ('high') | 414.50 | **2.66** |
| 2 | | | ref, bf16 autocast (aggregator; heads fp32/TF32) | 349.11 | **3.16** |
| 2 | | | ref, fp16 autocast (aggregator; heads fp32/TF32) | 342.27 | **3.23** |
Reading: on the device-forward definition (upload + network + readback of the five outputs) the RTX 5090 running the port's fp32 reference eagerly in strict fp32 is **3.13x faster than the p150a's fused bf16 trace at S=1 (134.1 vs 420.1 ms) and 3.62x at S=2 (245.0 vs 887.3 ms)**; with TF32 4.52x / 5.34x; in the p150a's own precision class (bf16 autocast) **6.41x / 8.58x** (65.6 / 103.4 ms); fp16 autocast 6.63x / 9.27x. The gap widens with S because the p150a scales almost linearly in S (420 -> 887 ms) while the GPU's S=2 forward costs only 1.58x its S=1 forward under bf16 autocast (the 1374-token frame blocks fill the GPU better at S=2 and the per-frame DPT heads are cheap). PCIe transfers are a small part of the GPU number (incl vs excl within ~1-2 ms; input 3.2 MB per view, readback 6.4 / 12.9 MB). Converting the whole module to fp16 once (informational) gives 52.3 / 76.8 ms (8.03x / 11.55x) at min PCC 0.99993 / 0.99994. `torch.compile` (inductor, torch 2.11.0+cu128, triton 3.6.0, sm_120) was **slower** than eager for this module in every variant tried (83.6 ms compiled vs 63.3 ms eager at S=1 fp16 autocast; compile 18-27 s per (S, mode), peak memory ~9.2-10.5 GB vs 5.4-5.8 GB eager): the compiled rows are reported as measured, not as a tuned deployment.
On accuracy, every GPU precision tracks the fp32 reference more closely (min PCC 0.9992 bf16 autocast S=1, >= 0.99999 fp16 autocast / TF32) than the p150a's bf16 graph does (served S=2 depth 0.99927 / world_points 0.99911 vs the same reference; synthetic min-PCC 0.9952 / 0.9981). bf16 is the least accurate GPU option here (min PCC 0.9992 / 0.9993, depth max |d| 0.25-0.28 on a ~1.0 median depth): the 8-bit bf16 mantissa costs more than fp16's range restriction on this network, whose exp/expm1 head activations stay fp32 under autocast.
End-to-end (the `server/app.py` interval): **180-248 ms vs 528.5 ms at S=1 (2.13-2.97x) and 342-493 ms vs 1104.6 ms at S=2 (2.24-3.23x)**. The ratio compresses because ~110 ms (S=1) / ~224 ms (S=2) of every request is the same single-threaded host `np.savez_compressed` of the float16/float32 dense arrays on both sides (the p150a server logs 105-125 ms of `encode` at S=1 on this host), plus 4 / 21 ms of PNG decode + bicubic pad: on the GPU the response encoding is now the largest stage of a served request at every precision.
Not measured / not claimed: p150a power (not measured in any pass -> no power or efficiency comparison; the GPU power figures above are nvidia-smi means over each timed loop, 31.2 W idle). p150a numbers were not re-measured. The GPU numbers exclude HTTP/JSON framing, as do the p150a `timing_ms` keys. S=3 / S=4 were not run on the GPU (the p150a card has only best-of-N `test_vggt.py` numbers for them, no served medians). The p150a's `timing_ms.forward` includes its host prelude (ImageNet-normalise + DINOv2 patch-embed on the host, fp32 torch) and the host `activate_head`; the GPU incl_h2d runs the same stages on the GPU inside the module, so both sides cover the whole `images -> outputs` path.
## Reproduce
```bash
cd /home/deepgadget/experiments/tt-models/logs/gpu-vs-p150/vggt-1b
# upstream package (pinned, no __init__.py -> namespace package): git clone https://github.com/facebookresearch/vggt vggt_upstream && git -C vggt_upstream checkout 44b3afbd1869d8bde4894dd8ea1e293112dd5eba
HF_HUB_OFFLINE=1 /home/deepgadget/experiments/tt-models/.venv-gpu/main/bin/python bench_vggt_gpu.py --iters 50 --warmup 10 --skip-cpu | tee full_run.log
/home/deepgadget/experiments/tt-models/.venv-gpu/main/bin/python render_report.py
# outputs: reports/gpu-vs-p150/vggt-1b.json + .md, ./result.json, ./cpu_fp32_reference.pt (written by the first run without --skip-cpu, see smoke.log)
# --no-compile skips the torch.compile rows; --no-served / --no-lowprec-weights skip those sections; --seqs 1 runs S=1 only
```
GPU released after the run: `nvidia-smi --query-compute-apps=pid --format=csv,noheader` -> empty.
## Update 2026-10-04
The GPU numbers above did not change. This section compares them with the optimized TT build of 2026-10-04. The TT numbers come from [`OPT_REPORT.md`](OPT_REPORT.md) (round 15) and from the final verification run. See [`VERIFICATION_2026-10-04.md`](VERIFICATION_2026-10-04.md).
TT setup: one Galaxy Blackhole chip (chip 4), 12×10 compute grid, stock worker dispatch with 2 command queues, tt-metal `8b98410e730`, warm, traced, batch 1, 518×518, all outputs. The device programs are the same in the 2026-10-04 Python API commit.
### Like-for-like: forward with upload and readback
TT row: `bench_vggt.py` served (upload, trace, readback and host finish). GPU row: incl_h2d of the table above. Ratio = GPU ms / TT ms (> 1 means that TT is faster).
| S | TT ms | GPU precision | GPU ms | ratio GPU / TT |
|---|---:|---|---:|---:|
| 1 | 49.3 | fp32 strict (no TF32) | 134.07 | **2.72** |
| 1 | 49.3 | fp32 + TF32 | 93.04 | **1.89** |
| 1 | 49.3 | bf16 autocast | 65.58 | **1.33** |
| 1 | 49.3 | fp16 autocast | 63.32 | **1.28** |
| 1 | 49.3 | bf16_weights (informational) | 54.47 | 1.10 |
| 1 | 49.3 | fp16_weights (informational) | 52.32 | 1.06 |
| 2 | 91.6 | fp32 strict (no TF32) | 244.97 | **2.67** |
| 2 | 91.6 | fp32 + TF32 | 166.20 | **1.81** |
| 2 | 91.6 | bf16 autocast | 103.37 | **1.13** |
| 2 | 91.6 | fp16 autocast | 95.72 | **1.05** |
| 2 | 91.6 | bf16_weights (informational) | 84.59 | 0.92 |
| 2 | 91.6 | fp16_weights (informational) | 76.80 | 0.84 |
TT device trace time: S=1 46.64 ms, S=2 87.45 ms. TT S=3 / S=4: 171.8 / 233.4 ms served, 165.68 / 227.67 ms device. The GPU did not run S=3 or S=4.
Result: in the reference precision class (bf16 / fp16 autocast), TT is 1.28-1.33× faster at S=1 and 1.05-1.13× faster at S=2. The GPU is faster only when the whole module is converted to fp16 or bf16 weights at S=2 (informational rows).
### Served end to end (not like-for-like)
TT `/predict` `timing_ms.total` (demo media, npz float16): S=1 56.2 ms, S=2 108.8 ms. The GPU served-like totals are 180.34 / 349.11 ms (bf16 autocast) and 178.04 / 342.27 ms (fp16 autocast). The ratio is 3.21× at S=1 and S=2 for bf16 autocast. **This comparison is not like-for-like.** The GPU rows used single-threaded host stages (decode, `np.savez_compressed`) on another host. The TT server uses a multi-threaded host path, libspng PNG decode, zlib-ng deflate at the same level and pybase64. The like-for-like comparison is the table above.
### Python API
The Python `model()` call on the demo files (decode, forward and output conversion, float32, no npz encode) takes 53.8 ms at S=1 and 104.7 ms at S=2. It gives the same output bits as the server. No GPU row uses the same interval.
### Accuracy
The TT outputs did not change in the 2026-10-04 Python API commit. Against the fp32 torch reference: `test_vggt.py` S=1 min PCC 0.9971 (`world_points_conf`); bench min PCC (round 15) S=1 / 2 / 3 / 4 0.9964 / 0.9988 / 0.9979 / 0.9950; real-photo depth AbsRel mean 0.00480 (previous release 0.00479). The GPU bf16 / fp16 autocast rows above stay closer to the fp32 reference (min PCC 0.9992 or higher).
### Not measured
p150a and Galaxy chip power were not measured. Thus, no efficiency comparison is made. The GPU numbers were not measured again.
## Update 2026-10-05
The GPU numbers above did not change. The TT build now opens the chip in the p150a configuration: dispatch on ETH cores, 1 command queue and a 12×10 compute grid. The 2026-10-04 TT numbers above used stock worker dispatch with 2 command queues on a Galaxy Blackhole chip. The TT numbers in this section replace them. The TT outputs are bit-identical to the 2026-10-04 build. See [`VERIFICATION_2026-10-04.md`](VERIFICATION_2026-10-04.md), section "p150 ETH-dispatch compliance (2026-10-05)".
TT setup: one Galaxy Blackhole chip (chip 14), ETH dispatch, 1 command queue, 12×10 compute grid, tt-metal `8b98410e730` with `patches/tt-metal-eth-dispatch.patch`, warm, traced, batch 1, 518×518, all outputs. An independent check on a different chip (chip 1) gave values within 4 % of these.
### Like-for-like: forward with upload and readback
TT row: `bench_vggt.py` served (upload, trace, readback and host finish). GPU row: incl_h2d of the table above. Ratio = GPU ms / TT ms (> 1 means that TT is faster).
| S | TT ms | GPU precision | GPU ms | ratio GPU / TT |
|---|---:|---|---:|---:|
| 1 | 50.1 | fp32 strict (no TF32) | 134.07 | **2.68** |
| 1 | 50.1 | fp32 + TF32 | 93.04 | **1.86** |
| 1 | 50.1 | bf16 autocast | 65.58 | **1.31** |
| 1 | 50.1 | fp16 autocast | 63.32 | **1.26** |
| 1 | 50.1 | bf16_weights (informational) | 54.47 | 1.09 |
| 1 | 50.1 | fp16_weights (informational) | 52.32 | 1.04 |
| 2 | 91.4 | fp32 strict (no TF32) | 244.97 | **2.68** |
| 2 | 91.4 | fp32 + TF32 | 166.20 | **1.82** |
| 2 | 91.4 | bf16 autocast | 103.37 | **1.13** |
| 2 | 91.4 | fp16 autocast | 95.72 | **1.05** |
| 2 | 91.4 | bf16_weights (informational) | 84.59 | 0.93 |
| 2 | 91.4 | fp16_weights (informational) | 76.80 | 0.84 |
TT device trace time: S=1 46.80 ms, S=2 87.38 ms. TT S=3 / S=4: 173.9 / 236.4 ms served, 165.46 / 225.77 ms device. The GPU did not run S=3 or S=4.
Result: in the reference precision class (bf16 / fp16 autocast), TT is 1.26-1.31× faster at S=1 and 1.05-1.13× faster at S=2. The GPU is faster only when the whole module is converted to fp16 or bf16 weights at S=2 (informational rows).
### Served end to end (not like-for-like)
TT `/predict` `timing_ms.total` (demo media, npz float16): S=1 57.0 ms, S=2 109.7 ms. The GPU served-like totals are 180.34 / 349.11 ms (bf16 autocast). The ratio is 3.16× at S=1 and 3.18× at S=2. **This comparison is not like-for-like**, for the same reasons as in the 2026-10-04 update.
### Python API
The Python `model()` call on the demo files (decode, forward and output conversion, float32, no npz encode) takes 54.7 ms at S=1 and 105.8 ms at S=2.
### Accuracy
Unchanged. The outputs are bit-identical to the 2026-10-04 build. `test_vggt.py` S=1 min PCC 0.9971; bench min PCC S=1 / 2 / 3 / 4 0.9964 / 0.9988 / 0.9979 / 0.9950.
|