changh95 commited on
Commit
1428722
·
verified ·
1 Parent(s): 4b8f630

tt-model.yaml: RTX 5090 comparison row + caveat

Browse files
Files changed (1) hide show
  1. tt-model.yaml +2 -0
tt-model.yaml CHANGED
@@ -155,6 +155,7 @@ card:
155
  | Inference, served over HTTP (warm, S=1, 518×518, npz; 30 requests, median / min / max) | 411 / 408 / 435 ms forward · 526 / 521 / 548 ms end-to-end |
156
  | Inference, served over HTTP (warm, S=2; 10 requests) | 881 / 875 / 887 ms forward · 1091 / 1085 / 1097 ms end-to-end |
157
  | Inference, `test_vggt.py` best-of-N (S=3 / S=4) | 1945 ms / 2759 ms |
 
158
 
159
  ### Caveats
160
 
@@ -164,6 +165,7 @@ card:
164
  - Weights are CC-BY-NC-4.0 (non-commercial only); the 5.0 GB fp32 `model.safetensors` is fetched into your HF cache, and the server keeps the torch model (~5 GB host RAM) as the driver.
165
  - Dense outputs are npz-encoded by default (see Response); point tracking is not exposed (track head not ported).
166
  - Not an OpenAI-compatible API; `GET /v1/models` is a stub so the tt-model ready card does not 404. Validated on tt-metal `v0.78.0-dev20260820` (main `8b98410e730`), single p150a only.
 
167
 
168
  ### Licensing
169
 
 
155
  | Inference, served over HTTP (warm, S=1, 518×518, npz; 30 requests, median / min / max) | 411 / 408 / 435 ms forward · 526 / 521 / 548 ms end-to-end |
156
  | Inference, served over HTTP (warm, S=2; 10 requests) | 881 / 875 / 887 ms forward · 1091 / 1085 / 1097 ms end-to-end |
157
  | Inference, `test_vggt.py` best-of-N (S=3 / S=4) | 1945 ms / 2759 ms |
158
+ | Same forward on an RTX 5090 (same host, port's torch reference, eager PyTorch, batch 1, incl. H2D/D2H; bf16 / fp16 autocast) | S=1 65.6 / 63.3 ms, S=2 103.4 / 95.7 ms → GPU 6.4–6.6× (S=1) and 8.6–9.3× (S=2) faster than the p150a's 420 / 887 ms; fp32-strict 134 / 245 ms (3.1× / 3.6×); fp16 weights resident 52.3 / 76.8 ms; `torch.compile` was slower than eager here |
159
 
160
  ### Caveats
161
 
 
165
  - Weights are CC-BY-NC-4.0 (non-commercial only); the 5.0 GB fp32 `model.safetensors` is fetched into your HF cache, and the server keeps the torch model (~5 GB host RAM) as the driver.
166
  - Dense outputs are npz-encoded by default (see Response); point tracking is not exposed (track head not ported).
167
  - Not an OpenAI-compatible API; `GET /v1/models` is a stub so the tt-model ready card does not 404. Validated on tt-metal `v0.78.0-dev20260820` (main `8b98410e730`), single p150a only.
168
+ - GPU comparison: GPU bf16/fp16 is 6–9× faster: the 1B-parameter aggregator is compute-bound on the p150a (one metal-trace per S, no host syncs), so the gap reflects raw throughput, not dispatch overhead. RTX 5090 rows (2026-09-14): same host, the port's own torch reference (same weights) run eagerly in PyTorch 2.11 cu128 with fp32 weights + autocast unless stated, no TensorRT; medians of 50 iterations after warm-up, H2D/D2H included; the p150a rows are the served bf16 fused path incl. upload/readback. p150a power was not measured, so no efficiency comparison is made. Full table: [`GPU_COMPARISON.md`](GPU_COMPARISON.md).
169
 
170
  ### Licensing
171