tt-model.yaml: RTX 5090 comparison row + caveat
Browse files- tt-model.yaml +2 -0
tt-model.yaml
CHANGED
|
@@ -155,6 +155,7 @@ card:
|
|
| 155 |
| Inference, served over HTTP (warm, S=1, 518×518, npz; 30 requests, median / min / max) | 411 / 408 / 435 ms forward · 526 / 521 / 548 ms end-to-end |
|
| 156 |
| Inference, served over HTTP (warm, S=2; 10 requests) | 881 / 875 / 887 ms forward · 1091 / 1085 / 1097 ms end-to-end |
|
| 157 |
| Inference, `test_vggt.py` best-of-N (S=3 / S=4) | 1945 ms / 2759 ms |
|
|
|
|
| 158 |
|
| 159 |
### Caveats
|
| 160 |
|
|
@@ -164,6 +165,7 @@ card:
|
|
| 164 |
- Weights are CC-BY-NC-4.0 (non-commercial only); the 5.0 GB fp32 `model.safetensors` is fetched into your HF cache, and the server keeps the torch model (~5 GB host RAM) as the driver.
|
| 165 |
- Dense outputs are npz-encoded by default (see Response); point tracking is not exposed (track head not ported).
|
| 166 |
- Not an OpenAI-compatible API; `GET /v1/models` is a stub so the tt-model ready card does not 404. Validated on tt-metal `v0.78.0-dev20260820` (main `8b98410e730`), single p150a only.
|
|
|
|
| 167 |
|
| 168 |
### Licensing
|
| 169 |
|
|
|
|
| 155 |
| Inference, served over HTTP (warm, S=1, 518×518, npz; 30 requests, median / min / max) | 411 / 408 / 435 ms forward · 526 / 521 / 548 ms end-to-end |
|
| 156 |
| Inference, served over HTTP (warm, S=2; 10 requests) | 881 / 875 / 887 ms forward · 1091 / 1085 / 1097 ms end-to-end |
|
| 157 |
| Inference, `test_vggt.py` best-of-N (S=3 / S=4) | 1945 ms / 2759 ms |
|
| 158 |
+
| Same forward on an RTX 5090 (same host, port's torch reference, eager PyTorch, batch 1, incl. H2D/D2H; bf16 / fp16 autocast) | S=1 65.6 / 63.3 ms, S=2 103.4 / 95.7 ms → GPU 6.4–6.6× (S=1) and 8.6–9.3× (S=2) faster than the p150a's 420 / 887 ms; fp32-strict 134 / 245 ms (3.1× / 3.6×); fp16 weights resident 52.3 / 76.8 ms; `torch.compile` was slower than eager here |
|
| 159 |
|
| 160 |
### Caveats
|
| 161 |
|
|
|
|
| 165 |
- Weights are CC-BY-NC-4.0 (non-commercial only); the 5.0 GB fp32 `model.safetensors` is fetched into your HF cache, and the server keeps the torch model (~5 GB host RAM) as the driver.
|
| 166 |
- Dense outputs are npz-encoded by default (see Response); point tracking is not exposed (track head not ported).
|
| 167 |
- Not an OpenAI-compatible API; `GET /v1/models` is a stub so the tt-model ready card does not 404. Validated on tt-metal `v0.78.0-dev20260820` (main `8b98410e730`), single p150a only.
|
| 168 |
+
- GPU comparison: GPU bf16/fp16 is 6–9× faster: the 1B-parameter aggregator is compute-bound on the p150a (one metal-trace per S, no host syncs), so the gap reflects raw throughput, not dispatch overhead. RTX 5090 rows (2026-09-14): same host, the port's own torch reference (same weights) run eagerly in PyTorch 2.11 cu128 with fp32 weights + autocast unless stated, no TensorRT; medians of 50 iterations after warm-up, H2D/D2H included; the p150a rows are the served bf16 fused path incl. upload/readback. p150a power was not measured, so no efficiency comparison is made. Full table: [`GPU_COMPARISON.md`](GPU_COMPARISON.md).
|
| 169 |
|
| 170 |
### Licensing
|
| 171 |
|