tt-model.yaml: RTX 5090 comparison row + caveat
Browse files- tt-model.yaml +2 -0
tt-model.yaml
CHANGED
|
@@ -112,6 +112,7 @@ card:
|
|
| 112 |
| Detection-IoU agreement vs fp32 reference (demo image) | 98.70 (2 cat + 2 remote, scores within 0.008 of the reference) |
|
| 113 |
| Backbone feature-map PCC vs torch | 0.999–0.9999 |
|
| 114 |
| Inference, served over HTTP (warm, batch 1, 560×560) | ~8.6 ms device · ~15.5 ms server-side incl. PNG decode (~120 FPS device-bound) |
|
|
|
|
| 115 |
|
| 116 |
### Caveats
|
| 117 |
|
|
@@ -119,6 +120,7 @@ card:
|
|
| 119 |
- bf16 on device: scores differ slightly from the fp32 reference (see the IoU figure above).
|
| 120 |
- Not an OpenAI-compatible API; `GET /v1/models` is a stub so the tt-model ready card does not 404.
|
| 121 |
- Validated on tt-metal `v0.78.0-dev20260820` (main `8b98410e730`), single p150a only.
|
|
|
|
| 122 |
|
| 123 |
### Licensing
|
| 124 |
|
|
|
|
| 112 |
| Detection-IoU agreement vs fp32 reference (demo image) | 98.70 (2 cat + 2 remote, scores within 0.008 of the reference) |
|
| 113 |
| Backbone feature-map PCC vs torch | 0.999–0.9999 |
|
| 114 |
| Inference, served over HTTP (warm, batch 1, 560×560) | ~8.6 ms device · ~15.5 ms server-side incl. PNG decode (~120 FPS device-bound) |
|
| 115 |
+
| Same forward on an RTX 5090 (same host, port's torch reference, eager PyTorch, batch 1, incl. H2D/D2H; bf16 / fp16 autocast) | 5.9 / 5.9 ms → GPU 1.4× faster than the p150a's 8.2 ms; fp32-strict GPU 8.0 ms (parity); best `torch.compile` 2.2 ms |
|
| 116 |
|
| 117 |
### Caveats
|
| 118 |
|
|
|
|
| 120 |
- bf16 on device: scores differ slightly from the fp32 reference (see the IoU figure above).
|
| 121 |
- Not an OpenAI-compatible API; `GET /v1/models` is a stub so the tt-model ready card does not 404.
|
| 122 |
- Validated on tt-metal `v0.78.0-dev20260820` (main `8b98410e730`), single p150a only.
|
| 123 |
+
- GPU comparison: fp32-strict eager GPU 8.0 ms ≈ p150a 8.2 ms; GPU bf16/fp16 1.4× faster, compiled 3.7×. RTX 5090 rows (2026-09-14): same host, the port's own torch reference (same weights) run eagerly in PyTorch 2.11 cu128 with fp32 weights + autocast unless stated, no TensorRT; medians of 50 iterations after warm-up, H2D/D2H included; the p150a rows are the served bf16 fused path incl. upload/readback. p150a power was not measured, so no efficiency comparison is made. Full table: [`GPU_COMPARISON.md`](GPU_COMPARISON.md).
|
| 124 |
|
| 125 |
### Licensing
|
| 126 |
|