changh95 commited on
Commit
6718c9f
·
verified ·
1 Parent(s): 5db9e11

tt-model.yaml: RTX 5090 comparison row + caveat

Browse files
Files changed (1) hide show
  1. tt-model.yaml +2 -0
tt-model.yaml CHANGED
@@ -112,6 +112,7 @@ card:
112
  | Detection-IoU agreement vs fp32 reference (demo image) | 98.70 (2 cat + 2 remote, scores within 0.008 of the reference) |
113
  | Backbone feature-map PCC vs torch | 0.999–0.9999 |
114
  | Inference, served over HTTP (warm, batch 1, 560×560) | ~8.6 ms device · ~15.5 ms server-side incl. PNG decode (~120 FPS device-bound) |
 
115
 
116
  ### Caveats
117
 
@@ -119,6 +120,7 @@ card:
119
  - bf16 on device: scores differ slightly from the fp32 reference (see the IoU figure above).
120
  - Not an OpenAI-compatible API; `GET /v1/models` is a stub so the tt-model ready card does not 404.
121
  - Validated on tt-metal `v0.78.0-dev20260820` (main `8b98410e730`), single p150a only.
 
122
 
123
  ### Licensing
124
 
 
112
  | Detection-IoU agreement vs fp32 reference (demo image) | 98.70 (2 cat + 2 remote, scores within 0.008 of the reference) |
113
  | Backbone feature-map PCC vs torch | 0.999–0.9999 |
114
  | Inference, served over HTTP (warm, batch 1, 560×560) | ~8.6 ms device · ~15.5 ms server-side incl. PNG decode (~120 FPS device-bound) |
115
+ | Same forward on an RTX 5090 (same host, port's torch reference, eager PyTorch, batch 1, incl. H2D/D2H; bf16 / fp16 autocast) | 5.9 / 5.9 ms → GPU 1.4× faster than the p150a's 8.2 ms; fp32-strict GPU 8.0 ms (parity); best `torch.compile` 2.2 ms |
116
 
117
  ### Caveats
118
 
 
120
  - bf16 on device: scores differ slightly from the fp32 reference (see the IoU figure above).
121
  - Not an OpenAI-compatible API; `GET /v1/models` is a stub so the tt-model ready card does not 404.
122
  - Validated on tt-metal `v0.78.0-dev20260820` (main `8b98410e730`), single p150a only.
123
+ - GPU comparison: fp32-strict eager GPU 8.0 ms ≈ p150a 8.2 ms; GPU bf16/fp16 1.4× faster, compiled 3.7×. RTX 5090 rows (2026-09-14): same host, the port's own torch reference (same weights) run eagerly in PyTorch 2.11 cu128 with fp32 weights + autocast unless stated, no TensorRT; medians of 50 iterations after warm-up, H2D/D2H included; the p150a rows are the served bf16 fused path incl. upload/readback. p150a power was not measured, so no efficiency comparison is made. Full table: [`GPU_COMPARISON.md`](GPU_COMPARISON.md).
124
 
125
  ### Licensing
126