changh95 commited on
Commit
ccde6d4
·
verified ·
1 Parent(s): 4b6b847

tt-model.yaml: RTX 5090 comparison row + caveat

Browse files
Files changed (1) hide show
  1. tt-model.yaml +2 -0
tt-model.yaml CHANGED
@@ -124,6 +124,7 @@ card:
124
  | Keypoint set vs reference (natural image, top-500, 2 px) | recall 98.20% · precision 99.40% · F1 98.80% |
125
  | Inference, served over HTTP (warm, batch 1, 480×640, 1600×900 JPEG in; median of 50 requests) | 5.3 ms device (trace + device NMS) · 1.3 ms host post-processing · 18 ms JPEG decode/resize · 24.7 ms end-to-end (~40 FPS) |
126
  | Same, legacy path (`TT_FUSED=0`: untraced, host NMS) | 12.4 ms device · 26.5 ms host NMS · 56.8 ms end-to-end (~18 FPS) |
 
127
 
128
  ### Caveats
129
 
@@ -133,6 +134,7 @@ card:
133
  - Weights are research-only: the Magic Leap SuperPoint licence allows academic or non-profit organisation NONCOMMERCIAL research use.
134
  - Not an OpenAI-compatible API; `GET /v1/models` is a stub so the tt-model ready card does not 404.
135
  - Validated on tt-metal `v0.78.0-dev20260820` (main `8b98410e730`), single p150a only; numbers above measured 2026-09-13 through this container image (`DEVICE_VALIDATION.md`).
 
136
 
137
  ### Licensing
138
 
 
124
  | Keypoint set vs reference (natural image, top-500, 2 px) | recall 98.20% · precision 99.40% · F1 98.80% |
125
  | Inference, served over HTTP (warm, batch 1, 480×640, 1600×900 JPEG in; median of 50 requests) | 5.3 ms device (trace + device NMS) · 1.3 ms host post-processing · 18 ms JPEG decode/resize · 24.7 ms end-to-end (~40 FPS) |
126
  | Same, legacy path (`TT_FUSED=0`: untraced, host NMS) | 12.4 ms device · 26.5 ms host NMS · 56.8 ms end-to-end (~18 FPS) |
127
+ | Same forward on an RTX 5090 (same host, port's torch reference, eager PyTorch, batch 1, incl. H2D/D2H; bf16 / fp16 autocast) | 1.6 / 1.5 ms → GPU 3.2–3.4× faster than the p150a's 5.1 ms device forward (incl. device NMS); fp32-strict 2.9 ms (1.7×); best `torch.compile` 1.2 ms. End-to-end both sides are bound by the ~18 ms JPEG decode/resize (GPU 21.5 vs p150a 23.9 ms) |
128
 
129
  ### Caveats
130
 
 
134
  - Weights are research-only: the Magic Leap SuperPoint licence allows academic or non-profit organisation NONCOMMERCIAL research use.
135
  - Not an OpenAI-compatible API; `GET /v1/models` is a stub so the tt-model ready card does not 404.
136
  - Validated on tt-metal `v0.78.0-dev20260820` (main `8b98410e730`), single p150a only; numbers above measured 2026-09-13 through this container image (`DEVICE_VALIDATION.md`).
137
+ - GPU comparison: GPU bf16/fp16 3.2–3.4× faster on the device forward, but the served request is host-bound on both sides (JPEG decode/resize ~18 ms), so end-to-end the GPU is only 1.1× faster. RTX 5090 rows (2026-09-14): same host, the port's own torch reference (same weights) run eagerly in PyTorch 2.11 cu128 with fp32 weights + autocast unless stated, no TensorRT; medians of 50 iterations after warm-up, H2D/D2H included; the p150a rows are the served bf16 fused path incl. upload/readback. p150a power was not measured, so no efficiency comparison is made. Full table: [`GPU_COMPARISON.md`](GPU_COMPARISON.md).
138
 
139
  ### Licensing
140