tt-model authoring files (fused default, megakernel pass)
Browse files- SERVING.md +42 -8
- tt-model.yaml +12 -5
SERVING.md
CHANGED
|
@@ -14,8 +14,12 @@ the manifest, `code/` the port, `code/models/server/app.py` the ASGI app uvicorn
|
|
| 14 |
|
| 15 |
## What the server does
|
| 16 |
|
| 17 |
-
|
| 18 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
| 19 |
|
| 20 |
1. base64 -> PIL RGB -> bilinear resize to 640x480 -> /255 -> fp32 `(1, 3, 480, 640)`
|
| 21 |
(HF `SuperPointImageProcessor` defaults; the model reads channel 0 like
|
|
@@ -35,9 +39,10 @@ One `threading.Lock` serialises every device call; handlers are sync and run und
|
|
| 35 |
frame) happens in the ASGI lifespan, so `Application startup complete` means warm. Shutdown
|
| 36 |
deallocates the device tensors and `ttnn.close_device`s inside the 120 s SIGTERM budget.
|
| 37 |
|
| 38 |
-
|
| 39 |
-
|
| 40 |
-
patched tt-metal) --
|
|
|
|
| 41 |
|
| 42 |
## Environment the app reads (lifespan only, never at import)
|
| 43 |
|
|
@@ -48,6 +53,9 @@ patched tt-metal) -- deliberately not what this image runs.
|
|
| 48 |
| `SP_WEIGHTS_DIR` | you (host/offline) | local dir with `config.json` + `model.safetensors`; overrides the two above |
|
| 49 |
| `TT_MESH_SHAPE` | launcher (`runtime.mesh_shape_env`) | `1x1` (also `(1, 1)` / `1,1`); any other shape -> `RuntimeError` at startup |
|
| 50 |
| `TT_DEVICE_ID` | you | chip to open, default `0` |
|
|
|
|
|
|
|
|
|
|
| 51 |
| `MESH_DEVICE` | launcher | `P150` (informational) |
|
| 52 |
| `TT_METAL_VISIBLE_DEVICES` | `serve.env` | `0` |
|
| 53 |
|
|
@@ -60,7 +68,7 @@ The launcher exports no revision, which is why `serve.env.TT_WEIGHTS_REVISION` r
|
|
| 60 |
| route | response |
|
| 61 |
|---|---|
|
| 62 |
| `GET /health` | `{"status": "ok" \| "starting", "model": "superpoint-p150", "device": {"arch", "id", "open"}}` (always 200) |
|
| 63 |
-
| `GET /info` | model/task/io, hardware, `weights {repo, revision, local_dir, loaded}`, `source {repo, commit}`, `input` (480x640, batch 1, preprocessing), `defaults`, `limits`, `serving_path` (traced=false, device_nms=false), `warmup_ms`, `descriptors` encoding, `license` |
|
| 64 |
| `GET /v1/models` | `{"object": "list", "data": [{"id": "<weights repo>", "object": "model", "owned_by": "changh95"}]}` (so OpenAI-shaped probes do not 404; not a chat API) |
|
| 65 |
| `POST /predict` | see below |
|
| 66 |
|
|
@@ -139,8 +147,10 @@ kernels ...)`, `Warmup complete: first forward <ms> (compile), second <ms>`, the
|
|
| 139 |
|
| 140 |
`PYTHONPATH` must start with `code/` so that `models` resolves to this repo's regular package
|
| 141 |
(tt-metal's own `models/` has no `__init__.py` and is shadowed on the host; in the image it is
|
| 142 |
-
excluded). The
|
| 143 |
-
|
|
|
|
|
|
|
| 144 |
|
| 145 |
## Package / serve / push (Blackhole host, rootless Docker)
|
| 146 |
|
|
@@ -182,6 +192,30 @@ worth keeping lives in `card.description` / `card.quickstart`); `media/`, `SERVI
|
|
| 182 |
`.gitattributes` and `tt-model.yaml` at the root survive. The orchestrator restores
|
| 183 |
`license`/`pipeline_tag` front matter after push.
|
| 184 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 185 |
## Caveats
|
| 186 |
|
| 187 |
- **Licence**: the weights are Magic Leap "academic or non-profit organisation noncommercial
|
|
|
|
| 14 |
|
| 15 |
## What the server does
|
| 16 |
|
| 17 |
+
Default since 2026-09-13 (`TT_FUSED` unset or `1`, also pinned in `serve.env`): pure ttnn, the
|
| 18 |
+
whole device graph as **one metal trace** per request with a standard-op device NMS, no custom
|
| 19 |
+
kernel -- see *Fused path* below for the per-request sequence and the measured numbers.
|
| 20 |
+
`TT_FUSED=0` restores the legacy path this section describes: pure-ttnn, **untraced**, host
|
| 21 |
+
NMS -- the port's `SP_TRACE_NMS=0 SP_NO_TRACE=1` configuration. Both paths: fixed 480x640
|
| 22 |
+
network input; every image is resized server-side. Per request (legacy):
|
| 23 |
|
| 24 |
1. base64 -> PIL RGB -> bilinear resize to 640x480 -> /255 -> fp32 `(1, 3, 480, 640)`
|
| 25 |
(HF `SuperPointImageProcessor` defaults; the model reads channel 0 like
|
|
|
|
| 39 |
frame) happens in the ASGI lifespan, so `Application startup complete` means warm. Shutdown
|
| 40 |
deallocates the device tensors and `ttnn.close_device`s inside the 120 s SIGTERM budget.
|
| 41 |
|
| 42 |
+
Measured 2026-09-13 through the container image (50 warm requests, medians; legacy path):
|
| 43 |
+
device_forward 12.4 ms, host NMS + post 26.5 ms, total 56.8 ms. The README's 40.7 fps requires
|
| 44 |
+
the fused `sp_eq_mul_mask` C++ kernel (`code/kernels/`, needs a patched tt-metal) -- not what
|
| 45 |
+
this image runs; the fused path below gets to ~40 fps with standard ops instead.
|
| 46 |
|
| 47 |
## Environment the app reads (lifespan only, never at import)
|
| 48 |
|
|
|
|
| 53 |
| `SP_WEIGHTS_DIR` | you (host/offline) | local dir with `config.json` + `model.safetensors`; overrides the two above |
|
| 54 |
| `TT_MESH_SHAPE` | launcher (`runtime.mesh_shape_env`) | `1x1` (also `(1, 1)` / `1,1`); any other shape -> `RuntimeError` at startup |
|
| 55 |
| `TT_DEVICE_ID` | you | chip to open, default `0` |
|
| 56 |
+
| `TT_FUSED` | `serve.env` (`"1"`; the code default when unset/empty is also fused) | unset/`1` = fused serving path: ONE metal trace per request (64-byte-page input upload, encoder + heads, `rms_norm` L2-norm, standard-op device NMS at radius 4, row-major outputs), captured during warm-up before READY. `0` = the legacy untraced path above, byte-identical to the 2026-09-12 image. See `DEVICE_VALIDATION.md` |
|
| 57 |
+
| `TT_FUSED_STAGES` | you (device A/B only) | comma list of fused stages, default all (`wide,nms,rms,rm`); `""` = trace-only |
|
| 58 |
+
| `SP_TRACE_REGION` | you | `trace_region_size` bytes for `ttnn.CreateDevice` on the fused path (default 32 MiB) |
|
| 59 |
| `MESH_DEVICE` | launcher | `P150` (informational) |
|
| 60 |
| `TT_METAL_VISIBLE_DEVICES` | `serve.env` | `0` |
|
| 61 |
|
|
|
|
| 68 |
| route | response |
|
| 69 |
|---|---|
|
| 70 |
| `GET /health` | `{"status": "ok" \| "starting", "model": "superpoint-p150", "device": {"arch", "id", "open"}}` (always 200) |
|
| 71 |
+
| `GET /info` | model/task/io, hardware, `weights {repo, revision, local_dir, loaded}`, `source {repo, commit}`, `input` (480x640, batch 1, preprocessing), `defaults`, `limits`, `serving_path` (fused: `traced=true, device_nms=true, nms_radius_traced=4, fused_stages`; legacy: `traced=false, device_nms=false`), `warmup_ms`, `descriptors` encoding, `license` |
|
| 72 |
| `GET /v1/models` | `{"object": "list", "data": [{"id": "<weights repo>", "object": "model", "owned_by": "changh95"}]}` (so OpenAI-shaped probes do not 404; not a chat API) |
|
| 73 |
| `POST /predict` | see below |
|
| 74 |
|
|
|
|
| 147 |
|
| 148 |
`PYTHONPATH` must start with `code/` so that `models` resolves to this repo's regular package
|
| 149 |
(tt-metal's own `models/` has no `__init__.py` and is shadowed on the host; in the image it is
|
| 150 |
+
excluded). The device tests in `code/models/tests/test_superpoint.py` get their `device` /
|
| 151 |
+
`device_params` fixtures and `--device-id` from the repo-local `code/conftest.py` (run pytest
|
| 152 |
+
from `code/`; `code/pytest.ini` pins the rootdir). tt-metal's own conftest cannot be used next
|
| 153 |
+
to this repo: it imports `models.demos...`, which the shadowing above breaks.
|
| 154 |
|
| 155 |
## Package / serve / push (Blackhole host, rootless Docker)
|
| 156 |
|
|
|
|
| 192 |
`.gitattributes` and `tt-model.yaml` at the root survive. The orchestrator restores
|
| 193 |
`license`/`pipeline_tag` front matter after push.
|
| 194 |
|
| 195 |
+
## Fused path (default; `TT_FUSED=0` = legacy) -- branch `opt/superpoint-p150-megakernel`
|
| 196 |
+
|
| 197 |
+
Device-validated on the p150a 2026-09-13 (`DEVICE_VALIDATION.md` "Results") and made the default
|
| 198 |
+
(code default + `serve.env.TT_FUSED: "1"`); `TT_FUSED=0` restores the legacy path above
|
| 199 |
+
byte-for-byte. On the fused path the lifespan opens the device with a trace region, runs the fused graph once
|
| 200 |
+
eagerly (kernel compile), captures it into a metal trace and replays it once -- all before
|
| 201 |
+
`Application startup complete` (boot log: `Warming up TT_FUSED path ...`, `Warmup complete:
|
| 202 |
+
compile forward ... trace capture ... traced forward ...`). Per request: H2D of the
|
| 203 |
+
`[1,1,9600,32]` bf16 input, one `execute_trace`, D2H of the row-major device NMS map (480x640)
|
| 204 |
+
and descriptors, then threshold / border / top-k / `grid_sample` on the host
|
| 205 |
+
(`postprocess_from_nms_map`). Requests with `nms_radius != 4` read the traced softmax scores
|
| 206 |
+
instead and run the legacy host NMS (same output, slower); `response.serving_path.device_nms`
|
| 207 |
+
and `/info.serving_path` report which path ran. Exactness: trace, upload, device NMS and
|
| 208 |
+
row-major outputs are bit-identical to the legacy path; the `rms_norm` descriptor L2-norm is
|
| 209 |
+
bf16-rounding-level (gate: descriptor PCC >= 0.999). Host proofs: `code/models/tests/test_fused_host.py`;
|
| 210 |
+
hardware plan, gates and results: `DEVICE_VALIDATION.md`. Measured 2026-09-13 through the
|
| 211 |
+
container image (`smoke_test.py` PASS, 50 warm requests, server `timing_ms` medians): device_forward
|
| 212 |
+
5.3 ms (min 5.1), post-processing 1.3 ms, preprocess 18 ms, total 24.7 ms (~40 fps) vs the legacy
|
| 213 |
+
path's 12.4 / 26.5 / 17.8 / 56.8 ms in the same session; keypoints and scores identical to the
|
| 214 |
+
legacy server for `nms_radius` 4 (device NMS) and 3 (host fallback), descriptors within bf16
|
| 215 |
+
rounding (max |diff| 8.5e-4, cosine >= 0.999996). Boot log landmarks: `Warming up TT_FUSED path
|
| 216 |
+
(stages nms,rm,rms,wide, traced nms_radius 4)`, `Warmup complete: compile forward ... trace capture
|
| 217 |
+
... traced forward ...`.
|
| 218 |
+
|
| 219 |
## Caveats
|
| 220 |
|
| 221 |
- **Licence**: the weights are Magic Leap "academic or non-profit organisation noncommercial
|
tt-model.yaml
CHANGED
|
@@ -60,6 +60,10 @@ serve:
|
|
| 60 |
env:
|
| 61 |
TT_WEIGHTS_REVISION: "734450e9ffe229074f5998494ddc615475cdb20a"
|
| 62 |
TT_METAL_VISIBLE_DEVICES: "0"
|
|
|
|
|
|
|
|
|
|
|
|
|
| 63 |
|
| 64 |
# Build-time assertions run INSIDE the finished image as uid 1000, no device, no weights.
|
| 65 |
verify:
|
|
@@ -99,7 +103,8 @@ card:
|
|
| 99 |
"scores": [0.609375, 0.589844, 0.582031],
|
| 100 |
"original_size": {"height": 900, "width": 1600}, "image_size": {"height": 480, "width": 640}, "scale": {"x": 2.5, "y": 1.875},
|
| 101 |
"descriptors": {"format": "npz", "key": "descriptors", "dtype": "float16", "shape": [539, 256], "data": "..."},
|
| 102 |
-
"
|
|
|
|
| 103 |
```
|
| 104 |
|
| 105 |
- `keypoints` are `[x, y]` in original image pixels, sorted by descending `scores`; the network frame is 480×640 and `scale` = original / network.
|
|
@@ -116,16 +121,18 @@ card:
|
|
| 116 |
| Metric | Value |
|
| 117 |
|---|---:|
|
| 118 |
| Pre-NMS score map · descriptor map PCC vs fp32 torch reference | 0.9971 · 0.9991 |
|
| 119 |
-
| Keypoint set vs reference (natural image, top-500, 2 px) | recall 98.
|
| 120 |
-
| Inference, served over HTTP (warm, batch 1, 480×640, 1600×900 JPEG in) |
|
|
|
|
| 121 |
|
| 122 |
### Caveats
|
| 123 |
|
| 124 |
- Every image is resized to 480×640 (bilinear, /255, channel 0); one image per request, batch 1, requests serialised on the chip.
|
| 125 |
-
- This server runs the
|
|
|
|
| 126 |
- Weights are research-only: the Magic Leap SuperPoint licence allows academic or non-profit organisation NONCOMMERCIAL research use.
|
| 127 |
- Not an OpenAI-compatible API; `GET /v1/models` is a stub so the tt-model ready card does not 404.
|
| 128 |
-
- Validated on tt-metal `v0.78.0-dev20260820` (main `8b98410e730`), single p150a only.
|
| 129 |
|
| 130 |
### Licensing
|
| 131 |
|
|
|
|
| 60 |
env:
|
| 61 |
TT_WEIGHTS_REVISION: "734450e9ffe229074f5998494ddc615475cdb20a"
|
| 62 |
TT_METAL_VISIBLE_DEVICES: "0"
|
| 63 |
+
# Fused serving path (one metal trace per request, standard-op device NMS), device-validated
|
| 64 |
+
# 2026-09-13 (DEVICE_VALIDATION.md "Results"); it is also the code default. "0" = the legacy
|
| 65 |
+
# untraced host-NMS path of the 2026-09-12 image.
|
| 66 |
+
TT_FUSED: "1"
|
| 67 |
|
| 68 |
# Build-time assertions run INSIDE the finished image as uid 1000, no device, no weights.
|
| 69 |
verify:
|
|
|
|
| 103 |
"scores": [0.609375, 0.589844, 0.582031],
|
| 104 |
"original_size": {"height": 900, "width": 1600}, "image_size": {"height": 480, "width": 640}, "scale": {"x": 2.5, "y": 1.875},
|
| 105 |
"descriptors": {"format": "npz", "key": "descriptors", "dtype": "float16", "shape": [539, 256], "data": "..."},
|
| 106 |
+
"serving_path": {"traced": true, "device_nms": true},
|
| 107 |
+
"timing_ms": {"preprocess": 17.7, "device_forward": 5.3, "postprocess": 1.3, "total": 24.3}}
|
| 108 |
```
|
| 109 |
|
| 110 |
- `keypoints` are `[x, y]` in original image pixels, sorted by descending `scores`; the network frame is 480×640 and `scale` = original / network.
|
|
|
|
| 121 |
| Metric | Value |
|
| 122 |
|---|---:|
|
| 123 |
| Pre-NMS score map · descriptor map PCC vs fp32 torch reference | 0.9971 · 0.9991 |
|
| 124 |
+
| Keypoint set vs reference (natural image, top-500, 2 px) | recall 98.20% · precision 99.40% · F1 98.80% |
|
| 125 |
+
| Inference, served over HTTP (warm, batch 1, 480×640, 1600×900 JPEG in; median of 50 requests) | 5.3 ms device (trace + device NMS) · 1.3 ms host post-processing · 18 ms JPEG decode/resize · 24.7 ms end-to-end (~40 FPS) |
|
| 126 |
+
| Same, legacy path (`TT_FUSED=0`: untraced, host NMS) | 12.4 ms device · 26.5 ms host NMS · 56.8 ms end-to-end (~18 FPS) |
|
| 127 |
|
| 128 |
### Caveats
|
| 129 |
|
| 130 |
- Every image is resized to 480×640 (bilinear, /255, channel 0); one image per request, batch 1, requests serialised on the chip.
|
| 131 |
+
- This server runs the whole device graph as one metal trace per request with a standard-op device NMS (radius 4; bit-identical to the host single-pass NMS) and an `rms_norm` descriptor L2-norm (bf16-rounding-level vs the legacy chain; PCC 0.9991 either way). bf16 + HiFi2 + fp32 accumulate throughout (bfloat8/LoFi drop score PCC to ~0.91). No custom kernel: the port README's `sp_eq_mul_mask` path is not built into this image.
|
| 132 |
+
- `nms_radius` other than 4 falls back to the host NMS on the traced scores (same keypoints, ~25 ms slower); `TT_FUSED=0` in the environment restores the untraced legacy path (validated 2026-09-12).
|
| 133 |
- Weights are research-only: the Magic Leap SuperPoint licence allows academic or non-profit organisation NONCOMMERCIAL research use.
|
| 134 |
- Not an OpenAI-compatible API; `GET /v1/models` is a stub so the tt-model ready card does not 404.
|
| 135 |
+
- Validated on tt-metal `v0.78.0-dev20260820` (main `8b98410e730`), single p150a only; numbers above measured 2026-09-13 through this container image (`DEVICE_VALIDATION.md`).
|
| 136 |
|
| 137 |
### Licensing
|
| 138 |
|