changh95 commited on
Commit
d2f174c
·
verified ·
1 Parent(s): eee3601

tt-model authoring files (fused default, megakernel pass)

Browse files
Files changed (2) hide show
  1. SERVING.md +42 -8
  2. tt-model.yaml +12 -5
SERVING.md CHANGED
@@ -14,8 +14,12 @@ the manifest, `code/` the port, `code/models/server/app.py` the ASGI app uvicorn
14
 
15
  ## What the server does
16
 
17
- Pure-ttnn, **untraced**, host NMS -- the port's `SP_TRACE_NMS=0 SP_NO_TRACE=1` configuration,
18
- no custom kernel. Fixed 480x640 network input; every image is resized server-side. Per request:
 
 
 
 
19
 
20
  1. base64 -> PIL RGB -> bilinear resize to 640x480 -> /255 -> fp32 `(1, 3, 480, 640)`
21
  (HF `SuperPointImageProcessor` defaults; the model reads channel 0 like
@@ -35,9 +39,10 @@ One `threading.Lock` serialises every device call; handlers are sync and run und
35
  frame) happens in the ASGI lifespan, so `Application startup complete` means warm. Shutdown
36
  deallocates the device tensors and `ttnn.close_device`s inside the 120 s SIGTERM budget.
37
 
38
- Expected speed: ~6 fps device forward (untraced) + ~36 ms host NMS. The README's 40.7 fps
39
- requires trace capture and the fused `sp_eq_mul_mask` C++ kernel (`code/kernels/`, needs a
40
- patched tt-metal) -- deliberately not what this image runs.
 
41
 
42
  ## Environment the app reads (lifespan only, never at import)
43
 
@@ -48,6 +53,9 @@ patched tt-metal) -- deliberately not what this image runs.
48
  | `SP_WEIGHTS_DIR` | you (host/offline) | local dir with `config.json` + `model.safetensors`; overrides the two above |
49
  | `TT_MESH_SHAPE` | launcher (`runtime.mesh_shape_env`) | `1x1` (also `(1, 1)` / `1,1`); any other shape -> `RuntimeError` at startup |
50
  | `TT_DEVICE_ID` | you | chip to open, default `0` |
 
 
 
51
  | `MESH_DEVICE` | launcher | `P150` (informational) |
52
  | `TT_METAL_VISIBLE_DEVICES` | `serve.env` | `0` |
53
 
@@ -60,7 +68,7 @@ The launcher exports no revision, which is why `serve.env.TT_WEIGHTS_REVISION` r
60
  | route | response |
61
  |---|---|
62
  | `GET /health` | `{"status": "ok" \| "starting", "model": "superpoint-p150", "device": {"arch", "id", "open"}}` (always 200) |
63
- | `GET /info` | model/task/io, hardware, `weights {repo, revision, local_dir, loaded}`, `source {repo, commit}`, `input` (480x640, batch 1, preprocessing), `defaults`, `limits`, `serving_path` (traced=false, device_nms=false), `warmup_ms`, `descriptors` encoding, `license` |
64
  | `GET /v1/models` | `{"object": "list", "data": [{"id": "<weights repo>", "object": "model", "owned_by": "changh95"}]}` (so OpenAI-shaped probes do not 404; not a chat API) |
65
  | `POST /predict` | see below |
66
 
@@ -139,8 +147,10 @@ kernels ...)`, `Warmup complete: first forward <ms> (compile), second <ms>`, the
139
 
140
  `PYTHONPATH` must start with `code/` so that `models` resolves to this repo's regular package
141
  (tt-metal's own `models/` has no `__init__.py` and is shadowed on the host; in the image it is
142
- excluded). The benchmark `code/models/tests/test_superpoint.py` still needs a tt-metal checkout as
143
- pytest rootdir for its `device` fixture -- it is unchanged by the packaging work.
 
 
144
 
145
  ## Package / serve / push (Blackhole host, rootless Docker)
146
 
@@ -182,6 +192,30 @@ worth keeping lives in `card.description` / `card.quickstart`); `media/`, `SERVI
182
  `.gitattributes` and `tt-model.yaml` at the root survive. The orchestrator restores
183
  `license`/`pipeline_tag` front matter after push.
184
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
185
  ## Caveats
186
 
187
  - **Licence**: the weights are Magic Leap "academic or non-profit organisation noncommercial
 
14
 
15
  ## What the server does
16
 
17
+ Default since 2026-09-13 (`TT_FUSED` unset or `1`, also pinned in `serve.env`): pure ttnn, the
18
+ whole device graph as **one metal trace** per request with a standard-op device NMS, no custom
19
+ kernel -- see *Fused path* below for the per-request sequence and the measured numbers.
20
+ `TT_FUSED=0` restores the legacy path this section describes: pure-ttnn, **untraced**, host
21
+ NMS -- the port's `SP_TRACE_NMS=0 SP_NO_TRACE=1` configuration. Both paths: fixed 480x640
22
+ network input; every image is resized server-side. Per request (legacy):
23
 
24
  1. base64 -> PIL RGB -> bilinear resize to 640x480 -> /255 -> fp32 `(1, 3, 480, 640)`
25
  (HF `SuperPointImageProcessor` defaults; the model reads channel 0 like
 
39
  frame) happens in the ASGI lifespan, so `Application startup complete` means warm. Shutdown
40
  deallocates the device tensors and `ttnn.close_device`s inside the 120 s SIGTERM budget.
41
 
42
+ Measured 2026-09-13 through the container image (50 warm requests, medians; legacy path):
43
+ device_forward 12.4 ms, host NMS + post 26.5 ms, total 56.8 ms. The README's 40.7 fps requires
44
+ the fused `sp_eq_mul_mask` C++ kernel (`code/kernels/`, needs a patched tt-metal) -- not what
45
+ this image runs; the fused path below gets to ~40 fps with standard ops instead.
46
 
47
  ## Environment the app reads (lifespan only, never at import)
48
 
 
53
  | `SP_WEIGHTS_DIR` | you (host/offline) | local dir with `config.json` + `model.safetensors`; overrides the two above |
54
  | `TT_MESH_SHAPE` | launcher (`runtime.mesh_shape_env`) | `1x1` (also `(1, 1)` / `1,1`); any other shape -> `RuntimeError` at startup |
55
  | `TT_DEVICE_ID` | you | chip to open, default `0` |
56
+ | `TT_FUSED` | `serve.env` (`"1"`; the code default when unset/empty is also fused) | unset/`1` = fused serving path: ONE metal trace per request (64-byte-page input upload, encoder + heads, `rms_norm` L2-norm, standard-op device NMS at radius 4, row-major outputs), captured during warm-up before READY. `0` = the legacy untraced path above, byte-identical to the 2026-09-12 image. See `DEVICE_VALIDATION.md` |
57
+ | `TT_FUSED_STAGES` | you (device A/B only) | comma list of fused stages, default all (`wide,nms,rms,rm`); `""` = trace-only |
58
+ | `SP_TRACE_REGION` | you | `trace_region_size` bytes for `ttnn.CreateDevice` on the fused path (default 32 MiB) |
59
  | `MESH_DEVICE` | launcher | `P150` (informational) |
60
  | `TT_METAL_VISIBLE_DEVICES` | `serve.env` | `0` |
61
 
 
68
  | route | response |
69
  |---|---|
70
  | `GET /health` | `{"status": "ok" \| "starting", "model": "superpoint-p150", "device": {"arch", "id", "open"}}` (always 200) |
71
+ | `GET /info` | model/task/io, hardware, `weights {repo, revision, local_dir, loaded}`, `source {repo, commit}`, `input` (480x640, batch 1, preprocessing), `defaults`, `limits`, `serving_path` (fused: `traced=true, device_nms=true, nms_radius_traced=4, fused_stages`; legacy: `traced=false, device_nms=false`), `warmup_ms`, `descriptors` encoding, `license` |
72
  | `GET /v1/models` | `{"object": "list", "data": [{"id": "<weights repo>", "object": "model", "owned_by": "changh95"}]}` (so OpenAI-shaped probes do not 404; not a chat API) |
73
  | `POST /predict` | see below |
74
 
 
147
 
148
  `PYTHONPATH` must start with `code/` so that `models` resolves to this repo's regular package
149
  (tt-metal's own `models/` has no `__init__.py` and is shadowed on the host; in the image it is
150
+ excluded). The device tests in `code/models/tests/test_superpoint.py` get their `device` /
151
+ `device_params` fixtures and `--device-id` from the repo-local `code/conftest.py` (run pytest
152
+ from `code/`; `code/pytest.ini` pins the rootdir). tt-metal's own conftest cannot be used next
153
+ to this repo: it imports `models.demos...`, which the shadowing above breaks.
154
 
155
  ## Package / serve / push (Blackhole host, rootless Docker)
156
 
 
192
  `.gitattributes` and `tt-model.yaml` at the root survive. The orchestrator restores
193
  `license`/`pipeline_tag` front matter after push.
194
 
195
+ ## Fused path (default; `TT_FUSED=0` = legacy) -- branch `opt/superpoint-p150-megakernel`
196
+
197
+ Device-validated on the p150a 2026-09-13 (`DEVICE_VALIDATION.md` "Results") and made the default
198
+ (code default + `serve.env.TT_FUSED: "1"`); `TT_FUSED=0` restores the legacy path above
199
+ byte-for-byte. On the fused path the lifespan opens the device with a trace region, runs the fused graph once
200
+ eagerly (kernel compile), captures it into a metal trace and replays it once -- all before
201
+ `Application startup complete` (boot log: `Warming up TT_FUSED path ...`, `Warmup complete:
202
+ compile forward ... trace capture ... traced forward ...`). Per request: H2D of the
203
+ `[1,1,9600,32]` bf16 input, one `execute_trace`, D2H of the row-major device NMS map (480x640)
204
+ and descriptors, then threshold / border / top-k / `grid_sample` on the host
205
+ (`postprocess_from_nms_map`). Requests with `nms_radius != 4` read the traced softmax scores
206
+ instead and run the legacy host NMS (same output, slower); `response.serving_path.device_nms`
207
+ and `/info.serving_path` report which path ran. Exactness: trace, upload, device NMS and
208
+ row-major outputs are bit-identical to the legacy path; the `rms_norm` descriptor L2-norm is
209
+ bf16-rounding-level (gate: descriptor PCC >= 0.999). Host proofs: `code/models/tests/test_fused_host.py`;
210
+ hardware plan, gates and results: `DEVICE_VALIDATION.md`. Measured 2026-09-13 through the
211
+ container image (`smoke_test.py` PASS, 50 warm requests, server `timing_ms` medians): device_forward
212
+ 5.3 ms (min 5.1), post-processing 1.3 ms, preprocess 18 ms, total 24.7 ms (~40 fps) vs the legacy
213
+ path's 12.4 / 26.5 / 17.8 / 56.8 ms in the same session; keypoints and scores identical to the
214
+ legacy server for `nms_radius` 4 (device NMS) and 3 (host fallback), descriptors within bf16
215
+ rounding (max |diff| 8.5e-4, cosine >= 0.999996). Boot log landmarks: `Warming up TT_FUSED path
216
+ (stages nms,rm,rms,wide, traced nms_radius 4)`, `Warmup complete: compile forward ... trace capture
217
+ ... traced forward ...`.
218
+
219
  ## Caveats
220
 
221
  - **Licence**: the weights are Magic Leap "academic or non-profit organisation noncommercial
tt-model.yaml CHANGED
@@ -60,6 +60,10 @@ serve:
60
  env:
61
  TT_WEIGHTS_REVISION: "734450e9ffe229074f5998494ddc615475cdb20a"
62
  TT_METAL_VISIBLE_DEVICES: "0"
 
 
 
 
63
 
64
  # Build-time assertions run INSIDE the finished image as uid 1000, no device, no weights.
65
  verify:
@@ -99,7 +103,8 @@ card:
99
  "scores": [0.609375, 0.589844, 0.582031],
100
  "original_size": {"height": 900, "width": 1600}, "image_size": {"height": 480, "width": 640}, "scale": {"x": 2.5, "y": 1.875},
101
  "descriptors": {"format": "npz", "key": "descriptors", "dtype": "float16", "shape": [539, 256], "data": "..."},
102
- "timing_ms": {"preprocess": 19.4, "device_forward": 12.5, "postprocess": 32.5, "total": 64.4}}
 
103
  ```
104
 
105
  - `keypoints` are `[x, y]` in original image pixels, sorted by descending `scores`; the network frame is 480×640 and `scale` = original / network.
@@ -116,16 +121,18 @@ card:
116
  | Metric | Value |
117
  |---|---:|
118
  | Pre-NMS score map · descriptor map PCC vs fp32 torch reference | 0.9971 · 0.9991 |
119
- | Keypoint set vs reference (natural image, top-500, 2 px) | recall 98.80% · precision 98.80% · F1 98.80% |
120
- | Inference, served over HTTP (warm, batch 1, 480×640, 1600×900 JPEG in) | ~12–20 ms device · ~65 ms end-to-end (~15 FPS) |
 
121
 
122
  ### Caveats
123
 
124
  - Every image is resized to 480×640 (bilinear, /255, channel 0); one image per request, batch 1, requests serialised on the chip.
125
- - This server runs the untraced pure-ttnn path with host single-pass NMS (bf16 + HiFi2 + fp32 accumulate; bfloat8/LoFi drop score PCC to ~0.91). The port README's 40.7 fps needs `ttnn.trace` plus the fused `sp_eq_mul_mask` C++ kernel in `code/kernels/`, which is not built into this image.
 
126
  - Weights are research-only: the Magic Leap SuperPoint licence allows academic or non-profit organisation NONCOMMERCIAL research use.
127
  - Not an OpenAI-compatible API; `GET /v1/models` is a stub so the tt-model ready card does not 404.
128
- - Validated on tt-metal `v0.78.0-dev20260820` (main `8b98410e730`), single p150a only.
129
 
130
  ### Licensing
131
 
 
60
  env:
61
  TT_WEIGHTS_REVISION: "734450e9ffe229074f5998494ddc615475cdb20a"
62
  TT_METAL_VISIBLE_DEVICES: "0"
63
+ # Fused serving path (one metal trace per request, standard-op device NMS), device-validated
64
+ # 2026-09-13 (DEVICE_VALIDATION.md "Results"); it is also the code default. "0" = the legacy
65
+ # untraced host-NMS path of the 2026-09-12 image.
66
+ TT_FUSED: "1"
67
 
68
  # Build-time assertions run INSIDE the finished image as uid 1000, no device, no weights.
69
  verify:
 
103
  "scores": [0.609375, 0.589844, 0.582031],
104
  "original_size": {"height": 900, "width": 1600}, "image_size": {"height": 480, "width": 640}, "scale": {"x": 2.5, "y": 1.875},
105
  "descriptors": {"format": "npz", "key": "descriptors", "dtype": "float16", "shape": [539, 256], "data": "..."},
106
+ "serving_path": {"traced": true, "device_nms": true},
107
+ "timing_ms": {"preprocess": 17.7, "device_forward": 5.3, "postprocess": 1.3, "total": 24.3}}
108
  ```
109
 
110
  - `keypoints` are `[x, y]` in original image pixels, sorted by descending `scores`; the network frame is 480×640 and `scale` = original / network.
 
121
  | Metric | Value |
122
  |---|---:|
123
  | Pre-NMS score map · descriptor map PCC vs fp32 torch reference | 0.9971 · 0.9991 |
124
+ | Keypoint set vs reference (natural image, top-500, 2 px) | recall 98.20% · precision 99.40% · F1 98.80% |
125
+ | Inference, served over HTTP (warm, batch 1, 480×640, 1600×900 JPEG in; median of 50 requests) | 5.3 ms device (trace + device NMS) · 1.3 ms host post-processing · 18 ms JPEG decode/resize · 24.7 ms end-to-end (~40 FPS) |
126
+ | Same, legacy path (`TT_FUSED=0`: untraced, host NMS) | 12.4 ms device · 26.5 ms host NMS · 56.8 ms end-to-end (~18 FPS) |
127
 
128
  ### Caveats
129
 
130
  - Every image is resized to 480×640 (bilinear, /255, channel 0); one image per request, batch 1, requests serialised on the chip.
131
+ - This server runs the whole device graph as one metal trace per request with a standard-op device NMS (radius 4; bit-identical to the host single-pass NMS) and an `rms_norm` descriptor L2-norm (bf16-rounding-level vs the legacy chain; PCC 0.9991 either way). bf16 + HiFi2 + fp32 accumulate throughout (bfloat8/LoFi drop score PCC to ~0.91). No custom kernel: the port README's `sp_eq_mul_mask` path is not built into this image.
132
+ - `nms_radius` other than 4 falls back to the host NMS on the traced scores (same keypoints, ~25 ms slower); `TT_FUSED=0` in the environment restores the untraced legacy path (validated 2026-09-12).
133
  - Weights are research-only: the Magic Leap SuperPoint licence allows academic or non-profit organisation NONCOMMERCIAL research use.
134
  - Not an OpenAI-compatible API; `GET /v1/models` is a stub so the tt-model ready card does not 404.
135
+ - Validated on tt-metal `v0.78.0-dev20260820` (main `8b98410e730`), single p150a only; numbers above measured 2026-09-13 through this container image (`DEVICE_VALIDATION.md`).
136
 
137
  ### Licensing
138