changh95 commited on
Commit
d471346
·
verified ·
1 Parent(s): bd29268

tt-model authoring file: compact card

Browse files
Files changed (1) hide show
  1. tt-model.yaml +39 -103
tt-model.yaml CHANGED
@@ -72,126 +72,62 @@ verify:
72
 
73
  card:
74
  description: >
75
- SuperPoint (magic-leap-community/superpoint) keypoint detection and description on a
76
- single Tenstorrent Blackhole p150a via tt-nn: a base64 image in, keypoints in original
77
- pixel coordinates, scores and 256-d descriptors out, at a fixed 480x640 network input.
78
- Pre-NMS score-map PCC 0.9971 and descriptor PCC 0.9991 vs the fp32 torch reference,
79
- keypoint F1 98.8% (top-500, 2 px); this server runs the pure-ttnn untraced path with host
80
- NMS (roughly 6 fps; the README's 40.7 fps needs trace plus the fused C++ NMS kernel in
81
- kernels/, which is not built into this image). Weights are under the Magic Leap
82
- SuperPoint licence: academic or non-profit organisation NONCOMMERCIAL research use only.
83
- Port source: github.com/changh95/tt-superpoint @ e1eab66e29ff424bc9af6b1118671d9bc08e899e.
84
  quickstart: |
85
- ### With tt-cli
86
 
87
  ```bash
88
- tt serve changh95/superpoint-p150 # pulls image + weights, boots, prints the port
 
 
89
  tt model stop changh95/superpoint-p150
90
  ```
91
 
92
- The port is the one `serve` printed (20000, or the next free one). This is **not** an
93
- OpenAI-style server: `tt-model curl` / chat clients do not apply. `GET /v1/models` exists
94
- only so generic probes do not 404; the real routes are below.
95
 
96
- ### Call it
97
 
98
- ```bash
99
- PORT=20000 # the port serve printed
100
- curl -s localhost:$PORT/health # {"status":"ok","model":"superpoint-p150","device":{...}}
101
- curl -s localhost:$PORT/info # weights repo+revision, canonical input, defaults, limits, licence
102
-
103
- # One image -> keypoints. Any PNG/JPEG; it is resized to 640x480 server-side and the
104
- # keypoints come back in YOUR image's pixel coordinates (see "scale").
105
- IMG=code/sample_data/house_in_field_1080p.jpg # 1600x900 natural image
106
- python3 - "$IMG" "$PORT" <<'EOF'
107
- import base64, json, sys, urllib.request
108
- img, port = sys.argv[1], sys.argv[2]
109
- req = {"image": base64.b64encode(open(img, "rb").read()).decode(),
110
- "max_keypoints": 1024, # -1 = all above threshold
111
- "keypoint_threshold": 0.005,
112
- "nms_radius": 4,
113
- "return_descriptors": True}
114
- r = urllib.request.Request(f"http://127.0.0.1:{port}/predict", json.dumps(req).encode(),
115
- {"Content-Type": "application/json"})
116
- out = json.load(urllib.request.urlopen(r, timeout=300))
117
- print(out["num_keypoints"], out["keypoints"][:3], out["scores"][:3], out["timing_ms"])
118
- EOF
119
  ```
120
 
121
- Response fields: `num_keypoints`; `keypoints` (N x [x, y] floats, original-image pixels);
122
- `scores` (N, post-NMS softmax scores); `original_size`, `image_size` (480x640) and
123
- `scale` (x = W/640, y = H/480; divide to get network-frame coordinates);
124
- `params` echoed; `timing_ms` (`preprocess`, `device_forward`, `postprocess`, `total`);
125
- when `return_descriptors` is true, `descriptors` = `{format: "npz", key: "descriptors",
126
- dtype: "float16", shape: [N, 256], data: <base64 NPZ>}` -- decode with
127
- `numpy.load(io.BytesIO(base64.b64decode(d["data"])))["descriptors"]`; rows are
128
- L2-normalised. Errors: 400 undecodable image / bad field, 503 while starting, 500 with
129
- the exception text. One image per request; requests are serialised on the chip.
130
-
131
- A ready-made check: `python code/models/server/smoke_test.py --url http://127.0.0.1:$PORT`
132
- prints one PASS/FAIL line with the keypoint count and timings.
133
-
134
- ### First boot
135
-
136
- Weights are 5 MB (`config.json`, `model.safetensors`, `preprocessor_config.json` at the
137
- pinned revision) and land in your HF cache. The first start JIT-compiles the conv /
138
- pool / softmax kernels (a few minutes, cached under
139
- `~/.cache/tt-model/superpoint-p150/cache`); the server logs `Loading weights`,
140
- `Warming up`, `Warmup complete` and is ready at uvicorn's `Application startup complete`.
141
- Later boots reuse the kernel cache. The weights repo is public and ungated (no token).
142
-
143
- ### What this server runs
144
 
145
- Fixed 480x640 input (bilinear resize, /255, channel 0 -- the HF
146
- `SuperPointImageProcessor` defaults), 8 encoder convs + 3 max-pools + score and
147
- descriptor heads in bfloat16 activations / bfloat16 weights / HiFi2 / fp32 accumulate,
148
- softmax and descriptor L2-norm on device, then host single-pass NMS, threshold, border
149
- removal, top-k and bilinear descriptor sampling. Untraced, host NMS: this is the
150
- port's `SP_TRACE_NMS=0`, `SP_NO_TRACE=1` configuration. The benchmark numbers below
151
- that need trace or the fused `sp_eq_mul_mask` kernel are **not** what this server does.
152
 
153
- ### Results from the port (Blackhole p150b, 480x640, batch 1, natural image)
 
 
154
 
155
- PCC vs the Hugging Face fp32 reference:
156
 
157
- | Tensor | PCC |
158
  |---|---:|
159
- | Pre-NMS score map | **0.9971** |
160
- | Descriptor map (post L2-norm) | **0.9991** |
161
-
162
- Keypoint set vs the reference (`sample_data/house_in_field_1080p.jpg`, top-500, 2 px):
163
- recall **98.80%**, precision **98.80%**, F1 **98.80%** (synthetic `torch.rand` input: F1 97.80%).
164
-
165
- Throughput measured by `models/tests/test_superpoint.py` (needs a tt-metal checkout as
166
- pytest rootdir for its `device` fixture):
167
-
168
- | Path | Device forward (input resident) | Traced fwd incl. H2D | Full e2e |
169
- |---|---:|---:|---:|
170
- | Forward-only trace, host NMS (`SP_TRACE_NMS=0`) | 355 fps (2.81 ms) | 73.6 fps | 17.0 fps |
171
- | Forward + device NMS, fused kernel (`SP_TRACE_NMS=1`) | 85.6 fps | 44.6 fps | **40.7 fps** |
172
- | Untraced, host NMS (**this server**) | ~6 fps | -- | ~5 fps |
173
-
174
- The fused path needs `kernels/sp_eq_mul_mask/` compiled into tt-metal (see its README);
175
- a pure-ttnn equivalent (`ttnn.eq` + `ttnn.multiply`) is measured in
176
- `kernels/sp_eq_mul_mask/bench.py`. `code/results.tsv` is the full experiment log
177
- (precision cliff: bfloat8/LoFi drop score PCC to 0.70-0.91; trace was a 10.8x unlock).
178
-
179
- Sample output (top-500 keypoints on the resized frame): ![keypoints](media/sample.png)
180
 
181
- ### Layout of `code/`
182
 
183
- `models/tt/superpoint_ttnn.py` (tt-nn model), `models/tt/postprocess.py` (validated host
184
- post-processing), `models/server/app.py` + `smoke_test.py`, `models/reference/`
185
- (HF reference loader, needs torchvision), `models/tests/test_superpoint.py` (benchmark +
186
- PCC + F1), `models/visualize.py`, `kernels/sp_eq_mul_mask/` (fused C++ Tensix NMS
187
- kernel, ~450 LoC), `sample_data/`, `results.tsv`, `run_benchmark.sh`.
188
- `models/common/lightweightmodule.py` is a schema-required filler from tt-metal.
189
 
190
  ### Licensing
191
 
192
- The upstream **weights** (`magic-leap-community/superpoint`) carry the Magic Leap
193
- SuperPoint licence: *academic or non-profit organisation noncommercial research use
194
- only* -- see https://huggingface.co/magic-leap-community/superpoint. The port code
195
- (Apache-2.0 headers, by Hyunggi Chang) is published under the same terms, since a port
196
- cannot grant more than its upstream does. Weights are not redistributed here; they are
197
- fetched from the upstream repo at the pinned revision.
 
72
 
73
  card:
74
  description: >
75
+ SuperPoint (Magic Leap's self-supervised interest-point detector and descriptor) running entirely on one Tenstorrent Blackhole p150a via tt-nn: image in, keypoints with scores and 256-d descriptors out.
76
+
77
+ Weights: [magic-leap-community/superpoint](https://huggingface.co/magic-leap-community/superpoint) ·
78
+ Paper: [arXiv:1712.07629](https://arxiv.org/abs/1712.07629) ·
79
+ Upstream code: [magicleap/SuperPointPretrainedNetwork](https://github.com/magicleap/SuperPointPretrainedNetwork) ·
80
+ Port: [changh95/tt-superpoint](https://github.com/changh95/tt-superpoint)
 
 
 
81
  quickstart: |
82
+ ### Run with tt-cli
83
 
84
  ```bash
85
+ tt serve changh95/superpoint-p150
86
+ printf '{"image":"%s"}' "$(base64 -w0 code/sample_data/house_in_field_1080p.jpg)" > req.json
87
+ curl -s localhost:20000/predict -H 'Content-Type: application/json' -d @req.json
88
  tt model stop changh95/superpoint-p150
89
  ```
90
 
91
+ - `POST /predict`: `image` (base64 PNG/JPEG); optional `max_keypoints` (1024, `-1` = all above threshold), `keypoint_threshold` (0.005), `nms_radius` (4), `return_descriptors` (true).
92
+ - `GET /health`, `GET /info`.
 
93
 
94
+ ### Response
95
 
96
+ ```json
97
+ {"num_keypoints": 539,
98
+ "keypoints": [[610.0, 703.125], [1042.5, 446.25], [1122.5, 442.5]],
99
+ "scores": [0.609375, 0.589844, 0.582031],
100
+ "original_size": {"height": 900, "width": 1600}, "image_size": {"height": 480, "width": 640}, "scale": {"x": 2.5, "y": 1.875},
101
+ "descriptors": {"format": "npz", "key": "descriptors", "dtype": "float16", "shape": [539, 256], "data": "..."},
102
+ "timing_ms": {"preprocess": 19.4, "device_forward": 12.5, "postprocess": 32.5, "total": 64.4}}
 
 
 
 
 
 
 
 
 
 
 
 
 
 
103
  ```
104
 
105
+ - `keypoints` are `[x, y]` in original image pixels, sorted by descending `scores`; the network frame is 480×640 and `scale` = original / network.
106
+ - `descriptors.data` is a base64 NPZ: `np.load(io.BytesIO(base64.b64decode(data)))["descriptors"]` gives `(N, 256)` float16 rows, L2-normalised, in keypoint order.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
107
 
108
+ ### Demo
 
 
 
 
 
 
109
 
110
+ | Top-500 keypoints on the 480×640 network frame of `code/sample_data/house_in_field_1080p.jpg` (`media/sample.png`, natural image) |
111
+ |:---:|
112
+ | ![](media/sample.png) |
113
 
114
+ ### Accuracy and speed
115
 
116
+ | Metric | Value |
117
  |---|---:|
118
+ | Pre-NMS score map · descriptor map PCC vs fp32 torch reference | 0.9971 · 0.9991 |
119
+ | Keypoint set vs reference (natural image, top-500, 2 px) | recall 98.80% · precision 98.80% · F1 98.80% |
120
+ | Inference, served over HTTP (warm, batch 1, 480×640, 1600×900 JPEG in) | ~12–20 ms device · ~65 ms end-to-end (~15 FPS) |
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
121
 
122
+ ### Caveats
123
 
124
+ - Every image is resized to 480×640 (bilinear, /255, channel 0); one image per request, batch 1, requests serialised on the chip.
125
+ - This server runs the untraced pure-ttnn path with host single-pass NMS (bf16 + HiFi2 + fp32 accumulate; bfloat8/LoFi drop score PCC to ~0.91). The port README's 40.7 fps needs `ttnn.trace` plus the fused `sp_eq_mul_mask` C++ kernel in `code/kernels/`, which is not built into this image.
126
+ - Weights are research-only: the Magic Leap SuperPoint licence allows academic or non-profit organisation NONCOMMERCIAL research use.
127
+ - Not an OpenAI-compatible API; `GET /v1/models` is a stub so the tt-model ready card does not 404.
128
+ - Validated on tt-metal `v0.78.0-dev20260820` (main `8b98410e730`), single p150a only.
 
129
 
130
  ### Licensing
131
 
132
+ - Weights: [magic-leap-community/superpoint](https://huggingface.co/magic-leap-community/superpoint), `other` ([Magic Leap SuperPoint licence](https://huggingface.co/magic-leap-community/superpoint/blob/main/LICENSE), noncommercial research use only); not redistributed here.
133
+ - Port and serving code (`code/`): Apache-2.0 (SPDX headers on the modules), from [changh95/tt-superpoint](https://github.com/changh95/tt-superpoint), distributed under the same upstream terms since a port cannot grant more than its upstream does.