Add files using upload-large-folder tool
Browse files- README.md +36 -109
- tt_kernel_manifest.json +2 -2
README.md
CHANGED
|
@@ -1,28 +1,16 @@
|
|
| 1 |
---
|
| 2 |
tags:
|
| 3 |
- blackhole
|
| 4 |
-
- keypoint-detection
|
| 5 |
- p150
|
| 6 |
-
- superpoint
|
| 7 |
-
- tenstorrent
|
| 8 |
- tt-dit-server
|
| 9 |
-
- tt-metal
|
| 10 |
- tt-model-cache
|
| 11 |
-
- tt-model-catalog
|
| 12 |
- tt-model-container
|
| 13 |
-
- tt-nn
|
| 14 |
-
- ttnn
|
| 15 |
-
base_model:
|
| 16 |
-
- magic-leap-community/superpoint
|
| 17 |
-
license: other
|
| 18 |
-
license_name: magic-leap-superpoint
|
| 19 |
-
license_link: https://huggingface.co/magic-leap-community/superpoint
|
| 20 |
-
pipeline_tag: keypoint-detection
|
| 21 |
---
|
| 22 |
|
| 23 |
# superpoint-p150
|
| 24 |
|
| 25 |
-
SuperPoint (
|
|
|
|
| 26 |
|
| 27 |
Runs on **p150** (mesh `P150`).
|
| 28 |
|
|
@@ -37,119 +25,58 @@ tt-model serve changh95/superpoint-p150
|
|
| 37 |
|
| 38 |
`pull --with-weights` downloads the Docker image and the [`magic-leap-community/superpoint`](https://huggingface.co/magic-leap-community/superpoint) weights at `734450e9ffe229074f5998494ddc615475cdb20a` (into your HF cache; they are not in the image). `serve` starts the model's own HTTP server on port 20000 (or the next free port, if that one is busy); the first start compiles kernels for your device, which takes several minutes, and the server is ready when it logs `Application startup complete`.
|
| 39 |
|
| 40 |
-
###
|
| 41 |
|
| 42 |
```bash
|
| 43 |
-
tt serve changh95/superpoint-p150
|
|
|
|
|
|
|
| 44 |
tt model stop changh95/superpoint-p150
|
| 45 |
```
|
| 46 |
|
| 47 |
-
|
| 48 |
-
|
| 49 |
-
only so generic probes do not 404; the real routes are below.
|
| 50 |
|
| 51 |
-
###
|
| 52 |
|
| 53 |
-
```
|
| 54 |
-
|
| 55 |
-
|
| 56 |
-
|
| 57 |
-
|
| 58 |
-
|
| 59 |
-
|
| 60 |
-
IMG=code/sample_data/house_in_field_1080p.jpg # 1600x900 natural image
|
| 61 |
-
python3 - "$IMG" "$PORT" <<'EOF'
|
| 62 |
-
import base64, json, sys, urllib.request
|
| 63 |
-
img, port = sys.argv[1], sys.argv[2]
|
| 64 |
-
req = {"image": base64.b64encode(open(img, "rb").read()).decode(),
|
| 65 |
-
"max_keypoints": 1024, # -1 = all above threshold
|
| 66 |
-
"keypoint_threshold": 0.005,
|
| 67 |
-
"nms_radius": 4,
|
| 68 |
-
"return_descriptors": True}
|
| 69 |
-
r = urllib.request.Request(f"http://127.0.0.1:{port}/predict", json.dumps(req).encode(),
|
| 70 |
-
{"Content-Type": "application/json"})
|
| 71 |
-
out = json.load(urllib.request.urlopen(r, timeout=300))
|
| 72 |
-
print(out["num_keypoints"], out["keypoints"][:3], out["scores"][:3], out["timing_ms"])
|
| 73 |
-
EOF
|
| 74 |
```
|
| 75 |
|
| 76 |
-
|
| 77 |
-
`
|
| 78 |
-
`scale` (x = W/640, y = H/480; divide to get network-frame coordinates);
|
| 79 |
-
`params` echoed; `timing_ms` (`preprocess`, `device_forward`, `postprocess`, `total`);
|
| 80 |
-
when `return_descriptors` is true, `descriptors` = `{format: "npz", key: "descriptors",
|
| 81 |
-
dtype: "float16", shape: [N, 256], data: <base64 NPZ>}` -- decode with
|
| 82 |
-
`numpy.load(io.BytesIO(base64.b64decode(d["data"])))["descriptors"]`; rows are
|
| 83 |
-
L2-normalised. Errors: 400 undecodable image / bad field, 503 while starting, 500 with
|
| 84 |
-
the exception text. One image per request; requests are serialised on the chip.
|
| 85 |
-
|
| 86 |
-
A ready-made check: `python code/models/server/smoke_test.py --url http://127.0.0.1:$PORT`
|
| 87 |
-
prints one PASS/FAIL line with the keypoint count and timings.
|
| 88 |
-
|
| 89 |
-
### First boot
|
| 90 |
-
|
| 91 |
-
Weights are 5 MB (`config.json`, `model.safetensors`, `preprocessor_config.json` at the
|
| 92 |
-
pinned revision) and land in your HF cache. The first start JIT-compiles the conv /
|
| 93 |
-
pool / softmax kernels (a few minutes, cached under
|
| 94 |
-
`~/.cache/tt-model/superpoint-p150/cache`); the server logs `Loading weights`,
|
| 95 |
-
`Warming up`, `Warmup complete` and is ready at uvicorn's `Application startup complete`.
|
| 96 |
-
Later boots reuse the kernel cache. The weights repo is public and ungated (no token).
|
| 97 |
-
|
| 98 |
-
### What this server runs
|
| 99 |
|
| 100 |
-
|
| 101 |
-
`SuperPointImageProcessor` defaults), 8 encoder convs + 3 max-pools + score and
|
| 102 |
-
descriptor heads in bfloat16 activations / bfloat16 weights / HiFi2 / fp32 accumulate,
|
| 103 |
-
softmax and descriptor L2-norm on device, then host single-pass NMS, threshold, border
|
| 104 |
-
removal, top-k and bilinear descriptor sampling. Untraced, host NMS: this is the
|
| 105 |
-
port's `SP_TRACE_NMS=0`, `SP_NO_TRACE=1` configuration. The benchmark numbers below
|
| 106 |
-
that need trace or the fused `sp_eq_mul_mask` kernel are **not** what this server does.
|
| 107 |
|
| 108 |
-
|
|
|
|
|
|
|
| 109 |
|
| 110 |
-
|
| 111 |
|
| 112 |
-
|
|
| 113 |
|---|---:|
|
| 114 |
-
| Pre-NMS score map |
|
| 115 |
-
|
|
| 116 |
-
|
| 117 |
-
Keypoint set vs the reference (`sample_data/house_in_field_1080p.jpg`, top-500, 2 px):
|
| 118 |
-
recall **98.80%**, precision **98.80%**, F1 **98.80%** (synthetic `torch.rand` input: F1 97.80%).
|
| 119 |
-
|
| 120 |
-
Throughput measured by `models/tests/test_superpoint.py` (needs a tt-metal checkout as
|
| 121 |
-
pytest rootdir for its `device` fixture):
|
| 122 |
-
|
| 123 |
-
| Path | Device forward (input resident) | Traced fwd incl. H2D | Full e2e |
|
| 124 |
-
|---|---:|---:|---:|
|
| 125 |
-
| Forward-only trace, host NMS (`SP_TRACE_NMS=0`) | 355 fps (2.81 ms) | 73.6 fps | 17.0 fps |
|
| 126 |
-
| Forward + device NMS, fused kernel (`SP_TRACE_NMS=1`) | 85.6 fps | 44.6 fps | **40.7 fps** |
|
| 127 |
-
| Untraced, host NMS (**this server**) | ~6 fps | -- | ~5 fps |
|
| 128 |
-
|
| 129 |
-
The fused path needs `kernels/sp_eq_mul_mask/` compiled into tt-metal (see its README);
|
| 130 |
-
a pure-ttnn equivalent (`ttnn.eq` + `ttnn.multiply`) is measured in
|
| 131 |
-
`kernels/sp_eq_mul_mask/bench.py`. `code/results.tsv` is the full experiment log
|
| 132 |
-
(precision cliff: bfloat8/LoFi drop score PCC to 0.70-0.91; trace was a 10.8x unlock).
|
| 133 |
-
|
| 134 |
-
Sample output (top-500 keypoints on the resized frame): 
|
| 135 |
|
| 136 |
-
###
|
| 137 |
|
| 138 |
-
|
| 139 |
-
|
| 140 |
-
|
| 141 |
-
|
| 142 |
-
|
| 143 |
-
`models/common/lightweightmodule.py` is a schema-required filler from tt-metal.
|
| 144 |
|
| 145 |
### Licensing
|
| 146 |
|
| 147 |
-
|
| 148 |
-
|
| 149 |
-
only* -- see https://huggingface.co/magic-leap-community/superpoint. The port code
|
| 150 |
-
(Apache-2.0 headers, by Hyunggi Chang) is published under the same terms, since a port
|
| 151 |
-
cannot grant more than its upstream does. Weights are not redistributed here; they are
|
| 152 |
-
fetched from the upstream repo at the pinned revision.
|
| 153 |
|
| 154 |
## Provenance
|
| 155 |
|
|
@@ -159,5 +86,5 @@ The exact sources the image was built from — `code/` in this repo is byte-iden
|
|
| 159 |
| --- | --- |
|
| 160 |
| tt-metal | [`8b98410e730bb504fea43a88609756e34821d91d`](https://github.com/tenstorrent/tt-metal/commit/8b98410e730bb504fea43a88609756e34821d91d) |
|
| 161 |
| `code/` digest | `006ddc1eb4e69a41` (sha256, first 16 hex digits) |
|
| 162 |
-
| built | 2026-09-
|
| 163 |
|
|
|
|
| 1 |
---
|
| 2 |
tags:
|
| 3 |
- blackhole
|
|
|
|
| 4 |
- p150
|
|
|
|
|
|
|
| 5 |
- tt-dit-server
|
|
|
|
| 6 |
- tt-model-cache
|
|
|
|
| 7 |
- tt-model-container
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 8 |
---
|
| 9 |
|
| 10 |
# superpoint-p150
|
| 11 |
|
| 12 |
+
SuperPoint (Magic Leap's self-supervised interest-point detector and descriptor) running entirely on one Tenstorrent Blackhole p150a via tt-nn: image in, keypoints with scores and 256-d descriptors out.
|
| 13 |
+
Weights: [magic-leap-community/superpoint](https://huggingface.co/magic-leap-community/superpoint) · Paper: [arXiv:1712.07629](https://arxiv.org/abs/1712.07629) · Upstream code: [magicleap/SuperPointPretrainedNetwork](https://github.com/magicleap/SuperPointPretrainedNetwork) · Port: [changh95/tt-superpoint](https://github.com/changh95/tt-superpoint)
|
| 14 |
|
| 15 |
Runs on **p150** (mesh `P150`).
|
| 16 |
|
|
|
|
| 25 |
|
| 26 |
`pull --with-weights` downloads the Docker image and the [`magic-leap-community/superpoint`](https://huggingface.co/magic-leap-community/superpoint) weights at `734450e9ffe229074f5998494ddc615475cdb20a` (into your HF cache; they are not in the image). `serve` starts the model's own HTTP server on port 20000 (or the next free port, if that one is busy); the first start compiles kernels for your device, which takes several minutes, and the server is ready when it logs `Application startup complete`.
|
| 27 |
|
| 28 |
+
### Run with tt-cli
|
| 29 |
|
| 30 |
```bash
|
| 31 |
+
tt serve changh95/superpoint-p150
|
| 32 |
+
printf '{"image":"%s"}' "$(base64 -w0 code/sample_data/house_in_field_1080p.jpg)" > req.json
|
| 33 |
+
curl -s localhost:20000/predict -H 'Content-Type: application/json' -d @req.json
|
| 34 |
tt model stop changh95/superpoint-p150
|
| 35 |
```
|
| 36 |
|
| 37 |
+
- `POST /predict`: `image` (base64 PNG/JPEG); optional `max_keypoints` (1024, `-1` = all above threshold), `keypoint_threshold` (0.005), `nms_radius` (4), `return_descriptors` (true).
|
| 38 |
+
- `GET /health`, `GET /info`.
|
|
|
|
| 39 |
|
| 40 |
+
### Response
|
| 41 |
|
| 42 |
+
```json
|
| 43 |
+
{"num_keypoints": 539,
|
| 44 |
+
"keypoints": [[610.0, 703.125], [1042.5, 446.25], [1122.5, 442.5]],
|
| 45 |
+
"scores": [0.609375, 0.589844, 0.582031],
|
| 46 |
+
"original_size": {"height": 900, "width": 1600}, "image_size": {"height": 480, "width": 640}, "scale": {"x": 2.5, "y": 1.875},
|
| 47 |
+
"descriptors": {"format": "npz", "key": "descriptors", "dtype": "float16", "shape": [539, 256], "data": "..."},
|
| 48 |
+
"timing_ms": {"preprocess": 19.4, "device_forward": 12.5, "postprocess": 32.5, "total": 64.4}}
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 49 |
```
|
| 50 |
|
| 51 |
+
- `keypoints` are `[x, y]` in original image pixels, sorted by descending `scores`; the network frame is 480×640 and `scale` = original / network.
|
| 52 |
+
- `descriptors.data` is a base64 NPZ: `np.load(io.BytesIO(base64.b64decode(data)))["descriptors"]` gives `(N, 256)` float16 rows, L2-normalised, in keypoint order.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 53 |
|
| 54 |
+
### Demo
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 55 |
|
| 56 |
+
| Top-500 keypoints on the 480×640 network frame of `code/sample_data/house_in_field_1080p.jpg` (`media/sample.png`, natural image) |
|
| 57 |
+
|:---:|
|
| 58 |
+
|  |
|
| 59 |
|
| 60 |
+
### Accuracy and speed
|
| 61 |
|
| 62 |
+
| Metric | Value |
|
| 63 |
|---|---:|
|
| 64 |
+
| Pre-NMS score map · descriptor map PCC vs fp32 torch reference | 0.9971 · 0.9991 |
|
| 65 |
+
| Keypoint set vs reference (natural image, top-500, 2 px) | recall 98.80% · precision 98.80% · F1 98.80% |
|
| 66 |
+
| Inference, served over HTTP (warm, batch 1, 480×640, 1600×900 JPEG in) | ~12–20 ms device · ~65 ms end-to-end (~15 FPS) |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 67 |
|
| 68 |
+
### Caveats
|
| 69 |
|
| 70 |
+
- Every image is resized to 480×640 (bilinear, /255, channel 0); one image per request, batch 1, requests serialised on the chip.
|
| 71 |
+
- This server runs the untraced pure-ttnn path with host single-pass NMS (bf16 + HiFi2 + fp32 accumulate; bfloat8/LoFi drop score PCC to ~0.91). The port README's 40.7 fps needs `ttnn.trace` plus the fused `sp_eq_mul_mask` C++ kernel in `code/kernels/`, which is not built into this image.
|
| 72 |
+
- Weights are research-only: the Magic Leap SuperPoint licence allows academic or non-profit organisation NONCOMMERCIAL research use.
|
| 73 |
+
- Not an OpenAI-compatible API; `GET /v1/models` is a stub so the tt-model ready card does not 404.
|
| 74 |
+
- Validated on tt-metal `v0.78.0-dev20260820` (main `8b98410e730`), single p150a only.
|
|
|
|
| 75 |
|
| 76 |
### Licensing
|
| 77 |
|
| 78 |
+
- Weights: [magic-leap-community/superpoint](https://huggingface.co/magic-leap-community/superpoint), `other` ([Magic Leap SuperPoint licence](https://huggingface.co/magic-leap-community/superpoint/blob/main/LICENSE), noncommercial research use only); not redistributed here.
|
| 79 |
+
- Port and serving code (`code/`): Apache-2.0 (SPDX headers on the modules), from [changh95/tt-superpoint](https://github.com/changh95/tt-superpoint), distributed under the same upstream terms since a port cannot grant more than its upstream does.
|
|
|
|
|
|
|
|
|
|
|
|
|
| 80 |
|
| 81 |
## Provenance
|
| 82 |
|
|
|
|
| 86 |
| --- | --- |
|
| 87 |
| tt-metal | [`8b98410e730bb504fea43a88609756e34821d91d`](https://github.com/tenstorrent/tt-metal/commit/8b98410e730bb504fea43a88609756e34821d91d) |
|
| 88 |
| `code/` digest | `006ddc1eb4e69a41` (sha256, first 16 hex digits) |
|
| 89 |
+
| built | 2026-09-12T13:22:54+00:00 by tt-model 0.1.0 |
|
| 90 |
|
tt_kernel_manifest.json
CHANGED
|
@@ -6,7 +6,7 @@
|
|
| 6 |
"device_count": 1,
|
| 7 |
"producer": {
|
| 8 |
"tt_kernel_version": "0.1.0",
|
| 9 |
-
"created_at": "2026-09-
|
| 10 |
"hostname": "deepgadget"
|
| 11 |
},
|
| 12 |
"weights": {
|
|
@@ -93,7 +93,7 @@
|
|
| 93 |
"image": "tt-model/superpoint-p150:0544890bca09",
|
| 94 |
"repo": "changh95/superpoint-p150",
|
| 95 |
"tt_model_version": "0.1.0",
|
| 96 |
-
"created_at": "2026-09-
|
| 97 |
"tt_metal": {
|
| 98 |
"sha": "8b98410e730bb504fea43a88609756e34821d91d",
|
| 99 |
"describe": "v0.78.0-dev20260820-25-g8b98410e73",
|
|
|
|
| 6 |
"device_count": 1,
|
| 7 |
"producer": {
|
| 8 |
"tt_kernel_version": "0.1.0",
|
| 9 |
+
"created_at": "2026-09-12T13:23:20.741264+00:00",
|
| 10 |
"hostname": "deepgadget"
|
| 11 |
},
|
| 12 |
"weights": {
|
|
|
|
| 93 |
"image": "tt-model/superpoint-p150:0544890bca09",
|
| 94 |
"repo": "changh95/superpoint-p150",
|
| 95 |
"tt_model_version": "0.1.0",
|
| 96 |
+
"created_at": "2026-09-12T13:22:54+00:00",
|
| 97 |
"tt_metal": {
|
| 98 |
"sha": "8b98410e730bb504fea43a88609756e34821d91d",
|
| 99 |
"describe": "v0.78.0-dev20260820-25-g8b98410e73",
|