File size: 7,949 Bytes
f64e572 3696119 4c179e2 f64e572 42b7595 f64e572 c699c4c 736060c f64e572 3696119 f64e572 3696119 f64e572 3696119 f64e572 3696119 42b7595 3696119 f64e572 c699c4c eee3601 f64e572 736060c f64e572 3696119 736060c 42b7595 3696119 f64e572 736060c f64e572 736060c f64e572 736060c d69d96a f64e572 736060c f64e572 736060c f64e572 736060c f64e572 c699c4c f64e572 c699c4c f64e572 c699c4c f64e572 736060c f64e572 c699c4c 736060c c699c4c 736060c f64e572 3696119 736060c c699c4c f64e572 3696119 f64e572 c699c4c f64e572 3696119 c699c4c d69d96a | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 | ---
tags:
- blackhole
- p150
- tt-dit-server
- tt-model-cache
- tt-model-container
- tenstorrent
- ttnn
- tt-metal
- tt-nn
- keypoint-detection
- superpoint
- tt-model-catalog
pipeline_tag: keypoint-detection
license_link: https://huggingface.co/magic-leap-community/superpoint
license: other
license_name: magic-leap-superpoint
base_model:
- magic-leap-community/superpoint
---
# superpoint-p150
Magic Leap SuperPoint port on one Tenstorrent Blackhole p150a.
Weights: [magic-leap-community/superpoint](https://huggingface.co/magic-leap-community/superpoint) · Paper: [arXiv:1712.07629](https://arxiv.org/abs/1712.07629) · Upstream code: [magicleap/SuperPointPretrainedNetwork](https://github.com/magicleap/SuperPointPretrainedNetwork) · Port: [changh95/tt-superpoint](https://github.com/changh95/tt-superpoint)
Runs on **p150** (mesh `P150`).
Packaged and published with [tt-model-manager](https://github.com/tenstorrent/tt-model-manager) 0.1.0 (manifest schema 5.1).
## Quickstart
```bash
tt-model pull changh95/superpoint-p150 --with-weights
tt-model serve changh95/superpoint-p150
```
- Weights [`magic-leap-community/superpoint`](https://huggingface.co/magic-leap-community/superpoint) go to your HF cache; the image does not contain them.
- Serves on port 20000 (or the next free port); ready when the log says `Application startup complete`.
### Run with tt-cli
```bash
tt serve changh95/superpoint-p150
printf '{"image":"%s"}' "$(base64 -w0 code/sample_data/house_in_field_1080p.jpg)" > req.json
curl -s localhost:20000/predict -H 'Content-Type: application/json' -d @req.json
tt model stop changh95/superpoint-p150
```
- `POST /predict`: `image` (base64 PNG/JPEG); optional `max_keypoints` (1024, `-1` = all above threshold), `keypoint_threshold` (0.005), `nms_radius` (4), `return_descriptors` (true).
- `GET /health`, `GET /info`.
### Response
```json
{"num_keypoints": 539,
"keypoints": [[610.0, 703.125], [1042.5, 446.25], [1122.5, 442.5]],
"scores": [0.609375, 0.589844, 0.582031],
"original_size": {"height": 900, "width": 1600}, "image_size": {"height": 480, "width": 640}, "scale": {"x": 2.5, "y": 1.875},
"descriptors": {"format": "npz", "key": "descriptors", "dtype": "float16", "shape": [539, 256], "data": "..."},
"serving_path": {"traced": true, "device_nms": true},
"timing_ms": {"preprocess": 17.7, "device_forward": 5.3, "postprocess": 1.3, "total": 24.3}}
```
- `keypoints` are `[x, y]` in original image pixels, sorted by descending `scores`; the network frame is 480×640 and `scale` = original / network.
- `descriptors.data` is a base64 NPZ: `np.load(io.BytesIO(base64.b64decode(data)))["descriptors"]` gives `(N, 256)` float16 rows, L2-normalised, in keypoint order.
### Demo
| Top-500 keypoints on the 480×640 network frame of `code/sample_data/house_in_field_1080p.jpg` (`media/sample.png`, natural image) |
|:---:|
|  |
### Demo & Performances
Warm, batch 1, 480×640 network frame, 1600×900 JPEG for served requests. We measured on a Blackhole chip with a 12×10 compute grid and dispatch on ETH cores (2026-10-03, `code/` at the optimized build).
| Metric | Performance |
|---|---:|
| End-to-end `predict()` (base64 + JPEG decode, device, post-processing; 100 requests) | **17.5–18.0 ms median** |
| `device_forward` in `predict()` (full-size plane upload, device resize, one trace, D2H) | **1.55–1.58 ms median** · 1.49 ms min |
| Request from the 8-bit 480×640 plane (H2D + trace + D2H + host decode) | **0.91 ms median** (0.90–0.92) · 0.876 ms min |
| Trace replay, input resident (network + NMS + keypoint list + descriptor sampling) | **0.50 ms median** · 0.48 ms min |
| Device time per trace replay (device profiler, 23 ops) | **462.8 µs** |
RTX 5090 reference measurements (2026-09-14) are unchanged. They use the port's torch reference in eager PyTorch 2.11, batch 1, with the same host pre/post-processing. The GPU host is a different machine. Full table: [`GPU_COMPARISON.md`](GPU_COMPARISON.md).
| RTX 5090 precision | GPU served-like total | vs ours (18.0 ms) | GPU upload + forward + readback vs our request (0.91 ms) | GPU forward, input resident, vs our trace (0.50 ms) |
|---|---:|---:|---:|---:|
| fp32 strict | 22.75 ms | **Blackhole 1.26× faster** | 2.94 ms: Blackhole 3.23× faster | 2.26 ms: Blackhole 4.52× faster |
| bf16 autocast | 21.45 ms | **Blackhole 1.19× faster** | 1.59 ms: Blackhole 1.74× faster | 0.91 ms: Blackhole 1.83× faster |
| fp16 autocast | 21.90 ms | **Blackhole 1.22× faster** | 1.53 ms: Blackhole 1.68× faster | 0.86 ms: Blackhole 1.71× faster |
| bf16 autocast + `torch.compile` (CUDA graphs) | not measured | – | 1.29 ms: Blackhole 1.42× faster | 0.59 ms: Blackhole 1.18× faster |
| fp16 autocast + `torch.compile` (CUDA graphs) | not measured | – | 1.21 ms: Blackhole 1.33× faster | 0.51 ms: parity (1.01×) |
The Blackhole trace also does the keypoint list and the descriptor sampling. The GPU does these steps on the host in a further 2.2 ms. Our request uploads only the 8-bit image plane and reads back only the keypoints and their descriptors. The GPU uploads an fp32 tensor and reads back the full score and descriptor maps. The host JPEG decode is the largest part of the served total on both sides. The previous release was 3.2× slower than the bf16 GPU on the device forward.
### Caveats
- Does not scale to multiple p150a in a mesh configuration. Current build enforces 12x10 tensix cores for compute, which is enabled by moving dispatch functions that were originally in 10 tensix cores to ETH cores. This implementation therefore assumes that chip-to-chip ethernet communication is not required by the user. The tt-metal change is in [`patches/tt-metal-eth-dispatch.patch`](patches/tt-metal-eth-dispatch.patch).
- Every image is resized to 480×640 (bilinear, /255, channel 0); one image per request, batch 1, requests serialised on the chip.
- Accuracy against the fp32 torch reference (natural image): score PCC 0.998583, descriptor PCC 0.999280, F1 0.9890. Some custom kernels sum in a different order than the ttnn conv. Thus the outputs are not bit-identical to the previous release. See [`OPT_REPORT.md`](OPT_REPORT.md).
- `nms_radius` 4 runs in the main trace. Radii 1–8 use a per-radius device NMS trace, which compiles on first use. Radius 0 or above 8 uses the host NMS.
- Weights are research-only: the Magic Leap SuperPoint licence allows academic or non-profit organisation NONCOMMERCIAL research use.
- Not an OpenAI-compatible API; `GET /v1/models` is a stub so the tt-model ready card does not 404.
### Licensing
- Weights: [magic-leap-community/superpoint](https://huggingface.co/magic-leap-community/superpoint), `other` ([Magic Leap SuperPoint licence](https://huggingface.co/magic-leap-community/superpoint/blob/main/LICENSE), noncommercial research use only); not redistributed here.
- Port and serving code (`code/`): Apache-2.0 (SPDX headers on the modules), from [changh95/tt-superpoint](https://github.com/changh95/tt-superpoint), distributed under the same upstream terms since a port cannot grant more than its upstream does.
- `patches/tt-metal-eth-dispatch.patch` changes [tt-metal](https://github.com/tenstorrent/tt-metal), which is Apache-2.0.
## Provenance
These are the exact sources the container image was built from. **`code/` has since been updated (2026-10-03 optimized build; see [`OPT_REPORT.md`](OPT_REPORT.md)) and is newer than the image.** `tt-model serve` runs the image's code until the image is rebuilt. `tt-model.yaml` and `SERVING.md` still describe the image:
| component | built from |
| --- | --- |
| tt-metal | [`8b98410e730bb504fea43a88609756e34821d91d`](https://github.com/tenstorrent/tt-metal/commit/8b98410e730bb504fea43a88609756e34821d91d) |
| `code/` digest (image) | `ae768681d4aa677d` (sha256, first 16 hex digits; the current `code/` differs) |
| built | 2026-09-13T15:21:44+00:00 by tt-model 0.1.0 |
|