File size: 7,949 Bytes
f64e572
 
 
 
3696119
 
 
4c179e2
 
 
 
 
 
 
 
 
 
 
 
 
f64e572
 
42b7595
f64e572
c699c4c
736060c
f64e572
3696119
f64e572
3696119
f64e572
3696119
f64e572
3696119
42b7595
 
3696119
f64e572
c699c4c
eee3601
f64e572
736060c
f64e572
3696119
736060c
 
 
42b7595
3696119
f64e572
736060c
 
f64e572
736060c
f64e572
736060c
 
 
 
 
 
d69d96a
 
f64e572
 
736060c
 
f64e572
736060c
f64e572
736060c
 
 
f64e572
c699c4c
f64e572
c699c4c
 
 
f64e572
c699c4c
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
f64e572
736060c
f64e572
c699c4c
736060c
c699c4c
 
736060c
 
f64e572
3696119
 
736060c
 
c699c4c
f64e572
3696119
f64e572
c699c4c
f64e572
3696119
 
 
c699c4c
d69d96a
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
---
tags:
- blackhole
- p150
- tt-dit-server
- tt-model-cache
- tt-model-container
- tenstorrent
- ttnn
- tt-metal
- tt-nn
- keypoint-detection
- superpoint
- tt-model-catalog
pipeline_tag: keypoint-detection
license_link: https://huggingface.co/magic-leap-community/superpoint
license: other
license_name: magic-leap-superpoint
base_model:
- magic-leap-community/superpoint
---

# superpoint-p150

Magic Leap SuperPoint port on one Tenstorrent Blackhole p150a.
Weights: [magic-leap-community/superpoint](https://huggingface.co/magic-leap-community/superpoint) · Paper: [arXiv:1712.07629](https://arxiv.org/abs/1712.07629) · Upstream code: [magicleap/SuperPointPretrainedNetwork](https://github.com/magicleap/SuperPointPretrainedNetwork) · Port: [changh95/tt-superpoint](https://github.com/changh95/tt-superpoint)

Runs on **p150** (mesh `P150`).

Packaged and published with [tt-model-manager](https://github.com/tenstorrent/tt-model-manager) 0.1.0 (manifest schema 5.1).

## Quickstart

```bash
tt-model pull  changh95/superpoint-p150 --with-weights
tt-model serve changh95/superpoint-p150
```

- Weights [`magic-leap-community/superpoint`](https://huggingface.co/magic-leap-community/superpoint) go to your HF cache; the image does not contain them.
- Serves on port 20000 (or the next free port); ready when the log says `Application startup complete`.

### Run with tt-cli

```bash
tt serve changh95/superpoint-p150
printf '{"image":"%s"}' "$(base64 -w0 code/sample_data/house_in_field_1080p.jpg)" > req.json
curl -s localhost:20000/predict -H 'Content-Type: application/json' -d @req.json
tt model stop changh95/superpoint-p150
```

- `POST /predict`: `image` (base64 PNG/JPEG); optional `max_keypoints` (1024, `-1` = all above threshold), `keypoint_threshold` (0.005), `nms_radius` (4), `return_descriptors` (true).
- `GET /health`, `GET /info`.

### Response

```json
{"num_keypoints": 539,
 "keypoints": [[610.0, 703.125], [1042.5, 446.25], [1122.5, 442.5]],
 "scores": [0.609375, 0.589844, 0.582031],
 "original_size": {"height": 900, "width": 1600}, "image_size": {"height": 480, "width": 640}, "scale": {"x": 2.5, "y": 1.875},
 "descriptors": {"format": "npz", "key": "descriptors", "dtype": "float16", "shape": [539, 256], "data": "..."},
 "serving_path": {"traced": true, "device_nms": true},
 "timing_ms": {"preprocess": 17.7, "device_forward": 5.3, "postprocess": 1.3, "total": 24.3}}
```

- `keypoints` are `[x, y]` in original image pixels, sorted by descending `scores`; the network frame is 480×640 and `scale` = original / network.
- `descriptors.data` is a base64 NPZ: `np.load(io.BytesIO(base64.b64decode(data)))["descriptors"]` gives `(N, 256)` float16 rows, L2-normalised, in keypoint order.

### Demo

| Top-500 keypoints on the 480×640 network frame of `code/sample_data/house_in_field_1080p.jpg` (`media/sample.png`, natural image) |
|:---:|
| ![](media/sample.png) |

### Demo & Performances

Warm, batch 1, 480×640 network frame, 1600×900 JPEG for served requests. We measured on a Blackhole chip with a 12×10 compute grid and dispatch on ETH cores (2026-10-03, `code/` at the optimized build).

| Metric | Performance |
|---|---:|
| End-to-end `predict()` (base64 + JPEG decode, device, post-processing; 100 requests) | **17.5–18.0 ms median** |
| `device_forward` in `predict()` (full-size plane upload, device resize, one trace, D2H) | **1.55–1.58 ms median** · 1.49 ms min |
| Request from the 8-bit 480×640 plane (H2D + trace + D2H + host decode) | **0.91 ms median** (0.90–0.92) · 0.876 ms min |
| Trace replay, input resident (network + NMS + keypoint list + descriptor sampling) | **0.50 ms median** · 0.48 ms min |
| Device time per trace replay (device profiler, 23 ops) | **462.8 µs** |

RTX 5090 reference measurements (2026-09-14) are unchanged. They use the port's torch reference in eager PyTorch 2.11, batch 1, with the same host pre/post-processing. The GPU host is a different machine. Full table: [`GPU_COMPARISON.md`](GPU_COMPARISON.md).

| RTX 5090 precision | GPU served-like total | vs ours (18.0 ms) | GPU upload + forward + readback vs our request (0.91 ms) | GPU forward, input resident, vs our trace (0.50 ms) |
|---|---:|---:|---:|---:|
| fp32 strict | 22.75 ms | **Blackhole 1.26× faster** | 2.94 ms: Blackhole 3.23× faster | 2.26 ms: Blackhole 4.52× faster |
| bf16 autocast | 21.45 ms | **Blackhole 1.19× faster** | 1.59 ms: Blackhole 1.74× faster | 0.91 ms: Blackhole 1.83× faster |
| fp16 autocast | 21.90 ms | **Blackhole 1.22× faster** | 1.53 ms: Blackhole 1.68× faster | 0.86 ms: Blackhole 1.71× faster |
| bf16 autocast + `torch.compile` (CUDA graphs) | not measured | – | 1.29 ms: Blackhole 1.42× faster | 0.59 ms: Blackhole 1.18× faster |
| fp16 autocast + `torch.compile` (CUDA graphs) | not measured | – | 1.21 ms: Blackhole 1.33× faster | 0.51 ms: parity (1.01×) |

The Blackhole trace also does the keypoint list and the descriptor sampling. The GPU does these steps on the host in a further 2.2 ms. Our request uploads only the 8-bit image plane and reads back only the keypoints and their descriptors. The GPU uploads an fp32 tensor and reads back the full score and descriptor maps. The host JPEG decode is the largest part of the served total on both sides. The previous release was 3.2× slower than the bf16 GPU on the device forward.

### Caveats

- Does not scale to multiple p150a in a mesh configuration. Current build enforces 12x10 tensix cores for compute, which is enabled by moving dispatch functions that were originally in 10 tensix cores to ETH cores. This implementation therefore assumes that chip-to-chip ethernet communication is not required by the user. The tt-metal change is in [`patches/tt-metal-eth-dispatch.patch`](patches/tt-metal-eth-dispatch.patch).
- Every image is resized to 480×640 (bilinear, /255, channel 0); one image per request, batch 1, requests serialised on the chip.
- Accuracy against the fp32 torch reference (natural image): score PCC 0.998583, descriptor PCC 0.999280, F1 0.9890. Some custom kernels sum in a different order than the ttnn conv. Thus the outputs are not bit-identical to the previous release. See [`OPT_REPORT.md`](OPT_REPORT.md).
- `nms_radius` 4 runs in the main trace. Radii 1–8 use a per-radius device NMS trace, which compiles on first use. Radius 0 or above 8 uses the host NMS.
- Weights are research-only: the Magic Leap SuperPoint licence allows academic or non-profit organisation NONCOMMERCIAL research use.
- Not an OpenAI-compatible API; `GET /v1/models` is a stub so the tt-model ready card does not 404.

### Licensing

- Weights: [magic-leap-community/superpoint](https://huggingface.co/magic-leap-community/superpoint), `other` ([Magic Leap SuperPoint licence](https://huggingface.co/magic-leap-community/superpoint/blob/main/LICENSE), noncommercial research use only); not redistributed here.
- Port and serving code (`code/`): Apache-2.0 (SPDX headers on the modules), from [changh95/tt-superpoint](https://github.com/changh95/tt-superpoint), distributed under the same upstream terms since a port cannot grant more than its upstream does.
- `patches/tt-metal-eth-dispatch.patch` changes [tt-metal](https://github.com/tenstorrent/tt-metal), which is Apache-2.0.

## Provenance

These are the exact sources the container image was built from. **`code/` has since been updated (2026-10-03 optimized build; see [`OPT_REPORT.md`](OPT_REPORT.md)) and is newer than the image.** `tt-model serve` runs the image's code until the image is rebuilt. `tt-model.yaml` and `SERVING.md` still describe the image:

| component | built from |
| --- | --- |
| tt-metal | [`8b98410e730bb504fea43a88609756e34821d91d`](https://github.com/tenstorrent/tt-metal/commit/8b98410e730bb504fea43a88609756e34821d91d) |
| `code/` digest (image) | `ae768681d4aa677d` (sha256, first 16 hex digits; the current `code/` differs) |
| built | 2026-09-13T15:21:44+00:00 by tt-model 0.1.0 |