changh95 commited on
Commit
736060c
·
verified ·
1 Parent(s): ff4d41f

Add files using upload-large-folder tool

Browse files
Files changed (2) hide show
  1. README.md +36 -109
  2. tt_kernel_manifest.json +2 -2
README.md CHANGED
@@ -1,28 +1,16 @@
1
  ---
2
  tags:
3
  - blackhole
4
- - keypoint-detection
5
  - p150
6
- - superpoint
7
- - tenstorrent
8
  - tt-dit-server
9
- - tt-metal
10
  - tt-model-cache
11
- - tt-model-catalog
12
  - tt-model-container
13
- - tt-nn
14
- - ttnn
15
- base_model:
16
- - magic-leap-community/superpoint
17
- license: other
18
- license_name: magic-leap-superpoint
19
- license_link: https://huggingface.co/magic-leap-community/superpoint
20
- pipeline_tag: keypoint-detection
21
  ---
22
 
23
  # superpoint-p150
24
 
25
- SuperPoint (magic-leap-community/superpoint) keypoint detection and description on a single Tenstorrent Blackhole p150a via tt-nn: a base64 image in, keypoints in original pixel coordinates, scores and 256-d descriptors out, at a fixed 480x640 network input. Pre-NMS score-map PCC 0.9971 and descriptor PCC 0.9991 vs the fp32 torch reference, keypoint F1 98.8% (top-500, 2 px); this server runs the pure-ttnn untraced path with host NMS (roughly 6 fps; the README's 40.7 fps needs trace plus the fused C++ NMS kernel in kernels/, which is not built into this image). Weights are under the Magic Leap SuperPoint licence: academic or non-profit organisation NONCOMMERCIAL research use only. Port source: github.com/changh95/tt-superpoint @ e1eab66e29ff424bc9af6b1118671d9bc08e899e.
 
26
 
27
  Runs on **p150** (mesh `P150`).
28
 
@@ -37,119 +25,58 @@ tt-model serve changh95/superpoint-p150
37
 
38
  `pull --with-weights` downloads the Docker image and the [`magic-leap-community/superpoint`](https://huggingface.co/magic-leap-community/superpoint) weights at `734450e9ffe229074f5998494ddc615475cdb20a` (into your HF cache; they are not in the image). `serve` starts the model's own HTTP server on port 20000 (or the next free port, if that one is busy); the first start compiles kernels for your device, which takes several minutes, and the server is ready when it logs `Application startup complete`.
39
 
40
- ### With tt-cli
41
 
42
  ```bash
43
- tt serve changh95/superpoint-p150 # pulls image + weights, boots, prints the port
 
 
44
  tt model stop changh95/superpoint-p150
45
  ```
46
 
47
- The port is the one `serve` printed (20000, or the next free one). This is **not** an
48
- OpenAI-style server: `tt-model curl` / chat clients do not apply. `GET /v1/models` exists
49
- only so generic probes do not 404; the real routes are below.
50
 
51
- ### Call it
52
 
53
- ```bash
54
- PORT=20000 # the port serve printed
55
- curl -s localhost:$PORT/health # {"status":"ok","model":"superpoint-p150","device":{...}}
56
- curl -s localhost:$PORT/info # weights repo+revision, canonical input, defaults, limits, licence
57
-
58
- # One image -> keypoints. Any PNG/JPEG; it is resized to 640x480 server-side and the
59
- # keypoints come back in YOUR image's pixel coordinates (see "scale").
60
- IMG=code/sample_data/house_in_field_1080p.jpg # 1600x900 natural image
61
- python3 - "$IMG" "$PORT" <<'EOF'
62
- import base64, json, sys, urllib.request
63
- img, port = sys.argv[1], sys.argv[2]
64
- req = {"image": base64.b64encode(open(img, "rb").read()).decode(),
65
- "max_keypoints": 1024, # -1 = all above threshold
66
- "keypoint_threshold": 0.005,
67
- "nms_radius": 4,
68
- "return_descriptors": True}
69
- r = urllib.request.Request(f"http://127.0.0.1:{port}/predict", json.dumps(req).encode(),
70
- {"Content-Type": "application/json"})
71
- out = json.load(urllib.request.urlopen(r, timeout=300))
72
- print(out["num_keypoints"], out["keypoints"][:3], out["scores"][:3], out["timing_ms"])
73
- EOF
74
  ```
75
 
76
- Response fields: `num_keypoints`; `keypoints` (N x [x, y] floats, original-image pixels);
77
- `scores` (N, post-NMS softmax scores); `original_size`, `image_size` (480x640) and
78
- `scale` (x = W/640, y = H/480; divide to get network-frame coordinates);
79
- `params` echoed; `timing_ms` (`preprocess`, `device_forward`, `postprocess`, `total`);
80
- when `return_descriptors` is true, `descriptors` = `{format: "npz", key: "descriptors",
81
- dtype: "float16", shape: [N, 256], data: <base64 NPZ>}` -- decode with
82
- `numpy.load(io.BytesIO(base64.b64decode(d["data"])))["descriptors"]`; rows are
83
- L2-normalised. Errors: 400 undecodable image / bad field, 503 while starting, 500 with
84
- the exception text. One image per request; requests are serialised on the chip.
85
-
86
- A ready-made check: `python code/models/server/smoke_test.py --url http://127.0.0.1:$PORT`
87
- prints one PASS/FAIL line with the keypoint count and timings.
88
-
89
- ### First boot
90
-
91
- Weights are 5 MB (`config.json`, `model.safetensors`, `preprocessor_config.json` at the
92
- pinned revision) and land in your HF cache. The first start JIT-compiles the conv /
93
- pool / softmax kernels (a few minutes, cached under
94
- `~/.cache/tt-model/superpoint-p150/cache`); the server logs `Loading weights`,
95
- `Warming up`, `Warmup complete` and is ready at uvicorn's `Application startup complete`.
96
- Later boots reuse the kernel cache. The weights repo is public and ungated (no token).
97
-
98
- ### What this server runs
99
 
100
- Fixed 480x640 input (bilinear resize, /255, channel 0 -- the HF
101
- `SuperPointImageProcessor` defaults), 8 encoder convs + 3 max-pools + score and
102
- descriptor heads in bfloat16 activations / bfloat16 weights / HiFi2 / fp32 accumulate,
103
- softmax and descriptor L2-norm on device, then host single-pass NMS, threshold, border
104
- removal, top-k and bilinear descriptor sampling. Untraced, host NMS: this is the
105
- port's `SP_TRACE_NMS=0`, `SP_NO_TRACE=1` configuration. The benchmark numbers below
106
- that need trace or the fused `sp_eq_mul_mask` kernel are **not** what this server does.
107
 
108
- ### Results from the port (Blackhole p150b, 480x640, batch 1, natural image)
 
 
109
 
110
- PCC vs the Hugging Face fp32 reference:
111
 
112
- | Tensor | PCC |
113
  |---|---:|
114
- | Pre-NMS score map | **0.9971** |
115
- | Descriptor map (post L2-norm) | **0.9991** |
116
-
117
- Keypoint set vs the reference (`sample_data/house_in_field_1080p.jpg`, top-500, 2 px):
118
- recall **98.80%**, precision **98.80%**, F1 **98.80%** (synthetic `torch.rand` input: F1 97.80%).
119
-
120
- Throughput measured by `models/tests/test_superpoint.py` (needs a tt-metal checkout as
121
- pytest rootdir for its `device` fixture):
122
-
123
- | Path | Device forward (input resident) | Traced fwd incl. H2D | Full e2e |
124
- |---|---:|---:|---:|
125
- | Forward-only trace, host NMS (`SP_TRACE_NMS=0`) | 355 fps (2.81 ms) | 73.6 fps | 17.0 fps |
126
- | Forward + device NMS, fused kernel (`SP_TRACE_NMS=1`) | 85.6 fps | 44.6 fps | **40.7 fps** |
127
- | Untraced, host NMS (**this server**) | ~6 fps | -- | ~5 fps |
128
-
129
- The fused path needs `kernels/sp_eq_mul_mask/` compiled into tt-metal (see its README);
130
- a pure-ttnn equivalent (`ttnn.eq` + `ttnn.multiply`) is measured in
131
- `kernels/sp_eq_mul_mask/bench.py`. `code/results.tsv` is the full experiment log
132
- (precision cliff: bfloat8/LoFi drop score PCC to 0.70-0.91; trace was a 10.8x unlock).
133
-
134
- Sample output (top-500 keypoints on the resized frame): ![keypoints](media/sample.png)
135
 
136
- ### Layout of `code/`
137
 
138
- `models/tt/superpoint_ttnn.py` (tt-nn model), `models/tt/postprocess.py` (validated host
139
- post-processing), `models/server/app.py` + `smoke_test.py`, `models/reference/`
140
- (HF reference loader, needs torchvision), `models/tests/test_superpoint.py` (benchmark +
141
- PCC + F1), `models/visualize.py`, `kernels/sp_eq_mul_mask/` (fused C++ Tensix NMS
142
- kernel, ~450 LoC), `sample_data/`, `results.tsv`, `run_benchmark.sh`.
143
- `models/common/lightweightmodule.py` is a schema-required filler from tt-metal.
144
 
145
  ### Licensing
146
 
147
- The upstream **weights** (`magic-leap-community/superpoint`) carry the Magic Leap
148
- SuperPoint licence: *academic or non-profit organisation noncommercial research use
149
- only* -- see https://huggingface.co/magic-leap-community/superpoint. The port code
150
- (Apache-2.0 headers, by Hyunggi Chang) is published under the same terms, since a port
151
- cannot grant more than its upstream does. Weights are not redistributed here; they are
152
- fetched from the upstream repo at the pinned revision.
153
 
154
  ## Provenance
155
 
@@ -159,5 +86,5 @@ The exact sources the image was built from — `code/` in this repo is byte-iden
159
  | --- | --- |
160
  | tt-metal | [`8b98410e730bb504fea43a88609756e34821d91d`](https://github.com/tenstorrent/tt-metal/commit/8b98410e730bb504fea43a88609756e34821d91d) |
161
  | `code/` digest | `006ddc1eb4e69a41` (sha256, first 16 hex digits) |
162
- | built | 2026-09-12T10:56:19+00:00 by tt-model 0.1.0 |
163
 
 
1
  ---
2
  tags:
3
  - blackhole
 
4
  - p150
 
 
5
  - tt-dit-server
 
6
  - tt-model-cache
 
7
  - tt-model-container
 
 
 
 
 
 
 
 
8
  ---
9
 
10
  # superpoint-p150
11
 
12
+ SuperPoint (Magic Leap's self-supervised interest-point detector and descriptor) running entirely on one Tenstorrent Blackhole p150a via tt-nn: image in, keypoints with scores and 256-d descriptors out.
13
+ Weights: [magic-leap-community/superpoint](https://huggingface.co/magic-leap-community/superpoint) · Paper: [arXiv:1712.07629](https://arxiv.org/abs/1712.07629) · Upstream code: [magicleap/SuperPointPretrainedNetwork](https://github.com/magicleap/SuperPointPretrainedNetwork) · Port: [changh95/tt-superpoint](https://github.com/changh95/tt-superpoint)
14
 
15
  Runs on **p150** (mesh `P150`).
16
 
 
25
 
26
  `pull --with-weights` downloads the Docker image and the [`magic-leap-community/superpoint`](https://huggingface.co/magic-leap-community/superpoint) weights at `734450e9ffe229074f5998494ddc615475cdb20a` (into your HF cache; they are not in the image). `serve` starts the model's own HTTP server on port 20000 (or the next free port, if that one is busy); the first start compiles kernels for your device, which takes several minutes, and the server is ready when it logs `Application startup complete`.
27
 
28
+ ### Run with tt-cli
29
 
30
  ```bash
31
+ tt serve changh95/superpoint-p150
32
+ printf '{"image":"%s"}' "$(base64 -w0 code/sample_data/house_in_field_1080p.jpg)" > req.json
33
+ curl -s localhost:20000/predict -H 'Content-Type: application/json' -d @req.json
34
  tt model stop changh95/superpoint-p150
35
  ```
36
 
37
+ - `POST /predict`: `image` (base64 PNG/JPEG); optional `max_keypoints` (1024, `-1` = all above threshold), `keypoint_threshold` (0.005), `nms_radius` (4), `return_descriptors` (true).
38
+ - `GET /health`, `GET /info`.
 
39
 
40
+ ### Response
41
 
42
+ ```json
43
+ {"num_keypoints": 539,
44
+ "keypoints": [[610.0, 703.125], [1042.5, 446.25], [1122.5, 442.5]],
45
+ "scores": [0.609375, 0.589844, 0.582031],
46
+ "original_size": {"height": 900, "width": 1600}, "image_size": {"height": 480, "width": 640}, "scale": {"x": 2.5, "y": 1.875},
47
+ "descriptors": {"format": "npz", "key": "descriptors", "dtype": "float16", "shape": [539, 256], "data": "..."},
48
+ "timing_ms": {"preprocess": 19.4, "device_forward": 12.5, "postprocess": 32.5, "total": 64.4}}
 
 
 
 
 
 
 
 
 
 
 
 
 
 
49
  ```
50
 
51
+ - `keypoints` are `[x, y]` in original image pixels, sorted by descending `scores`; the network frame is 480×640 and `scale` = original / network.
52
+ - `descriptors.data` is a base64 NPZ: `np.load(io.BytesIO(base64.b64decode(data)))["descriptors"]` gives `(N, 256)` float16 rows, L2-normalised, in keypoint order.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
53
 
54
+ ### Demo
 
 
 
 
 
 
55
 
56
+ | Top-500 keypoints on the 480×640 network frame of `code/sample_data/house_in_field_1080p.jpg` (`media/sample.png`, natural image) |
57
+ |:---:|
58
+ | ![](media/sample.png) |
59
 
60
+ ### Accuracy and speed
61
 
62
+ | Metric | Value |
63
  |---|---:|
64
+ | Pre-NMS score map · descriptor map PCC vs fp32 torch reference | 0.9971 · 0.9991 |
65
+ | Keypoint set vs reference (natural image, top-500, 2 px) | recall 98.80% · precision 98.80% · F1 98.80% |
66
+ | Inference, served over HTTP (warm, batch 1, 480×640, 1600×900 JPEG in) | ~12–20 ms device · ~65 ms end-to-end (~15 FPS) |
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
67
 
68
+ ### Caveats
69
 
70
+ - Every image is resized to 480×640 (bilinear, /255, channel 0); one image per request, batch 1, requests serialised on the chip.
71
+ - This server runs the untraced pure-ttnn path with host single-pass NMS (bf16 + HiFi2 + fp32 accumulate; bfloat8/LoFi drop score PCC to ~0.91). The port README's 40.7 fps needs `ttnn.trace` plus the fused `sp_eq_mul_mask` C++ kernel in `code/kernels/`, which is not built into this image.
72
+ - Weights are research-only: the Magic Leap SuperPoint licence allows academic or non-profit organisation NONCOMMERCIAL research use.
73
+ - Not an OpenAI-compatible API; `GET /v1/models` is a stub so the tt-model ready card does not 404.
74
+ - Validated on tt-metal `v0.78.0-dev20260820` (main `8b98410e730`), single p150a only.
 
75
 
76
  ### Licensing
77
 
78
+ - Weights: [magic-leap-community/superpoint](https://huggingface.co/magic-leap-community/superpoint), `other` ([Magic Leap SuperPoint licence](https://huggingface.co/magic-leap-community/superpoint/blob/main/LICENSE), noncommercial research use only); not redistributed here.
79
+ - Port and serving code (`code/`): Apache-2.0 (SPDX headers on the modules), from [changh95/tt-superpoint](https://github.com/changh95/tt-superpoint), distributed under the same upstream terms since a port cannot grant more than its upstream does.
 
 
 
 
80
 
81
  ## Provenance
82
 
 
86
  | --- | --- |
87
  | tt-metal | [`8b98410e730bb504fea43a88609756e34821d91d`](https://github.com/tenstorrent/tt-metal/commit/8b98410e730bb504fea43a88609756e34821d91d) |
88
  | `code/` digest | `006ddc1eb4e69a41` (sha256, first 16 hex digits) |
89
+ | built | 2026-09-12T13:22:54+00:00 by tt-model 0.1.0 |
90
 
tt_kernel_manifest.json CHANGED
@@ -6,7 +6,7 @@
6
  "device_count": 1,
7
  "producer": {
8
  "tt_kernel_version": "0.1.0",
9
- "created_at": "2026-09-12T10:59:21.045020+00:00",
10
  "hostname": "deepgadget"
11
  },
12
  "weights": {
@@ -93,7 +93,7 @@
93
  "image": "tt-model/superpoint-p150:0544890bca09",
94
  "repo": "changh95/superpoint-p150",
95
  "tt_model_version": "0.1.0",
96
- "created_at": "2026-09-12T10:56:19+00:00",
97
  "tt_metal": {
98
  "sha": "8b98410e730bb504fea43a88609756e34821d91d",
99
  "describe": "v0.78.0-dev20260820-25-g8b98410e73",
 
6
  "device_count": 1,
7
  "producer": {
8
  "tt_kernel_version": "0.1.0",
9
+ "created_at": "2026-09-12T13:23:20.741264+00:00",
10
  "hostname": "deepgadget"
11
  },
12
  "weights": {
 
93
  "image": "tt-model/superpoint-p150:0544890bca09",
94
  "repo": "changh95/superpoint-p150",
95
  "tt_model_version": "0.1.0",
96
+ "created_at": "2026-09-12T13:22:54+00:00",
97
  "tt_metal": {
98
  "sha": "8b98410e730bb504fea43a88609756e34821d91d",
99
  "describe": "v0.78.0-dev20260820-25-g8b98410e73",