File size: 7,031 Bytes
dc14621 6575fcf dc14621 6575fcf 9c95d27 dc14621 9c95d27 dc14621 6575fcf dc14621 6575fcf dc14621 6575fcf dc14621 6575fcf dc14621 6575fcf dc14621 6575fcf dc14621 6575fcf dc14621 6575fcf dc14621 6575fcf dc14621 d471346 dc14621 d471346 6575fcf d471346 9c95d27 6575fcf d471346 6575fcf d471346 dc14621 d471346 dc14621 6575fcf d471346 6575fcf d471346 6575fcf d471346 6575fcf d471346 6575fcf d471346 6575fcf d471346 6575fcf d471346 6575fcf d471346 6575fcf d471346 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 | # SPDX-License-Identifier: Apache-2.0
# tt-model-manager container manifest (schema 5.1): SuperPoint on one Blackhole p150.
#
# Run EVERY tt-model command from this directory (extra_code.root and --out are CWD-relative):
# tt-model package --container tt-model.yaml --out $ROOT/build # amd64 + Docker>=25, 2.5-4 h cold
# tt-model serve $ROOT/build/superpoint-p150/tt_kernel_manifest.json
# tt-model push $ROOT/build/superpoint-p150 --publish
schema: "5.1"
repo: changh95/superpoint-p150
name: superpoint-p150
# POINTER to the weights (never baked into the image). Pinned to the snapshot the port
# was validated against; the same sha is handed to the app as TT_WEIGHTS_REVISION below
# because the launcher only exports HF_MODEL. pytorch_model.bin is a 5 MB duplicate.
weights:
repo: magic-leap-community/superpoint
revision: 734450e9ffe229074f5998494ddc615475cdb20a
allow_patterns: ["config.json", "model.safetensors", "preprocessor_config.json"]
kind: tt-dit-server
arch: blackhole
source:
tt_metal: /home/deepgadget/experiments/gbp-tt/tt-metal # v0.78.0-dev20260820-25, torch 2.11.0 pin
# The port imports nothing from tt-metal's models/ tree; one small real file satisfies
# the schema's at-least-one rule (staged to code/models/common/lightweightmodule.py).
code:
- models/common/lightweightmodule.py
# The port itself, from this repo. Everything under code/ that should stay on the Hub
# must be listed: push makes code/ exactly this set (kernels/ is the fused C++ NMS
# kernel's source, not built into this image; sample_data/ is the smoke-test frame).
extra_code:
- root: code
paths:
- models
- kernels
- sample_data
- results.tsv
- run_benchmark.sh
ubuntu: "22.04"
python: "3.12"
runtime:
app: models.server.app:app
mesh_shape_env: TT_MESH_SHAPE
# On top of the kind defaults (fastapi, uvicorn, pydantic>=2, pillow) and the auto-pinned
# torch==2.11.0+cpu. transformers provides SuperPointForKeypointDetection (present in
# 4.53 .. 5.x). No torchvision on the serve path (resize/normalise done with PIL+torch).
packages:
- "numpy>=1.24.4,<2"
- "transformers>=4.53,<6"
- huggingface_hub
- safetensors
serve:
port: 20000
hardware: p150
mesh_device: P150
env:
TT_WEIGHTS_REVISION: "734450e9ffe229074f5998494ddc615475cdb20a"
TT_METAL_VISIBLE_DEVICES: "0"
# Build-time assertions run INSIDE the finished image as uid 1000, no device, no weights.
verify:
- "import models.server.app as a; assert a.app"
- "from models.tt.superpoint_ttnn import TtSuperPoint, device_outputs_to_host; assert TtSuperPoint"
- "from models.tt.postprocess import postprocess_keypoints; assert postprocess_keypoints"
- "import transformers, huggingface_hub, safetensors; from transformers import SuperPointForKeypointDetection; assert SuperPointForKeypointDetection"
- "import numpy; assert numpy.__version__.startswith('1.'), numpy.__version__"
- "from pathlib import Path; assert Path('/opt/tt-metal/sample_data/house_in_field_1080p.jpg').is_file()"
card:
description: >
SuperPoint (Magic Leap's self-supervised interest-point detector and descriptor) running entirely on one Tenstorrent Blackhole p150a via tt-nn: image in, keypoints with scores and 256-d descriptors out.
Weights: [magic-leap-community/superpoint](https://huggingface.co/magic-leap-community/superpoint) ·
Paper: [arXiv:1712.07629](https://arxiv.org/abs/1712.07629) ·
Upstream code: [magicleap/SuperPointPretrainedNetwork](https://github.com/magicleap/SuperPointPretrainedNetwork) ·
Port: [changh95/tt-superpoint](https://github.com/changh95/tt-superpoint)
quickstart: |
### Run with tt-cli
```bash
tt serve changh95/superpoint-p150
printf '{"image":"%s"}' "$(base64 -w0 code/sample_data/house_in_field_1080p.jpg)" > req.json
curl -s localhost:20000/predict -H 'Content-Type: application/json' -d @req.json
tt model stop changh95/superpoint-p150
```
- `POST /predict`: `image` (base64 PNG/JPEG); optional `max_keypoints` (1024, `-1` = all above threshold), `keypoint_threshold` (0.005), `nms_radius` (4), `return_descriptors` (true).
- `GET /health`, `GET /info`.
### Response
```json
{"num_keypoints": 539,
"keypoints": [[610.0, 703.125], [1042.5, 446.25], [1122.5, 442.5]],
"scores": [0.609375, 0.589844, 0.582031],
"original_size": {"height": 900, "width": 1600}, "image_size": {"height": 480, "width": 640}, "scale": {"x": 2.5, "y": 1.875},
"descriptors": {"format": "npz", "key": "descriptors", "dtype": "float16", "shape": [539, 256], "data": "..."},
"timing_ms": {"preprocess": 19.4, "device_forward": 12.5, "postprocess": 32.5, "total": 64.4}}
```
- `keypoints` are `[x, y]` in original image pixels, sorted by descending `scores`; the network frame is 480×640 and `scale` = original / network.
- `descriptors.data` is a base64 NPZ: `np.load(io.BytesIO(base64.b64decode(data)))["descriptors"]` gives `(N, 256)` float16 rows, L2-normalised, in keypoint order.
### Demo
| Top-500 keypoints on the 480×640 network frame of `code/sample_data/house_in_field_1080p.jpg` (`media/sample.png`, natural image) |
|:---:|
|  |
### Accuracy and speed
| Metric | Value |
|---|---:|
| Pre-NMS score map · descriptor map PCC vs fp32 torch reference | 0.9971 · 0.9991 |
| Keypoint set vs reference (natural image, top-500, 2 px) | recall 98.80% · precision 98.80% · F1 98.80% |
| Inference, served over HTTP (warm, batch 1, 480×640, 1600×900 JPEG in) | ~12–20 ms device · ~65 ms end-to-end (~15 FPS) |
### Caveats
- Every image is resized to 480×640 (bilinear, /255, channel 0); one image per request, batch 1, requests serialised on the chip.
- This server runs the untraced pure-ttnn path with host single-pass NMS (bf16 + HiFi2 + fp32 accumulate; bfloat8/LoFi drop score PCC to ~0.91). The port README's 40.7 fps needs `ttnn.trace` plus the fused `sp_eq_mul_mask` C++ kernel in `code/kernels/`, which is not built into this image.
- Weights are research-only: the Magic Leap SuperPoint licence allows academic or non-profit organisation NONCOMMERCIAL research use.
- Not an OpenAI-compatible API; `GET /v1/models` is a stub so the tt-model ready card does not 404.
- Validated on tt-metal `v0.78.0-dev20260820` (main `8b98410e730`), single p150a only.
### Licensing
- Weights: [magic-leap-community/superpoint](https://huggingface.co/magic-leap-community/superpoint), `other` ([Magic Leap SuperPoint licence](https://huggingface.co/magic-leap-community/superpoint/blob/main/LICENSE), noncommercial research use only); not redistributed here.
- Port and serving code (`code/`): Apache-2.0 (SPDX headers on the modules), from [changh95/tt-superpoint](https://github.com/changh95/tt-superpoint), distributed under the same upstream terms since a port cannot grant more than its upstream does.
|