superpoint-p150 / tt-model.yaml
changh95's picture
tt-model authoring file: compact card
d471346 verified
Raw History Blame
7.03 kB
# SPDX-License-Identifier: Apache-2.0
# tt-model-manager container manifest (schema 5.1): SuperPoint on one Blackhole p150.
#
# Run EVERY tt-model command from this directory (extra_code.root and --out are CWD-relative):
# tt-model package --container tt-model.yaml --out $ROOT/build # amd64 + Docker>=25, 2.5-4 h cold
# tt-model serve $ROOT/build/superpoint-p150/tt_kernel_manifest.json
# tt-model push $ROOT/build/superpoint-p150 --publish
schema: "5.1"
repo: changh95/superpoint-p150
name: superpoint-p150
# POINTER to the weights (never baked into the image). Pinned to the snapshot the port
# was validated against; the same sha is handed to the app as TT_WEIGHTS_REVISION below
# because the launcher only exports HF_MODEL. pytorch_model.bin is a 5 MB duplicate.
weights:
repo: magic-leap-community/superpoint
revision: 734450e9ffe229074f5998494ddc615475cdb20a
allow_patterns: ["config.json", "model.safetensors", "preprocessor_config.json"]
kind: tt-dit-server
arch: blackhole
source:
tt_metal: /home/deepgadget/experiments/gbp-tt/tt-metal # v0.78.0-dev20260820-25, torch 2.11.0 pin
# The port imports nothing from tt-metal's models/ tree; one small real file satisfies
# the schema's at-least-one rule (staged to code/models/common/lightweightmodule.py).
code:
- models/common/lightweightmodule.py
# The port itself, from this repo. Everything under code/ that should stay on the Hub
# must be listed: push makes code/ exactly this set (kernels/ is the fused C++ NMS
# kernel's source, not built into this image; sample_data/ is the smoke-test frame).
extra_code:
- root: code
paths:
- models
- kernels
- sample_data
- results.tsv
- run_benchmark.sh
ubuntu: "22.04"
python: "3.12"
runtime:
app: models.server.app:app
mesh_shape_env: TT_MESH_SHAPE
# On top of the kind defaults (fastapi, uvicorn, pydantic>=2, pillow) and the auto-pinned
# torch==2.11.0+cpu. transformers provides SuperPointForKeypointDetection (present in
# 4.53 .. 5.x). No torchvision on the serve path (resize/normalise done with PIL+torch).
packages:
- "numpy>=1.24.4,<2"
- "transformers>=4.53,<6"
- huggingface_hub
- safetensors
serve:
port: 20000
hardware: p150
mesh_device: P150
env:
TT_WEIGHTS_REVISION: "734450e9ffe229074f5998494ddc615475cdb20a"
TT_METAL_VISIBLE_DEVICES: "0"
# Build-time assertions run INSIDE the finished image as uid 1000, no device, no weights.
verify:
- "import models.server.app as a; assert a.app"
- "from models.tt.superpoint_ttnn import TtSuperPoint, device_outputs_to_host; assert TtSuperPoint"
- "from models.tt.postprocess import postprocess_keypoints; assert postprocess_keypoints"
- "import transformers, huggingface_hub, safetensors; from transformers import SuperPointForKeypointDetection; assert SuperPointForKeypointDetection"
- "import numpy; assert numpy.__version__.startswith('1.'), numpy.__version__"
- "from pathlib import Path; assert Path('/opt/tt-metal/sample_data/house_in_field_1080p.jpg').is_file()"
card:
description: >
SuperPoint (Magic Leap's self-supervised interest-point detector and descriptor) running entirely on one Tenstorrent Blackhole p150a via tt-nn: image in, keypoints with scores and 256-d descriptors out.
Weights: [magic-leap-community/superpoint](https://huggingface.co/magic-leap-community/superpoint) ·
Paper: [arXiv:1712.07629](https://arxiv.org/abs/1712.07629) ·
Upstream code: [magicleap/SuperPointPretrainedNetwork](https://github.com/magicleap/SuperPointPretrainedNetwork) ·
Port: [changh95/tt-superpoint](https://github.com/changh95/tt-superpoint)
quickstart: |
### Run with tt-cli
```bash
tt serve changh95/superpoint-p150
printf '{"image":"%s"}' "$(base64 -w0 code/sample_data/house_in_field_1080p.jpg)" > req.json
curl -s localhost:20000/predict -H 'Content-Type: application/json' -d @req.json
tt model stop changh95/superpoint-p150
```
- `POST /predict`: `image` (base64 PNG/JPEG); optional `max_keypoints` (1024, `-1` = all above threshold), `keypoint_threshold` (0.005), `nms_radius` (4), `return_descriptors` (true).
- `GET /health`, `GET /info`.
### Response
```json
{"num_keypoints": 539,
"keypoints": [[610.0, 703.125], [1042.5, 446.25], [1122.5, 442.5]],
"scores": [0.609375, 0.589844, 0.582031],
"original_size": {"height": 900, "width": 1600}, "image_size": {"height": 480, "width": 640}, "scale": {"x": 2.5, "y": 1.875},
"descriptors": {"format": "npz", "key": "descriptors", "dtype": "float16", "shape": [539, 256], "data": "..."},
"timing_ms": {"preprocess": 19.4, "device_forward": 12.5, "postprocess": 32.5, "total": 64.4}}
```
- `keypoints` are `[x, y]` in original image pixels, sorted by descending `scores`; the network frame is 480×640 and `scale` = original / network.
- `descriptors.data` is a base64 NPZ: `np.load(io.BytesIO(base64.b64decode(data)))["descriptors"]` gives `(N, 256)` float16 rows, L2-normalised, in keypoint order.
### Demo
| Top-500 keypoints on the 480×640 network frame of `code/sample_data/house_in_field_1080p.jpg` (`media/sample.png`, natural image) |
|:---:|
| ![](media/sample.png) |
### Accuracy and speed
| Metric | Value |
|---|---:|
| Pre-NMS score map · descriptor map PCC vs fp32 torch reference | 0.9971 · 0.9991 |
| Keypoint set vs reference (natural image, top-500, 2 px) | recall 98.80% · precision 98.80% · F1 98.80% |
| Inference, served over HTTP (warm, batch 1, 480×640, 1600×900 JPEG in) | ~12–20 ms device · ~65 ms end-to-end (~15 FPS) |
### Caveats
- Every image is resized to 480×640 (bilinear, /255, channel 0); one image per request, batch 1, requests serialised on the chip.
- This server runs the untraced pure-ttnn path with host single-pass NMS (bf16 + HiFi2 + fp32 accumulate; bfloat8/LoFi drop score PCC to ~0.91). The port README's 40.7 fps needs `ttnn.trace` plus the fused `sp_eq_mul_mask` C++ kernel in `code/kernels/`, which is not built into this image.
- Weights are research-only: the Magic Leap SuperPoint licence allows academic or non-profit organisation NONCOMMERCIAL research use.
- Not an OpenAI-compatible API; `GET /v1/models` is a stub so the tt-model ready card does not 404.
- Validated on tt-metal `v0.78.0-dev20260820` (main `8b98410e730`), single p150a only.
### Licensing
- Weights: [magic-leap-community/superpoint](https://huggingface.co/magic-leap-community/superpoint), `other` ([Magic Leap SuperPoint licence](https://huggingface.co/magic-leap-community/superpoint/blob/main/LICENSE), noncommercial research use only); not redistributed here.
- Port and serving code (`code/`): Apache-2.0 (SPDX headers on the modules), from [changh95/tt-superpoint](https://github.com/changh95/tt-superpoint), distributed under the same upstream terms since a port cannot grant more than its upstream does.