changh95 commited on
Commit
8eab111
·
verified ·
1 Parent(s): 6575fcf

tt-model authoring files

Browse files
Files changed (1) hide show
  1. SERVING.md +187 -93
SERVING.md CHANGED
@@ -1,112 +1,206 @@
1
- # Serving this port with tt-model-manager
2
-
3
- This repo carries everything needed to build a **tt-model container bundle** from it, so
4
- `tt-model pull` / `tt-model serve` and `tt-cli` can run it. Do this on a
5
- **Blackhole host** (amd64 Linux + Docker >= 25).
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
6
 
7
- ## Why not on a Mac
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
8
 
9
- `tt-model package --container` builds from `ghcr.io/.../ubuntu-22.04-dev-amd64` and a cold
10
- build is 2.5-4 hours. Packaging and pushing need **no card** (`require_host(need_devices=False)`),
11
- but they do need amd64. `tt-model serve` needs `/dev/tenstorrent/*`.
12
 
13
- ## Steps
14
 
15
  ```bash
16
- # 0. get tt-model (not on PyPI)
17
- git clone https://github.com/tenstorrent/tt-model-manager && cd tt-model-manager
18
- python -m venv .venv && source .venv/bin/activate && pip install -e .
19
-
20
- # 1. pull THIS repo
21
- hf download changh95/superpoint-blackhole --local-dir superpoint-blackhole && cd superpoint-blackhole
22
-
23
- # 2. point the manifest at your tt-metal checkout
24
- $EDITOR tt-model.yaml # source.tt_metal: /path/to/tt-metal
25
 
26
- # 3. finish the serving adapter
27
- $EDITOR code/models/server/app.py # see "VERIFY ON HARDWARE" markers
28
 
29
- # 4. validate with no hardware and no build
30
- python -c "from tt_kernel.container_manifest import ContainerManifest as M; M.load('tt-model.yaml')"
 
31
 
32
- # 5. build, prove, publish
33
- tt-model package --container tt-model.yaml
34
- tt-model serve build/superpoint-blackhole/tt_kernel_manifest.json
35
- tt-model push build/superpoint-blackhole --publish
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
36
  ```
37
 
38
- Step 5's `--publish` implies `--public` and adds the `tt-model-catalog` tag, which is what
39
- lists it in the community catalog.
 
40
 
41
- > **Run every `tt-model` command from this repo's root.** `source.extra_code[].root: code`
42
- > is resolved against the **current working directory**, not against the YAML's location --
43
- > verified: loading the same manifest from a parent directory fails with
44
- > `source.extra_code lists 1 path(s) that do not exist under code`. `cd` into the repo first.
45
 
46
- ## What the bundle push does to this repo
47
 
48
- - **`code/` and `image/` are replaced wholesale** (`_prune_removed` in `tt_kernel/hub.py`).
49
- The port code staged here is exactly what `extra_code` re-stages, so it is replaced, not
50
- duplicated. Files under `code/` that the allowlist does not ship get pruned.
51
- - **`README.md` is overwritten** by the generated card. The text you want to keep belongs in
52
- `card.description` / `card.quickstart` in `tt-model.yaml` -- it is already seeded there.
53
- - **Root files survive**, which is why `media/` and this file live at the root.
54
 
55
- ## The serving contract
 
 
 
56
 
57
- `kind: tt-dit-server` requires only `hardware` + `mesh_device` (no `max_num_seqs` /
58
- `block_size` -- those are vLLM engine settings). `runtime.app` is an ASGI target
59
- `"module:attr"` whose **top-level package must appear in an allowlist entry**, which is why
60
- `runtime.app` and `source.extra_code[0].paths` have to stay in sync. The kind installs
61
- fastapi / uvicorn / pydantic / pillow; `runtime.packages` adds to that set.
62
 
63
- Despite the name, this kind is not diffusion-only -- `jashansinghTT/s2-pro-blackhole` is a
64
- TTS model using it. It simply means "the model's own ASGI app under uvicorn".
65
 
66
- ## The catalog tags are automatic -- do not add them by hand
 
 
 
67
 
68
- `tt-model-cache`, `tt-model-container` and `tt-model-catalog` are written by the real flow,
69
- not by editing this card:
 
70
 
71
- | tag | written by | source |
72
- |---|---|---|
73
- | `tt-model-cache` | `package` | `TT_MODEL_TAG` in `tt_kernel/__init__.py:9` |
74
- | `tt-model-container` | `package` | `build.py:996` -- `{TT_MODEL_TAG, m.arch, m.kind, "tt-model-container"}` |
75
- | `blackhole`, `tt-dit-server`, `p150` | `package` | same line, from `arch` / `kind` / each profile's `hardware` |
76
- | `tt-model-catalog` | `--publish` | `cli.py:620` / `cli.py:820` |
77
-
78
- So finishing step 5 produces exactly the tag set the catalog models carry. Adding them to
79
- this repo now would be actively harmful, because **nothing validates them**:
80
-
81
- - `tt-model-cache` is what `tt-model search` filters on (`hub.py:268`), so the repo would
82
- show up as an installable bundle and `tt-model pull` would then fail -- there is no
83
- `tt_kernel_manifest.json` here, only the authoring `tt-model.yaml`.
84
- - `tt-model-container` asserts a v5.1 OCI image in `image/`, which does not exist yet.
85
- - `tt-model-catalog` lists it in the public community index, which renders each entry by
86
- reading `tt_kernel_manifest.json`.
87
- - `tt-model publish <id>` does **not** check that the target is a real bundle -- it only
88
- flips visibility and sets the tag -- so the guard rail you might expect is not there.
89
-
90
- Until the bundle is built, the honest tag set is the one this card already carries
91
- (`tenstorrent`, `blackhole`, `p150`, `ttnn`, `tt-metal`, plus the task tags).
92
-
93
- ## Known sharp edges (offline-validated, not yet hardware-validated)
94
-
95
- The manifest loads clean (`load_container_manifest(..., check_sources=False)` passes, which
96
- includes the launcher's own `validate()`), but two things can only be settled on the box:
97
-
98
- 1. **`code/` becomes the image's ONLY `models` package.** `source.code: [models/common]` is
99
- staged from tt-metal into `code/models/common`, and an `extra_code` path of `models`
100
- merges the port's own tree into the same `code/models/`. tt-metal's `models/` has no
101
- `__init__.py`, so this relies on namespace packages resolving. Check the import in
102
- `verify.sh` output on the first build.
103
- 2. **Generically-named top-level packages.** Repos whose package is `models`, `tt`, or
104
- `common` put a very common name at the root of the image's import path. If an import
105
- collides, rename the package in the port repo (and update `runtime.app` +
106
- `source.extra_code[].paths` together -- the launcher cross-checks them).
107
-
108
- `source.code` needs at least one tt-metal-relative entry even when all real code comes from
109
- `extra_code`; `models/common` is used as that minimum. If your port genuinely imports more
110
- from tt-metal (`locate-anything` needs `models/tt_transformers` and
111
- `models/demos/qwen25_vl`), list it there instead -- under-listing fails the image's own
112
- build-time import check on your machine, which is the cheap place to find out.
 
1
+ # Serving SuperPoint on Blackhole with tt-model-manager
2
+
3
+ This repo is the authoring source of the **tt-model container package**
4
+ `changh95/superpoint-blackhole` (`kind: tt-dit-server`, schema 5.1). `tt-model.yaml` is
5
+ the manifest, `code/` the port, `code/models/server/app.py` the ASGI app uvicorn runs.
6
+
7
+ | item | value |
8
+ |---|---|
9
+ | tt-metal tree | `/home/deepgadget/experiments/gbp-tt/tt-metal` (main 8b98410e730, v0.78.0-dev20260820-25, torch 2.11.0 pin) |
10
+ | weights | `magic-leap-community/superpoint` @ `734450e9ffe229074f5998494ddc615475cdb20a` (`config.json`, `model.safetensors`, `preprocessor_config.json`; 5 MB; public, ungated) |
11
+ | port source | github.com/changh95/tt-superpoint @ `e1eab66e29ff424bc9af6b1118671d9bc08e899e` (+ `code/models/server/`, `code/models/tt/postprocess.py` added here) |
12
+ | hardware | one Blackhole p150 (`hardware: p150`, `mesh_device: P150`, `TT_MESH_SHAPE=1x1`) |
13
+ | app | `models.server.app:app` |
14
+
15
+ ## What the server does
16
+
17
+ Pure-ttnn, **untraced**, host NMS -- the port's `SP_TRACE_NMS=0 SP_NO_TRACE=1` configuration,
18
+ no custom kernel. Fixed 480x640 network input; every image is resized server-side. Per request:
19
+
20
+ 1. base64 -> PIL RGB -> bilinear resize to 640x480 -> /255 -> fp32 `(1, 3, 480, 640)`
21
+ (HF `SuperPointImageProcessor` defaults; the model reads channel 0 like
22
+ `SuperPointForKeypointDetection.extract_one_channel_pixel_values`).
23
+ 2. `TtSuperPoint.run_untraced(tt_in, pixel_values)`: H2D into the persistent device input,
24
+ `run_device_compute(trace_nms=False)` (8 convs + 3 max-pools + heads, bf16/HiFi2/fp32-acc,
25
+ softmax and descriptor L2-norm on device), `device_outputs_to_host` (D2H, **no second
26
+ softmax**), deallocate.
27
+ 3. `models.tt.postprocess.postprocess_keypoints`: fold 8x8 cells, single-pass NMS
28
+ (`nms_radius`), threshold, border removal (4 px), top-k (`max_keypoints`), bilinear
29
+ descriptor sampling + L2-norm. Identical to the sequence validated in
30
+ `code/models/tests/test_superpoint.py` (score PCC 0.9971, descriptor PCC 0.9991, F1 98.8%).
31
+ 4. Keypoints are scaled back to the client's pixel frame (`scale = {W/640, H/480}`).
32
+
33
+ One `threading.Lock` serialises every device call; handlers are sync and run under
34
+ `torch.inference_mode()`. Startup (weights -> device -> model -> 2 warm-up forwards on a zero
35
+ frame) happens in the ASGI lifespan, so `Application startup complete` means warm. Shutdown
36
+ deallocates the device tensors and `ttnn.close_device`s inside the 120 s SIGTERM budget.
37
+
38
+ Expected speed: ~6 fps device forward (untraced) + ~36 ms host NMS. The README's 40.7 fps
39
+ requires trace capture and the fused `sp_eq_mul_mask` C++ kernel (`code/kernels/`, needs a
40
+ patched tt-metal) -- deliberately not what this image runs.
41
+
42
+ ## Environment the app reads (lifespan only, never at import)
43
+
44
+ | var | set by | meaning / default |
45
+ |---|---|---|
46
+ | `HF_MODEL` | launcher (`weights.repo`) | weights repo id; default `magic-leap-community/superpoint` |
47
+ | `TT_WEIGHTS_REVISION` | `serve.env` | commit sha passed to `from_pretrained(revision=)`; default: repo default branch |
48
+ | `SP_WEIGHTS_DIR` | you (host/offline) | local dir with `config.json` + `model.safetensors`; overrides the two above |
49
+ | `TT_MESH_SHAPE` | launcher (`runtime.mesh_shape_env`) | `1x1` (also `(1, 1)` / `1,1`); any other shape -> `RuntimeError` at startup |
50
+ | `TT_DEVICE_ID` | you | chip to open, default `0` |
51
+ | `MESH_DEVICE` | launcher | `P150` (informational) |
52
+ | `TT_METAL_VISIBLE_DEVICES` | `serve.env` | `0` |
53
+
54
+ The launcher exports no revision, which is why `serve.env.TT_WEIGHTS_REVISION` repeats
55
+ `weights.revision`. A sha-pinned snapshot has no `refs/main`, so the app must pass the sha
56
+ (and falls back to `local_files_only=True` if the Hub is unreachable but the snapshot is cached).
57
+
58
+ ## HTTP contract
59
+
60
+ | route | response |
61
+ |---|---|
62
+ | `GET /health` | `{"status": "ok" \| "starting", "model": "superpoint-blackhole", "device": {"arch", "id", "open"}}` (always 200) |
63
+ | `GET /info` | model/task/io, hardware, `weights {repo, revision, local_dir, loaded}`, `source {repo, commit}`, `input` (480x640, batch 1, preprocessing), `defaults`, `limits`, `serving_path` (traced=false, device_nms=false), `warmup_ms`, `descriptors` encoding, `license` |
64
+ | `GET /v1/models` | `{"object": "list", "data": [{"id": "<weights repo>", "object": "model", "owned_by": "changh95"}]}` (so OpenAI-shaped probes do not 404; not a chat API) |
65
+ | `POST /predict` | see below |
66
+
67
+ Request (`application/json`):
68
+
69
+ ```json
70
+ {
71
+ "image": "<base64 PNG/JPEG>", // required; RGB or grayscale; any size (resized to 640x480)
72
+ "max_keypoints": 1024, // optional; -1 = all above threshold; cap 307200
73
+ "keypoint_threshold": 0.005, // optional; [0, 1]
74
+ "nms_radius": 4, // optional; 0..32 canonical-frame pixels; 0 = no NMS
75
+ "return_descriptors": true // optional
76
+ }
77
+ ```
78
 
79
+ Response:
80
+
81
+ ```json
82
+ {
83
+ "num_keypoints": N,
84
+ "keypoints": [[x, y], ...], // N x 2 floats, ORIGINAL image pixel coordinates
85
+ "scores": [...], // N floats, always descending (keypoints/descriptors share the order)
86
+ "original_size": {"height": H, "width": W},
87
+ "image_size": {"height": 480, "width": 640},
88
+ "scale": {"x": W/640, "y": H/480}, // divide keypoints by this to get network-frame coords
89
+ "params": {"max_keypoints", "keypoint_threshold", "nms_radius", "border_removal_distance"},
90
+ "timing_ms": {"preprocess", "device_forward", "postprocess", "total"},
91
+ "descriptors": { // only when return_descriptors is true
92
+ "format": "npz", "key": "descriptors", "dtype": "float16", "shape": [N, 256],
93
+ "data": "<base64 NPZ>" // numpy.load(io.BytesIO(base64.b64decode(data)))["descriptors"]
94
+ }
95
+ }
96
+ ```
97
 
98
+ Errors: `400` undecodable image or invalid field (pydantic errors are mapped to 400),
99
+ `503` while starting, `500` with `"<ExceptionType>: <message>"` on an inference failure.
100
+ One image per request.
101
 
102
+ Smoke test (the hardware phase runs it unchanged):
103
 
104
  ```bash
105
+ python code/models/server/smoke_test.py --url http://127.0.0.1:<port>
106
+ # PASS superpoint-blackhole: <N> keypoints on house_in_field_1080p.jpg (1600x900), ...
107
+ ```
 
 
 
 
 
 
108
 
109
+ ## Running on the HOST for validation (no Docker)
 
110
 
111
+ Uses the tree's own venv (`python_env`, Python 3.10, torch 2.11.0+cpu, transformers 5.12.1)
112
+ plus fastapi/uvicorn, which that venv lacks -- install them into a throwaway venv and append
113
+ its `site-packages` rather than touching the tree venv:
114
 
115
+ ```bash
116
+ ROOT=/home/deepgadget/experiments/tt-models
117
+ T=/home/deepgadget/experiments/gbp-tt/tt-metal
118
+ export PATH=$HOME/.local/bin:$PATH
119
+ uv venv --python 3.10 /tmp/sp-http -q && uv pip install --python /tmp/sp-http/bin/python -q fastapi uvicorn
120
+
121
+ export PYTHONPATH=$ROOT/models/superpoint-blackhole/code:$T:$T/ttnn:$T/tools:/tmp/sp-http/lib/python3.10/site-packages
122
+ export TT_METAL_HOME=$T ARCH_NAME=blackhole
123
+ export HF_MODEL=magic-leap-community/superpoint
124
+ export TT_WEIGHTS_REVISION=734450e9ffe229074f5998494ddc615475cdb20a
125
+ export TT_MESH_SHAPE=1x1 TT_DEVICE_ID=0 TT_METAL_VISIBLE_DEVICES=0
126
+
127
+ # import check with NO device (what the image's verify.sh does):
128
+ cd /tmp && $T/python_env/bin/python -c "import models.server.app as a; assert a.app"
129
+
130
+ # serve (opens the chip; hardware phase only):
131
+ cd $ROOT/models/superpoint-blackhole && $T/python_env/bin/python -m uvicorn --host 0.0.0.0 --port 20000 --lifespan on models.server.app:app
132
+ # then, from another shell:
133
+ python code/models/server/smoke_test.py --url http://127.0.0.1:20000
134
  ```
135
 
136
+ Boot log landmarks: `Loading weights ...`, `Opening device 0`, `Warming up (compiling
137
+ kernels ...)`, `Warmup complete: first forward <ms> (compile), second <ms>`, then uvicorn's
138
+ `Application startup complete`. Ctrl-C / SIGTERM closes the device.
139
 
140
+ `PYTHONPATH` must start with `code/` so that `models` resolves to this repo's regular package
141
+ (tt-metal's own `models/` has no `__init__.py` and is shadowed on the host; in the image it is
142
+ excluded). The benchmark `code/models/tests/test_superpoint.py` still needs a tt-metal checkout as
143
+ pytest rootdir for its `device` fixture -- it is unchanged by the packaging work.
144
 
145
+ ## Package / serve / push (Blackhole host, rootless Docker)
146
 
147
+ Every command from **this directory**: `extra_code.root: code` and `--out` resolve against the
148
+ CWD. `source $ROOT/bin/docker-env.sh` first (rootless Docker 28 + buildx on this box; the bare
149
+ `docker` on PATH is podman).
 
 
 
150
 
151
+ ```bash
152
+ ROOT=/home/deepgadget/experiments/tt-models
153
+ cd $ROOT/models/superpoint-blackhole
154
+ source $ROOT/bin/docker-env.sh
155
 
156
+ # offline validation (no docker, no device) -- must print VALID
157
+ $ROOT/.venv/bin/python -c "from tt_kernel.container_manifest import load_container_manifest; m = load_container_manifest('tt-model.yaml', check_sources=True); p = m.resolve_profile(); print('VALID', m.name, m.kind, p.hardware, p.mesh_device, m.weights_ref)"
 
 
 
158
 
159
+ # build (2.5-4 h cold; verify.sh imports the app + the verify: lines inside the image)
160
+ $ROOT/.venv/bin/tt-model package --container tt-model.yaml --out $ROOT/build
161
 
162
+ # serve + smoke (hardware phase)
163
+ $ROOT/.venv/bin/tt-model serve $ROOT/build/superpoint-blackhole/tt_kernel_manifest.json
164
+ python code/models/server/smoke_test.py --url http://127.0.0.1:<port serve printed>
165
+ $ROOT/.venv/bin/tt-model stop changh95/superpoint-blackhole
166
 
167
+ # publish (after validation only)
168
+ $ROOT/.venv/bin/tt-model push $ROOT/build/superpoint-blackhole --publish
169
+ ```
170
 
171
+ tt-cli users: `tt serve changh95/superpoint-blackhole` / `tt model stop changh95/superpoint-blackhole`
172
+ (`tt config set tools.override.tt-model $ROOT/.venv/bin/tt-model` is already set here because the
173
+ pinned tt-model breaks under rootless Docker).
174
+
175
+ ## What push does to this repo
176
+
177
+ `code/` and `image/` on the Hub become exactly the staged trees, so `extra_code.paths` lists
178
+ everything under `code/` worth keeping: `models`, `kernels` (fused NMS kernel source, not built
179
+ into this image), `sample_data`, `results.tsv`, `run_benchmark.sh`. `code/.gitignore` is not
180
+ listed and will be pruned. `README.md` is replaced by the generated card (all README content
181
+ worth keeping lives in `card.description` / `card.quickstart`); `media/`, `SERVING.md`,
182
+ `.gitattributes` and `tt-model.yaml` at the root survive. The orchestrator restores
183
+ `license`/`pipeline_tag` front matter after push.
184
+
185
+ ## Caveats
186
+
187
+ - **Licence**: the weights are Magic Leap "academic or non-profit organisation noncommercial
188
+ research use only". Stated in the card; publishing to the public catalog inherits it.
189
+ - **Package name `models`** collides with tt-metal's `models/` tree. It works because the
190
+ image excludes tt-metal's `models/` and the port ships `models/__init__.py`; the filler
191
+ `models/common/lightweightmodule.py` becomes a namespace subpackage nobody imports.
192
+ - **API drift**: the port was validated on a tt-metal of ~v0.71 (spring 2026); this package
193
+ builds against v0.78 (Aug 2026). `import ttnn` and every symbol the port touches
194
+ (`Conv2dSliceConfig`, `Conv2dDRAMSliceHeight`, `UnaryWithParam`, `CreateDevice`,
195
+ `copy_host_to_device_tensor(cq_id=)`, `max_pool2d` kwargs) exist in the v0.78 tree, but
196
+ conv/pool behaviour (L1 budgets of the `(4, 2, 1, 1)` DRAM slicing) is only proven on the
197
+ chip -- re-run the PCC benchmark in the hardware phase.
198
+ - **Device open**: `ttnn.CreateDevice(device_id, l1_small_size=32768)` -- the recipe of the
199
+ port's untraced script (`models/visualize.py`); no trace region, one command queue.
200
+ - **Thread model**: uvicorn runs the sync `predict` in a worker thread; the lock serialises
201
+ ttnn calls, and the lifespan opened the device on the event-loop thread (same pattern as the
202
+ template and other tt-dit servers).
203
+ - **Host venv is Python 3.10**, the image is 3.12 -- the uv dry-run of the runtime packages
204
+ on 3.12 resolves to torch 2.11.0+cpu, numpy 1.26.4, transformers 5.17.0, no torchvision.
205
+ - `transformers` in the image resolves to 5.x (5.17.0 at authoring time); `SuperPointForKeypointDetection`
206
+ exists in 4.53 .. 5.17 (`>=4.53,<6` pin).