Download SERVING.md from changh95/superpoint-p150: direct link, hf CLI and curl.
- Browser
- Download file 4.14 kB
-
https://huggingface.co/changh95/superpoint-p150/resolve/dc1462131f6ccf95ce1c9d9015f45e1cea8b3c07/SERVING.md
- Command line
-
hf download hf://changh95/superpoint-p150@dc1462131f6ccf95ce1c9d9015f45e1cea8b3c07/SERVING.md
-
curl -L -o SERVING.md https://huggingface.co/changh95/superpoint-p150/resolve/dc1462131f6ccf95ce1c9d9015f45e1cea8b3c07/SERVING.md
Serving this port with tt-model-manager
This repo carries everything needed to build a tt-model container bundle from it, so
tt-model pull / tt-model serve and tt-cli can run it. Do this on a
Blackhole host (amd64 Linux + Docker >= 25).
Why not on a Mac
tt-model package --container builds from ghcr.io/.../ubuntu-22.04-dev-amd64 and a cold
build is 2.5-4 hours. Packaging and pushing need no card (require_host(need_devices=False)),
but they do need amd64. tt-model serve needs /dev/tenstorrent/*.
Steps
# 0. get tt-model (not on PyPI)
git clone https://github.com/tenstorrent/tt-model-manager && cd tt-model-manager
python -m venv .venv && source .venv/bin/activate && pip install -e .
# 1. pull THIS repo
hf download changh95/superpoint-blackhole --local-dir superpoint-blackhole && cd superpoint-blackhole
# 2. point the manifest at your tt-metal checkout
$EDITOR tt-model.yaml # source.tt_metal: /path/to/tt-metal
# 3. finish the serving adapter
$EDITOR code/models/server/app.py # see "VERIFY ON HARDWARE" markers
# 4. validate with no hardware and no build
python -c "from tt_kernel.container_manifest import ContainerManifest as M; M.load('tt-model.yaml')"
# 5. build, prove, publish
tt-model package --container tt-model.yaml
tt-model serve build/superpoint-blackhole/tt_kernel_manifest.json
tt-model push build/superpoint-blackhole --publish
Step 5's --publish implies --public and adds the tt-model-catalog tag, which is what
lists it in the community catalog.
What the bundle push does to this repo
code/andimage/are replaced wholesale (_prune_removedintt_kernel/hub.py). The port code staged here is exactly whatextra_codere-stages, so it is replaced, not duplicated. Files undercode/that the allowlist does not ship get pruned.README.mdis overwritten by the generated card. The text you want to keep belongs incard.description/card.quickstartintt-model.yaml-- it is already seeded there.- Root files survive, which is why
media/and this file live at the root.
The serving contract
kind: tt-dit-server requires only hardware + mesh_device (no max_num_seqs /
block_size -- those are vLLM engine settings). runtime.app is an ASGI target
"module:attr" whose top-level package must appear in an allowlist entry, which is why
runtime.app and source.extra_code[0].paths have to stay in sync. The kind installs
fastapi / uvicorn / pydantic / pillow; runtime.packages adds to that set.
Despite the name, this kind is not diffusion-only -- jashansinghTT/s2-pro-blackhole is a
TTS model using it. It simply means "the model's own ASGI app under uvicorn".
Known sharp edges (offline-validated, not yet hardware-validated)
The manifest loads clean (load_container_manifest(..., check_sources=False) passes, which
includes the launcher's own validate()), but two things can only be settled on the box:
code/becomes the image's ONLYmodelspackage.source.code: [models/common]is staged from tt-metal intocode/models/common, and anextra_codepath ofmodelsmerges the port's own tree into the samecode/models/. tt-metal'smodels/has no__init__.py, so this relies on namespace packages resolving. Check the import inverify.shoutput on the first build.- Generically-named top-level packages. Repos whose package is
models,tt, orcommonput a very common name at the root of the image's import path. If an import collides, rename the package in the port repo (and updateruntime.app+source.extra_code[].pathstogether -- the launcher cross-checks them).
source.code needs at least one tt-metal-relative entry even when all real code comes from
extra_code; models/common is used as that minimum. If your port genuinely imports more
from tt-metal (locate-anything needs models/tt_transformers and
models/demos/qwen25_vl), list it there instead -- under-listing fails the image's own
build-time import check on your machine, which is the cheap place to find out.