superpoint-p150 / SERVING.md
changh95's picture
Add Tenstorrent Blackhole tt-nn port
dc14621 verified
|
Raw History Blame
4.14 kB

Serving this port with tt-model-manager

This repo carries everything needed to build a tt-model container bundle from it, so tt-model pull / tt-model serve and tt-cli can run it. Do this on a Blackhole host (amd64 Linux + Docker >= 25).

Why not on a Mac

tt-model package --container builds from ghcr.io/.../ubuntu-22.04-dev-amd64 and a cold build is 2.5-4 hours. Packaging and pushing need no card (require_host(need_devices=False)), but they do need amd64. tt-model serve needs /dev/tenstorrent/*.

Steps

# 0. get tt-model (not on PyPI)
git clone https://github.com/tenstorrent/tt-model-manager && cd tt-model-manager
python -m venv .venv && source .venv/bin/activate && pip install -e .

# 1. pull THIS repo
hf download changh95/superpoint-blackhole --local-dir superpoint-blackhole && cd superpoint-blackhole

# 2. point the manifest at your tt-metal checkout
$EDITOR tt-model.yaml          # source.tt_metal: /path/to/tt-metal

# 3. finish the serving adapter
$EDITOR code/models/server/app.py   # see "VERIFY ON HARDWARE" markers

# 4. validate with no hardware and no build
python -c "from tt_kernel.container_manifest import ContainerManifest as M; M.load('tt-model.yaml')"

# 5. build, prove, publish
tt-model package --container tt-model.yaml
tt-model serve   build/superpoint-blackhole/tt_kernel_manifest.json
tt-model push    build/superpoint-blackhole --publish

Step 5's --publish implies --public and adds the tt-model-catalog tag, which is what lists it in the community catalog.

What the bundle push does to this repo

  • code/ and image/ are replaced wholesale (_prune_removed in tt_kernel/hub.py). The port code staged here is exactly what extra_code re-stages, so it is replaced, not duplicated. Files under code/ that the allowlist does not ship get pruned.
  • README.md is overwritten by the generated card. The text you want to keep belongs in card.description / card.quickstart in tt-model.yaml -- it is already seeded there.
  • Root files survive, which is why media/ and this file live at the root.

The serving contract

kind: tt-dit-server requires only hardware + mesh_device (no max_num_seqs / block_size -- those are vLLM engine settings). runtime.app is an ASGI target "module:attr" whose top-level package must appear in an allowlist entry, which is why runtime.app and source.extra_code[0].paths have to stay in sync. The kind installs fastapi / uvicorn / pydantic / pillow; runtime.packages adds to that set.

Despite the name, this kind is not diffusion-only -- jashansinghTT/s2-pro-blackhole is a TTS model using it. It simply means "the model's own ASGI app under uvicorn".

Known sharp edges (offline-validated, not yet hardware-validated)

The manifest loads clean (load_container_manifest(..., check_sources=False) passes, which includes the launcher's own validate()), but two things can only be settled on the box:

  1. code/ becomes the image's ONLY models package. source.code: [models/common] is staged from tt-metal into code/models/common, and an extra_code path of models merges the port's own tree into the same code/models/. tt-metal's models/ has no __init__.py, so this relies on namespace packages resolving. Check the import in verify.sh output on the first build.
  2. Generically-named top-level packages. Repos whose package is models, tt, or common put a very common name at the root of the image's import path. If an import collides, rename the package in the port repo (and update runtime.app + source.extra_code[].paths together -- the launcher cross-checks them).

source.code needs at least one tt-metal-relative entry even when all real code comes from extra_code; models/common is used as that minimum. If your port genuinely imports more from tt-metal (locate-anything needs models/tt_transformers and models/demos/qwen25_vl), list it there instead -- under-listing fails the image's own build-time import check on your machine, which is the cheap place to find out.