# Serving RF-DETR-base on Blackhole with tt-model-manager
This repo is a **tt-model container package source**: `tt-model.yaml` + `code/rf_detr`
(the tt-nn port, with the FastAPI server under `code/rf_detr/server/`). It packages into
an OCI image that `tt-model serve` / `tt serve changh95/rf-detr-p150` run on a
Blackhole p150a host. Weights (`Roboflow/rf-detr-base`, Apache-2.0, ~129 MB) are a pinned
pointer, never baked into the image.
| | |
|---|---|
| tt-metal tree | `/home/deepgadget/experiments/gbp-tt/tt-metal` (main `8b98410e730`, v0.78.0-dev20260820-25; torch 2.11.0+cpu) |
| weights | `Roboflow/rf-detr-base` @ `7b95b089788e6c7db56d5ea9b0a07ca08ea6ac0a` (`model.safetensors`, `config.json`) |
| app | `rf_detr.server.app:app` (uvicorn, ASGI lifespan does weights -> device -> model -> warm-up) |
| device recipe | `ttnn.open_device(device_id, l1_small_size=32768, trace_region_size=90_000_000, num_command_queues=1)` — the port's validated params (`conftest.py`, `scripts/make_demo.py`, `benchmark.py`) |
| port source | github.com/changh95/tt-RF-DETR @ `5b0149869465892477141b4c1dfc4622db0d1ec5` (+ `rf_detr/server/` added here) |
Below, `$ROOT=/home/deepgadget/experiments/tt-models`.
## 1. Run on the HOST (hardware validation, no Docker)
The tree venv `$TT_METAL_HOME/python_env` (py3.10, torch 2.11.0+cpu, ttnn editable) has
everything except `fastapi`/`uvicorn`. Do not install into the tree venv; put the HTTP
stack in a side directory and prepend it to `PYTHONPATH`:
```bash
export TT_METAL_HOME=/home/deepgadget/experiments/gbp-tt/tt-metal
export PATH=$HOME/.local/bin:$PATH # uv
mkdir -p /tmp/rfdetr-http && uv pip install --python $TT_METAL_HOME/python_env/bin/python \
--target /tmp/rfdetr-http fastapi uvicorn
cd $ROOT/models/rf-detr-p150
export PYTHONPATH=$PWD/code:$TT_METAL_HOME:$TT_METAL_HOME/ttnn:$TT_METAL_HOME/tools:/tmp/rfdetr-http
export HF_MODEL=Roboflow/rf-detr-base
export TT_WEIGHTS_REVISION=7b95b089788e6c7db56d5ea9b0a07ca08ea6ac0a # already in ~/.cache/huggingface
export TT_MESH_SHAPE=1x1
export TT_DEVICE_ID=0 # optional (default 0)
export TT_METAL_VISIBLE_DEVICES=0 # optional, one-card box
$TT_METAL_HOME/python_env/bin/python -m uvicorn --host 0.0.0.0 --port 20000 --lifespan on rf_detr.server.app:app
```
Boot log landmarks: `Loading weights` -> `Opening device` -> `TT model built` ->
`Warming up: capturing trace` -> `Warmup complete` -> uvicorn `Application startup complete`.
The first boot compiles kernels (minutes on a cold `~/.cache/ttnn` / `TT_METAL_CACHE`);
later boots take seconds. Startup failures raise and uvicorn exits non-zero (no CPU
fallback). Stop with Ctrl-C / SIGTERM: the lifespan releases the trace, drops the device
tensors and calls `ttnn.close_device`.
Then, from a second shell (stdlib only, any Python):
```bash
python code/rf_detr/server/smoke_test.py --url http://127.0.0.1:20000 # PASS/FAIL one-liner
python code/rf_detr/server/smoke_test.py --url http://127.0.0.1:20000 --wait 1800 --out /tmp/rfdetr.json
```
Expected on the demo image (`media/demo_source.png`, 640x480): `cat` 0.96 / 0.94,
`remote` 0.90 / 0.73, `couch` 0.67 — `PASS n=5 counts={"cat": 2, "couch": 1, "remote": 2} ...`.
The CPU reference gives cat 0.960/0.933, remote 0.898/0.728, couch 0.673 for the same input.
Offline override (no Hub, no HF cache): `TT_RF_DETR_WEIGHTS=
`.
Offline checks that need no device (what the implement phase ran):
```bash
# manifest + launcher preview (from the repo dir — `root: code` is CWD-relative)
$ROOT/.venv/bin/python -c "from tt_kernel.container_manifest import load_container_manifest; m = load_container_manifest('tt-model.yaml', check_sources=True); p = m.resolve_profile(); print('VALID', m.name, m.kind, p.hardware, p.mesh_device, m.weights_ref); from tt_kernel.launchers import launcher_for; L = launcher_for(m.kind); print(L.install_lines(m)); print(L.verify_lines(m))"
# import the app with no device
PYTHONPATH=$PWD/code:$TT_METAL_HOME:$TT_METAL_HOME/ttnn:$TT_METAL_HOME/tools:/tmp/rfdetr-http $TT_METAL_HOME/python_env/bin/python -c "import rf_detr.server.app as a; assert a.app"
```
## 2. Package, serve, push (Docker)
Rootless Docker on this box needs `source $ROOT/bin/docker-env.sh` first (sets PATH and
`DOCKER_HOST`; the bare `docker` on PATH is podman). **Run every `tt-model` command from
this directory**: `source.tt_metal` and `source.extra_code[].root: code` resolve against
the process CWD, not against the YAML. `--out` points outside the git checkout because
`stage()` deletes `/rf-detr-p150` before rebuilding.
```bash
source $ROOT/bin/docker-env.sh
cd $ROOT/models/rf-detr-p150
tt-model package --container tt-model.yaml --out $ROOT/build # builds the image (hours cold), runs verify.sh
tt-model serve $ROOT/build/rf-detr-p150/tt_kernel_manifest.json # pre-downloads the pinned weights, boots, waits for READY
python code/rf_detr/server/smoke_test.py --url http://127.0.0.1:
tt-model stop changh95/rf-detr-p150
tt-model push $ROOT/build/rf-detr-p150 --publish # uploads code/, image/, README (generated card)
```
`serve` publishes the first free port from 20000 and exports into the container exactly:
`HF_MODEL=Roboflow/rf-detr-base`, `MESH_DEVICE=P150`, `TT_MESH_SHAPE=1x1`, then `serve.env`
(`TT_WEIGHTS_REVISION=7b95b0…`, `TT_METAL_VISIBLE_DEVICES=0`). The HF cache is mounted at
`/hf` (rw), kernels at `/cache`, and `/weight-cache` is unused by this port (weights are
converted to bf16/bf8 device tensors at boot in seconds). tt-cli users:
`tt serve changh95/rf-detr-p150` / `tt model stop changh95/rf-detr-p150`.
`push` makes `code/` and `image/` on the Hub exactly the staged trees (`rf_detr`,
`scripts`, `conftest.py` and the `models/common/lightweightmodule.py` filler) and replaces
`README.md` with the generated card (everything worth keeping lives in
`card.description` / `card.quickstart`). `media/`, `SERVING.md`, `.gitattributes` and
`tt-model.yaml` at the repo root survive; the orchestrator restores the `license` /
`pipeline_tag` front matter afterwards.
## 3. The request / response contract
| route | purpose |
|---|---|
| `GET /health` | `{"status": "ok" \| "starting", "model": "RF-DETR-base", "device": "blackhole p150 (device_id=0, mesh 1x1)" \| null}` — 200 always; `ok` only after warm-up |
| `GET /info` | model, task, hardware, weights repo + revision + resolved files, port source + commit, input limits (560x560, 1 image, batch 1), COCO labels, device params |
| `GET /v1/models` | `{"object": "list", "data": [{"id": "Roboflow/rf-detr-base", "object": "model", "owned_by": "changh95"}]}` — only so the OpenAI-shaped ready card / `tt-model curl` do not 404 |
| `POST /predict` | one image -> detections (below) |
`POST /predict` request (JSON):
| field | type | default | meaning |
|---|---|---|---|
| `image` | str | required | base64 PNG/JPEG, any size (max side 8192, max 32 MB); a `data:image/...;base64,` prefix and wrapped base64 are tolerated |
| `threshold` | float 0..1 | `0.5` | keep detections with `sigmoid(logit) > threshold` |
| `max_detections` | int 1..300 | `100` | keep at most this many (highest scores first) |
| `return_raw` | bool | `false` | also return `raw.logits` `[300][91]` and `raw.pred_boxes` `[300][4]` (normalised cxcywh) |
Response (200):
```json
{
"detections": [
{"label": "cat", "label_id": 17, "score": 0.96,
"box": [7.35, 54.64, 318.45, 472.14], "box_cxcywh_norm": [0.2545, 0.5487, 0.4861, 0.8698]}
],
"num_detections": 5,
"image_size": {"width": 640, "height": 480},
"input_size": [560, 560],
"threshold": 0.5,
"timing_ms": {"decode": 3.1, "preprocess": 4.0, "inference": 47.0, "postprocess": 0.4, "total": 54.5}
}
```
`box` is `[x1, y1, x2, y2]` in pixels of the ORIGINAL image (the preprocessor squashes the
whole image to 560x560, so the model's normalised cxcywh is relative to the original too;
clamped to the image bounds), sorted by score descending. `label_id` is the COCO-91 index
from `config.json`'s `id2label`. Errors: **400** undecodable/oversized image, **422**
schema violation (missing `image`, `threshold` out of range), **503** while starting,
**500** with `inference failed: : ` if the device call raises.
Server-side pipeline (the port's validated recipe): PIL RGB -> bilinear+antialias resize
to 560x560 -> /255 -> ImageNet mean/std -> `TtRfDetr(pixel_values)` under one global
`threading.Lock` and `torch.inference_mode()` -> `sigmoid(logits).max(-1)` -> threshold ->
sort -> scale cxcywh by (W, H).
## 4. Caveats
- **Batch 1, one image per request.** The TT graph is shape-locked to 560x560
(16 windows x 101 tokens, 40x40 feature grid, 300 queries, 91 classes). Requests are
serialised by a lock; concurrent clients queue.
- **Warm-up captures the metal-trace** of the projector+transformer tail inside the
lifespan (two forwards on a synthetic 560x560 image), so READY means warm; the first
cold boot pays the ttnn JIT (minutes), later boots reuse `/cache` (`TT_METAL_CACHE`).
- **Weights are pinned by sha** in both `weights.revision` and `serve.env.TT_WEIGHTS_REVISION`;
`hf_hub_download(..., revision=)` is a pure cache hit after `serve`'s pre-download
and falls back to `local_files_only=True` if the Hub is unreachable.
- **Precision knobs stay off.** `RFDETR_BB_FIDELITY` / `RFDETR_BB_FP32ACC` / `RFDETR_BB_L1`
are read by `TtRfDetr` but every setting other than the default regressed accuracy or hung
(port README); do not put them in `serve.env`.
- **ttnn API drift** is the main hardware-phase risk: the port was validated on a
~June-2026 tree; the packaged tree is 2026-08-20. `Conv2dConfig.shard_layout`,
`ttnn.topk` index dtype, `execute_trace` kwargs and `grid_sample` kwargs all exist in
this tree (checked offline); numerics/gate (98.5 detection-IoU) must be re-confirmed
with `pytest code/rf_detr/tests/test_pretrained_eval.py` or the smoke test on the box.
- The `tt_metal/python_env` venv on this host is Python 3.10 (the image uses 3.12); the
code is version-agnostic (3.9+ syntax, `from __future__ import annotations`).
- `tt-model curl` and the ready card's `/v1/models` hint are OpenAI-shaped and are not this
API; use the routes above.