# Serving RF-DETR-base on Blackhole with tt-model-manager This repo is a **tt-model container package source**: `tt-model.yaml` + `code/rf_detr` (the tt-nn port, with the FastAPI server under `code/rf_detr/server/`). It packages into an OCI image that `tt-model serve` / `tt serve changh95/rf-detr-p150` run on a Blackhole p150a host. Weights (`Roboflow/rf-detr-base`, Apache-2.0, ~129 MB) are a pinned pointer, never baked into the image. | | | |---|---| | tt-metal tree | `/home/deepgadget/experiments/gbp-tt/tt-metal` (main `8b98410e730`, v0.78.0-dev20260820-25; torch 2.11.0+cpu) | | weights | `Roboflow/rf-detr-base` @ `7b95b089788e6c7db56d5ea9b0a07ca08ea6ac0a` (`model.safetensors`, `config.json`) | | app | `rf_detr.server.app:app` (uvicorn, ASGI lifespan does weights -> device -> model -> warm-up) | | device recipe | `ttnn.open_device(device_id, l1_small_size=32768, trace_region_size=90_000_000, num_command_queues=1)` — the port's validated params (`conftest.py`, `scripts/make_demo.py`, `benchmark.py`) | | port source | github.com/changh95/tt-RF-DETR @ `5b0149869465892477141b4c1dfc4622db0d1ec5` (+ `rf_detr/server/` added here) | Below, `$ROOT=/home/deepgadget/experiments/tt-models`. ## 1. Run on the HOST (hardware validation, no Docker) The tree venv `$TT_METAL_HOME/python_env` (py3.10, torch 2.11.0+cpu, ttnn editable) has everything except `fastapi`/`uvicorn`. Do not install into the tree venv; put the HTTP stack in a side directory and prepend it to `PYTHONPATH`: ```bash export TT_METAL_HOME=/home/deepgadget/experiments/gbp-tt/tt-metal export PATH=$HOME/.local/bin:$PATH # uv mkdir -p /tmp/rfdetr-http && uv pip install --python $TT_METAL_HOME/python_env/bin/python \ --target /tmp/rfdetr-http fastapi uvicorn cd $ROOT/models/rf-detr-p150 export PYTHONPATH=$PWD/code:$TT_METAL_HOME:$TT_METAL_HOME/ttnn:$TT_METAL_HOME/tools:/tmp/rfdetr-http export HF_MODEL=Roboflow/rf-detr-base export TT_WEIGHTS_REVISION=7b95b089788e6c7db56d5ea9b0a07ca08ea6ac0a # already in ~/.cache/huggingface export TT_MESH_SHAPE=1x1 export TT_DEVICE_ID=0 # optional (default 0) export TT_METAL_VISIBLE_DEVICES=0 # optional, one-card box $TT_METAL_HOME/python_env/bin/python -m uvicorn --host 0.0.0.0 --port 20000 --lifespan on rf_detr.server.app:app ``` Boot log landmarks: `Loading weights` -> `Opening device` -> `TT model built` -> `Warming up: capturing trace` -> `Warmup complete` -> uvicorn `Application startup complete`. The first boot compiles kernels (minutes on a cold `~/.cache/ttnn` / `TT_METAL_CACHE`); later boots take seconds. Startup failures raise and uvicorn exits non-zero (no CPU fallback). Stop with Ctrl-C / SIGTERM: the lifespan releases the trace, drops the device tensors and calls `ttnn.close_device`. Then, from a second shell (stdlib only, any Python): ```bash python code/rf_detr/server/smoke_test.py --url http://127.0.0.1:20000 # PASS/FAIL one-liner python code/rf_detr/server/smoke_test.py --url http://127.0.0.1:20000 --wait 1800 --out /tmp/rfdetr.json ``` Expected on the demo image (`media/demo_source.png`, 640x480): `cat` 0.96 / 0.94, `remote` 0.90 / 0.73, `couch` 0.67 — `PASS n=5 counts={"cat": 2, "couch": 1, "remote": 2} ...`. The CPU reference gives cat 0.960/0.933, remote 0.898/0.728, couch 0.673 for the same input. Offline override (no Hub, no HF cache): `TT_RF_DETR_WEIGHTS=`. Offline checks that need no device (what the implement phase ran): ```bash # manifest + launcher preview (from the repo dir — `root: code` is CWD-relative) $ROOT/.venv/bin/python -c "from tt_kernel.container_manifest import load_container_manifest; m = load_container_manifest('tt-model.yaml', check_sources=True); p = m.resolve_profile(); print('VALID', m.name, m.kind, p.hardware, p.mesh_device, m.weights_ref); from tt_kernel.launchers import launcher_for; L = launcher_for(m.kind); print(L.install_lines(m)); print(L.verify_lines(m))" # import the app with no device PYTHONPATH=$PWD/code:$TT_METAL_HOME:$TT_METAL_HOME/ttnn:$TT_METAL_HOME/tools:/tmp/rfdetr-http $TT_METAL_HOME/python_env/bin/python -c "import rf_detr.server.app as a; assert a.app" ``` ## 2. Package, serve, push (Docker) Rootless Docker on this box needs `source $ROOT/bin/docker-env.sh` first (sets PATH and `DOCKER_HOST`; the bare `docker` on PATH is podman). **Run every `tt-model` command from this directory**: `source.tt_metal` and `source.extra_code[].root: code` resolve against the process CWD, not against the YAML. `--out` points outside the git checkout because `stage()` deletes `/rf-detr-p150` before rebuilding. ```bash source $ROOT/bin/docker-env.sh cd $ROOT/models/rf-detr-p150 tt-model package --container tt-model.yaml --out $ROOT/build # builds the image (hours cold), runs verify.sh tt-model serve $ROOT/build/rf-detr-p150/tt_kernel_manifest.json # pre-downloads the pinned weights, boots, waits for READY python code/rf_detr/server/smoke_test.py --url http://127.0.0.1: tt-model stop changh95/rf-detr-p150 tt-model push $ROOT/build/rf-detr-p150 --publish # uploads code/, image/, README (generated card) ``` `serve` publishes the first free port from 20000 and exports into the container exactly: `HF_MODEL=Roboflow/rf-detr-base`, `MESH_DEVICE=P150`, `TT_MESH_SHAPE=1x1`, then `serve.env` (`TT_WEIGHTS_REVISION=7b95b0…`, `TT_METAL_VISIBLE_DEVICES=0`). The HF cache is mounted at `/hf` (rw), kernels at `/cache`, and `/weight-cache` is unused by this port (weights are converted to bf16/bf8 device tensors at boot in seconds). tt-cli users: `tt serve changh95/rf-detr-p150` / `tt model stop changh95/rf-detr-p150`. `push` makes `code/` and `image/` on the Hub exactly the staged trees (`rf_detr`, `scripts`, `conftest.py` and the `models/common/lightweightmodule.py` filler) and replaces `README.md` with the generated card (everything worth keeping lives in `card.description` / `card.quickstart`). `media/`, `SERVING.md`, `.gitattributes` and `tt-model.yaml` at the repo root survive; the orchestrator restores the `license` / `pipeline_tag` front matter afterwards. ## 3. The request / response contract | route | purpose | |---|---| | `GET /health` | `{"status": "ok" \| "starting", "model": "RF-DETR-base", "device": "blackhole p150 (device_id=0, mesh 1x1)" \| null}` — 200 always; `ok` only after warm-up | | `GET /info` | model, task, hardware, weights repo + revision + resolved files, port source + commit, input limits (560x560, 1 image, batch 1), COCO labels, device params | | `GET /v1/models` | `{"object": "list", "data": [{"id": "Roboflow/rf-detr-base", "object": "model", "owned_by": "changh95"}]}` — only so the OpenAI-shaped ready card / `tt-model curl` do not 404 | | `POST /predict` | one image -> detections (below) | `POST /predict` request (JSON): | field | type | default | meaning | |---|---|---|---| | `image` | str | required | base64 PNG/JPEG, any size (max side 8192, max 32 MB); a `data:image/...;base64,` prefix and wrapped base64 are tolerated | | `threshold` | float 0..1 | `0.5` | keep detections with `sigmoid(logit) > threshold` | | `max_detections` | int 1..300 | `100` | keep at most this many (highest scores first) | | `return_raw` | bool | `false` | also return `raw.logits` `[300][91]` and `raw.pred_boxes` `[300][4]` (normalised cxcywh) | Response (200): ```json { "detections": [ {"label": "cat", "label_id": 17, "score": 0.96, "box": [7.35, 54.64, 318.45, 472.14], "box_cxcywh_norm": [0.2545, 0.5487, 0.4861, 0.8698]} ], "num_detections": 5, "image_size": {"width": 640, "height": 480}, "input_size": [560, 560], "threshold": 0.5, "timing_ms": {"decode": 3.1, "preprocess": 4.0, "inference": 47.0, "postprocess": 0.4, "total": 54.5} } ``` `box` is `[x1, y1, x2, y2]` in pixels of the ORIGINAL image (the preprocessor squashes the whole image to 560x560, so the model's normalised cxcywh is relative to the original too; clamped to the image bounds), sorted by score descending. `label_id` is the COCO-91 index from `config.json`'s `id2label`. Errors: **400** undecodable/oversized image, **422** schema violation (missing `image`, `threshold` out of range), **503** while starting, **500** with `inference failed: : ` if the device call raises. Server-side pipeline (the port's validated recipe): PIL RGB -> bilinear+antialias resize to 560x560 -> /255 -> ImageNet mean/std -> `TtRfDetr(pixel_values)` under one global `threading.Lock` and `torch.inference_mode()` -> `sigmoid(logits).max(-1)` -> threshold -> sort -> scale cxcywh by (W, H). ## 4. Caveats - **Batch 1, one image per request.** The TT graph is shape-locked to 560x560 (16 windows x 101 tokens, 40x40 feature grid, 300 queries, 91 classes). Requests are serialised by a lock; concurrent clients queue. - **Warm-up captures the metal-trace** of the projector+transformer tail inside the lifespan (two forwards on a synthetic 560x560 image), so READY means warm; the first cold boot pays the ttnn JIT (minutes), later boots reuse `/cache` (`TT_METAL_CACHE`). - **Weights are pinned by sha** in both `weights.revision` and `serve.env.TT_WEIGHTS_REVISION`; `hf_hub_download(..., revision=)` is a pure cache hit after `serve`'s pre-download and falls back to `local_files_only=True` if the Hub is unreachable. - **Precision knobs stay off.** `RFDETR_BB_FIDELITY` / `RFDETR_BB_FP32ACC` / `RFDETR_BB_L1` are read by `TtRfDetr` but every setting other than the default regressed accuracy or hung (port README); do not put them in `serve.env`. - **ttnn API drift** is the main hardware-phase risk: the port was validated on a ~June-2026 tree; the packaged tree is 2026-08-20. `Conv2dConfig.shard_layout`, `ttnn.topk` index dtype, `execute_trace` kwargs and `grid_sample` kwargs all exist in this tree (checked offline); numerics/gate (98.5 detection-IoU) must be re-confirmed with `pytest code/rf_detr/tests/test_pretrained_eval.py` or the smoke test on the box. - The `tt_metal/python_env` venv on this host is Python 3.10 (the image uses 3.12); the code is version-agnostic (3.9+ syntax, `from __future__ import annotations`). - `tt-model curl` and the ready card's `/v1/models` hint are OpenAI-shaped and are not this API; use the routes above.