rf-detr-p150 / SERVING.md
changh95's picture
tt-model authoring files
bfb7f89 verified
|
Raw History Blame Contribute Delete
10.4 kB

Serving RF-DETR-base on Blackhole with tt-model-manager

This repo is a tt-model container package source: tt-model.yaml + code/rf_detr (the tt-nn port, with the FastAPI server under code/rf_detr/server/). It packages into an OCI image that tt-model serve / tt serve changh95/rf-detr-p150 run on a Blackhole p150a host. Weights (Roboflow/rf-detr-base, Apache-2.0, ~129 MB) are a pinned pointer, never baked into the image.

tt-metal tree /home/deepgadget/experiments/gbp-tt/tt-metal (main 8b98410e730, v0.78.0-dev20260820-25; torch 2.11.0+cpu)
weights Roboflow/rf-detr-base @ 7b95b089788e6c7db56d5ea9b0a07ca08ea6ac0a (model.safetensors, config.json)
app rf_detr.server.app:app (uvicorn, ASGI lifespan does weights -> device -> model -> warm-up)
device recipe ttnn.open_device(device_id, l1_small_size=32768, trace_region_size=90_000_000, num_command_queues=1) — the port's validated params (conftest.py, scripts/make_demo.py, benchmark.py)
port source github.com/changh95/tt-RF-DETR @ 5b0149869465892477141b4c1dfc4622db0d1ec5 (+ rf_detr/server/ added here)

Below, $ROOT=/home/deepgadget/experiments/tt-models.

1. Run on the HOST (hardware validation, no Docker)

The tree venv $TT_METAL_HOME/python_env (py3.10, torch 2.11.0+cpu, ttnn editable) has everything except fastapi/uvicorn. Do not install into the tree venv; put the HTTP stack in a side directory and prepend it to PYTHONPATH:

export TT_METAL_HOME=/home/deepgadget/experiments/gbp-tt/tt-metal
export PATH=$HOME/.local/bin:$PATH                 # uv
mkdir -p /tmp/rfdetr-http && uv pip install --python $TT_METAL_HOME/python_env/bin/python \
    --target /tmp/rfdetr-http fastapi uvicorn

cd $ROOT/models/rf-detr-p150
export PYTHONPATH=$PWD/code:$TT_METAL_HOME:$TT_METAL_HOME/ttnn:$TT_METAL_HOME/tools:/tmp/rfdetr-http
export HF_MODEL=Roboflow/rf-detr-base
export TT_WEIGHTS_REVISION=7b95b089788e6c7db56d5ea9b0a07ca08ea6ac0a   # already in ~/.cache/huggingface
export TT_MESH_SHAPE=1x1
export TT_DEVICE_ID=0                                                  # optional (default 0)
export TT_METAL_VISIBLE_DEVICES=0                                      # optional, one-card box

$TT_METAL_HOME/python_env/bin/python -m uvicorn --host 0.0.0.0 --port 20000 --lifespan on rf_detr.server.app:app

Boot log landmarks: Loading weights -> Opening device -> TT model built -> Warming up: capturing trace -> Warmup complete -> uvicorn Application startup complete. The first boot compiles kernels (minutes on a cold ~/.cache/ttnn / TT_METAL_CACHE); later boots take seconds. Startup failures raise and uvicorn exits non-zero (no CPU fallback). Stop with Ctrl-C / SIGTERM: the lifespan releases the trace, drops the device tensors and calls ttnn.close_device.

Then, from a second shell (stdlib only, any Python):

python code/rf_detr/server/smoke_test.py --url http://127.0.0.1:20000        # PASS/FAIL one-liner
python code/rf_detr/server/smoke_test.py --url http://127.0.0.1:20000 --wait 1800 --out /tmp/rfdetr.json

Expected on the demo image (media/demo_source.png, 640x480): cat 0.96 / 0.94, remote 0.90 / 0.73, couch 0.67 — PASS n=5 counts={"cat": 2, "couch": 1, "remote": 2} .... The CPU reference gives cat 0.960/0.933, remote 0.898/0.728, couch 0.673 for the same input.

Offline override (no Hub, no HF cache): TT_RF_DETR_WEIGHTS=<dir holding model.safetensors + config.json>.

Offline checks that need no device (what the implement phase ran):

# manifest + launcher preview (from the repo dir — `root: code` is CWD-relative)
$ROOT/.venv/bin/python -c "from tt_kernel.container_manifest import load_container_manifest; m = load_container_manifest('tt-model.yaml', check_sources=True); p = m.resolve_profile(); print('VALID', m.name, m.kind, p.hardware, p.mesh_device, m.weights_ref); from tt_kernel.launchers import launcher_for; L = launcher_for(m.kind); print(L.install_lines(m)); print(L.verify_lines(m))"
# import the app with no device
PYTHONPATH=$PWD/code:$TT_METAL_HOME:$TT_METAL_HOME/ttnn:$TT_METAL_HOME/tools:/tmp/rfdetr-http $TT_METAL_HOME/python_env/bin/python -c "import rf_detr.server.app as a; assert a.app"

2. Package, serve, push (Docker)

Rootless Docker on this box needs source $ROOT/bin/docker-env.sh first (sets PATH and DOCKER_HOST; the bare docker on PATH is podman). Run every tt-model command from this directory: source.tt_metal and source.extra_code[].root: code resolve against the process CWD, not against the YAML. --out points outside the git checkout because stage() deletes <out>/rf-detr-p150 before rebuilding.

source $ROOT/bin/docker-env.sh
cd $ROOT/models/rf-detr-p150

tt-model package --container tt-model.yaml --out $ROOT/build       # builds the image (hours cold), runs verify.sh
tt-model serve   $ROOT/build/rf-detr-p150/tt_kernel_manifest.json   # pre-downloads the pinned weights, boots, waits for READY
python code/rf_detr/server/smoke_test.py --url http://127.0.0.1:<port serve printed>
tt-model stop    changh95/rf-detr-p150
tt-model push    $ROOT/build/rf-detr-p150 --publish            # uploads code/, image/, README (generated card)

serve publishes the first free port from 20000 and exports into the container exactly: HF_MODEL=Roboflow/rf-detr-base, MESH_DEVICE=P150, TT_MESH_SHAPE=1x1, then serve.env (TT_WEIGHTS_REVISION=7b95b0…, TT_METAL_VISIBLE_DEVICES=0). The HF cache is mounted at /hf (rw), kernels at /cache, and /weight-cache is unused by this port (weights are converted to bf16/bf8 device tensors at boot in seconds). tt-cli users: tt serve changh95/rf-detr-p150 / tt model stop changh95/rf-detr-p150.

push makes code/ and image/ on the Hub exactly the staged trees (rf_detr, scripts, conftest.py and the models/common/lightweightmodule.py filler) and replaces README.md with the generated card (everything worth keeping lives in card.description / card.quickstart). media/, SERVING.md, .gitattributes and tt-model.yaml at the repo root survive; the orchestrator restores the license / pipeline_tag front matter afterwards.

3. The request / response contract

route purpose
GET /health {"status": "ok" | "starting", "model": "RF-DETR-base", "device": "blackhole p150 (device_id=0, mesh 1x1)" | null} — 200 always; ok only after warm-up
GET /info model, task, hardware, weights repo + revision + resolved files, port source + commit, input limits (560x560, 1 image, batch 1), COCO labels, device params
GET /v1/models {"object": "list", "data": [{"id": "Roboflow/rf-detr-base", "object": "model", "owned_by": "changh95"}]} — only so the OpenAI-shaped ready card / tt-model curl do not 404
POST /predict one image -> detections (below)

POST /predict request (JSON):

field type default meaning
image str required base64 PNG/JPEG, any size (max side 8192, max 32 MB); a data:image/...;base64, prefix and wrapped base64 are tolerated
threshold float 0..1 0.5 keep detections with sigmoid(logit) > threshold
max_detections int 1..300 100 keep at most this many (highest scores first)
return_raw bool false also return raw.logits [300][91] and raw.pred_boxes [300][4] (normalised cxcywh)

Response (200):

{
  "detections": [
    {"label": "cat", "label_id": 17, "score": 0.96,
     "box": [7.35, 54.64, 318.45, 472.14], "box_cxcywh_norm": [0.2545, 0.5487, 0.4861, 0.8698]}
  ],
  "num_detections": 5,
  "image_size": {"width": 640, "height": 480},
  "input_size": [560, 560],
  "threshold": 0.5,
  "timing_ms": {"decode": 3.1, "preprocess": 4.0, "inference": 47.0, "postprocess": 0.4, "total": 54.5}
}

box is [x1, y1, x2, y2] in pixels of the ORIGINAL image (the preprocessor squashes the whole image to 560x560, so the model's normalised cxcywh is relative to the original too; clamped to the image bounds), sorted by score descending. label_id is the COCO-91 index from config.json's id2label. Errors: 400 undecodable/oversized image, 422 schema violation (missing image, threshold out of range), 503 while starting, 500 with inference failed: <ExceptionType>: <text> if the device call raises.

Server-side pipeline (the port's validated recipe): PIL RGB -> bilinear+antialias resize to 560x560 -> /255 -> ImageNet mean/std -> TtRfDetr(pixel_values) under one global threading.Lock and torch.inference_mode() -> sigmoid(logits).max(-1) -> threshold -> sort -> scale cxcywh by (W, H).

4. Caveats

  • Batch 1, one image per request. The TT graph is shape-locked to 560x560 (16 windows x 101 tokens, 40x40 feature grid, 300 queries, 91 classes). Requests are serialised by a lock; concurrent clients queue.
  • Warm-up captures the metal-trace of the projector+transformer tail inside the lifespan (two forwards on a synthetic 560x560 image), so READY means warm; the first cold boot pays the ttnn JIT (minutes), later boots reuse /cache (TT_METAL_CACHE).
  • Weights are pinned by sha in both weights.revision and serve.env.TT_WEIGHTS_REVISION; hf_hub_download(..., revision=<sha>) is a pure cache hit after serve's pre-download and falls back to local_files_only=True if the Hub is unreachable.
  • Precision knobs stay off. RFDETR_BB_FIDELITY / RFDETR_BB_FP32ACC / RFDETR_BB_L1 are read by TtRfDetr but every setting other than the default regressed accuracy or hung (port README); do not put them in serve.env.
  • ttnn API drift is the main hardware-phase risk: the port was validated on a ~June-2026 tree; the packaged tree is 2026-08-20. Conv2dConfig.shard_layout, ttnn.topk index dtype, execute_trace kwargs and grid_sample kwargs all exist in this tree (checked offline); numerics/gate (98.5 detection-IoU) must be re-confirmed with pytest code/rf_detr/tests/test_pretrained_eval.py or the smoke test on the box.
  • The tt_metal/python_env venv on this host is Python 3.10 (the image uses 3.12); the code is version-agnostic (3.9+ syntax, from __future__ import annotations).
  • tt-model curl and the ready card's /v1/models hint are OpenAI-shaped and are not this API; use the routes above.