Download SERVING.md from changh95/rf-detr-p150: direct link, hf CLI and curl.
- Browser
- Download file 10.4 kB
-
https://huggingface.co/changh95/rf-detr-p150/resolve/main/SERVING.md
- Command line
-
hf download hf://changh95/rf-detr-p150/SERVING.md
-
curl -L -o SERVING.md https://huggingface.co/changh95/rf-detr-p150/resolve/main/SERVING.md
Serving RF-DETR-base on Blackhole with tt-model-manager
This repo is a tt-model container package source: tt-model.yaml + code/rf_detr
(the tt-nn port, with the FastAPI server under code/rf_detr/server/). It packages into
an OCI image that tt-model serve / tt serve changh95/rf-detr-p150 run on a
Blackhole p150a host. Weights (Roboflow/rf-detr-base, Apache-2.0, ~129 MB) are a pinned
pointer, never baked into the image.
| tt-metal tree | /home/deepgadget/experiments/gbp-tt/tt-metal (main 8b98410e730, v0.78.0-dev20260820-25; torch 2.11.0+cpu) |
| weights | Roboflow/rf-detr-base @ 7b95b089788e6c7db56d5ea9b0a07ca08ea6ac0a (model.safetensors, config.json) |
| app | rf_detr.server.app:app (uvicorn, ASGI lifespan does weights -> device -> model -> warm-up) |
| device recipe | ttnn.open_device(device_id, l1_small_size=32768, trace_region_size=90_000_000, num_command_queues=1) — the port's validated params (conftest.py, scripts/make_demo.py, benchmark.py) |
| port source | github.com/changh95/tt-RF-DETR @ 5b0149869465892477141b4c1dfc4622db0d1ec5 (+ rf_detr/server/ added here) |
Below, $ROOT=/home/deepgadget/experiments/tt-models.
1. Run on the HOST (hardware validation, no Docker)
The tree venv $TT_METAL_HOME/python_env (py3.10, torch 2.11.0+cpu, ttnn editable) has
everything except fastapi/uvicorn. Do not install into the tree venv; put the HTTP
stack in a side directory and prepend it to PYTHONPATH:
export TT_METAL_HOME=/home/deepgadget/experiments/gbp-tt/tt-metal
export PATH=$HOME/.local/bin:$PATH # uv
mkdir -p /tmp/rfdetr-http && uv pip install --python $TT_METAL_HOME/python_env/bin/python \
--target /tmp/rfdetr-http fastapi uvicorn
cd $ROOT/models/rf-detr-p150
export PYTHONPATH=$PWD/code:$TT_METAL_HOME:$TT_METAL_HOME/ttnn:$TT_METAL_HOME/tools:/tmp/rfdetr-http
export HF_MODEL=Roboflow/rf-detr-base
export TT_WEIGHTS_REVISION=7b95b089788e6c7db56d5ea9b0a07ca08ea6ac0a # already in ~/.cache/huggingface
export TT_MESH_SHAPE=1x1
export TT_DEVICE_ID=0 # optional (default 0)
export TT_METAL_VISIBLE_DEVICES=0 # optional, one-card box
$TT_METAL_HOME/python_env/bin/python -m uvicorn --host 0.0.0.0 --port 20000 --lifespan on rf_detr.server.app:app
Boot log landmarks: Loading weights -> Opening device -> TT model built ->
Warming up: capturing trace -> Warmup complete -> uvicorn Application startup complete.
The first boot compiles kernels (minutes on a cold ~/.cache/ttnn / TT_METAL_CACHE);
later boots take seconds. Startup failures raise and uvicorn exits non-zero (no CPU
fallback). Stop with Ctrl-C / SIGTERM: the lifespan releases the trace, drops the device
tensors and calls ttnn.close_device.
Then, from a second shell (stdlib only, any Python):
python code/rf_detr/server/smoke_test.py --url http://127.0.0.1:20000 # PASS/FAIL one-liner
python code/rf_detr/server/smoke_test.py --url http://127.0.0.1:20000 --wait 1800 --out /tmp/rfdetr.json
Expected on the demo image (media/demo_source.png, 640x480): cat 0.96 / 0.94,
remote 0.90 / 0.73, couch 0.67 — PASS n=5 counts={"cat": 2, "couch": 1, "remote": 2} ....
The CPU reference gives cat 0.960/0.933, remote 0.898/0.728, couch 0.673 for the same input.
Offline override (no Hub, no HF cache): TT_RF_DETR_WEIGHTS=<dir holding model.safetensors + config.json>.
Offline checks that need no device (what the implement phase ran):
# manifest + launcher preview (from the repo dir — `root: code` is CWD-relative)
$ROOT/.venv/bin/python -c "from tt_kernel.container_manifest import load_container_manifest; m = load_container_manifest('tt-model.yaml', check_sources=True); p = m.resolve_profile(); print('VALID', m.name, m.kind, p.hardware, p.mesh_device, m.weights_ref); from tt_kernel.launchers import launcher_for; L = launcher_for(m.kind); print(L.install_lines(m)); print(L.verify_lines(m))"
# import the app with no device
PYTHONPATH=$PWD/code:$TT_METAL_HOME:$TT_METAL_HOME/ttnn:$TT_METAL_HOME/tools:/tmp/rfdetr-http $TT_METAL_HOME/python_env/bin/python -c "import rf_detr.server.app as a; assert a.app"
2. Package, serve, push (Docker)
Rootless Docker on this box needs source $ROOT/bin/docker-env.sh first (sets PATH and
DOCKER_HOST; the bare docker on PATH is podman). Run every tt-model command from
this directory: source.tt_metal and source.extra_code[].root: code resolve against
the process CWD, not against the YAML. --out points outside the git checkout because
stage() deletes <out>/rf-detr-p150 before rebuilding.
source $ROOT/bin/docker-env.sh
cd $ROOT/models/rf-detr-p150
tt-model package --container tt-model.yaml --out $ROOT/build # builds the image (hours cold), runs verify.sh
tt-model serve $ROOT/build/rf-detr-p150/tt_kernel_manifest.json # pre-downloads the pinned weights, boots, waits for READY
python code/rf_detr/server/smoke_test.py --url http://127.0.0.1:<port serve printed>
tt-model stop changh95/rf-detr-p150
tt-model push $ROOT/build/rf-detr-p150 --publish # uploads code/, image/, README (generated card)
serve publishes the first free port from 20000 and exports into the container exactly:
HF_MODEL=Roboflow/rf-detr-base, MESH_DEVICE=P150, TT_MESH_SHAPE=1x1, then serve.env
(TT_WEIGHTS_REVISION=7b95b0…, TT_METAL_VISIBLE_DEVICES=0). The HF cache is mounted at
/hf (rw), kernels at /cache, and /weight-cache is unused by this port (weights are
converted to bf16/bf8 device tensors at boot in seconds). tt-cli users:
tt serve changh95/rf-detr-p150 / tt model stop changh95/rf-detr-p150.
push makes code/ and image/ on the Hub exactly the staged trees (rf_detr,
scripts, conftest.py and the models/common/lightweightmodule.py filler) and replaces
README.md with the generated card (everything worth keeping lives in
card.description / card.quickstart). media/, SERVING.md, .gitattributes and
tt-model.yaml at the repo root survive; the orchestrator restores the license /
pipeline_tag front matter afterwards.
3. The request / response contract
| route | purpose |
|---|---|
GET /health |
{"status": "ok" | "starting", "model": "RF-DETR-base", "device": "blackhole p150 (device_id=0, mesh 1x1)" | null} — 200 always; ok only after warm-up |
GET /info |
model, task, hardware, weights repo + revision + resolved files, port source + commit, input limits (560x560, 1 image, batch 1), COCO labels, device params |
GET /v1/models |
{"object": "list", "data": [{"id": "Roboflow/rf-detr-base", "object": "model", "owned_by": "changh95"}]} — only so the OpenAI-shaped ready card / tt-model curl do not 404 |
POST /predict |
one image -> detections (below) |
POST /predict request (JSON):
| field | type | default | meaning |
|---|---|---|---|
image |
str | required | base64 PNG/JPEG, any size (max side 8192, max 32 MB); a data:image/...;base64, prefix and wrapped base64 are tolerated |
threshold |
float 0..1 | 0.5 |
keep detections with sigmoid(logit) > threshold |
max_detections |
int 1..300 | 100 |
keep at most this many (highest scores first) |
return_raw |
bool | false |
also return raw.logits [300][91] and raw.pred_boxes [300][4] (normalised cxcywh) |
Response (200):
{
"detections": [
{"label": "cat", "label_id": 17, "score": 0.96,
"box": [7.35, 54.64, 318.45, 472.14], "box_cxcywh_norm": [0.2545, 0.5487, 0.4861, 0.8698]}
],
"num_detections": 5,
"image_size": {"width": 640, "height": 480},
"input_size": [560, 560],
"threshold": 0.5,
"timing_ms": {"decode": 3.1, "preprocess": 4.0, "inference": 47.0, "postprocess": 0.4, "total": 54.5}
}
box is [x1, y1, x2, y2] in pixels of the ORIGINAL image (the preprocessor squashes the
whole image to 560x560, so the model's normalised cxcywh is relative to the original too;
clamped to the image bounds), sorted by score descending. label_id is the COCO-91 index
from config.json's id2label. Errors: 400 undecodable/oversized image, 422
schema violation (missing image, threshold out of range), 503 while starting,
500 with inference failed: <ExceptionType>: <text> if the device call raises.
Server-side pipeline (the port's validated recipe): PIL RGB -> bilinear+antialias resize
to 560x560 -> /255 -> ImageNet mean/std -> TtRfDetr(pixel_values) under one global
threading.Lock and torch.inference_mode() -> sigmoid(logits).max(-1) -> threshold ->
sort -> scale cxcywh by (W, H).
4. Caveats
- Batch 1, one image per request. The TT graph is shape-locked to 560x560 (16 windows x 101 tokens, 40x40 feature grid, 300 queries, 91 classes). Requests are serialised by a lock; concurrent clients queue.
- Warm-up captures the metal-trace of the projector+transformer tail inside the
lifespan (two forwards on a synthetic 560x560 image), so READY means warm; the first
cold boot pays the ttnn JIT (minutes), later boots reuse
/cache(TT_METAL_CACHE). - Weights are pinned by sha in both
weights.revisionandserve.env.TT_WEIGHTS_REVISION;hf_hub_download(..., revision=<sha>)is a pure cache hit afterserve's pre-download and falls back tolocal_files_only=Trueif the Hub is unreachable. - Precision knobs stay off.
RFDETR_BB_FIDELITY/RFDETR_BB_FP32ACC/RFDETR_BB_L1are read byTtRfDetrbut every setting other than the default regressed accuracy or hung (port README); do not put them inserve.env. - ttnn API drift is the main hardware-phase risk: the port was validated on a
~June-2026 tree; the packaged tree is 2026-08-20.
Conv2dConfig.shard_layout,ttnn.topkindex dtype,execute_tracekwargs andgrid_samplekwargs all exist in this tree (checked offline); numerics/gate (98.5 detection-IoU) must be re-confirmed withpytest code/rf_detr/tests/test_pretrained_eval.pyor the smoke test on the box. - The
tt_metal/python_envvenv on this host is Python 3.10 (the image uses 3.12); the code is version-agnostic (3.9+ syntax,from __future__ import annotations). tt-model curland the ready card's/v1/modelshint are OpenAI-shaped and are not this API; use the routes above.