|
Download SERVING.md from changh95/rf-detr-p150: direct link, hf CLI and curl.
- Browser
- Download file 10.4 kB
-
https://huggingface.co/changh95/rf-detr-p150/resolve/main/SERVING.md
- Command line
-
hf download hf://changh95/rf-detr-p150/SERVING.md
-
curl -L -o SERVING.md https://huggingface.co/changh95/rf-detr-p150/resolve/main/SERVING.md
10.4 kB
| # Serving RF-DETR-base on Blackhole with tt-model-manager | |
| This repo is a **tt-model container package source**: `tt-model.yaml` + `code/rf_detr` | |
| (the tt-nn port, with the FastAPI server under `code/rf_detr/server/`). It packages into | |
| an OCI image that `tt-model serve` / `tt serve changh95/rf-detr-p150` run on a | |
| Blackhole p150a host. Weights (`Roboflow/rf-detr-base`, Apache-2.0, ~129 MB) are a pinned | |
| pointer, never baked into the image. | |
| | | | | |
| |---|---| | |
| | tt-metal tree | `/home/deepgadget/experiments/gbp-tt/tt-metal` (main `8b98410e730`, v0.78.0-dev20260820-25; torch 2.11.0+cpu) | | |
| | weights | `Roboflow/rf-detr-base` @ `7b95b089788e6c7db56d5ea9b0a07ca08ea6ac0a` (`model.safetensors`, `config.json`) | | |
| | app | `rf_detr.server.app:app` (uvicorn, ASGI lifespan does weights -> device -> model -> warm-up) | | |
| | device recipe | `ttnn.open_device(device_id, l1_small_size=32768, trace_region_size=90_000_000, num_command_queues=1)` — the port's validated params (`conftest.py`, `scripts/make_demo.py`, `benchmark.py`) | | |
| | port source | github.com/changh95/tt-RF-DETR @ `5b0149869465892477141b4c1dfc4622db0d1ec5` (+ `rf_detr/server/` added here) | | |
| Below, `$ROOT=/home/deepgadget/experiments/tt-models`. | |
| ## 1. Run on the HOST (hardware validation, no Docker) | |
| The tree venv `$TT_METAL_HOME/python_env` (py3.10, torch 2.11.0+cpu, ttnn editable) has | |
| everything except `fastapi`/`uvicorn`. Do not install into the tree venv; put the HTTP | |
| stack in a side directory and prepend it to `PYTHONPATH`: | |
| ```bash | |
| export TT_METAL_HOME=/home/deepgadget/experiments/gbp-tt/tt-metal | |
| export PATH=$HOME/.local/bin:$PATH # uv | |
| mkdir -p /tmp/rfdetr-http && uv pip install --python $TT_METAL_HOME/python_env/bin/python \ | |
| --target /tmp/rfdetr-http fastapi uvicorn | |
| cd $ROOT/models/rf-detr-p150 | |
| export PYTHONPATH=$PWD/code:$TT_METAL_HOME:$TT_METAL_HOME/ttnn:$TT_METAL_HOME/tools:/tmp/rfdetr-http | |
| export HF_MODEL=Roboflow/rf-detr-base | |
| export TT_WEIGHTS_REVISION=7b95b089788e6c7db56d5ea9b0a07ca08ea6ac0a # already in ~/.cache/huggingface | |
| export TT_MESH_SHAPE=1x1 | |
| export TT_DEVICE_ID=0 # optional (default 0) | |
| export TT_METAL_VISIBLE_DEVICES=0 # optional, one-card box | |
| $TT_METAL_HOME/python_env/bin/python -m uvicorn --host 0.0.0.0 --port 20000 --lifespan on rf_detr.server.app:app | |
| ``` | |
| Boot log landmarks: `Loading weights` -> `Opening device` -> `TT model built` -> | |
| `Warming up: capturing trace` -> `Warmup complete` -> uvicorn `Application startup complete`. | |
| The first boot compiles kernels (minutes on a cold `~/.cache/ttnn` / `TT_METAL_CACHE`); | |
| later boots take seconds. Startup failures raise and uvicorn exits non-zero (no CPU | |
| fallback). Stop with Ctrl-C / SIGTERM: the lifespan releases the trace, drops the device | |
| tensors and calls `ttnn.close_device`. | |
| Then, from a second shell (stdlib only, any Python): | |
| ```bash | |
| python code/rf_detr/server/smoke_test.py --url http://127.0.0.1:20000 # PASS/FAIL one-liner | |
| python code/rf_detr/server/smoke_test.py --url http://127.0.0.1:20000 --wait 1800 --out /tmp/rfdetr.json | |
| ``` | |
| Expected on the demo image (`media/demo_source.png`, 640x480): `cat` 0.96 / 0.94, | |
| `remote` 0.90 / 0.73, `couch` 0.67 — `PASS n=5 counts={"cat": 2, "couch": 1, "remote": 2} ...`. | |
| The CPU reference gives cat 0.960/0.933, remote 0.898/0.728, couch 0.673 for the same input. | |
| Offline override (no Hub, no HF cache): `TT_RF_DETR_WEIGHTS=<dir holding model.safetensors + config.json>`. | |
| Offline checks that need no device (what the implement phase ran): | |
| ```bash | |
| # manifest + launcher preview (from the repo dir — `root: code` is CWD-relative) | |
| $ROOT/.venv/bin/python -c "from tt_kernel.container_manifest import load_container_manifest; m = load_container_manifest('tt-model.yaml', check_sources=True); p = m.resolve_profile(); print('VALID', m.name, m.kind, p.hardware, p.mesh_device, m.weights_ref); from tt_kernel.launchers import launcher_for; L = launcher_for(m.kind); print(L.install_lines(m)); print(L.verify_lines(m))" | |
| # import the app with no device | |
| PYTHONPATH=$PWD/code:$TT_METAL_HOME:$TT_METAL_HOME/ttnn:$TT_METAL_HOME/tools:/tmp/rfdetr-http $TT_METAL_HOME/python_env/bin/python -c "import rf_detr.server.app as a; assert a.app" | |
| ``` | |
| ## 2. Package, serve, push (Docker) | |
| Rootless Docker on this box needs `source $ROOT/bin/docker-env.sh` first (sets PATH and | |
| `DOCKER_HOST`; the bare `docker` on PATH is podman). **Run every `tt-model` command from | |
| this directory**: `source.tt_metal` and `source.extra_code[].root: code` resolve against | |
| the process CWD, not against the YAML. `--out` points outside the git checkout because | |
| `stage()` deletes `<out>/rf-detr-p150` before rebuilding. | |
| ```bash | |
| source $ROOT/bin/docker-env.sh | |
| cd $ROOT/models/rf-detr-p150 | |
| tt-model package --container tt-model.yaml --out $ROOT/build # builds the image (hours cold), runs verify.sh | |
| tt-model serve $ROOT/build/rf-detr-p150/tt_kernel_manifest.json # pre-downloads the pinned weights, boots, waits for READY | |
| python code/rf_detr/server/smoke_test.py --url http://127.0.0.1:<port serve printed> | |
| tt-model stop changh95/rf-detr-p150 | |
| tt-model push $ROOT/build/rf-detr-p150 --publish # uploads code/, image/, README (generated card) | |
| ``` | |
| `serve` publishes the first free port from 20000 and exports into the container exactly: | |
| `HF_MODEL=Roboflow/rf-detr-base`, `MESH_DEVICE=P150`, `TT_MESH_SHAPE=1x1`, then `serve.env` | |
| (`TT_WEIGHTS_REVISION=7b95b0…`, `TT_METAL_VISIBLE_DEVICES=0`). The HF cache is mounted at | |
| `/hf` (rw), kernels at `/cache`, and `/weight-cache` is unused by this port (weights are | |
| converted to bf16/bf8 device tensors at boot in seconds). tt-cli users: | |
| `tt serve changh95/rf-detr-p150` / `tt model stop changh95/rf-detr-p150`. | |
| `push` makes `code/` and `image/` on the Hub exactly the staged trees (`rf_detr`, | |
| `scripts`, `conftest.py` and the `models/common/lightweightmodule.py` filler) and replaces | |
| `README.md` with the generated card (everything worth keeping lives in | |
| `card.description` / `card.quickstart`). `media/`, `SERVING.md`, `.gitattributes` and | |
| `tt-model.yaml` at the repo root survive; the orchestrator restores the `license` / | |
| `pipeline_tag` front matter afterwards. | |
| ## 3. The request / response contract | |
| | route | purpose | | |
| |---|---| | |
| | `GET /health` | `{"status": "ok" \| "starting", "model": "RF-DETR-base", "device": "blackhole p150 (device_id=0, mesh 1x1)" \| null}` — 200 always; `ok` only after warm-up | | |
| | `GET /info` | model, task, hardware, weights repo + revision + resolved files, port source + commit, input limits (560x560, 1 image, batch 1), COCO labels, device params | | |
| | `GET /v1/models` | `{"object": "list", "data": [{"id": "Roboflow/rf-detr-base", "object": "model", "owned_by": "changh95"}]}` — only so the OpenAI-shaped ready card / `tt-model curl` do not 404 | | |
| | `POST /predict` | one image -> detections (below) | | |
| `POST /predict` request (JSON): | |
| | field | type | default | meaning | | |
| |---|---|---|---| | |
| | `image` | str | required | base64 PNG/JPEG, any size (max side 8192, max 32 MB); a `data:image/...;base64,` prefix and wrapped base64 are tolerated | | |
| | `threshold` | float 0..1 | `0.5` | keep detections with `sigmoid(logit) > threshold` | | |
| | `max_detections` | int 1..300 | `100` | keep at most this many (highest scores first) | | |
| | `return_raw` | bool | `false` | also return `raw.logits` `[300][91]` and `raw.pred_boxes` `[300][4]` (normalised cxcywh) | | |
| Response (200): | |
| ```json | |
| { | |
| "detections": [ | |
| {"label": "cat", "label_id": 17, "score": 0.96, | |
| "box": [7.35, 54.64, 318.45, 472.14], "box_cxcywh_norm": [0.2545, 0.5487, 0.4861, 0.8698]} | |
| ], | |
| "num_detections": 5, | |
| "image_size": {"width": 640, "height": 480}, | |
| "input_size": [560, 560], | |
| "threshold": 0.5, | |
| "timing_ms": {"decode": 3.1, "preprocess": 4.0, "inference": 47.0, "postprocess": 0.4, "total": 54.5} | |
| } | |
| ``` | |
| `box` is `[x1, y1, x2, y2]` in pixels of the ORIGINAL image (the preprocessor squashes the | |
| whole image to 560x560, so the model's normalised cxcywh is relative to the original too; | |
| clamped to the image bounds), sorted by score descending. `label_id` is the COCO-91 index | |
| from `config.json`'s `id2label`. Errors: **400** undecodable/oversized image, **422** | |
| schema violation (missing `image`, `threshold` out of range), **503** while starting, | |
| **500** with `inference failed: <ExceptionType>: <text>` if the device call raises. | |
| Server-side pipeline (the port's validated recipe): PIL RGB -> bilinear+antialias resize | |
| to 560x560 -> /255 -> ImageNet mean/std -> `TtRfDetr(pixel_values)` under one global | |
| `threading.Lock` and `torch.inference_mode()` -> `sigmoid(logits).max(-1)` -> threshold -> | |
| sort -> scale cxcywh by (W, H). | |
| ## 4. Caveats | |
| - **Batch 1, one image per request.** The TT graph is shape-locked to 560x560 | |
| (16 windows x 101 tokens, 40x40 feature grid, 300 queries, 91 classes). Requests are | |
| serialised by a lock; concurrent clients queue. | |
| - **Warm-up captures the metal-trace** of the projector+transformer tail inside the | |
| lifespan (two forwards on a synthetic 560x560 image), so READY means warm; the first | |
| cold boot pays the ttnn JIT (minutes), later boots reuse `/cache` (`TT_METAL_CACHE`). | |
| - **Weights are pinned by sha** in both `weights.revision` and `serve.env.TT_WEIGHTS_REVISION`; | |
| `hf_hub_download(..., revision=<sha>)` is a pure cache hit after `serve`'s pre-download | |
| and falls back to `local_files_only=True` if the Hub is unreachable. | |
| - **Precision knobs stay off.** `RFDETR_BB_FIDELITY` / `RFDETR_BB_FP32ACC` / `RFDETR_BB_L1` | |
| are read by `TtRfDetr` but every setting other than the default regressed accuracy or hung | |
| (port README); do not put them in `serve.env`. | |
| - **ttnn API drift** is the main hardware-phase risk: the port was validated on a | |
| ~June-2026 tree; the packaged tree is 2026-08-20. `Conv2dConfig.shard_layout`, | |
| `ttnn.topk` index dtype, `execute_trace` kwargs and `grid_sample` kwargs all exist in | |
| this tree (checked offline); numerics/gate (98.5 detection-IoU) must be re-confirmed | |
| with `pytest code/rf_detr/tests/test_pretrained_eval.py` or the smoke test on the box. | |
| - The `tt_metal/python_env` venv on this host is Python 3.10 (the image uses 3.12); the | |
| code is version-agnostic (3.9+ syntax, `from __future__ import annotations`). | |
| - `tt-model curl` and the ready card's `/v1/models` hint are OpenAI-shaped and are not this | |
| API; use the routes above. | |