# SPDX-License-Identifier: Apache-2.0 # tt-model-manager container manifest (schema 5.1) for RF-DETR-base on Blackhole. # # Run EVERY tt-model command from this directory: `source.tt_metal` and # `extra_code[].root: code` resolve against the process CWD, not this file. # # source $ROOT/bin/docker-env.sh # rootless Docker env on this box # tt-model package --container tt-model.yaml --out $ROOT/build # tt-model serve $ROOT/build/rf-detr-p150/tt_kernel_manifest.json # tt-model push $ROOT/build/rf-detr-p150 --publish schema: "5.1" repo: changh95/rf-detr-p150 name: rf-detr-p150 # A POINTER, pinned: `serve` pre-downloads exactly these two files at this sha into # the host HF cache; the app loads the same sha via TT_WEIGHTS_REVISION below. weights: repo: Roboflow/rf-detr-base revision: 7b95b089788e6c7db56d5ea9b0a07ca08ea6ac0a allow_patterns: ["model.safetensors", "config.json"] kind: tt-dit-server arch: blackhole source: # The tt-metal tree the image is built from (v0.78.0-dev20260820-25, main 8b98410e730); # its torch pin (2.11.0) is read from tt_metal/python_env/requirements-dev.txt. tt_metal: /home/deepgadget/experiments/gbp-tt/tt-metal # The port imports nothing from tt-metal's models/ tree; the schema needs one entry. code: - models/common/lightweightmodule.py # This repo's own code (lands at /opt/tt-metal/; PYTHONPATH=/opt/tt-metal). extra_code: - root: code paths: - rf_detr - scripts - conftest.py ubuntu: "22.04" python: "3.12" runtime: app: rf_detr.server.app:app mesh_shape_env: TT_MESH_SHAPE # On top of the kind defaults (fastapi, uvicorn, pydantic>=2, pillow) and the # auto-pinned torch==+cpu. ttnn needs numpy<2; say so explicitly. packages: - "numpy>=1.24.4,<2" - safetensors - huggingface_hub serve: port: 20000 hardware: p150 mesh_device: P150 env: TT_WEIGHTS_REVISION: "7b95b089788e6c7db56d5ea9b0a07ca08ea6ac0a" TT_METAL_VISIBLE_DEVICES: "0" # Build-time assertions, run INSIDE the finished image as uid 1000, no device, no weights. verify: - "import rf_detr.server.app as a; assert a.app" - "import safetensors, huggingface_hub, numpy, PIL; assert int(numpy.__version__.split('.')[0]) < 2, numpy.__version__" - "from rf_detr.reference.weights import load_rf_detr_base, get_preprocessor, resolve_weight_files; from rf_detr.tt import TtRfDetr" - "import ttnn; assert hasattr(ttnn, 'grid_sample') and hasattr(ttnn, 'topk') and hasattr(ttnn, 'copy_host_to_device_tensor') and hasattr(ttnn, 'release_trace')" - "from rf_detr.server.app import parse_mesh_shape as p; assert p('1x1') == p('(1, 1)') == p('1,1') == (1, 1)" card: description: > RF-DETR-base (Roboflow's real-time DETR, COCO-91) running entirely on one Tenstorrent Blackhole p150a via tt-nn: image in, labelled boxes out. Weights: [Roboflow/rf-detr-base](https://huggingface.co/Roboflow/rf-detr-base) · Paper: [arXiv:2511.09554](https://arxiv.org/abs/2511.09554) · Upstream code: [roboflow/rf-detr](https://github.com/roboflow/rf-detr) · Port: [changh95/tt-RF-DETR](https://github.com/changh95/tt-RF-DETR) quickstart: | ### Run with tt-cli ```bash tt serve changh95/rf-detr-p150 printf '{"image":"%s"}' "$(base64 -w0 media/demo_source.png)" > req.json curl -s localhost:20000/predict -H 'Content-Type: application/json' -d @req.json tt model stop changh95/rf-detr-p150 ``` - `POST /predict`: `image` (base64 PNG/JPEG); optional `threshold` (0.5), `max_detections` (100). - `GET /health`, `GET /info`. ### Response ```json {"detections": [ {"label": "cat", "label_id": 17, "score": 0.960, "box": [7.5, 54.4, 317.5, 470.6]}, {"label": "remote", "label_id": 75, "score": 0.903, "box": [40.9, 72.7, 176.6, 117.7]}], "num_detections": 5, "image_size": {"width": 640, "height": 480}, "input_size": [560, 560], "timing_ms": {"decode": 5.8, "preprocess": 0.9, "inference": 8.6, "postprocess": 0.2, "total": 15.4}} ``` - `box` is `[x1, y1, x2, y2]` in original image pixels; detections are sorted by score. ### Demo | Input (`media/demo_source.png`) | Detections on p150a (`media/demo_detections.png`) | |:---:|:---:| | ![](media/demo_source.png) | ![](media/demo_detections.png) | ### Accuracy and speed | Metric | Value | |---|---:| | Detection-IoU agreement vs fp32 reference (demo image) | 98.70 (2 cat + 2 remote, scores within 0.008 of the reference) | | Backbone feature-map PCC vs torch | 0.999–0.9999 | | Inference, served over HTTP (warm, batch 1, 560×560) | ~8.6 ms device · ~15.5 ms server-side incl. PNG decode (~120 FPS device-bound) | | Same forward on an RTX 5090 (same host, port's torch reference, eager PyTorch, batch 1, incl. H2D/D2H; bf16 / fp16 autocast) | 5.9 / 5.9 ms → GPU 1.4× faster than the p150a's 8.2 ms; fp32-strict GPU 8.0 ms (parity); best `torch.compile` 2.2 ms | ### Caveats - Every image is squashed to 560×560; one image per request, batch 1. - bf16 on device: scores differ slightly from the fp32 reference (see the IoU figure above). - Not an OpenAI-compatible API; `GET /v1/models` is a stub so the tt-model ready card does not 404. - Validated on tt-metal `v0.78.0-dev20260820` (main `8b98410e730`), single p150a only. - GPU comparison: fp32-strict eager GPU 8.0 ms ≈ p150a 8.2 ms; GPU bf16/fp16 1.4× faster, compiled 3.7×. RTX 5090 rows (2026-09-14): same host, the port's own torch reference (same weights) run eagerly in PyTorch 2.11 cu128 with fp32 weights + autocast unless stated, no TensorRT; medians of 50 iterations after warm-up, H2D/D2H included; the p150a rows are the served bf16 fused path incl. upload/readback. p150a power was not measured, so no efficiency comparison is made. Full table: [`GPU_COMPARISON.md`](GPU_COMPARISON.md). ### Licensing - Weights: [Roboflow/rf-detr-base](https://huggingface.co/Roboflow/rf-detr-base), Apache-2.0. - Port and serving code (`code/`): Apache-2.0, from [changh95/tt-RF-DETR](https://github.com/changh95/tt-RF-DETR).