rf-detr-p150 / tt-model.yaml
changh95's picture
tt-model.yaml: RTX 5090 comparison row + caveat
6718c9f verified
Raw History Blame Contribute Delete
6.18 kB
# SPDX-License-Identifier: Apache-2.0
# tt-model-manager container manifest (schema 5.1) for RF-DETR-base on Blackhole.
#
# Run EVERY tt-model command from this directory: `source.tt_metal` and
# `extra_code[].root: code` resolve against the process CWD, not this file.
#
# source $ROOT/bin/docker-env.sh # rootless Docker env on this box
# tt-model package --container tt-model.yaml --out $ROOT/build
# tt-model serve $ROOT/build/rf-detr-p150/tt_kernel_manifest.json
# tt-model push $ROOT/build/rf-detr-p150 --publish
schema: "5.1"
repo: changh95/rf-detr-p150
name: rf-detr-p150
# A POINTER, pinned: `serve` pre-downloads exactly these two files at this sha into
# the host HF cache; the app loads the same sha via TT_WEIGHTS_REVISION below.
weights:
repo: Roboflow/rf-detr-base
revision: 7b95b089788e6c7db56d5ea9b0a07ca08ea6ac0a
allow_patterns: ["model.safetensors", "config.json"]
kind: tt-dit-server
arch: blackhole
source:
# The tt-metal tree the image is built from (v0.78.0-dev20260820-25, main 8b98410e730);
# its torch pin (2.11.0) is read from tt_metal/python_env/requirements-dev.txt.
tt_metal: /home/deepgadget/experiments/gbp-tt/tt-metal
# The port imports nothing from tt-metal's models/ tree; the schema needs one entry.
code:
- models/common/lightweightmodule.py
# This repo's own code (lands at /opt/tt-metal/<path>; PYTHONPATH=/opt/tt-metal).
extra_code:
- root: code
paths:
- rf_detr
- scripts
- conftest.py
ubuntu: "22.04"
python: "3.12"
runtime:
app: rf_detr.server.app:app
mesh_shape_env: TT_MESH_SHAPE
# On top of the kind defaults (fastapi, uvicorn, pydantic>=2, pillow) and the
# auto-pinned torch==<tree pin>+cpu. ttnn needs numpy<2; say so explicitly.
packages:
- "numpy>=1.24.4,<2"
- safetensors
- huggingface_hub
serve:
port: 20000
hardware: p150
mesh_device: P150
env:
TT_WEIGHTS_REVISION: "7b95b089788e6c7db56d5ea9b0a07ca08ea6ac0a"
TT_METAL_VISIBLE_DEVICES: "0"
# Build-time assertions, run INSIDE the finished image as uid 1000, no device, no weights.
verify:
- "import rf_detr.server.app as a; assert a.app"
- "import safetensors, huggingface_hub, numpy, PIL; assert int(numpy.__version__.split('.')[0]) < 2, numpy.__version__"
- "from rf_detr.reference.weights import load_rf_detr_base, get_preprocessor, resolve_weight_files; from rf_detr.tt import TtRfDetr"
- "import ttnn; assert hasattr(ttnn, 'grid_sample') and hasattr(ttnn, 'topk') and hasattr(ttnn, 'copy_host_to_device_tensor') and hasattr(ttnn, 'release_trace')"
- "from rf_detr.server.app import parse_mesh_shape as p; assert p('1x1') == p('(1, 1)') == p('1,1') == (1, 1)"
card:
description: >
RF-DETR-base (Roboflow's real-time DETR, COCO-91) running entirely on one Tenstorrent Blackhole p150a via tt-nn: image in, labelled boxes out.
Weights: [Roboflow/rf-detr-base](https://huggingface.co/Roboflow/rf-detr-base) ·
Paper: [arXiv:2511.09554](https://arxiv.org/abs/2511.09554) ·
Upstream code: [roboflow/rf-detr](https://github.com/roboflow/rf-detr) ·
Port: [changh95/tt-RF-DETR](https://github.com/changh95/tt-RF-DETR)
quickstart: |
### Run with tt-cli
```bash
tt serve changh95/rf-detr-p150
printf '{"image":"%s"}' "$(base64 -w0 media/demo_source.png)" > req.json
curl -s localhost:20000/predict -H 'Content-Type: application/json' -d @req.json
tt model stop changh95/rf-detr-p150
```
- `POST /predict`: `image` (base64 PNG/JPEG); optional `threshold` (0.5), `max_detections` (100).
- `GET /health`, `GET /info`.
### Response
```json
{"detections": [
{"label": "cat", "label_id": 17, "score": 0.960, "box": [7.5, 54.4, 317.5, 470.6]},
{"label": "remote", "label_id": 75, "score": 0.903, "box": [40.9, 72.7, 176.6, 117.7]}],
"num_detections": 5, "image_size": {"width": 640, "height": 480}, "input_size": [560, 560],
"timing_ms": {"decode": 5.8, "preprocess": 0.9, "inference": 8.6, "postprocess": 0.2, "total": 15.4}}
```
- `box` is `[x1, y1, x2, y2]` in original image pixels; detections are sorted by score.
### Demo
| Input (`media/demo_source.png`) | Detections on p150a (`media/demo_detections.png`) |
|:---:|:---:|
| ![](media/demo_source.png) | ![](media/demo_detections.png) |
### Accuracy and speed
| Metric | Value |
|---|---:|
| Detection-IoU agreement vs fp32 reference (demo image) | 98.70 (2 cat + 2 remote, scores within 0.008 of the reference) |
| Backbone feature-map PCC vs torch | 0.999–0.9999 |
| Inference, served over HTTP (warm, batch 1, 560×560) | ~8.6 ms device · ~15.5 ms server-side incl. PNG decode (~120 FPS device-bound) |
| Same forward on an RTX 5090 (same host, port's torch reference, eager PyTorch, batch 1, incl. H2D/D2H; bf16 / fp16 autocast) | 5.9 / 5.9 ms → GPU 1.4× faster than the p150a's 8.2 ms; fp32-strict GPU 8.0 ms (parity); best `torch.compile` 2.2 ms |
### Caveats
- Every image is squashed to 560×560; one image per request, batch 1.
- bf16 on device: scores differ slightly from the fp32 reference (see the IoU figure above).
- Not an OpenAI-compatible API; `GET /v1/models` is a stub so the tt-model ready card does not 404.
- Validated on tt-metal `v0.78.0-dev20260820` (main `8b98410e730`), single p150a only.
- GPU comparison: fp32-strict eager GPU 8.0 ms ≈ p150a 8.2 ms; GPU bf16/fp16 1.4× faster, compiled 3.7×. RTX 5090 rows (2026-09-14): same host, the port's own torch reference (same weights) run eagerly in PyTorch 2.11 cu128 with fp32 weights + autocast unless stated, no TensorRT; medians of 50 iterations after warm-up, H2D/D2H included; the p150a rows are the served bf16 fused path incl. upload/readback. p150a power was not measured, so no efficiency comparison is made. Full table: [`GPU_COMPARISON.md`](GPU_COMPARISON.md).
### Licensing
- Weights: [Roboflow/rf-detr-base](https://huggingface.co/Roboflow/rf-detr-base), Apache-2.0.
- Port and serving code (`code/`): Apache-2.0, from [changh95/tt-RF-DETR](https://github.com/changh95/tt-RF-DETR).