Download tt-model.yaml from changh95/rf-detr-p150: direct link, hf CLI and curl.
- Browser
- Download file 6.18 kB
-
https://huggingface.co/changh95/rf-detr-p150/resolve/main/tt-model.yaml
- Command line
-
hf download hf://changh95/rf-detr-p150/tt-model.yaml
-
curl -L -o tt-model.yaml https://huggingface.co/changh95/rf-detr-p150/resolve/main/tt-model.yaml
6.18 kB
| # SPDX-License-Identifier: Apache-2.0 | |
| # tt-model-manager container manifest (schema 5.1) for RF-DETR-base on Blackhole. | |
| # | |
| # Run EVERY tt-model command from this directory: `source.tt_metal` and | |
| # `extra_code[].root: code` resolve against the process CWD, not this file. | |
| # | |
| # source $ROOT/bin/docker-env.sh # rootless Docker env on this box | |
| # tt-model package --container tt-model.yaml --out $ROOT/build | |
| # tt-model serve $ROOT/build/rf-detr-p150/tt_kernel_manifest.json | |
| # tt-model push $ROOT/build/rf-detr-p150 --publish | |
| schema: "5.1" | |
| repo: changh95/rf-detr-p150 | |
| name: rf-detr-p150 | |
| # A POINTER, pinned: `serve` pre-downloads exactly these two files at this sha into | |
| # the host HF cache; the app loads the same sha via TT_WEIGHTS_REVISION below. | |
| weights: | |
| repo: Roboflow/rf-detr-base | |
| revision: 7b95b089788e6c7db56d5ea9b0a07ca08ea6ac0a | |
| allow_patterns: ["model.safetensors", "config.json"] | |
| kind: tt-dit-server | |
| arch: blackhole | |
| source: | |
| # The tt-metal tree the image is built from (v0.78.0-dev20260820-25, main 8b98410e730); | |
| # its torch pin (2.11.0) is read from tt_metal/python_env/requirements-dev.txt. | |
| tt_metal: /home/deepgadget/experiments/gbp-tt/tt-metal | |
| # The port imports nothing from tt-metal's models/ tree; the schema needs one entry. | |
| code: | |
| - models/common/lightweightmodule.py | |
| # This repo's own code (lands at /opt/tt-metal/<path>; PYTHONPATH=/opt/tt-metal). | |
| extra_code: | |
| - root: code | |
| paths: | |
| - rf_detr | |
| - scripts | |
| - conftest.py | |
| ubuntu: "22.04" | |
| python: "3.12" | |
| runtime: | |
| app: rf_detr.server.app:app | |
| mesh_shape_env: TT_MESH_SHAPE | |
| # On top of the kind defaults (fastapi, uvicorn, pydantic>=2, pillow) and the | |
| # auto-pinned torch==<tree pin>+cpu. ttnn needs numpy<2; say so explicitly. | |
| packages: | |
| - "numpy>=1.24.4,<2" | |
| - safetensors | |
| - huggingface_hub | |
| serve: | |
| port: 20000 | |
| hardware: p150 | |
| mesh_device: P150 | |
| env: | |
| TT_WEIGHTS_REVISION: "7b95b089788e6c7db56d5ea9b0a07ca08ea6ac0a" | |
| TT_METAL_VISIBLE_DEVICES: "0" | |
| # Build-time assertions, run INSIDE the finished image as uid 1000, no device, no weights. | |
| verify: | |
| - "import rf_detr.server.app as a; assert a.app" | |
| - "import safetensors, huggingface_hub, numpy, PIL; assert int(numpy.__version__.split('.')[0]) < 2, numpy.__version__" | |
| - "from rf_detr.reference.weights import load_rf_detr_base, get_preprocessor, resolve_weight_files; from rf_detr.tt import TtRfDetr" | |
| - "import ttnn; assert hasattr(ttnn, 'grid_sample') and hasattr(ttnn, 'topk') and hasattr(ttnn, 'copy_host_to_device_tensor') and hasattr(ttnn, 'release_trace')" | |
| - "from rf_detr.server.app import parse_mesh_shape as p; assert p('1x1') == p('(1, 1)') == p('1,1') == (1, 1)" | |
| card: | |
| description: > | |
| RF-DETR-base (Roboflow's real-time DETR, COCO-91) running entirely on one Tenstorrent Blackhole p150a via tt-nn: image in, labelled boxes out. | |
| Weights: [Roboflow/rf-detr-base](https://huggingface.co/Roboflow/rf-detr-base) · | |
| Paper: [arXiv:2511.09554](https://arxiv.org/abs/2511.09554) · | |
| Upstream code: [roboflow/rf-detr](https://github.com/roboflow/rf-detr) · | |
| Port: [changh95/tt-RF-DETR](https://github.com/changh95/tt-RF-DETR) | |
| quickstart: | | |
| ### Run with tt-cli | |
| ```bash | |
| tt serve changh95/rf-detr-p150 | |
| printf '{"image":"%s"}' "$(base64 -w0 media/demo_source.png)" > req.json | |
| curl -s localhost:20000/predict -H 'Content-Type: application/json' -d @req.json | |
| tt model stop changh95/rf-detr-p150 | |
| ``` | |
| - `POST /predict`: `image` (base64 PNG/JPEG); optional `threshold` (0.5), `max_detections` (100). | |
| - `GET /health`, `GET /info`. | |
| ### Response | |
| ```json | |
| {"detections": [ | |
| {"label": "cat", "label_id": 17, "score": 0.960, "box": [7.5, 54.4, 317.5, 470.6]}, | |
| {"label": "remote", "label_id": 75, "score": 0.903, "box": [40.9, 72.7, 176.6, 117.7]}], | |
| "num_detections": 5, "image_size": {"width": 640, "height": 480}, "input_size": [560, 560], | |
| "timing_ms": {"decode": 5.8, "preprocess": 0.9, "inference": 8.6, "postprocess": 0.2, "total": 15.4}} | |
| ``` | |
| - `box` is `[x1, y1, x2, y2]` in original image pixels; detections are sorted by score. | |
| ### Demo | |
| | Input (`media/demo_source.png`) | Detections on p150a (`media/demo_detections.png`) | | |
| |:---:|:---:| | |
| |  |  | | |
| ### Accuracy and speed | |
| | Metric | Value | | |
| |---|---:| | |
| | Detection-IoU agreement vs fp32 reference (demo image) | 98.70 (2 cat + 2 remote, scores within 0.008 of the reference) | | |
| | Backbone feature-map PCC vs torch | 0.999–0.9999 | | |
| | Inference, served over HTTP (warm, batch 1, 560×560) | ~8.6 ms device · ~15.5 ms server-side incl. PNG decode (~120 FPS device-bound) | | |
| | Same forward on an RTX 5090 (same host, port's torch reference, eager PyTorch, batch 1, incl. H2D/D2H; bf16 / fp16 autocast) | 5.9 / 5.9 ms → GPU 1.4× faster than the p150a's 8.2 ms; fp32-strict GPU 8.0 ms (parity); best `torch.compile` 2.2 ms | | |
| ### Caveats | |
| - Every image is squashed to 560×560; one image per request, batch 1. | |
| - bf16 on device: scores differ slightly from the fp32 reference (see the IoU figure above). | |
| - Not an OpenAI-compatible API; `GET /v1/models` is a stub so the tt-model ready card does not 404. | |
| - Validated on tt-metal `v0.78.0-dev20260820` (main `8b98410e730`), single p150a only. | |
| - GPU comparison: fp32-strict eager GPU 8.0 ms ≈ p150a 8.2 ms; GPU bf16/fp16 1.4× faster, compiled 3.7×. RTX 5090 rows (2026-09-14): same host, the port's own torch reference (same weights) run eagerly in PyTorch 2.11 cu128 with fp32 weights + autocast unless stated, no TensorRT; medians of 50 iterations after warm-up, H2D/D2H included; the p150a rows are the served bf16 fused path incl. upload/readback. p150a power was not measured, so no efficiency comparison is made. Full table: [`GPU_COMPARISON.md`](GPU_COMPARISON.md). | |
| ### Licensing | |
| - Weights: [Roboflow/rf-detr-base](https://huggingface.co/Roboflow/rf-detr-base), Apache-2.0. | |
| - Port and serving code (`code/`): Apache-2.0, from [changh95/tt-RF-DETR](https://github.com/changh95/tt-RF-DETR). | |