rf-detr-p150

RF-DETR-base (Roboflow's real-time DETR, COCO-91) running entirely on one Tenstorrent Blackhole p150a via tt-nn: image in, labelled boxes out. Weights: Roboflow/rf-detr-base · Paper: arXiv:2511.09554 · Upstream code: roboflow/rf-detr · Port: changh95/tt-RF-DETR

Runs on p150 (mesh P150).

Packaged and published with tt-model-manager 0.1.0 (manifest schema 5.1).

Quickstart

tt-model pull  changh95/rf-detr-p150 --with-weights
tt-model serve changh95/rf-detr-p150
  • Weights Roboflow/rf-detr-base at 7b95b089788e go to your HF cache; the image does not contain them.
  • Serves on port 20000 (or the next free port); ready when the log says Application startup complete.

Run with tt-cli

tt serve changh95/rf-detr-p150
printf '{"image":"%s"}' "$(base64 -w0 media/demo_source.png)" > req.json
curl -s localhost:20000/predict -H 'Content-Type: application/json' -d @req.json
tt model stop changh95/rf-detr-p150
  • POST /predict: image (base64 PNG/JPEG); optional threshold (0.5), max_detections (100).
  • GET /health, GET /info.

Response

{"detections": [
   {"label": "cat",    "label_id": 17, "score": 0.960, "box": [7.5, 54.4, 317.5, 470.6]},
   {"label": "remote", "label_id": 75, "score": 0.903, "box": [40.9, 72.7, 176.6, 117.7]}],
 "num_detections": 5, "image_size": {"width": 640, "height": 480}, "input_size": [560, 560],
 "timing_ms": {"decode": 5.8, "preprocess": 0.9, "inference": 8.6, "postprocess": 0.2, "total": 15.4}}
  • box is [x1, y1, x2, y2] in original image pixels; detections are sorted by score.

Demo

Input (media/demo_source.png) Detections on p150a (media/demo_detections.png)

Accuracy and speed

Metric Value
Detection-IoU agreement vs fp32 reference (demo image) 98.70 (2 cat + 2 remote, scores within 0.008 of the reference)
Backbone feature-map PCC vs torch 0.999–0.9999
Inference, served over HTTP (warm, batch 1, 560×560) 8.6 ms device · ~15.5 ms server-side incl. PNG decode (120 FPS device-bound)
Same forward on an RTX 5090 (same host, port's torch reference, eager PyTorch, batch 1, incl. H2D/D2H; bf16 / fp16 autocast) 5.9 / 5.9 ms → GPU 1.4× faster than the p150a's 8.2 ms; fp32-strict GPU 8.0 ms (parity); best torch.compile 2.2 ms

Caveats

  • Every image is squashed to 560×560; one image per request, batch 1.
  • bf16 on device: scores differ slightly from the fp32 reference (see the IoU figure above).
  • Not an OpenAI-compatible API; GET /v1/models is a stub so the tt-model ready card does not 404.
  • Validated on tt-metal v0.78.0-dev20260820 (main 8b98410e730), single p150a only.
  • GPU comparison: fp32-strict eager GPU 8.0 ms ≈ p150a 8.2 ms; GPU bf16/fp16 1.4× faster, compiled 3.7×. RTX 5090 rows (2026-09-14): same host, the port's own torch reference (same weights) run eagerly in PyTorch 2.11 cu128 with fp32 weights + autocast unless stated, no TensorRT; medians of 50 iterations after warm-up, H2D/D2H included; the p150a rows are the served bf16 fused path incl. upload/readback. p150a power was not measured, so no efficiency comparison is made. Full table: GPU_COMPARISON.md.

Licensing

Provenance

The exact sources the image was built from — code/ in this repo is byte-identical to the model code inside the image:

component built from
tt-metal 8b98410e730bb504fea43a88609756e34821d91d
code/ digest 906212f3590bf302 (sha256, first 16 hex digits)
built 2026-09-12T20:08:45+00:00 by tt-model 0.1.0
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for changh95/rf-detr-p150

Finetuned
(8)
this model

Collection including changh95/rf-detr-p150

Paper for changh95/rf-detr-p150