Download README.md from changh95/locate-anything-3b-p150: direct link, hf CLI and curl.
- Browser
- Download file 14.4 kB
-
https://huggingface.co/changh95/locate-anything-3b-p150/resolve/main/README.md
- Command line
-
hf download hf://changh95/locate-anything-3b-p150/README.md
-
curl -L -o README.md https://huggingface.co/changh95/locate-anything-3b-p150/resolve/main/README.md
tags:
- blackhole
- p150
- tt-dit-server
- tt-model-cache
- tt-model-container
- tenstorrent
- ttnn
- tt-metal
- tt-nn
- visual-grounding
- open-vocabulary
- vlm
- qwen2.5
- tt-model-catalog
license: other
base_model:
- nvidia/LocateAnything-3B
license_link: https://huggingface.co/nvidia/LocateAnything-3B/blob/main/LICENSE
pipeline_tag: object-detection
license_name: nvidia-license
locate-anything-3b-p150
NVIDIA LocateAnything-3B port on one Tenstorrent Blackhole p150a. Weights: nvidia/LocateAnything-3B · Paper: arXiv:2605.27365 · Upstream code: NVlabs/Eagle (Embodied) · Port: changh95/tt-locate-anything
Runs on p150 (mesh P150).
Device configuration: dispatch on ETH cores, 1 command queue, 12x10 compute grid. This is the p150 target and the default of the Python API and the HTTP server (LA_DISPATCH=auto). It needs tt-metal with patches/tt-metal-eth-dispatch.patch. All performance numbers in this card were measured in this configuration, on Blackhole chips of a Galaxy host, except the container-image response in Serving (HTTP).
Packaged and published with tt-model-manager 0.1.0 (manifest schema 5.1).
Quickstart (Python)
Prerequisite: a tt-metal / ttnn environment at tt-metal 8b98410e730 with patches/tt-metal-eth-dispatch.patch applied. ttnn is not on PyPI.
hf download changh95/locate-anything-3b-p150 --exclude "image/*" --local-dir locate-anything-3b-p150 && cd locate-anything-3b-p150
pip install -e code/ # adds numpy<2, pillow, huggingface_hub, safetensors, loguru, tqdm, pyyaml
pip install -e "code/[server,test]" # also adds the HTTP server (fastapi, uvicorn) and the tests (pytest, httpx)
from locate_anything import LocateAnything
with LocateAnything.from_pretrained(device_id=0) as model: # weights from the HF cache, opens the chip, warms up
result = model("media/demo_input.png", "car") # path, PIL image, numpy or torch array
for label, box in zip(result.labels, result.boxes):
print(label, box.round(1)) # car [ 541.4 447.1 1165.4 860.8] (x1 y1 x2 y2, pixels)
result.draw().save("out.png")
- The first
from_pretraineddownloads thenvidia/LocateAnything-3Bweights (about 7.7 GB) to your HF cache. It also converts the weights to BFP8. This takes some minutes. - With warm caches,
from_pretrainedtakes 20-35 s. Each later call takes about 115 ms on the demo image. from_pretrainedwarms up the model before it returns. The first call is as fast as the later calls: ~115 ms on the demo image. The default warm-up prepares prompts up to 1024 tokens. To prepare prompts up to 2048 tokens, usefrom_pretrained(..., warmup_variants="all").- The path
media/demo_input.pngis relative. Run the snippet from the root of this repo, or give an absolute path. - The
withblock releases the traces and closes the chip at the end. Withoutwith, callmodel.close().
Input (image) |
A file path, PNG / JPEG bytes, a PIL.Image, a numpy array or a torch tensor (HxWx3 RGB uint8, 3xHxW, grayscale, or float in 0..1). Any size. |
Input (query) |
Free text ("the red mug on the left") or a list of categories. The list ["person", "car"] becomes "person</c>car". |
| Options | model(image, query, max_new_tokens=128) (1..1024). from_pretrained(device_id=0, device=None, dispatch="auto", weights_dir=None, cache_dir=..., precision="bfp8w"). |
Output (LocateResult) |
boxes float64 [N, 4] (x1, y1, x2, y2 in original image pixels), labels, boxes_norm (0..1000 model output), points, raw_text, stopped_on_eos, timing_ms. |
| Methods | result.draw() gives a PIL.Image with the boxes. result.to_dict() gives the same JSON body as the HTTP /predict. for det in result: gives each Detection. |
- Throughput:
model.predict_batch(images, queries)runs a list of images on the chip, one at a time. A host thread decodes the next image while the chip runs the current image. model.detect(image, ["person", "car"])is the upstreamdetectname.LocateAnything.parse_boxesandparse_pointsgive the upstream output format.- The model has state. A lock serializes the calls. Make all calls from the thread that built the model.
model(image, query)runs the same code as the HTTPPOST /predict. It gives the same tokens and boxes.- The model squashes each image to 616×336. Images from OpenCV are BGR, so give
img[..., ::-1]. - Full reference (all options, input forms, fields, tests):
code/PYTHON.md. Runnable example:examples/quickstart.py.
Serving (HTTP)
tt-model pull changh95/locate-anything-3b-p150 --with-weights
tt-model serve changh95/locate-anything-3b-p150 # or with tt-cli: tt serve changh95/locate-anything-3b-p150
{ printf '{"query":"car","image":"'; base64 -w0 media/demo_input.png; printf '"}'; } > req.json
curl -s localhost:20000/predict -H 'Content-Type: application/json' -d @req.json
tt model stop changh95/locate-anything-3b-p150
- Weights
nvidia/LocateAnything-3Bgo to your HF cache; the image does not contain them. - Serves on port 20000 (or the next free port); ready when the log says
Application startup complete. POST /predict:image(base64 PNG/JPEG, one image),query(free text, 1-1000 chars; join categories with</c>, e.g.person</c>car); optionalmax_new_tokens(128, cap 1024),return_overlay(false). AlsoGET /health,GET /info.
{"query": "car", "width": 1920, "height": 1080, "canonical_size": [616, 336], "grid_hw": [24, 44],
"raw_text": "<ref>car</ref><box><282><414><606><794></box><|im_end|>",
"detections": [{"label": "car", "box": [541.44, 447.12, 1163.52, 857.52], "box_norm": [282, 414, 606, 794]}],
"points": [], "num_generated_tokens": 10, "stopped_on_eos": true, "decode_mode": "ar_greedy_trace_device_sampling",
"timing_ms": {"vision": 1.4, "prefill": 86.3, "decode": 202.1, "total": 303.9, "decode_tok_s": 44.53}}
boxis[x1, y1, x2, y2]in original image pixels;box_normis the model's own 0..1000 output over the squashedcanonical_sizeview.return_overlay: trueaddsoverlay_png_b64.- This response comes from the container image (2026-09-13, real p150a). The image opens the chip with the stock Tensix-column dispatch, which gives an 11x10 grid on a p150a. It is not the target configuration.
- The current
code/serves with ETH dispatch, 1 command queue and a 12x10 grid. For the same request it gives<282><414><607><797>andtiming_ms{"vision": 0.4, "prefill": 0.3, "decode": 123.6, "total": 127.1, "decode_tok_s": 72.82}(chip 7, median of 100 requests, 2026-10-05; wall ~0.19 s). On chips where ETH dispatch is as fast as Tensix dispatch,pipeline.run()takes 115 ms (see Demo & Performances and Provenance).
Demo
Demo & Performances
Warm, batch 1, demo image + car, 10 tokens. ETH dispatch, 1 command queue, 12x10 compute grid for all rows.
The speed of ETH dispatch changes between the chips of the Galaxy host. On chips 9, 12, 14 and 16, ETH dispatch gave about 115 ms (column 2). On chips 9 and 16, Tensix dispatch gave the same speed. On chip 7 (2026-10-05 re-check, independently verified), ETH dispatch was about 12 ms slower than Tensix dispatch on the same chip (115.1 ms), all in the traced programs (column 3). The outputs were bit-identical. We expect column 2 to be the better p150 estimate, because a p150 has all its ETH cores free for dispatch. Column 3 is the slowest result that we measured.
| Metric | Performance | 2026-10-05 re-check, chip 7 |
|---|---|---|
End-to-end pipeline.run() |
115 ms median (115.0–115.1) · 114.2 ms min | 126.9 ms median · 126.2 ms min (verifier: 128.0 / 126.9) |
Python model() call, model(pil_image, "car") |
114.7 ms median · 113.7 ms min (file path input with PNG decode: 149.0 ms median) | 127.2 ms median · 126.2 ms min (file path input: 164.7 ms median) |
| Host preprocessing · pixel upload | 1.7–2.0 ms (byte-exact C reimplementation of PIL bicubic + fused normalize/patchify) · 0.6 ms | 1.8 ms · 0.55 ms |
| Vision trace (MoonViT, 1056 patches) | 22.0 ms | 22.4 ms |
| Prefill (trace, first token in-trace) | 17.7 ms | 18.9 ms |
| Decode | 8.3 ms/token (~121 tok/s) · 9 steps 74.2 ms | 9.5 ms/token (~106 tok/s) · 9 steps 84.8 ms |
/predict served over HTTP, 100 requests, timing_ms.total |
not measured | 127.1 ms median · 126.2 ms min · wall ~0.19 s |
The Python API (2026-10-04) does not change the model, the kernels or the device setup, so the other rows in column 2 are unchanged. The model() row comes from chip 12, ETH dispatch, 12x10 grid and 30 warm calls (test_api_device.py --runs 30). In the same process, pipeline.run() took 114.6 ms median. The API adds no measurable time.
RTX 5090 reference measurements (2026-09-14) are unchanged. They use the port's torch reference in eager PyTorch 2.11, batch 1, with H2D/D2H included, plus the same host pre/post-processing; "served-like" adds 13.1 ms of PIL preprocessing. Full table: GPU_COMPARISON.md.
| RTX 5090 precision | GPU served-like | vs current build (115 ms) | GPU vision+prefill vs ours (39.7 ms) | GPU 9-step decode vs ours (74.2 ms) |
|---|---|---|---|---|
| fp32 strict | 229.0 ms | Blackhole 1.99× faster | 92.0 ms: Blackhole 2.32× faster | 123.5 ms: Blackhole 1.66× faster |
| fp32 + TF32 | 198.9 ms | Blackhole 1.73× faster | 61.9 ms: Blackhole 1.56× faster | 123.5 ms: Blackhole 1.66× faster |
| bf16 native weights | 171.6 ms | Blackhole 1.49× faster | 49.1 ms: Blackhole 1.24× faster | 108.7 ms: Blackhole 1.46× faster |
bf16 native + torch.compile |
149.7 ms | Blackhole 1.30× faster | 39.9 ms: parity (1.01×) | 96.3 ms: Blackhole 1.30× faster |
With the chip-7 numbers (e2e 126.9 ms, vision+prefill 41.3 ms, 9-step decode 84.8 ms), the ratios are lower:
| RTX 5090 precision | GPU served-like | vs chip 7 (126.9 ms) | GPU vision+prefill vs chip 7 (41.3 ms) | GPU 9-step decode vs chip 7 (84.8 ms) |
|---|---|---|---|---|
| fp32 strict | 229.0 ms | Blackhole 1.80× faster | 92.0 ms: Blackhole 2.23× faster | 123.5 ms: Blackhole 1.46× faster |
| fp32 + TF32 | 198.9 ms | Blackhole 1.57× faster | 61.9 ms: Blackhole 1.50× faster | 123.5 ms: Blackhole 1.46× faster |
| bf16 native weights | 171.6 ms | Blackhole 1.35× faster | 49.1 ms: Blackhole 1.19× faster | 108.7 ms: Blackhole 1.28× faster |
bf16 native + torch.compile |
149.7 ms | Blackhole 1.18× faster | 39.9 ms: GPU 1.03× faster | 96.3 ms: Blackhole 1.14× faster |
Part of the end-to-end lead comes from host preprocessing: 1.7–2.0 ms here vs 13.1 ms for PIL on the GPU host. On device time alone (vision + prefill + decode), Blackhole is 1.39× faster than bf16 eager GPU (113.9 ms vs 157.8 ms). On chip 7 the device time is 126.1 ms, and Blackhole is 1.25× faster than bf16 eager GPU. The previous release was 1.78× slower than bf16 GPU.
Caveats
- Does not scale to multiple p150a in a mesh configuration. Current build enforces 12x10 tensix cores for compute, which is enabled by moving dispatch functions that were originally in 10 tensix cores to ETH cores. This implementation therefore assumes that chip-to-chip ethernet communication is not required by the user.
dispatch="worker"(Python API) andLA_DISPATCH=worker(server) are opt-in modes with the stock Tensix-column dispatch. They give only an 11x10 grid on a p150. On a tt-metal without the patch,"auto"also uses this mode and shows a warning. No number in this card uses this mode, except the container-image response above.- One image + one query per request, batch 1, requests are serialized. Every image is squash-resized onto a fixed 24×44-patch grid (616×336, 16:9;
LA_IN_TOKEN_LIMIT=1024, upstream default 25600): non-16:9 images are distorted before the model sees them, boxes still map back to original pixels. - Not an OpenAI-compatible API;
GET /v1/modelsis a stub so the tt-model ready card does not 404.
Licensing
- Weights: nvidia/LocateAnything-3B,
other(NVIDIA license, not OSI); fetched from upstream, not redistributed here. - Port, Python API and serving code (
code/locate_anything,code/scripts,examples/): Apache-2.0 SPDX headers, from changh95/tt-locate-anything, published under the same upstream terms; the vendored tt-metal overlaycode/models/tt_transformers/tt/mlp.pyis Apache-2.0.
Provenance
These are the exact sources the container image was built from. code/ has since been updated (2026-10-03 optimized build; see OPT_REPORT.md; 2026-10-04 Python API, see code/PYTHON.md; 2026-10-05 ETH dispatch, 1 command queue and 12x10 grid as the default of the server and the Python API, see VERIFICATION_2026-10-03.md) and is newer than the image. tt-model serve runs the image's code until the image is rebuilt. tt-model.yaml and SERVING.md still describe the image:
| component | built from |
|---|---|
| tt-metal | 8b98410e730bb504fea43a88609756e34821d91d |
code/ digest (image) |
27dc2b154aa86da9 (sha256, first 16 hex digits; the current code/ differs) |
| built | 2026-09-13T15:59:45+00:00 by tt-model 0.1.0 |

