--- tags: - blackhole - p150 - tt-dit-server - tt-model-cache - tt-model-container - tenstorrent - ttnn - tt-metal - tt-nn - visual-grounding - open-vocabulary - vlm - qwen2.5 - tt-model-catalog license: other base_model: - nvidia/LocateAnything-3B license_link: https://huggingface.co/nvidia/LocateAnything-3B/blob/main/LICENSE pipeline_tag: object-detection license_name: nvidia-license --- # locate-anything-3b-p150 NVIDIA LocateAnything-3B port on one Tenstorrent Blackhole p150a. Weights: [nvidia/LocateAnything-3B](https://huggingface.co/nvidia/LocateAnything-3B) · Paper: [arXiv:2605.27365](https://arxiv.org/abs/2605.27365) · Upstream code: [NVlabs/Eagle (Embodied)](https://github.com/NVlabs/Eagle/tree/main/Embodied) · Port: [changh95/tt-locate-anything](https://github.com/changh95/tt-locate-anything) Runs on **p150** (mesh `P150`). Device configuration: dispatch on ETH cores, 1 command queue, 12x10 compute grid. This is the p150 target and the default of the Python API and the HTTP server (`LA_DISPATCH=auto`). It needs tt-metal with [`patches/tt-metal-eth-dispatch.patch`](patches/tt-metal-eth-dispatch.patch). All performance numbers in this card were measured in this configuration, on Blackhole chips of a Galaxy host, except the container-image response in [Serving (HTTP)](#serving-http). Packaged and published with [tt-model-manager](https://github.com/tenstorrent/tt-model-manager) 0.1.0 (manifest schema 5.1). ## Quickstart (Python) Prerequisite: a tt-metal / ttnn environment at tt-metal [`8b98410e730`](https://github.com/tenstorrent/tt-metal/commit/8b98410e730bb504fea43a88609756e34821d91d) with [`patches/tt-metal-eth-dispatch.patch`](patches/tt-metal-eth-dispatch.patch) applied. ttnn is not on PyPI. ```bash hf download changh95/locate-anything-3b-p150 --exclude "image/*" --local-dir locate-anything-3b-p150 && cd locate-anything-3b-p150 pip install -e code/ # adds numpy<2, pillow, huggingface_hub, safetensors, loguru, tqdm, pyyaml pip install -e "code/[server,test]" # also adds the HTTP server (fastapi, uvicorn) and the tests (pytest, httpx) ``` ```python from locate_anything import LocateAnything with LocateAnything.from_pretrained(device_id=0) as model: # weights from the HF cache, opens the chip, warms up result = model("media/demo_input.png", "car") # path, PIL image, numpy or torch array for label, box in zip(result.labels, result.boxes): print(label, box.round(1)) # car [ 541.4 447.1 1165.4 860.8] (x1 y1 x2 y2, pixels) result.draw().save("out.png") ``` - The first `from_pretrained` downloads the [`nvidia/LocateAnything-3B`](https://huggingface.co/nvidia/LocateAnything-3B) weights (about 7.7 GB) to your HF cache. It also converts the weights to BFP8. This takes some minutes. - With warm caches, `from_pretrained` takes 20-35 s. Each later call takes about 115 ms on the demo image. - `from_pretrained` warms up the model before it returns. The first call is as fast as the later calls: ~115 ms on the demo image. The default warm-up prepares prompts up to 1024 tokens. To prepare prompts up to 2048 tokens, use `from_pretrained(..., warmup_variants="all")`. - The path `media/demo_input.png` is relative. Run the snippet from the root of this repo, or give an absolute path. - The `with` block releases the traces and closes the chip at the end. Without `with`, call `model.close()`. | | | |---|---| | **Input** (`image`) | A file path, PNG / JPEG `bytes`, a `PIL.Image`, a numpy array or a torch tensor (`HxWx3` RGB `uint8`, `3xHxW`, grayscale, or float in 0..1). Any size. | | **Input** (`query`) | Free text (`"the red mug on the left"`) or a list of categories. The list `["person", "car"]` becomes `"personcar"`. | | **Options** | `model(image, query, max_new_tokens=128)` (1..1024). `from_pretrained(device_id=0, device=None, dispatch="auto", weights_dir=None, cache_dir=..., precision="bfp8w")`. | | **Output** (`LocateResult`) | `boxes` `float64 [N, 4]` (`x1, y1, x2, y2` in original image pixels), `labels`, `boxes_norm` (0..1000 model output), `points`, `raw_text`, `stopped_on_eos`, `timing_ms`. | | **Methods** | `result.draw()` gives a `PIL.Image` with the boxes. `result.to_dict()` gives the same JSON body as the HTTP `/predict`. `for det in result:` gives each `Detection`. | - Throughput: `model.predict_batch(images, queries)` runs a list of images on the chip, one at a time. A host thread decodes the next image while the chip runs the current image. - `model.detect(image, ["person", "car"])` is the upstream `detect` name. `LocateAnything.parse_boxes` and `parse_points` give the upstream output format. - The model has state. A lock serializes the calls. Make all calls from the thread that built the model. - `model(image, query)` runs the same code as the HTTP `POST /predict`. It gives the same tokens and boxes. - The model squashes each image to 616×336. Images from OpenCV are BGR, so give `img[..., ::-1]`. - Full reference (all options, input forms, fields, tests): [`code/PYTHON.md`](code/PYTHON.md). Runnable example: [`examples/quickstart.py`](examples/quickstart.py). ## Serving (HTTP) ```bash tt-model pull changh95/locate-anything-3b-p150 --with-weights tt-model serve changh95/locate-anything-3b-p150 # or with tt-cli: tt serve changh95/locate-anything-3b-p150 { printf '{"query":"car","image":"'; base64 -w0 media/demo_input.png; printf '"}'; } > req.json curl -s localhost:20000/predict -H 'Content-Type: application/json' -d @req.json tt model stop changh95/locate-anything-3b-p150 ``` - Weights [`nvidia/LocateAnything-3B`](https://huggingface.co/nvidia/LocateAnything-3B) go to your HF cache; the image does not contain them. - Serves on port 20000 (or the next free port); ready when the log says `Application startup complete`. - `POST /predict`: `image` (base64 PNG/JPEG, one image), `query` (free text, 1-1000 chars; join categories with ``, e.g. `personcar`); optional `max_new_tokens` (128, cap 1024), `return_overlay` (false). Also `GET /health`, `GET /info`. ```json {"query": "car", "width": 1920, "height": 1080, "canonical_size": [616, 336], "grid_hw": [24, 44], "raw_text": "car<282><414><606><794><|im_end|>", "detections": [{"label": "car", "box": [541.44, 447.12, 1163.52, 857.52], "box_norm": [282, 414, 606, 794]}], "points": [], "num_generated_tokens": 10, "stopped_on_eos": true, "decode_mode": "ar_greedy_trace_device_sampling", "timing_ms": {"vision": 1.4, "prefill": 86.3, "decode": 202.1, "total": 303.9, "decode_tok_s": 44.53}} ``` - `box` is `[x1, y1, x2, y2]` in original image pixels; `box_norm` is the model's own 0..1000 output over the squashed `canonical_size` view. `return_overlay: true` adds `overlay_png_b64`. - This response comes from the container image (2026-09-13, real p150a). The image opens the chip with the stock Tensix-column dispatch, which gives an 11x10 grid on a p150a. It is not the target configuration. - The current `code/` serves with ETH dispatch, 1 command queue and a 12x10 grid. For the same request it gives `<282><414><607><797>` and `timing_ms` `{"vision": 0.4, "prefill": 0.3, "decode": 123.6, "total": 127.1, "decode_tok_s": 72.82}` (chip 7, median of 100 requests, 2026-10-05; wall ~0.19 s). On chips where ETH dispatch is as fast as Tensix dispatch, `pipeline.run()` takes 115 ms (see [Demo & Performances](#demo--performances) and [Provenance](#provenance)). ## Demo | Input (`media/demo_input.png`), query `car` | Greedy decode on p150a (`media/demo_ar.png`) | |:---:|:---:| | ![](media/demo_input.png) | ![](media/demo_ar.png) | ## Demo & Performances Warm, batch 1, demo image + `car`, 10 tokens. ETH dispatch, 1 command queue, 12x10 compute grid for all rows. The speed of ETH dispatch changes between the chips of the Galaxy host. On chips 9, 12, 14 and 16, ETH dispatch gave about 115 ms (column 2). On chips 9 and 16, Tensix dispatch gave the same speed. On chip 7 (2026-10-05 re-check, independently verified), ETH dispatch was about 12 ms slower than Tensix dispatch on the same chip (115.1 ms), all in the traced programs (column 3). The outputs were bit-identical. We expect column 2 to be the better p150 estimate, because a p150 has all its ETH cores free for dispatch. Column 3 is the slowest result that we measured. | Metric | Performance | 2026-10-05 re-check, chip 7 | |---|---:|---:| | End-to-end `pipeline.run()` | **115 ms median** (115.0–115.1) · 114.2 ms min | 126.9 ms median · 126.2 ms min (verifier: 128.0 / 126.9) | | Python `model()` call, `model(pil_image, "car")` | **114.7 ms median** · 113.7 ms min (file path input with PNG decode: 149.0 ms median) | 127.2 ms median · 126.2 ms min (file path input: 164.7 ms median) | | Host preprocessing · pixel upload | 1.7–2.0 ms (byte-exact C reimplementation of PIL bicubic + fused normalize/patchify) · 0.6 ms | 1.8 ms · 0.55 ms | | Vision trace (MoonViT, 1056 patches) | **22.0 ms** | 22.4 ms | | Prefill (trace, first token in-trace) | **17.7 ms** | 18.9 ms | | Decode | **8.3 ms/token (~121 tok/s)** · 9 steps 74.2 ms | 9.5 ms/token (~106 tok/s) · 9 steps 84.8 ms | | `/predict` served over HTTP, 100 requests, `timing_ms.total` | not measured | 127.1 ms median · 126.2 ms min · wall ~0.19 s | The Python API (2026-10-04) does not change the model, the kernels or the device setup, so the other rows in column 2 are unchanged. The `model()` row comes from chip 12, ETH dispatch, 12x10 grid and 30 warm calls (`test_api_device.py --runs 30`). In the same process, `pipeline.run()` took 114.6 ms median. The API adds no measurable time. RTX 5090 reference measurements (2026-09-14) are unchanged. They use the port's torch reference in eager PyTorch 2.11, batch 1, with H2D/D2H included, plus the same host pre/post-processing; "served-like" adds 13.1 ms of PIL preprocessing. Full table: [`GPU_COMPARISON.md`](GPU_COMPARISON.md). | RTX 5090 precision | GPU served-like | vs current build (115 ms) | GPU vision+prefill vs ours (39.7 ms) | GPU 9-step decode vs ours (74.2 ms) | |---|---:|---:|---:|---:| | fp32 strict | 229.0 ms | **Blackhole 1.99× faster** | 92.0 ms: Blackhole 2.32× faster | 123.5 ms: Blackhole 1.66× faster | | fp32 + TF32 | 198.9 ms | **Blackhole 1.73× faster** | 61.9 ms: Blackhole 1.56× faster | 123.5 ms: Blackhole 1.66× faster | | bf16 native weights | 171.6 ms | **Blackhole 1.49× faster** | 49.1 ms: Blackhole 1.24× faster | 108.7 ms: Blackhole 1.46× faster | | bf16 native + `torch.compile` | 149.7 ms | **Blackhole 1.30× faster** | 39.9 ms: parity (1.01×) | 96.3 ms: Blackhole 1.30× faster | With the chip-7 numbers (e2e 126.9 ms, vision+prefill 41.3 ms, 9-step decode 84.8 ms), the ratios are lower: | RTX 5090 precision | GPU served-like | vs chip 7 (126.9 ms) | GPU vision+prefill vs chip 7 (41.3 ms) | GPU 9-step decode vs chip 7 (84.8 ms) | |---|---:|---:|---:|---:| | fp32 strict | 229.0 ms | **Blackhole 1.80× faster** | 92.0 ms: Blackhole 2.23× faster | 123.5 ms: Blackhole 1.46× faster | | fp32 + TF32 | 198.9 ms | **Blackhole 1.57× faster** | 61.9 ms: Blackhole 1.50× faster | 123.5 ms: Blackhole 1.46× faster | | bf16 native weights | 171.6 ms | **Blackhole 1.35× faster** | 49.1 ms: Blackhole 1.19× faster | 108.7 ms: Blackhole 1.28× faster | | bf16 native + `torch.compile` | 149.7 ms | **Blackhole 1.18× faster** | 39.9 ms: GPU 1.03× faster | 96.3 ms: Blackhole 1.14× faster | Part of the end-to-end lead comes from host preprocessing: 1.7–2.0 ms here vs 13.1 ms for PIL on the GPU host. On device time alone (vision + prefill + decode), Blackhole is 1.39× faster than bf16 eager GPU (113.9 ms vs 157.8 ms). On chip 7 the device time is 126.1 ms, and Blackhole is 1.25× faster than bf16 eager GPU. The previous release was 1.78× slower than bf16 GPU. ## Caveats - Does not scale to multiple p150a in a mesh configuration. Current build enforces 12x10 tensix cores for compute, which is enabled by moving dispatch functions that were originally in 10 tensix cores to ETH cores. This implementation therefore assumes that chip-to-chip ethernet communication is not required by the user. - `dispatch="worker"` (Python API) and `LA_DISPATCH=worker` (server) are opt-in modes with the stock Tensix-column dispatch. They give only an 11x10 grid on a p150. On a tt-metal without the patch, `"auto"` also uses this mode and shows a warning. No number in this card uses this mode, except the container-image response above. - One image + one query per request, batch 1, requests are serialized. Every image is squash-resized onto a fixed 24×44-patch grid (616×336, 16:9; `LA_IN_TOKEN_LIMIT=1024`, upstream default 25600): non-16:9 images are distorted before the model sees them, boxes still map back to original pixels. - Not an OpenAI-compatible API; `GET /v1/models` is a stub so the tt-model ready card does not 404. ## Licensing - Weights: [nvidia/LocateAnything-3B](https://huggingface.co/nvidia/LocateAnything-3B), `other` ([NVIDIA license](https://huggingface.co/nvidia/LocateAnything-3B/blob/main/LICENSE), not OSI); fetched from upstream, not redistributed here. - Port, Python API and serving code (`code/locate_anything`, `code/scripts`, `examples/`): Apache-2.0 SPDX headers, from [changh95/tt-locate-anything](https://github.com/changh95/tt-locate-anything), published under the same upstream terms; the vendored [tt-metal](https://github.com/tenstorrent/tt-metal) overlay `code/models/tt_transformers/tt/mlp.py` is Apache-2.0. ## Provenance These are the exact sources the container image was built from. **`code/` has since been updated (2026-10-03 optimized build; see [`OPT_REPORT.md`](OPT_REPORT.md); 2026-10-04 Python API, see [`code/PYTHON.md`](code/PYTHON.md); 2026-10-05 ETH dispatch, 1 command queue and 12x10 grid as the default of the server and the Python API, see [`VERIFICATION_2026-10-03.md`](VERIFICATION_2026-10-03.md)) and is newer than the image.** `tt-model serve` runs the image's code until the image is rebuilt. `tt-model.yaml` and `SERVING.md` still describe the image: | component | built from | | --- | --- | | tt-metal | [`8b98410e730bb504fea43a88609756e34821d91d`](https://github.com/tenstorrent/tt-metal/commit/8b98410e730bb504fea43a88609756e34821d91d) | | `code/` digest (image) | `27dc2b154aa86da9` (sha256, first 16 hex digits; the current `code/` differs) | | built | 2026-09-13T15:59:45+00:00 by tt-model 0.1.0 |