devstral-small-2-24b-instruct-2512

Devstral Small 2 is Mistral's 24B agentic coding model, tuned for software engineering tasks: exploring a repository, editing multiple files, and driving tools. This package serves the text decoder on Tenstorrent Blackhole through vLLM, with tool calling enabled. The Pixtral vision tower is not ported, and a single die is not supported -- two dies is the minimum.

Runs on p300 or p150x4 or p300x2 β€” see the serve profiles below.

Packaged and published with tt-model-manager 0.1.0 (manifest schema 5.1).

Quickstart

uv tool install tenstorrent   # once β€” the Tenstorrent CLI, `tt`
tt model pull anirud/devstral-small-2-24b-instruct-2512
tt serve anirud/devstral-small-2-24b-instruct-2512

tt model pull downloads the Docker image and the mistralai/Devstral-Small-2-24B-Instruct-2512 weights at 55c5b41e98c2dbd21b0c8afffc540dcfc9eb5128 into your HF cache; they are not baked into the image. There is no --with-weights flag on tt β€” a bundle's weights come down by default. tt serve starts an OpenAI-compatible server on port 20000 (or the next free port); the first start compiles kernels for your device, which takes several minutes, and the server is ready when it logs Application startup complete.

Without tt-cli β€” tt-model alone does the whole job:

tt-model pull  anirud/devstral-small-2-24b-instruct-2512 --with-weights
tt-model serve anirud/devstral-small-2-24b-instruct-2512

Point a client at it

Once the server reports ready, the OpenAI-compatible endpoint is on the port tt serve printed (20000 by default).

tt serve anirud/devstral-small-2-24b-instruct-2512 --profile p150x4
curl http://127.0.0.1:20000/v1/chat/completions -H 'Content-Type: application/json' -d '{
  "model": "mistralai/Devstral-Small-2-24B-Instruct-2512",
  "messages": [{"role": "user", "content": "Refactor this function to use a generator."}],
  "temperature": 0.15
}'

Tool calling is on: send an OpenAI tools payload and the reply carries finish_reason: tool_calls.

Profiles

profile dies notes
p300 2 one p300 card
p150x4 4 the default
p300x2 4 the same four dies under the QB2 label

There is no single-die profile. The weights and KV cache fit in one 32 GB die, but a prefill matmul overruns L1: tt_transformers picks the same MLP prefill grids on one die as on four, so one die's cores carry the whole MLP rather than a quarter of it.

Not ported

The Pixtral vision tower. This is a text-only port: image inputs are not supported.

Serve profiles

One image serves every profile below; pick one with --profile.

profile hardware mesh max_num_seqs max_model_len
p300 p300 P300 4 32768
p150x4 (default) p150x4 P150x4 8 65536
p300x2 p300x2 P300x2 32 65536

Provenance

The exact sources the image was built from β€” code/ in this repo is byte-identical to the model code inside the image:

component built from
tt-metal a local checkout β€” commit not published
vLLM v0.26.0
vllm-tt-plugin a local checkout β€” commit not published
code/ digest 71722ddade285f7f (sha256, first 16 hex digits)
built 2026-10-01T15:37:57+00:00 by tt-model 0.1.0
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support