clm-v0.1-8b-p150

CLM-v0.1-8B (Contrastive Language Model) is a System One verifier and ranker: a frozen Qwen3-8B encoder with last-token pooling plus two 9.4M-parameter projection heads trained with InfoNCE. This package serves its System One API (typed noul / choice / score questions, free-form ranking) and an OpenAI-compatible embeddings endpoint on one Blackhole p150 chip.

Runs on p150 or p150 or p150 or p150x4 โ€” see the serve profiles below.

Packaged and published with tt-model-manager 0.1.0 (manifest schema 5.1).

At a glance

Architecture Qwen3-8B encoder (36 layers, last-token pooling) + 2 x MLP heads (4096-1536-1536-512)
Hardware p150, p150x4
License apache-2.0
Status Experimental community bring-up

Intended use

Direct use: Zero-shot action selection, best-of-N ranking, typed decisions over text states.

Out-of-scope use: Text generation. The package produces scores and embeddings only.

Quickstart

uv tool install tenstorrent   # once โ€” the Tenstorrent CLI, `tt`
tt model pull tt-hous/clm-v0.1-8b-p150
tt serve tt-hous/clm-v0.1-8b-p150

tt model pull (or tt-model pull --with-weights) downloads the Docker image and the Qwen/Qwen3-8B weights at b968826d9c46dd6066d109eabc6255188de91218 (into your HF cache; they are not in the image). tt serve (or tt-model serve) starts the model's own HTTP server on port 20000 (or the next free port, if that one is busy); the first start compiles kernels for your device, which takes several minutes, and the server is ready when it logs Application startup complete.

Without tt-cli โ€” tt-model alone does the whole job:

tt-model pull  tt-hous/clm-v0.1-8b-p150 --with-weights
tt-model serve tt-hous/clm-v0.1-8b-p150

Serve profiles

One image serves every profile below; pick one with --profile.

profile hardware mesh
p150 (default) p150 P150
p150-accuracy p150 P150
p150-fast p150 P150
p150x4 p150x4 P150x4

Using it

This package exposes the model's own HTTP server on port 20000, not an OpenAI-compatible chat API โ€” its request and response shapes are the model's own. See the author's notes above for the payload it expects.

The server speaks the CLM System One wire format. POST /v1/systemone takes {"state": ..., "questions": {id: {"type": "noul"|"choice"|"score", "instructions": ..., "criteria": ...}}, "temperature": 1.0} and returns per-question probability distributions; POST /v1/rank takes {"context", "question", "answers": [...]}; POST /v1/embeddings is OpenAI-compatible (model qwen3-8b, 4096-d L2-normalized, last-token pooled Qwen3-8B). Every response carries X-CLM-Latency-Ms. The playground UI is served at /. The clm Python client from github.com/Contrastive-LM/CLM works unchanged with CLM_BASE_URL=http://<host>:<port>. A live demo of the CLM repository's T-Rex game is served at /demo: the game runs headless inside the container at 60 FPS, the model on the chip chooses jump, duck or run through this server's own POST /v1/systemone, and the page streams the game and every decision (the three option texts with probabilities, the executed action, shield interventions, round trip and X-CLM-Latency-Ms) over a WebSocket at /demo/ws; the table at the end of a run has the same fields as the T-Rex row in the performance section, with the authors' RTX 4090 numbers alongside.

Expected performance

Served from this package on one Blackhole p150 chip (two p300c boards, four chips, qb2 box), default profile (accuracy_lofi_mlp policy with the program-config overrides), client on the same host, p50 latencies. System One API: the README customer example (three questions, 98 tokens embedded) 220 ms with a cold cache and 0.1 ms warm (all vectors cached); a new state against a fixed action set 56 ms server-side (the CLM README's 58.1 ms on one RTX 4090 is the comparable call: 38 tokens embedded, option vectors cached; its vector-cache table reports 28.0 ms for a new state and 0.6 ms for a cached one, against 56 ms and 0.1 ms here). Embeddings endpoint: one 128-token text 56.6 ms, 512 tokens 83 ms, 1024 tokens 149 ms, 2048 tokens 272 ms; eight 128-token texts 149 ms, eight 512-token texts 530 ms; throughput up to 7.7k tokens/s at batch 32. Fidelity vs an fp32 CPU run of the same weights: cosine 0.9992 mean / 0.9960 min on 308 texts (host run of the same encoder code); from the served package, 96.0 percent identical argmax on 200 Typed Decisions subset decisions and 98.4 percent of the 188 with a reference margin above 0.10. Zero-shot on LocalLLaMA/typed-decisions (400 cases, 2,000 decisions): accuracy 0.345 (0.364 with the previous stock-accuracy default; the fp32 reference scores 0.370 on the 40-case subset against 0.355 here), KL 2.04, Brier 0.63, 258 ms p50 per 5-question case; T-Rex real-time game 3 of 5 courses survived, 2,195 decisions per course (authors on a 4090: 5 of 5, 3,342). Profiles: p150-accuracy (the previous default) 263 ms cold README example, 60 ms new state, about 9 percent slower than p150 at 128 tokens batch 1 and 19 to 22 percent slower in the other measured cells; p150-fast 234 ms and 56 ms with 97.3 / 95.7 percent confident-decision agreement (sweep / served); p150x4 (four chips, tensor parallel) 169 ms cold README example and 34 ms new state from the published image.

Limitations

Encoder-only; produces scores and embeddings, never text. Inputs longer than 2048 tokens are truncated to their last 2048 tokens (upstream clm-serve default); the encoder itself supports 40960. Outputs are deterministic for identical requests, but a text embedded alongside different batch-mates can differ at the cosine 0.999 level (batch-variant reduction order of the prefill kernels, tt-metal issue 47238): between a single-text and a batched embedding of the same 200 Typed Decisions subset decisions, 4 to 5 argmax answers change, all near ties. Against the fp32 CPU reference the default profile reverses 8 of those 200 decisions (5 of them where the reference's own top-2 margin is below 0.10, 3 of the 188 confident ones: 98.4 percent). The published README example probabilities (0.410 / 0.939 / 1.984) are not reproduced by the published head on any hardware; the default profile gives 0.844 / 0.991 / 2.000 on the single-text path (host run of the shipped policy) and 0.875 / 0.986 / 2.000 when the served request is embedded as one batch, against the CPU fp32 reference's 0.842 / 0.988 / 2.000. Zero-shot Typed Decisions accuracy is below the dataset's prior baseline on both CPU and TT; the CLM authors treat that benchmark as a fine-tuning target. The p150-fast profile fails this port's decision-agreement gate (97.3 percent of confident decisions in the standalone sweep, 95.7 percent measured from a served package) and is kept for consumers who prefer its latency. Only the p150, p150-accuracy, p150-fast and p150x4 profiles were run; a two-chip 1x2 mesh does not open on the build host (fabric router sync timeout). The host must set TT_METAL_PINNED_MEMORY_CACHE_LIMIT_BYTES=0 (the package does) because loading cached weight tensors from disk otherwise hangs on this tt-metal version (see the port's doc/probe/README.md). The branch hous/clm-v0.1-8b at github.com/housTT/tt-metal carries the port; the image is built from a clean checkout of it.

Risks and safety considerations

The matmul kernels use 32 of 110 cores at 128 tokens (model_config grid cap), so single-text latency is about 2x an RTX 4090; latency-sensitive loops like the T-Rex game lose decisions. The container serves one model process; concurrent requests are serialized on the device.

Licensing

Qwen3-8B weights under Apache-2.0; CLM-v0.1-8B head weights under Apache-2.0 (Contrastive-LM); the vendored clm serving code under Apache-2.0; the TTNN port code under Apache-2.0.

Feedback

Questions or problems with this package: open a discussion at https://huggingface.co/tt-hous/clm-v0.1-8b-p150/discussions โ€” that is what reaches its author. A problem with the tt tooling itself: tt report issue (collects your environment and opens a prefilled issue against tenstorrent/tt-cli). Product feedback: support@tenstorrent.com.

Provenance

The exact sources the image was built from โ€” code/ in this repo is byte-identical to the model code inside the image:

component built from
tt-metal 2641f871d7c031866dba01db8eb52946f2772861
code/ digest 949aa72e5600889b (sha256, first 16 hex digits)
built 2026-10-02T18:37:16+00:00 by tt-model 0.1.0
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for tt-hous/clm-v0.1-8b-p150

Finetuned
Qwen/Qwen3-8B
Finetuned
(5)
this model