Upload folder using huggingface_hub
Browse files- .gitattributes +3 -0
- README.md +106 -0
- base_config/config.json +140 -0
- base_config/preprocessor_config.json +21 -0
- base_config/video_preprocessor_config.json +21 -0
- install.sh +36 -0
- model.py +6 -0
- prepare_model_dir.py +83 -0
- requirements.txt +5 -0
- run.sh +87 -0
- tt_kernel_manifest.json +80 -0
- vllm-overrides.txt +6 -0
- vllm_models/hemmingway-1-p300/vllm_metadata.json +4 -0
- wheels/tt_metal_models-0.79.0.dev20260929+qwen36p150x2.g9f6b02d-py3-none-any.whl +3 -0
- wheels/ttnn-0.79.0.dev20260929+qwen36p150x2.g9f6b02d-cp312-cp312-linux_x86_64.whl +3 -0
- wheels/vllm_tt_plugin-0.1.0-py3-none-any.whl +3 -0
.gitattributes
CHANGED
|
@@ -33,3 +33,6 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
|
|
| 33 |
*.zip filter=lfs diff=lfs merge=lfs -text
|
| 34 |
*.zst filter=lfs diff=lfs merge=lfs -text
|
| 35 |
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
|
|
|
|
|
|
|
|
|
|
|
| 33 |
*.zip filter=lfs diff=lfs merge=lfs -text
|
| 34 |
*.zst filter=lfs diff=lfs merge=lfs -text
|
| 35 |
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
| 36 |
+
wheels/tt_metal_models-0.79.0.dev20260929+qwen36p150x2.g9f6b02d-py3-none-any.whl filter=lfs diff=lfs merge=lfs -text
|
| 37 |
+
wheels/ttnn-0.79.0.dev20260929+qwen36p150x2.g9f6b02d-cp312-cp312-linux_x86_64.whl filter=lfs diff=lfs merge=lfs -text
|
| 38 |
+
wheels/vllm_tt_plugin-0.1.0-py3-none-any.whl filter=lfs diff=lfs merge=lfs -text
|
README.md
ADDED
|
@@ -0,0 +1,106 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: cc-by-nc-4.0
|
| 3 |
+
base_model: Altworld/Hemmingway-1
|
| 4 |
+
pipeline_tag: text-generation
|
| 5 |
+
tags:
|
| 6 |
+
- tt-model-cache
|
| 7 |
+
- blackhole
|
| 8 |
+
- vllm
|
| 9 |
+
- thin
|
| 10 |
+
- p150x2
|
| 11 |
+
- chat
|
| 12 |
+
- creative-writing
|
| 13 |
+
---
|
| 14 |
+
|
| 15 |
+
# hemmingway-1-p300
|
| 16 |
+
|
| 17 |
+
> **License: CC BY-NC 4.0. Non-commercial use only.** The weights are by Altworld (https://huggingface.co/Altworld/Hemmingway-1). Commercial use needs a separate agreement with Altworld. This package inherits that license.
|
| 18 |
+
|
| 19 |
+
This is a tt-model v6 thin package that serves Altworld/Hemmingway-1 (revision `1a5f363a3dd2d1cc456c28b8abbb403b9555efaf`) on 2 Tenstorrent Blackhole chips (one P300 board, or a P150x2 mesh). Hemmingway-1 (`Altworld/Hemmingway-1`) is a 27B-parameter open-weights text model from Altworld, built on Qwen/Qwen3.8-27B for everyday writing: messages, emails and short pieces. Its upstream card says it gives you the text itself, without a preamble or a list of options. The upstream card also reports comparisons with other models. We did not reproduce them.
|
| 20 |
+
|
| 21 |
+
The package was produced by an automated bring-up harness (tt-orchard), and a person reviewed the results before release.
|
| 22 |
+
|
| 23 |
+
## How to serve
|
| 24 |
+
|
| 25 |
+
tt-model serve episod/hemmingway-1-p300
|
| 26 |
+
|
| 27 |
+
The server exposes an OpenAI-compatible endpoint (`/v1/chat/completions` and `/v1/completions`). Context length is 262144 tokens with up to 4 sequences at a time. Speculative decoding uses the drafter `incoai/Qwen3.8-27B-DFlash2` (K=7 draft tokens per step), which was trained on base Qwen3.8-27B. It was not trained on this model.
|
| 28 |
+
|
| 29 |
+
Notes that matter before you start:
|
| 30 |
+
|
| 31 |
+
- This package ships no weights. `tt-model` downloads them from the upstream repository `Altworld/Hemmingway-1` with your own Hugging Face account.
|
| 32 |
+
- Only greedy decoding is supported. Requests that ask for logprobs are rejected.
|
| 33 |
+
- The server needs a fresh, empty tensor cache for each model. Do not share a tensor cache between Qwen3.8-based models: the cache is keyed by layer name, so a cache built for another Qwen3.8-27B model would load that model's weights under this model's name.
|
| 34 |
+
- The first boot compiles kernels. Boot takes a long time (see the table below).
|
| 35 |
+
|
| 36 |
+
**Turn thinking off for direct answers.** Send `"chat_template_kwargs": {"enable_thinking": false}` with each chat request. With thinking on (the server default), the model spends its token budget on reasoning and may return no visible answer. In one test with thinking on, a creative prompt used all 700 tokens on reasoning and `content` was null.
|
| 37 |
+
|
| 38 |
+
## Evidence status of this card
|
| 39 |
+
|
| 40 |
+
All numbers below were measured on 2026-10-04 on one box, on this package installed in its own virtual environment, with the weights and tensor cache for this model only. Each is labelled measured and gives its conditions.
|
| 41 |
+
|
| 42 |
+
## Measured results (2 chips)
|
| 43 |
+
|
| 44 |
+
Conditions: greedy, thinking off, tensor cache already built for this model. Speed workload: the 80 first-turn `coding` prompts of nvidia/SPEED-Bench, output cap 2048 tokens (6 of 80 hit the cap). Per-user tok/s = 1000 / mean time per output token. Short answers distort that mean, so a token-weighted figure is also given. One run per cell.
|
| 45 |
+
|
| 46 |
+
| Item | Value |
|
| 47 |
+
|---|---|
|
| 48 |
+
| Boot, launch to `/health` OK, tensor cache present | 2025 s (33.8 min), of which 1261.7 s was the drafter and verify warm-up (alloc and compile). The verification install measured 2034.4 s. |
|
| 49 |
+
| Coding, 1 user, per-user speed | 54.1 tok/s (mean 18.50 ms per token, median 13.61 ms) |
|
| 50 |
+
| Coding, 1 user, token-weighted | 61.8 tok/s over 80 requests; 69.1 tok/s over the 71 requests with at least 64 output tokens |
|
| 51 |
+
| Coding, 1 user, TTFT | mean 232 ms, median 185 ms |
|
| 52 |
+
| Coding, 4 users, per-user speed | 28.5 tok/s (mean 35.13 ms per token, median 23.89 ms) |
|
| 53 |
+
| Coding, 4 users, token-weighted | 42.2 tok/s; 42.0 tok/s over the 71 requests of at least 64 tokens |
|
| 54 |
+
| Coding, 4 users, TTFT | mean 749 ms, median 728 ms |
|
| 55 |
+
| Coding, aggregate | 59.3 tok/s at 1 user (24,941 output tokens in 420.5 s), 130.4 tok/s at 4 users (191.3 s) |
|
| 56 |
+
|
| 57 |
+
Context sweep (`vllm bench serve`, random prompts, ignore EOS, 128 output tokens, 4 to 8 requests per cell):
|
| 58 |
+
|
| 59 |
+
| Prompt tokens / users | Per user incl. TTFT | TTFT | Time per token |
|
| 60 |
+
|---|---|---|---|
|
| 61 |
+
| 1024 / 1 | 51.5 tok/s | 318 ms | 17.1 ms (58 tok/s decode) |
|
| 62 |
+
| 16384 / 1 | 21.8 tok/s | 3788 ms | 16.3 ms (61 tok/s decode) |
|
| 63 |
+
| 1024 / 4 | 30.2 tok/s (120.9 aggregate) | 969 ms | 25.7 ms |
|
| 64 |
+
|
| 65 |
+
### Draft acceptance
|
| 66 |
+
|
| 67 |
+
The drafter was trained on base Qwen3.8-27B. Acceptance is the mean number of draft tokens accepted per step out of 7, weighted by steps.
|
| 68 |
+
|
| 69 |
+
| Workload | Accepted of 7 | Tokens committed per step |
|
| 70 |
+
|---|---|---|
|
| 71 |
+
| 80 coding prompts (4,644 steps) | 4.38 (62.6%) | 5.36 |
|
| 72 |
+
| 10 everyday-writing prompts (356 steps) | 1.72 (24.6%) | 2.69 |
|
| 73 |
+
|
| 74 |
+
Acceptance on writing prompts is much lower than on code. Speed on writing prompts was measured only as part of a 10-prompt run and not at scale. Expect lower speed on prose than the coding numbers suggest. No comparison with the base model's acceptance was run, so we cannot say whether the fine-tune lowered acceptance.
|
| 75 |
+
|
| 76 |
+
### Long context retrieval
|
| 77 |
+
|
| 78 |
+
One passkey in the middle of filler text, one trial per length, greedy. All were retrieved: 1063, 2075, 2944, 15783 (8.0 s) and 135567 tokens (44.5 s). The key was a 6-digit number in synthetic text. This is not a long-context quality evaluation.
|
| 79 |
+
|
| 80 |
+
### Agreement with the CPU bf16 reference
|
| 81 |
+
|
| 82 |
+
The chip generated 48 tokens for each of 8 prompts, and the CPU bf16 model scored each chip token given the chip's own prefix. The chip token matched at 365 of 384 positions (95.1 percent). Per prompt: 45, 42, 45, 47, 47, 46, 46, 47 of 48. The default chat template was used, so the tokens are mostly the model's thinking text. The 19 mismatches were not examined beyond their first tokens. An earlier single-prompt check gave 30 of 32 (0.9375). That is thin evidence, and the 8-prompt figure is the better one.
|
| 83 |
+
|
| 84 |
+
### Writing quality check
|
| 85 |
+
|
| 86 |
+
10 prompts we wrote for everyday messages and short writing (a note to a landlord, a raise request, declining a wedding invitation, a condolence message, a story opening and others), greedy, thinking off, 600-token cap. Result: 7 good, 3 partly, 0 poor. The grades were given by the AI model that ran the check, not by humans, on one answer per prompt. They support the claim that it writes the requested message directly, with no preamble or options. They say nothing about quality compared with another model, because no other model was run. The three "partly" answers had a literal placeholder ("[Manager's Name]"), an invented name ("Sarah") or a vague request.
|
| 87 |
+
|
| 88 |
+
### Sample
|
| 89 |
+
|
| 90 |
+
Prompt: "You are a building remembering your past. You were built in the 1930s. What do you remember?" Settings: greedy, thinking off, max_tokens 900, one sample, finished at a stop token after 830 tokens. This is not a quality measurement. Excerpt:
|
| 91 |
+
|
| 92 |
+
> I remember the first family. The Kowalskis. Mrs. Kowalski hung curtains in the front window within a week of moving in, and I felt the weight of them against my glass like a small, warm hand.
|
| 93 |
+
|
| 94 |
+
## What was not measured
|
| 95 |
+
|
| 96 |
+
- Accuracy benchmarks (MMLU and similar). None were run.
|
| 97 |
+
- A comparison with the base model Qwen3.8-27B, or with the unquantized upstream model on a GPU. None was run, so these cards make no claim that this package matches or beats either.
|
| 98 |
+
- A 4-chip package. None is published here.
|
| 99 |
+
- A fresh-machine pull test. Pulling this package from the Hub and booting it on a clean machine has not been done yet.
|
| 100 |
+
- Speed numbers use a smaller prompt subset and a lower token cap than the earlier Qwen3.8-27B DFlash2 card, so they are not comparable with that card.
|
| 101 |
+
- Sampling quality. The server accepts greedy decoding only and rejects logprobs.
|
| 102 |
+
- Boot with a cold kernel cache on a fresh machine. Warm boots in these measurements took about 28 to 34 minutes, so the first boot on a new machine can be expected to take at least that long, but it was not timed.
|
| 103 |
+
|
| 104 |
+
## Evidence
|
| 105 |
+
|
| 106 |
+
The measurements come from the run's benchmark summary and its quality-check write-up (RESULTS-bench.md and qualitative.md), with raw result files kept by the person who ran them. The risks and open items are in the run's risk list.
|
base_config/config.json
ADDED
|
@@ -0,0 +1,140 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"architectures": [
|
| 3 |
+
"Qwen3_5ForConditionalGeneration"
|
| 4 |
+
],
|
| 5 |
+
"image_token_id": 248056,
|
| 6 |
+
"language_model_only": false,
|
| 7 |
+
"model_type": "qwen3_5",
|
| 8 |
+
"text_config": {
|
| 9 |
+
"attention_bias": false,
|
| 10 |
+
"attention_dropout": 0.0,
|
| 11 |
+
"attn_output_gate": true,
|
| 12 |
+
"bos_token_id": 248044,
|
| 13 |
+
"dtype": "bfloat16",
|
| 14 |
+
"eos_token_id": 248044,
|
| 15 |
+
"full_attention_interval": 4,
|
| 16 |
+
"head_dim": 256,
|
| 17 |
+
"hidden_act": "silu",
|
| 18 |
+
"hidden_size": 5120,
|
| 19 |
+
"initializer_range": 0.02,
|
| 20 |
+
"intermediate_size": 17408,
|
| 21 |
+
"layer_types": [
|
| 22 |
+
"linear_attention",
|
| 23 |
+
"linear_attention",
|
| 24 |
+
"linear_attention",
|
| 25 |
+
"full_attention",
|
| 26 |
+
"linear_attention",
|
| 27 |
+
"linear_attention",
|
| 28 |
+
"linear_attention",
|
| 29 |
+
"full_attention",
|
| 30 |
+
"linear_attention",
|
| 31 |
+
"linear_attention",
|
| 32 |
+
"linear_attention",
|
| 33 |
+
"full_attention",
|
| 34 |
+
"linear_attention",
|
| 35 |
+
"linear_attention",
|
| 36 |
+
"linear_attention",
|
| 37 |
+
"full_attention",
|
| 38 |
+
"linear_attention",
|
| 39 |
+
"linear_attention",
|
| 40 |
+
"linear_attention",
|
| 41 |
+
"full_attention",
|
| 42 |
+
"linear_attention",
|
| 43 |
+
"linear_attention",
|
| 44 |
+
"linear_attention",
|
| 45 |
+
"full_attention",
|
| 46 |
+
"linear_attention",
|
| 47 |
+
"linear_attention",
|
| 48 |
+
"linear_attention",
|
| 49 |
+
"full_attention",
|
| 50 |
+
"linear_attention",
|
| 51 |
+
"linear_attention",
|
| 52 |
+
"linear_attention",
|
| 53 |
+
"full_attention",
|
| 54 |
+
"linear_attention",
|
| 55 |
+
"linear_attention",
|
| 56 |
+
"linear_attention",
|
| 57 |
+
"full_attention",
|
| 58 |
+
"linear_attention",
|
| 59 |
+
"linear_attention",
|
| 60 |
+
"linear_attention",
|
| 61 |
+
"full_attention",
|
| 62 |
+
"linear_attention",
|
| 63 |
+
"linear_attention",
|
| 64 |
+
"linear_attention",
|
| 65 |
+
"full_attention",
|
| 66 |
+
"linear_attention",
|
| 67 |
+
"linear_attention",
|
| 68 |
+
"linear_attention",
|
| 69 |
+
"full_attention",
|
| 70 |
+
"linear_attention",
|
| 71 |
+
"linear_attention",
|
| 72 |
+
"linear_attention",
|
| 73 |
+
"full_attention",
|
| 74 |
+
"linear_attention",
|
| 75 |
+
"linear_attention",
|
| 76 |
+
"linear_attention",
|
| 77 |
+
"full_attention",
|
| 78 |
+
"linear_attention",
|
| 79 |
+
"linear_attention",
|
| 80 |
+
"linear_attention",
|
| 81 |
+
"full_attention",
|
| 82 |
+
"linear_attention",
|
| 83 |
+
"linear_attention",
|
| 84 |
+
"linear_attention",
|
| 85 |
+
"full_attention"
|
| 86 |
+
],
|
| 87 |
+
"linear_conv_kernel_dim": 4,
|
| 88 |
+
"linear_key_head_dim": 128,
|
| 89 |
+
"linear_num_key_heads": 16,
|
| 90 |
+
"linear_num_value_heads": 48,
|
| 91 |
+
"linear_value_head_dim": 128,
|
| 92 |
+
"mamba_ssm_dtype": "float32",
|
| 93 |
+
"max_position_embeddings": 262144,
|
| 94 |
+
"model_type": "qwen3_5_text",
|
| 95 |
+
"mtp_num_hidden_layers": 1,
|
| 96 |
+
"mtp_use_dedicated_embeddings": false,
|
| 97 |
+
"num_attention_heads": 24,
|
| 98 |
+
"num_hidden_layers": 64,
|
| 99 |
+
"num_key_value_heads": 4,
|
| 100 |
+
"output_gate_type": "swish",
|
| 101 |
+
"pad_token_id": null,
|
| 102 |
+
"partial_rotary_factor": 0.25,
|
| 103 |
+
"rms_norm_eps": 1e-06,
|
| 104 |
+
"rope_parameters": {
|
| 105 |
+
"mrope_interleaved": true,
|
| 106 |
+
"mrope_section": [
|
| 107 |
+
11,
|
| 108 |
+
11,
|
| 109 |
+
10
|
| 110 |
+
],
|
| 111 |
+
"partial_rotary_factor": 0.25,
|
| 112 |
+
"rope_theta": 10000000,
|
| 113 |
+
"rope_type": "default"
|
| 114 |
+
},
|
| 115 |
+
"tie_word_embeddings": false,
|
| 116 |
+
"use_cache": true,
|
| 117 |
+
"vocab_size": 248320
|
| 118 |
+
},
|
| 119 |
+
"tie_word_embeddings": false,
|
| 120 |
+
"transformers_version": "5.8.0.dev0",
|
| 121 |
+
"video_token_id": 248057,
|
| 122 |
+
"vision_config": {
|
| 123 |
+
"deepstack_visual_indexes": [],
|
| 124 |
+
"depth": 27,
|
| 125 |
+
"hidden_act": "gelu_pytorch_tanh",
|
| 126 |
+
"hidden_size": 1152,
|
| 127 |
+
"in_channels": 3,
|
| 128 |
+
"initializer_range": 0.02,
|
| 129 |
+
"intermediate_size": 4304,
|
| 130 |
+
"model_type": "qwen3_5",
|
| 131 |
+
"num_heads": 16,
|
| 132 |
+
"num_position_embeddings": 2304,
|
| 133 |
+
"out_hidden_size": 5120,
|
| 134 |
+
"patch_size": 16,
|
| 135 |
+
"spatial_merge_size": 2,
|
| 136 |
+
"temporal_patch_size": 2
|
| 137 |
+
},
|
| 138 |
+
"vision_end_token_id": 248054,
|
| 139 |
+
"vision_start_token_id": 248053
|
| 140 |
+
}
|
base_config/preprocessor_config.json
ADDED
|
@@ -0,0 +1,21 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"size": {
|
| 3 |
+
"longest_edge": 16777216,
|
| 4 |
+
"shortest_edge": 65536
|
| 5 |
+
},
|
| 6 |
+
"patch_size": 16,
|
| 7 |
+
"temporal_patch_size": 2,
|
| 8 |
+
"merge_size": 2,
|
| 9 |
+
"image_mean": [
|
| 10 |
+
0.5,
|
| 11 |
+
0.5,
|
| 12 |
+
0.5
|
| 13 |
+
],
|
| 14 |
+
"image_std": [
|
| 15 |
+
0.5,
|
| 16 |
+
0.5,
|
| 17 |
+
0.5
|
| 18 |
+
],
|
| 19 |
+
"processor_class": "Qwen3VLProcessor",
|
| 20 |
+
"image_processor_type": "Qwen2VLImageProcessorFast"
|
| 21 |
+
}
|
base_config/video_preprocessor_config.json
ADDED
|
@@ -0,0 +1,21 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"size": {
|
| 3 |
+
"longest_edge": 25165824,
|
| 4 |
+
"shortest_edge": 4096
|
| 5 |
+
},
|
| 6 |
+
"patch_size": 16,
|
| 7 |
+
"temporal_patch_size": 2,
|
| 8 |
+
"merge_size": 2,
|
| 9 |
+
"image_mean": [
|
| 10 |
+
0.5,
|
| 11 |
+
0.5,
|
| 12 |
+
0.5
|
| 13 |
+
],
|
| 14 |
+
"image_std": [
|
| 15 |
+
0.5,
|
| 16 |
+
0.5,
|
| 17 |
+
0.5
|
| 18 |
+
],
|
| 19 |
+
"processor_class": "Qwen3VLProcessor",
|
| 20 |
+
"video_processor_type": "Qwen3VLVideoProcessor"
|
| 21 |
+
}
|
install.sh
ADDED
|
@@ -0,0 +1,36 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
#!/usr/bin/env bash
|
| 2 |
+
# Install this self-contained TT model package into an isolated, reproducible venv (via uv).
|
| 3 |
+
# Usage: ./install.sh [venv-path] (default: ./venv)
|
| 4 |
+
# Deps: v6 thin: ttnn/tt-metal-models (index) + empty-target vLLM + plugin/ops wheels (by path)
|
| 5 |
+
#
|
| 6 |
+
# HERMETIC INSTALL: everything the model needs to SERVE ends up UNDER this folder — the pinned
|
| 7 |
+
# interpreter (in .python/), the venv (with package contents copied in), and at serve time the
|
| 8 |
+
# caches/weights (run.sh points HF_HOME/TT_CACHE_PATH/... here). After this runs, serving depends
|
| 9 |
+
# on nothing outside the folder except the TT device + system libc. Only THIS install step reaches
|
| 10 |
+
# the network (to fetch the interpreter and, unless --vendor-deps, the pip deps).
|
| 11 |
+
set -euo pipefail
|
| 12 |
+
HERE="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
| 13 |
+
VENV="${1:-$HERE/venv}"
|
| 14 |
+
PYVER="3.12"
|
| 15 |
+
|
| 16 |
+
# uv gives us a pinned interpreter + deterministic installs, independent of the host Python.
|
| 17 |
+
if ! command -v uv >/dev/null 2>&1; then
|
| 18 |
+
export UV_INSTALL_DIR="$HERE/.uv"
|
| 19 |
+
curl -LsSf https://astral.sh/uv/install.sh | sh >/dev/null 2>&1
|
| 20 |
+
export PATH="$HERE/.uv:$PATH"
|
| 21 |
+
fi
|
| 22 |
+
|
| 23 |
+
# Keep the pinned interpreter INSIDE the bundle (not in uv's global ~/.local store), so the venv's
|
| 24 |
+
# python resolves within the folder wall. python-build-standalone (what uv provisions) is
|
| 25 |
+
# relocatable, so a --relocatable venv built against it stays self-contained.
|
| 26 |
+
export UV_PYTHON_INSTALL_DIR="$HERE/.python"
|
| 27 |
+
uv python install "$PYVER"
|
| 28 |
+
uv venv --relocatable --python "$PYVER" "$VENV"
|
| 29 |
+
uv pip install --python "$VENV/bin/python" --link-mode=copy --find-links "$HERE/wheels" --extra-index-url https://download.pytorch.org/whl/cpu -r "$HERE/requirements.txt"
|
| 30 |
+
VLLM_COMMON="$(mktemp)"
|
| 31 |
+
curl -fsSL "https://raw.githubusercontent.com/vllm-project/vllm/v0.26.0/requirements/common.txt" -o "$VLLM_COMMON"
|
| 32 |
+
uv pip install --python "$VENV/bin/python" --link-mode=copy --override "$HERE/vllm-overrides.txt" -r "$VLLM_COMMON"
|
| 33 |
+
rm -f "$VLLM_COMMON"
|
| 34 |
+
VLLM_TARGET_DEVICE=empty uv pip install --python "$VENV/bin/python" --link-mode=copy --no-deps --no-binary vllm vllm==0.26.0
|
| 35 |
+
uv pip install --python "$VENV/bin/python" --link-mode=copy --find-links "$HERE/wheels" "$HERE/wheels/vllm_tt_plugin-0.1.0-py3-none-any.whl"
|
| 36 |
+
echo "installed into $VENV (python $PYVER, interpreter under $HERE/.python)"
|
model.py
ADDED
|
@@ -0,0 +1,6 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Thin-bundle runner shim for Qwen3.8-27B on Blackhole (DFlash2 speculative decoding).
|
| 2 |
+
|
| 3 |
+
The model code itself ships in the `tt-metal-models` wheel (models.demos.blackhole.qwen36); this module only
|
| 4 |
+
re-exports the vLLM adapter class so `model:Qwen36DFlashForCausalLM` resolves with PYTHONPATH set to the bundle dir.
|
| 5 |
+
"""
|
| 6 |
+
from models.demos.blackhole.qwen36.tt.qwen36_vllm_dflash import Qwen36DFlashForCausalLM # noqa: F401
|
prepare_model_dir.py
ADDED
|
@@ -0,0 +1,83 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
#!/usr/bin/env python3
|
| 2 |
+
"""Build model-dir/ for this bundle before vLLM starts. run.sh runs it with the bundle's python.
|
| 3 |
+
|
| 4 |
+
Written by tt-orchard (stage 7) for a model whose weights are a fine-tune of a supported model.
|
| 5 |
+
The model code in this bundle is registered for the base model's architecture, so vLLM must see
|
| 6 |
+
the base model's config files. The tokenizer and weights must be the fine-tune's. This script
|
| 7 |
+
builds model-dir/ next to itself from both:
|
| 8 |
+
|
| 9 |
+
- COPIED from base_config/ (shipped in the bundle): every file there, for example config.json.
|
| 10 |
+
- LINKED from the fine-tune's Hugging Face snapshot (absolute links to the resolved files):
|
| 11 |
+
tokenizer.json, tokenizer_config.json, chat_template.jinja, generation_config.json,
|
| 12 |
+
model.safetensors.index.json and every *.safetensors file. tokenizer.json and at least one
|
| 13 |
+
*.safetensors file are required.
|
| 14 |
+
|
| 15 |
+
The snapshot is the manifest's `weights.repo_id` at `weights.revision` (which must be a pinned
|
| 16 |
+
40-character commit) under $HF_HUB_CACHE or $HF_HOME/hub (run.sh sets HF_HOME). When it is missing
|
| 17 |
+
and HF_HUB_OFFLINE is not set, it is downloaded with huggingface_hub from the bundle's venv. When it
|
| 18 |
+
is missing and HF_HUB_OFFLINE is set, the script exits 2 and names it.
|
| 19 |
+
|
| 20 |
+
model-dir/.weights records "<repo>@<revision>". A model-dir without that file was not built here,
|
| 21 |
+
and the script exits 2 without touching it. Exit 0 means model-dir is ready.
|
| 22 |
+
"""
|
| 23 |
+
from __future__ import annotations
|
| 24 |
+
|
| 25 |
+
import json
|
| 26 |
+
import os
|
| 27 |
+
import re
|
| 28 |
+
import shutil
|
| 29 |
+
import sys
|
| 30 |
+
from pathlib import Path
|
| 31 |
+
|
| 32 |
+
HERE = Path(__file__).resolve().parent
|
| 33 |
+
LINKED = ("tokenizer.json", "tokenizer_config.json", "chat_template.jinja", "generation_config.json",
|
| 34 |
+
"model.safetensors.index.json")
|
| 35 |
+
MARKER = ".weights"
|
| 36 |
+
|
| 37 |
+
|
| 38 |
+
def fail(message: str) -> None:
|
| 39 |
+
print(f"prepare_model_dir: {message}", file=sys.stderr)
|
| 40 |
+
sys.exit(2)
|
| 41 |
+
|
| 42 |
+
|
| 43 |
+
def snapshot(repo: str, rev: str) -> Path:
|
| 44 |
+
hub = os.environ.get("HF_HUB_CACHE") or os.path.join(
|
| 45 |
+
os.environ.get("HF_HOME") or os.path.expanduser("~/.cache/huggingface"), "hub")
|
| 46 |
+
org, name = repo.split("/", 1)
|
| 47 |
+
snap = Path(hub) / f"models--{org}--{name}" / "snapshots" / rev
|
| 48 |
+
if snap.is_dir():
|
| 49 |
+
return snap
|
| 50 |
+
if os.environ.get("HF_HUB_OFFLINE", "").lower() in ("1", "true", "yes", "on"):
|
| 51 |
+
fail(f"the weights snapshot {repo}@{rev} is not in {hub} and HF_HUB_OFFLINE is set. "
|
| 52 |
+
f"Download it first: hf download {repo} --revision {rev}")
|
| 53 |
+
from huggingface_hub import snapshot_download # in the bundle's venv
|
| 54 |
+
return Path(snapshot_download(repo_id=repo, revision=rev))
|
| 55 |
+
|
| 56 |
+
|
| 57 |
+
def main() -> int:
|
| 58 |
+
weights = json.loads((HERE / "tt_kernel_manifest.json").read_text(encoding="utf-8"))["weights"]
|
| 59 |
+
repo, rev = weights["repo_id"], weights.get("revision") or ""
|
| 60 |
+
if not re.fullmatch(r"[0-9a-f]{40}", rev):
|
| 61 |
+
fail(f"the manifest's weights revision {rev!r} is not a 40-character commit")
|
| 62 |
+
md = HERE / "model-dir"
|
| 63 |
+
if md.is_symlink() or (md.exists() and not (md / MARKER).is_file()):
|
| 64 |
+
fail(f"{md} exists and was not built by this script; move it aside")
|
| 65 |
+
snap = snapshot(repo, rev)
|
| 66 |
+
names = list(LINKED) + sorted(p.name for p in snap.glob("*.safetensors"))
|
| 67 |
+
present = [n for n in names if (snap / n).exists()]
|
| 68 |
+
if "tokenizer.json" not in present or not any(n.endswith(".safetensors") for n in present):
|
| 69 |
+
fail(f"{snap} needs tokenizer.json and at least one *.safetensors file")
|
| 70 |
+
if md.exists():
|
| 71 |
+
shutil.rmtree(md)
|
| 72 |
+
md.mkdir()
|
| 73 |
+
for f in sorted((HERE / "base_config").iterdir()):
|
| 74 |
+
shutil.copyfile(f, md / f.name)
|
| 75 |
+
for n in present:
|
| 76 |
+
os.symlink(os.path.realpath(snap / n), md / n)
|
| 77 |
+
(md / MARKER).write_text(f"{repo}@{rev}", encoding="utf-8")
|
| 78 |
+
print(f"prepare_model_dir: {md} holds the base config and {len(present)} files of {repo}@{rev}")
|
| 79 |
+
return 0
|
| 80 |
+
|
| 81 |
+
|
| 82 |
+
if __name__ == "__main__":
|
| 83 |
+
sys.exit(main())
|
requirements.txt
ADDED
|
@@ -0,0 +1,5 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
--extra-index-url https://download.pytorch.org/whl/cpu
|
| 2 |
+
ttnn==0.79.0.dev20260929+qwen36p150x2.g9f6b02d
|
| 3 |
+
tt-metal-models==0.79.0.dev20260929+qwen36p150x2.g9f6b02d
|
| 4 |
+
transformers==5.17.0
|
| 5 |
+
tokenizers==0.23.2
|
run.sh
ADDED
|
@@ -0,0 +1,87 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
#!/usr/bin/env bash
|
| 2 |
+
# Serve this model on TT hardware. Assumes ./install.sh has been run.
|
| 3 |
+
set -euo pipefail
|
| 4 |
+
HERE="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
| 5 |
+
VENV="${VENV:-$HERE/venv}"
|
| 6 |
+
PYBIN="$VENV/bin/python"
|
| 7 |
+
|
| 8 |
+
# Locate ttnn WITHOUT importing it — importing loads _ttnn.so, which is exactly what needs the
|
| 9 |
+
# LD_PRELOAD below (chicken-and-egg). find_spec resolves the path without executing the module.
|
| 10 |
+
TTNN_DIR="$("$PYBIN" -c 'import importlib.util,os;print(os.path.dirname(importlib.util.find_spec("ttnn").origin))')"
|
| 11 |
+
# _ttnncpp.so lives in ttnn.libs/ for an auditwheel-repaired (portable) wheel, or build/lib/ for a
|
| 12 |
+
# raw one; preload it to avoid the glibc "static TLS block" error on late dlopen.
|
| 13 |
+
# Prefer the auditwheel-vendored copy in *.libs/ (that's the one _ttnn.so actually loads via
|
| 14 |
+
# RPATH); fall back to build/lib for a raw (unrepaired) wheel — e.g. the plain `ttnn` PyPI wheel
|
| 15 |
+
# a v6 thin bundle installs, which isn't auditwheel-repaired at all.
|
| 16 |
+
# `|| true` on each probe: under `set -e -o pipefail`, a `ls <no-match> | head -1` pipe fails
|
| 17 |
+
# (pipefail surfaces ls's nonzero exit) and set -e would kill the script on THIS line, before the
|
| 18 |
+
# fallback below — or the deliberate `:?` error a few lines down — ever runs.
|
| 19 |
+
LD_PRELOAD="$(ls "$TTNN_DIR"/../*.libs/_ttnncpp*.so 2>/dev/null | head -1)" || true
|
| 20 |
+
[ -n "$LD_PRELOAD" ] || LD_PRELOAD="$(ls "$TTNN_DIR"/build/lib/_ttnncpp*.so 2>/dev/null | head -1)" || true
|
| 21 |
+
export LD_PRELOAD="${LD_PRELOAD:?could not locate _ttnncpp.so in the ttnn install}"
|
| 22 |
+
export TT_METAL_HOME="$TTNN_DIR"
|
| 23 |
+
# EXTRA_MODELS_DIR is a PARENT of per-model bundle folders; the plugin scans its children for
|
| 24 |
+
# each vllm_metadata.json (so the metadata lives in vllm_models/<model>/, not the bundle root).
|
| 25 |
+
export EXTRA_MODELS_DIR="$HERE/vllm_models"
|
| 26 |
+
export TT_VLLM_BUILTIN_MODELS=0
|
| 27 |
+
# Do NOT set VLLM_PLUGINS: it is an ALLOW-LIST — setting it suppresses the vllm.general_plugins
|
| 28 |
+
# group, so the model's tool/reasoning-parser overrides would silently not load. The TT platform
|
| 29 |
+
# + model registry load via entry points without it.
|
| 30 |
+
export PYTHONPATH="$HERE:${PYTHONPATH:-}" # resolves the adapter/model imports
|
| 31 |
+
export MESH_DEVICE="${MESH_DEVICE:-P150x2}"
|
| 32 |
+
NCHIPS=2
|
| 33 |
+
if [ -z "${TT_METAL_VISIBLE_DEVICES:-}" ]; then
|
| 34 |
+
if [ -n "${TT_VISIBLE_DEVICES:-}" ]; then
|
| 35 |
+
IFS=, read -ra _GRANT <<< "$TT_VISIBLE_DEVICES"
|
| 36 |
+
if [ "${#_GRANT[@]}" -lt "$NCHIPS" ]; then
|
| 37 |
+
echo "run.sh: this model needs $NCHIPS chip(s) but TT_VISIBLE_DEVICES grants ${#_GRANT[@]} ($TT_VISIBLE_DEVICES)" >&2
|
| 38 |
+
exit 1
|
| 39 |
+
fi
|
| 40 |
+
TT_VISIBLE_DEVICES="$(IFS=,; echo "${_GRANT[*]:0:$NCHIPS}")"
|
| 41 |
+
export TT_VISIBLE_DEVICES
|
| 42 |
+
TT_METAL_VISIBLE_DEVICES="0,1"
|
| 43 |
+
else
|
| 44 |
+
TT_METAL_VISIBLE_DEVICES="0,1"
|
| 45 |
+
fi
|
| 46 |
+
fi
|
| 47 |
+
export TT_METAL_VISIBLE_DEVICES
|
| 48 |
+
# HERMETIC RUNTIME: keep every cache/home INSIDE the folder wall, so serving writes and reads
|
| 49 |
+
# nothing outside it (the ttnn tensor cache even DEFAULTS to a hard-coded /mnt/... path upstream —
|
| 50 |
+
# a classic other-machine leak we must override). Each is overridable if the operator sets it.
|
| 51 |
+
export HF_HOME="${HF_HOME:-$HERE/.hf}" # HF weights + hub cache
|
| 52 |
+
export TT_CACHE_PATH="${TT_CACHE_PATH:-$HERE/.tt_cache}" # ttnn weight/tensor cache
|
| 53 |
+
export TT_CACHE_HOME="${TT_CACHE_HOME:-$HERE/.tt_cache}" # override upstream's /mnt/... default
|
| 54 |
+
export XDG_CACHE_HOME="${XDG_CACHE_HOME:-$HERE/.cache}" # generic catch-all (triton, etc.)
|
| 55 |
+
export TRITON_CACHE_DIR="${TRITON_CACHE_DIR:-$HERE/.cache/triton}"
|
| 56 |
+
export TORCHINDUCTOR_CACHE_DIR="${TORCHINDUCTOR_CACHE_DIR:-$HERE/.cache/inductor}"
|
| 57 |
+
export HF_MODEL="$HERE/model-dir"
|
| 58 |
+
export MODEL_WEIGHTS_DIR="$HERE/model-dir"
|
| 59 |
+
export TT_MODEL_WEIGHTS_REVISION="${TT_MODEL_WEIGHTS_REVISION:-1a5f363a3dd2d1cc456c28b8abbb403b9555efaf}"
|
| 60 |
+
export ARCH_NAME="blackhole"
|
| 61 |
+
export QWEN36_SKIP_VISION="1"
|
| 62 |
+
export QWEN_SDPA_BF8="1"
|
| 63 |
+
export TORCHDYNAMO_DISABLE="1"
|
| 64 |
+
export TT_QWEN35_TEXT_VER="qwen36_blackhole"
|
| 65 |
+
export VLLM_RPC_TIMEOUT="900000"
|
| 66 |
+
export VLLM_CONFIGURE_LOGGING="1"
|
| 67 |
+
export QWEN36_DRAFTER="dflash2"
|
| 68 |
+
export DFLASH_WEIGHTS="incoai/Qwen3.8-27B-DFlash2@dedf8df68adfb1afeaf7b7480c0a0243108177b4"
|
| 69 |
+
export QWEN36_DFLASH_TP="1"
|
| 70 |
+
export QWEN36_DFLASH_BLOCK="8"
|
| 71 |
+
export QWEN36_DFLASH_FOLD_SEED="1"
|
| 72 |
+
export QWEN36_DFLASH_SERVE_BLOCK="32"
|
| 73 |
+
export QWEN36_PREFILL_BUCKET_TRACE="0"
|
| 74 |
+
export TT_METAL_PINNED_MEMORY_CACHE_LIMIT_BYTES="0"
|
| 75 |
+
export QWEN36_GDN_SPEC_FUSED="1"
|
| 76 |
+
export QWEN36_MAX_TOKENS_ALL_USERS="262144"
|
| 77 |
+
CMD=("$PYBIN" -m vllm.entrypoints.openai.api_server --model "$HERE/model-dir" --max_num_seqs 4 --block_size 64 --max_model_len 262144 --additional-config '{"tt": {"l1_small_size": 24576, "fabric_config": "FABRIC_1D", "trace_region_size": 1073741824, "sample_on_device_mode": "decode_only"}}' --max-num-batched-tokens 65536 --enable-auto-tool-choice --tool-call-parser qwen3_coder --reasoning_parser qwen3 --no-async-scheduling "$@")
|
| 78 |
+
# TT_MODEL_PRINT=1 (set by `tt-model serve --print`) echoes the fully-resolved command+env
|
| 79 |
+
if [ "${TT_MODEL_PRINT:-0}" = "1" ]; then
|
| 80 |
+
printf 'LD_PRELOAD=%s TT_METAL_HOME=%s EXTRA_MODELS_DIR=%s MESH_DEVICE=%s HF_MODEL=%s
|
| 81 |
+
%s
|
| 82 |
+
' \
|
| 83 |
+
"$LD_PRELOAD" "$TT_METAL_HOME" "$EXTRA_MODELS_DIR" "$MESH_DEVICE" "${HF_MODEL:-}" "${CMD[*]}"
|
| 84 |
+
exit 0
|
| 85 |
+
fi
|
| 86 |
+
"$PYBIN" "$HERE/prepare_model_dir.py"
|
| 87 |
+
exec "${CMD[@]}"
|
tt_kernel_manifest.json
ADDED
|
@@ -0,0 +1,80 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"schema_version": "6",
|
| 3 |
+
"name": "hemmingway-1-p300",
|
| 4 |
+
"tt_metal_version": "0.79.0.dev20260929+qwen36p150x2.g9f6b02d",
|
| 5 |
+
"arch": "blackhole",
|
| 6 |
+
"device_count": 2,
|
| 7 |
+
"producer": {
|
| 8 |
+
"tt_kernel_version": "0.1.0",
|
| 9 |
+
"created_at": "2026-10-04T01:18:16.031583+00:00",
|
| 10 |
+
"hostname": "redacted"
|
| 11 |
+
},
|
| 12 |
+
"weights": {
|
| 13 |
+
"repo_id": "Altworld/Hemmingway-1",
|
| 14 |
+
"revision": "1a5f363a3dd2d1cc456c28b8abbb403b9555efaf",
|
| 15 |
+
"allow_patterns": null,
|
| 16 |
+
"ignore_patterns": null,
|
| 17 |
+
"repo_type": "model"
|
| 18 |
+
},
|
| 19 |
+
"mesh": {
|
| 20 |
+
"devices": 2,
|
| 21 |
+
"topology": "P150x2",
|
| 22 |
+
"fabric": null
|
| 23 |
+
},
|
| 24 |
+
"entrypoint": {
|
| 25 |
+
"cls": "models.demos.blackhole.qwen36.tt.qwen36_vllm_dflash:Qwen36DFlashForCausalLM",
|
| 26 |
+
"arch_name": "Qwen3_5ForConditionalGeneration"
|
| 27 |
+
},
|
| 28 |
+
"resources": {
|
| 29 |
+
"max_model_len": 262144,
|
| 30 |
+
"max_num_seqs": 4,
|
| 31 |
+
"block_size": 64,
|
| 32 |
+
"trace_region_bytes": null,
|
| 33 |
+
"extra_args": [],
|
| 34 |
+
"command_override": {}
|
| 35 |
+
},
|
| 36 |
+
"capabilities": null,
|
| 37 |
+
"env": {
|
| 38 |
+
"ARCH_NAME": "blackhole",
|
| 39 |
+
"QWEN36_SKIP_VISION": "1",
|
| 40 |
+
"QWEN_SDPA_BF8": "1",
|
| 41 |
+
"TORCHDYNAMO_DISABLE": "1",
|
| 42 |
+
"TT_QWEN35_TEXT_VER": "qwen36_blackhole",
|
| 43 |
+
"VLLM_RPC_TIMEOUT": "900000",
|
| 44 |
+
"VLLM_CONFIGURE_LOGGING": "1",
|
| 45 |
+
"QWEN36_DRAFTER": "dflash2",
|
| 46 |
+
"DFLASH_WEIGHTS": "incoai/Qwen3.8-27B-DFlash2@dedf8df68adfb1afeaf7b7480c0a0243108177b4",
|
| 47 |
+
"QWEN36_DFLASH_TP": "1",
|
| 48 |
+
"QWEN36_DFLASH_BLOCK": "8",
|
| 49 |
+
"QWEN36_DFLASH_FOLD_SEED": "1",
|
| 50 |
+
"QWEN36_DFLASH_SERVE_BLOCK": "32",
|
| 51 |
+
"QWEN36_PREFILL_BUCKET_TRACE": "0",
|
| 52 |
+
"TT_METAL_PINNED_MEMORY_CACHE_LIMIT_BYTES": "0",
|
| 53 |
+
"QWEN36_GDN_SPEC_FUSED": "1",
|
| 54 |
+
"QWEN36_MAX_TOKENS_ALL_USERS": "262144"
|
| 55 |
+
},
|
| 56 |
+
"bundled": null,
|
| 57 |
+
"deps": {
|
| 58 |
+
"python": "3.12",
|
| 59 |
+
"requirements": "requirements.txt",
|
| 60 |
+
"wheels": [
|
| 61 |
+
"wheels/vllm_tt_plugin-0.1.0-py3-none-any.whl"
|
| 62 |
+
],
|
| 63 |
+
"wheels_dir": "wheels",
|
| 64 |
+
"models_wheels": [
|
| 65 |
+
"wheels/ttnn-0.79.0.dev20260929+qwen36p150x2.g9f6b02d-cp312-cp312-linux_x86_64.whl",
|
| 66 |
+
"wheels/tt_metal_models-0.79.0.dev20260929+qwen36p150x2.g9f6b02d-py3-none-any.whl"
|
| 67 |
+
],
|
| 68 |
+
"vllm": {
|
| 69 |
+
"version": "0.26.0",
|
| 70 |
+
"target_device": "empty",
|
| 71 |
+
"overrides": "vllm-overrides.txt",
|
| 72 |
+
"common_requirements": null,
|
| 73 |
+
"wheel": null
|
| 74 |
+
},
|
| 75 |
+
"model_dir": ".",
|
| 76 |
+
"kind": "vllm",
|
| 77 |
+
"app": null
|
| 78 |
+
},
|
| 79 |
+
"container": null
|
| 80 |
+
}
|
vllm-overrides.txt
ADDED
|
@@ -0,0 +1,6 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# vLLM dependency overrides for the empty-target (TT) build — see tenstorrent/vllm-tt-plugin.
|
| 2 |
+
# ttnn needs numpy<2; vLLM's common.txt wants opencv-python-headless>=4.13 which needs numpy>=2.
|
| 3 |
+
# Pin opencv to the last numpy-1 release (its video path is unused by TT models) and hold numpy<2
|
| 4 |
+
# (fixed by the tt-metal/ttnn env this runs inside). Revisit on a vLLM bump or if a model gains video.
|
| 5 |
+
opencv-python-headless==4.11.0.86
|
| 6 |
+
numpy>=1.24.4,<2
|
vllm_models/hemmingway-1-p300/vllm_metadata.json
ADDED
|
@@ -0,0 +1,4 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"arch": "Qwen3_5ForConditionalGeneration",
|
| 3 |
+
"main_class": "models.demos.blackhole.qwen36.tt.qwen36_vllm_dflash:Qwen36DFlashForCausalLM"
|
| 4 |
+
}
|
wheels/tt_metal_models-0.79.0.dev20260929+qwen36p150x2.g9f6b02d-py3-none-any.whl
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:4528111a0415b9f5d0f5468b7645a1bf7d6d719cd4c5ff527fbc4434893b3390
|
| 3 |
+
size 9544741
|
wheels/ttnn-0.79.0.dev20260929+qwen36p150x2.g9f6b02d-cp312-cp312-linux_x86_64.whl
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:502ad84809ad51e70cfcab1f79372319ddbff1acf984ca9652cddf8bbd586cdd
|
| 3 |
+
size 89713784
|
wheels/vllm_tt_plugin-0.1.0-py3-none-any.whl
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:2fa641a876be2523d4cd7bd2f363643412e55060fc6b34b4d78ec42d0cd826ba
|
| 3 |
+
size 322371
|