hemmingway-1-p300
License: CC BY-NC 4.0. Non-commercial use only. The weights are by Altworld (https://huggingface.co/Altworld/Hemmingway-1). Commercial use needs a separate agreement with Altworld. This package inherits that license.
This is a tt-model v6 thin package that serves Altworld/Hemmingway-1 (revision 1a5f363a3dd2d1cc456c28b8abbb403b9555efaf) on 2 Tenstorrent Blackhole chips (one P300 board, or a P150x2 mesh). Hemmingway-1 (Altworld/Hemmingway-1) is a 27B-parameter open-weights text model from Altworld, built on Qwen/Qwen3.8-27B for everyday writing: messages, emails and short pieces. Its upstream card says it gives you the text itself, without a preamble or a list of options. The upstream card also reports comparisons with other models. We did not reproduce them.
The package was produced by an automated bring-up harness (tt-orchard), and a person reviewed the results before release.
How to serve
tt-model serve episod/hemmingway-1-p300
The server exposes an OpenAI-compatible endpoint (/v1/chat/completions and /v1/completions). Context length is 262144 tokens with up to 4 sequences at a time. Speculative decoding uses the drafter incoai/Qwen3.8-27B-DFlash2 (K=7 draft tokens per step), which was trained on base Qwen3.8-27B. It was not trained on this model.
Notes that matter before you start:
- This package ships no weights.
tt-modeldownloads them from the upstream repositoryAltworld/Hemmingway-1with your own Hugging Face account. - Only greedy decoding is supported. Requests that ask for logprobs are rejected.
- The server needs a fresh, empty tensor cache for each model. Do not share a tensor cache between Qwen3.8-based models: the cache is keyed by layer name, so a cache built for another Qwen3.8-27B model would load that model's weights under this model's name.
- The first boot compiles kernels. Boot takes a long time (see the table below).
Turn thinking off for direct answers. Send "chat_template_kwargs": {"enable_thinking": false} with each chat request. With thinking on (the server default), the model spends its token budget on reasoning and may return no visible answer. In one test with thinking on, a creative prompt used all 700 tokens on reasoning and content was null.
Evidence status of this card
All numbers below were measured on 2026-10-04 on one box, on this package installed in its own virtual environment, with the weights and tensor cache for this model only. Each is labelled measured and gives its conditions.
Measured results (2 chips)
Conditions: greedy, thinking off, tensor cache already built for this model. Speed workload: the 80 first-turn coding prompts of nvidia/SPEED-Bench, output cap 2048 tokens (6 of 80 hit the cap). Per-user tok/s = 1000 / mean time per output token. Short answers distort that mean, so a token-weighted figure is also given. One run per cell.
| Item | Value |
|---|---|
Boot, launch to /health OK, tensor cache present |
2025 s (33.8 min), of which 1261.7 s was the drafter and verify warm-up (alloc and compile). The verification install measured 2034.4 s. |
| Coding, 1 user, per-user speed | 54.1 tok/s (mean 18.50 ms per token, median 13.61 ms) |
| Coding, 1 user, token-weighted | 61.8 tok/s over 80 requests; 69.1 tok/s over the 71 requests with at least 64 output tokens |
| Coding, 1 user, TTFT | mean 232 ms, median 185 ms |
| Coding, 4 users, per-user speed | 28.5 tok/s (mean 35.13 ms per token, median 23.89 ms) |
| Coding, 4 users, token-weighted | 42.2 tok/s; 42.0 tok/s over the 71 requests of at least 64 tokens |
| Coding, 4 users, TTFT | mean 749 ms, median 728 ms |
| Coding, aggregate | 59.3 tok/s at 1 user (24,941 output tokens in 420.5 s), 130.4 tok/s at 4 users (191.3 s) |
Context sweep (vllm bench serve, random prompts, ignore EOS, 128 output tokens, 4 to 8 requests per cell):
| Prompt tokens / users | Per user incl. TTFT | TTFT | Time per token |
|---|---|---|---|
| 1024 / 1 | 51.5 tok/s | 318 ms | 17.1 ms (58 tok/s decode) |
| 16384 / 1 | 21.8 tok/s | 3788 ms | 16.3 ms (61 tok/s decode) |
| 1024 / 4 | 30.2 tok/s (120.9 aggregate) | 969 ms | 25.7 ms |
Draft acceptance
The drafter was trained on base Qwen3.8-27B. Acceptance is the mean number of draft tokens accepted per step out of 7, weighted by steps.
| Workload | Accepted of 7 | Tokens committed per step |
|---|---|---|
| 80 coding prompts (4,644 steps) | 4.38 (62.6%) | 5.36 |
| 10 everyday-writing prompts (356 steps) | 1.72 (24.6%) | 2.69 |
Acceptance on writing prompts is much lower than on code. Speed on writing prompts was measured only as part of a 10-prompt run and not at scale. Expect lower speed on prose than the coding numbers suggest. No comparison with the base model's acceptance was run, so we cannot say whether the fine-tune lowered acceptance.
Long context retrieval
One passkey in the middle of filler text, one trial per length, greedy. All were retrieved: 1063, 2075, 2944, 15783 (8.0 s) and 135567 tokens (44.5 s). The key was a 6-digit number in synthetic text. This is not a long-context quality evaluation.
Agreement with the CPU bf16 reference
The chip generated 48 tokens for each of 8 prompts, and the CPU bf16 model scored each chip token given the chip's own prefix. The chip token matched at 365 of 384 positions (95.1 percent). Per prompt: 45, 42, 45, 47, 47, 46, 46, 47 of 48. The default chat template was used, so the tokens are mostly the model's thinking text. The 19 mismatches were not examined beyond their first tokens. An earlier single-prompt check gave 30 of 32 (0.9375). That is thin evidence, and the 8-prompt figure is the better one.
Writing quality check
10 prompts we wrote for everyday messages and short writing (a note to a landlord, a raise request, declining a wedding invitation, a condolence message, a story opening and others), greedy, thinking off, 600-token cap. Result: 7 good, 3 partly, 0 poor. The grades were given by the AI model that ran the check, not by humans, on one answer per prompt. They support the claim that it writes the requested message directly, with no preamble or options. They say nothing about quality compared with another model, because no other model was run. The three "partly" answers had a literal placeholder ("[Manager's Name]"), an invented name ("Sarah") or a vague request.
Sample
Prompt: "You are a building remembering your past. You were built in the 1930s. What do you remember?" Settings: greedy, thinking off, max_tokens 900, one sample, finished at a stop token after 830 tokens. This is not a quality measurement. Excerpt:
I remember the first family. The Kowalskis. Mrs. Kowalski hung curtains in the front window within a week of moving in, and I felt the weight of them against my glass like a small, warm hand.
Pulled from the Hub and booted
On 2026-10-04 this repository was pulled with tt-model pull while it was private, which built a new virtual environment from the package's wheels and requirements. It was then started with tt-model serve on one Blackhole P300 board with an empty tensor cache, so the weights were converted on that boot. The server was ready after about 35 minutes (launch to /health OK, one boot, started at the same time as a second model's boot on the other board). It answered "What is 7 times 6?" with 42 and wrote a correct Python Fibonacci function. The weights were read from a local copy of the upstream repository, so a download from the upstream repository was not part of this test.
What was not measured
- Accuracy benchmarks (MMLU and similar). None were run.
- A comparison with the base model Qwen3.8-27B, or with the unquantized upstream model on a GPU. None was run, so these cards make no claim that this package matches or beats either.
- A 4-chip package. None is published here.
- A fresh download of the weights from the upstream repository. The Hub test used a local copy of the weights.
- Speed numbers use a smaller prompt subset and a lower token cap than the earlier Qwen3.8-27B DFlash2 card, so they are not comparable with that card.
- Sampling quality. The server accepts greedy decoding only and rejects logprobs.
- Boot with a cold kernel cache on a fresh machine. Warm boots in these measurements took about 28 to 34 minutes, so the first boot on a new machine can be expected to take at least that long, but it was not timed.
Evidence
The measurements come from the run's benchmark summary and its quality-check write-up (RESULTS-bench.md and qualitative.md), with raw result files kept by the person who ran them. The risks and open items are in the run's risk list.