episod commited on
Commit
aaf0ef5
·
verified ·
1 Parent(s): 278cf24

Card: Hub pull and boot check, measured results

Browse files
Files changed (1) hide show
  1. README.md +5 -1
README.md CHANGED
@@ -91,12 +91,16 @@ Prompt: "You are a building remembering your past. You were built in the 1930s.
91
 
92
  > I remember the first family. The Kowalskis. Mrs. Kowalski hung curtains in the front window within a week of moving in, and I felt the weight of them against my glass like a small, warm hand.
93
 
 
 
 
 
94
  ## What was not measured
95
 
96
  - Accuracy benchmarks (MMLU and similar). None were run.
97
  - A comparison with the base model Qwen3.8-27B, or with the unquantized upstream model on a GPU. None was run, so these cards make no claim that this package matches or beats either.
98
  - A 4-chip package. None is published here.
99
- - A fresh-machine pull test. Pulling this package from the Hub and booting it on a clean machine has not been done yet.
100
  - Speed numbers use a smaller prompt subset and a lower token cap than the earlier Qwen3.8-27B DFlash2 card, so they are not comparable with that card.
101
  - Sampling quality. The server accepts greedy decoding only and rejects logprobs.
102
  - Boot with a cold kernel cache on a fresh machine. Warm boots in these measurements took about 28 to 34 minutes, so the first boot on a new machine can be expected to take at least that long, but it was not timed.
 
91
 
92
  > I remember the first family. The Kowalskis. Mrs. Kowalski hung curtains in the front window within a week of moving in, and I felt the weight of them against my glass like a small, warm hand.
93
 
94
+ ## Pulled from the Hub and booted
95
+
96
+ On 2026-10-04 this repository was pulled with `tt-model pull` while it was private, which built a new virtual environment from the package's wheels and requirements. It was then started with `tt-model serve` on one Blackhole P300 board with an empty tensor cache, so the weights were converted on that boot. The server was ready after about 35 minutes (launch to `/health` OK, one boot, started at the same time as a second model's boot on the other board). It answered "What is 7 times 6?" with 42 and wrote a correct Python Fibonacci function. The weights were read from a local copy of the upstream repository, so a download from the upstream repository was not part of this test.
97
+
98
  ## What was not measured
99
 
100
  - Accuracy benchmarks (MMLU and similar). None were run.
101
  - A comparison with the base model Qwen3.8-27B, or with the unquantized upstream model on a GPU. None was run, so these cards make no claim that this package matches or beats either.
102
  - A 4-chip package. None is published here.
103
+ - A fresh download of the weights from the upstream repository. The Hub test used a local copy of the weights.
104
  - Speed numbers use a smaller prompt subset and a lower token cap than the earlier Qwen3.8-27B DFlash2 card, so they are not comparable with that card.
105
  - Sampling quality. The server accepts greedy decoding only and rejects logprobs.
106
  - Boot with a cold kernel cache on a fresh machine. Warm boots in these measurements took about 28 to 34 minutes, so the first boot on a new machine can be expected to take at least that long, but it was not timed.