infosave commited on
Commit
1a0eec0
Β·
verified Β·
1 Parent(s): 5f7b65c

0.7.1: 30B prompt ingest through the resident graph; Metal experts on the windowed arena; VRAM ladder

Browse files
Files changed (1) hide show
  1. README.md +29 -18
README.md CHANGED
@@ -177,7 +177,7 @@ tokenizes exactly like the 4-bit ones.
177
  ## Measured
178
 
179
  Steady-state decode, single stream, `cortiq bench --core --tokens 128
180
- --ignore-eos`, cortiq 0.7.0. The dense files are latency-bound on a
181
  discrete card (a 1 GB model needs ~8 submits per token), so the CPU
182
  matters as much as the GPU there; the MoE row is where the card counts.
183
 
@@ -186,7 +186,7 @@ matters as much as the GPU there; the MoE row is where the card counts.
186
  | `hy-mt2-1.8b-q4tp.cmf` | 137.7 | 29.1 | 72.8 | 56.9 |
187
  | `hy-mt2-1.8b-q1t.cmf` | 63.7 | 21.9 | 66.9 | 38.1 |
188
  | `hy-mt2-7b-q4tp.cmf` | 77.2 | 9.3 | 23.1 | 16.7 |
189
- | `hy-mt2-30b-a3b-q4tp.cmf` | 52.7 | 11.0 | —⁡ | β€” |
190
 
191
  ### The MoE on any card (dynamic loading)
192
 
@@ -198,25 +198,36 @@ auto-detect (the CPU alone: 10.5 tok/s on this 14-core Xeon):
198
 
199
  | VRAM budget | 4 GB | 6 GB | 8 GB | 12 GB | 16 GB | 24 GB |
200
  |---|---:|---:|---:|---:|---:|---:|
201
- | layers on the card | 7/48 | 14/48 | 20/48 | 33/48 | 45/48 | **48/48** |
202
- | decode, tok/s | 9.9 | 11.1 | 14.5 | 21.9 | 39.2 | **53.7** |
203
 
204
  Perplexity and the greedy continuation do not move along the ladder β€” the
205
  split changes where a layer runs, never what it computes. `CMF_GPU_VRAM_MB`
206
- overrides the auto-detected budget when you want to cap it by hand.
207
-
208
- ⁡ Not measured: the 15.8 GB file is at the edge of a 24 GB Mac (the 14.3 GB
209
- Qwen3.8-27B decodes at 5.7 tok/s there) and on Metal a sigmoid-routed,
210
- ungated-shared MoE layer still runs its experts on the CPU β€” the Metal
211
- select kernel is the next port.
212
-
213
- Prompt ingest (41-token prompt): on the RTX PRO 4000 the dense files take
214
- 180 (1.8B) and 118 (7B) tok/s; the 30B ingests at **8 tok/s** on this card
215
- β€” its batched prefill re-stages the expert buffers per 32-token chunk
216
- instead of sharing the decode graph's resident copy, so a long source
217
- paragraph costs seconds before the first token. Decode is unaffected;
218
- sharing the buffers is the next item on the MoE list. On the M4: 469 tok/s
219
- for the 1.8B q4tp, 218 for the ternary file, 122 for the 7B.
 
 
 
 
 
 
 
 
 
 
 
220
 
221
  **First answer vs. the rest.** On a discrete card the weights are uploaded
222
  when the whole-token graph is first built β€” 18.5 GB for the 30B, ~26 s on
 
177
  ## Measured
178
 
179
  Steady-state decode, single stream, `cortiq bench --core --tokens 128
180
+ --ignore-eos`, cortiq 0.7.1. The dense files are latency-bound on a
181
  discrete card (a 1 GB model needs ~8 submits per token), so the CPU
182
  matters as much as the GPU there; the MoE row is where the card counts.
183
 
 
186
  | `hy-mt2-1.8b-q4tp.cmf` | 137.7 | 29.1 | 72.8 | 56.9 |
187
  | `hy-mt2-1.8b-q1t.cmf` | 63.7 | 21.9 | 66.9 | 38.1 |
188
  | `hy-mt2-7b-q4tp.cmf` | 77.2 | 9.3 | 23.1 | 16.7 |
189
+ | `hy-mt2-30b-a3b-q4tp.cmf` | 57.4 | 11.0 | 32.5⁡ | 22.7 |
190
 
191
  ### The MoE on any card (dynamic loading)
192
 
 
198
 
199
  | VRAM budget | 4 GB | 6 GB | 8 GB | 12 GB | 16 GB | 24 GB |
200
  |---|---:|---:|---:|---:|---:|---:|
201
+ | layers on the card | 7/48 | 14/48 | 20/48 | 33/48 | **48/48** | **48/48** |
202
+ | decode, tok/s | 9.0 | 12.5 | 14.2 | 19.6 | 51.9 | **57.4** |
203
 
204
  Perplexity and the greedy continuation do not move along the ladder β€” the
205
  split changes where a layer runs, never what it computes. `CMF_GPU_VRAM_MB`
206
+ overrides the auto-detected budget when you want to cap it by hand. The
207
+ 16 GB point held 45 layers in 0.7.0; with the prompt on the graph (0.7.1)
208
+ the per-op prefill arena no longer competes for the budget and the whole
209
+ stack fits (16.8 GB resident), so a 16 GB card decodes at the full rate.
210
+
211
+ ⁡ The 15.8 GB file is larger than one Metal buffer (13.6 GB on a 24 GB
212
+ M4), so its weights map as two overlapping windows; 0.7.1 taught the
213
+ expert kernels to address them (before, every MoE layer of a windowed
214
+ file ran on the CPU: 17.9 tok/s, under the M4's own CPU). The experts themselves stream at the
215
+ card's bandwidth (~12 ms of the 31 ms token); the rest is the fixed cost
216
+ of the ~12 dispatches each of the 48 layers needs, which is why the
217
+ dense 7B, with a third of the layers' bytes per layer, is not faster.
218
+
219
+ Prompt ingest: on the RTX PRO 4000 the dense files take 180 (1.8B) and
220
+ 118 (7B) tok/s. The 30B ingests at **65 tok/s** (41-token prompt) and
221
+ 56 tok/s at 512 tokens β€” through the same resident graph that decodes,
222
+ one position at a time; `CMF_BATCH_K=32` switches the prompt to the
223
+ batched graph (32 positions per submit) for 80 / 71 tok/s. In 0.7.0 the
224
+ 30B's prompt went through the chunked host prefill, where every expert
225
+ ran on the CPU: 8 tok/s, ten seconds before the first token of a
226
+ paragraph. The graph route needs the whole stack on the card; on the
227
+ VRAM ladder below the chunked path stays (7–12 tok/s of ingest), because
228
+ walking a prompt through a device prefix finishes every position on the
229
+ host. On the M4: 469 tok/s for the 1.8B q4tp, 218 for the ternary file,
230
+ 122 for the 7B, 74 for the 30B.
231
 
232
  **First answer vs. the rest.** On a discrete card the weights are uploaded
233
  when the whole-token graph is first built β€” 18.5 GB for the 30B, ~26 s on