Instructions to use infosave/Hy-MT2-cmf with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- cortiq
How to use infosave/Hy-MT2-cmf with cortiq:
# one Rust binary, no additional dependencies cargo install cortiq-cli # or a prebuilt binary from github.com/infosave2007/cmf/releases hf download infosave/Hy-MT2-cmf --include "*.cmf" --local-dir . ls *.cmf # some repos ship more than one quantization
cortiq run FILE.cmf --prompt "What is the capital of France?"
cortiq serve FILE.cmf --port 8080 # OpenAI-compatible server
- Notebooks
- Google Colab
- Kaggle
0.7.1: 30B prompt ingest through the resident graph; Metal experts on the windowed arena; VRAM ladder
Browse files
README.md
CHANGED
|
@@ -177,7 +177,7 @@ tokenizes exactly like the 4-bit ones.
|
|
| 177 |
## Measured
|
| 178 |
|
| 179 |
Steady-state decode, single stream, `cortiq bench --core --tokens 128
|
| 180 |
-
--ignore-eos`, cortiq 0.7.
|
| 181 |
discrete card (a 1 GB model needs ~8 submits per token), so the CPU
|
| 182 |
matters as much as the GPU there; the MoE row is where the card counts.
|
| 183 |
|
|
@@ -186,7 +186,7 @@ matters as much as the GPU there; the MoE row is where the card counts.
|
|
| 186 |
| `hy-mt2-1.8b-q4tp.cmf` | 137.7 | 29.1 | 72.8 | 56.9 |
|
| 187 |
| `hy-mt2-1.8b-q1t.cmf` | 63.7 | 21.9 | 66.9 | 38.1 |
|
| 188 |
| `hy-mt2-7b-q4tp.cmf` | 77.2 | 9.3 | 23.1 | 16.7 |
|
| 189 |
-
| `hy-mt2-30b-a3b-q4tp.cmf` |
|
| 190 |
|
| 191 |
### The MoE on any card (dynamic loading)
|
| 192 |
|
|
@@ -198,25 +198,36 @@ auto-detect (the CPU alone: 10.5 tok/s on this 14-core Xeon):
|
|
| 198 |
|
| 199 |
| VRAM budget | 4 GB | 6 GB | 8 GB | 12 GB | 16 GB | 24 GB |
|
| 200 |
|---|---:|---:|---:|---:|---:|---:|
|
| 201 |
-
| layers on the card | 7/48 | 14/48 | 20/48 | 33/48 |
|
| 202 |
-
| decode, tok/s | 9.
|
| 203 |
|
| 204 |
Perplexity and the greedy continuation do not move along the ladder β the
|
| 205 |
split changes where a layer runs, never what it computes. `CMF_GPU_VRAM_MB`
|
| 206 |
-
overrides the auto-detected budget when you want to cap it by hand.
|
| 207 |
-
|
| 208 |
-
|
| 209 |
-
|
| 210 |
-
|
| 211 |
-
|
| 212 |
-
|
| 213 |
-
|
| 214 |
-
|
| 215 |
-
|
| 216 |
-
|
| 217 |
-
|
| 218 |
-
|
| 219 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 220 |
|
| 221 |
**First answer vs. the rest.** On a discrete card the weights are uploaded
|
| 222 |
when the whole-token graph is first built β 18.5 GB for the 30B, ~26 s on
|
|
|
|
| 177 |
## Measured
|
| 178 |
|
| 179 |
Steady-state decode, single stream, `cortiq bench --core --tokens 128
|
| 180 |
+
--ignore-eos`, cortiq 0.7.1. The dense files are latency-bound on a
|
| 181 |
discrete card (a 1 GB model needs ~8 submits per token), so the CPU
|
| 182 |
matters as much as the GPU there; the MoE row is where the card counts.
|
| 183 |
|
|
|
|
| 186 |
| `hy-mt2-1.8b-q4tp.cmf` | 137.7 | 29.1 | 72.8 | 56.9 |
|
| 187 |
| `hy-mt2-1.8b-q1t.cmf` | 63.7 | 21.9 | 66.9 | 38.1 |
|
| 188 |
| `hy-mt2-7b-q4tp.cmf` | 77.2 | 9.3 | 23.1 | 16.7 |
|
| 189 |
+
| `hy-mt2-30b-a3b-q4tp.cmf` | 57.4 | 11.0 | 32.5β΅ | 22.7 |
|
| 190 |
|
| 191 |
### The MoE on any card (dynamic loading)
|
| 192 |
|
|
|
|
| 198 |
|
| 199 |
| VRAM budget | 4 GB | 6 GB | 8 GB | 12 GB | 16 GB | 24 GB |
|
| 200 |
|---|---:|---:|---:|---:|---:|---:|
|
| 201 |
+
| layers on the card | 7/48 | 14/48 | 20/48 | 33/48 | **48/48** | **48/48** |
|
| 202 |
+
| decode, tok/s | 9.0 | 12.5 | 14.2 | 19.6 | 51.9 | **57.4** |
|
| 203 |
|
| 204 |
Perplexity and the greedy continuation do not move along the ladder β the
|
| 205 |
split changes where a layer runs, never what it computes. `CMF_GPU_VRAM_MB`
|
| 206 |
+
overrides the auto-detected budget when you want to cap it by hand. The
|
| 207 |
+
16 GB point held 45 layers in 0.7.0; with the prompt on the graph (0.7.1)
|
| 208 |
+
the per-op prefill arena no longer competes for the budget and the whole
|
| 209 |
+
stack fits (16.8 GB resident), so a 16 GB card decodes at the full rate.
|
| 210 |
+
|
| 211 |
+
β΅ The 15.8 GB file is larger than one Metal buffer (13.6 GB on a 24 GB
|
| 212 |
+
M4), so its weights map as two overlapping windows; 0.7.1 taught the
|
| 213 |
+
expert kernels to address them (before, every MoE layer of a windowed
|
| 214 |
+
file ran on the CPU: 17.9 tok/s, under the M4's own CPU). The experts themselves stream at the
|
| 215 |
+
card's bandwidth (~12 ms of the 31 ms token); the rest is the fixed cost
|
| 216 |
+
of the ~12 dispatches each of the 48 layers needs, which is why the
|
| 217 |
+
dense 7B, with a third of the layers' bytes per layer, is not faster.
|
| 218 |
+
|
| 219 |
+
Prompt ingest: on the RTX PRO 4000 the dense files take 180 (1.8B) and
|
| 220 |
+
118 (7B) tok/s. The 30B ingests at **65 tok/s** (41-token prompt) and
|
| 221 |
+
56 tok/s at 512 tokens β through the same resident graph that decodes,
|
| 222 |
+
one position at a time; `CMF_BATCH_K=32` switches the prompt to the
|
| 223 |
+
batched graph (32 positions per submit) for 80 / 71 tok/s. In 0.7.0 the
|
| 224 |
+
30B's prompt went through the chunked host prefill, where every expert
|
| 225 |
+
ran on the CPU: 8 tok/s, ten seconds before the first token of a
|
| 226 |
+
paragraph. The graph route needs the whole stack on the card; on the
|
| 227 |
+
VRAM ladder below the chunked path stays (7β12 tok/s of ingest), because
|
| 228 |
+
walking a prompt through a device prefix finishes every position on the
|
| 229 |
+
host. On the M4: 469 tok/s for the 1.8B q4tp, 218 for the ternary file,
|
| 230 |
+
122 for the 7B, 74 for the 30B.
|
| 231 |
|
| 232 |
**First answer vs. the rest.** On a discrete card the weights are uploaded
|
| 233 |
when the whole-token graph is first built β 18.5 GB for the 30B, ~26 s on
|