Instructions to use infosave/MiniMax-H3-Turbo-cmf with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- cortiq
How to use infosave/MiniMax-H3-Turbo-cmf with cortiq:
# one Rust binary, no additional dependencies cargo install cortiq-cli # or a prebuilt binary from github.com/infosave2007/cmf/releases hf download infosave/MiniMax-H3-Turbo-cmf --include "*.cmf" --local-dir . ls *.cmf # some repos ship more than one quantization
cortiq animate FILE.cmf --prompt "a corgi in a chef hat flipping a pancake" --out clip.avi
- Notebooks
- Google Colab
- Kaggle
Upload README.md with huggingface_hub
Browse files
README.md
CHANGED
|
@@ -58,56 +58,143 @@ and reference images, videos and audio (`ref2va`); those paths are not ported
|
|
| 58 |
and the vision tower is not packed. What is here is `t2va`: prompt in, video
|
| 59 |
and audio out.
|
| 60 |
|
| 61 |
-
##
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 62 |
|
| 63 |
```bash
|
| 64 |
-
|
| 65 |
-
|
| 66 |
-
cat parts-q4tp/part_* > mmh3-turbo-q4tp.cmf
|
| 67 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 68 |
cortiq animate mmh3-turbo-q4tp.cmf \
|
| 69 |
--prompt "A corgi in a chef hat flipping a pancake, sizzling sounds and a cheerful bark." \
|
| 70 |
-
--width 512 --height 288 --frames 39 --
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 71 |
```
|
| 72 |
|
| 73 |
-
|
| 74 |
-
|
| 75 |
-
|
| 76 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 77 |
|
| 78 |
-
|
| 79 |
-
|
| 80 |
-
|
| 81 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 82 |
|
| 83 |
## What it costs to run
|
| 84 |
|
| 85 |
The file is memory-mapped, so plan on RAM at least its size or every step
|
| 86 |
touches non-resident pages.
|
| 87 |
|
| 88 |
-
|
|
|
|
| 89 |
|
| 90 |
-
| |
|
| 91 |
-
|---|---|---|---|
|
| 92 |
-
|
|
| 93 |
-
|
|
| 94 |
|
| 95 |
-
|
| 96 |
|
| 97 |
-
|
| 98 |
-
|
|
|
|
| 99 |
512Γ288 second is five times the tokens of a 256Γ160 one.
|
| 100 |
|
| 101 |
-
**
|
| 102 |
-
|
| 103 |
-
|
| 104 |
-
|
| 105 |
-
|
| 106 |
-
|
| 107 |
-
|
| 108 |
-
|
| 109 |
-
|
| 110 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 111 |
|
| 112 |
## What the conversion did
|
| 113 |
|
|
|
|
| 58 |
and the vision tower is not packed. What is here is `t2va`: prompt in, video
|
| 59 |
and audio out.
|
| 60 |
|
| 61 |
+
## Running it
|
| 62 |
+
|
| 63 |
+
### 1. Get the runtime
|
| 64 |
+
|
| 65 |
+
`cortiq` is one Rust binary. Either install it β
|
| 66 |
+
|
| 67 |
+
```bash
|
| 68 |
+
cargo install cortiq-cli # needs Rust 1.85+; brings the GPU backend
|
| 69 |
+
```
|
| 70 |
+
|
| 71 |
+
β or take a prebuilt archive from the
|
| 72 |
+
[latest release](https://github.com/infosave2007/cmf/releases/latest)
|
| 73 |
+
(Linux x86-64, macOS on Apple Silicon and Intel, Windows x86-64 and ARM64;
|
| 74 |
+
each ships a `.sha256`). Nothing else is required: no Python, no PyTorch, no
|
| 75 |
+
CUDA toolkit, no ffmpeg.
|
| 76 |
+
|
| 77 |
+
Check it took:
|
| 78 |
|
| 79 |
```bash
|
| 80 |
+
cortiq --version
|
| 81 |
+
```
|
|
|
|
| 82 |
|
| 83 |
+
### 2. Get the weights
|
| 84 |
+
|
| 85 |
+
One file, 23.5 GB.
|
| 86 |
+
|
| 87 |
+
```bash
|
| 88 |
+
pip install -U "huggingface_hub[cli]" # only to fetch the file
|
| 89 |
+
hf download infosave/MiniMax-H3-Turbo-cmf mmh3-turbo-q4tp.cmf --local-dir .
|
| 90 |
+
```
|
| 91 |
+
|
| 92 |
+
Confirm it arrived whole β the container carries a hash per tensor:
|
| 93 |
+
|
| 94 |
+
```bash
|
| 95 |
+
cortiq verify mmh3-turbo-q4tp.cmf # β β all tensor hashes match
|
| 96 |
+
cortiq info mmh3-turbo-q4tp.cmf # β arch, layers, 47.83B params
|
| 97 |
+
```
|
| 98 |
+
|
| 99 |
+
### 3. Render
|
| 100 |
+
|
| 101 |
+
```bash
|
| 102 |
cortiq animate mmh3-turbo-q4tp.cmf \
|
| 103 |
--prompt "A corgi in a chef hat flipping a pancake, sizzling sounds and a cheerful bark." \
|
| 104 |
+
--width 512 --height 288 --frames 39 --steps 4 --seed 42 \
|
| 105 |
+
--out corgi.avi
|
| 106 |
+
```
|
| 107 |
+
|
| 108 |
+
That writes `corgi.avi` β MJPEG video with PCM stereo, playable in VLC, mpv,
|
| 109 |
+
QuickTime and Windows Media Player β and `corgi.wav` beside it. The JPEG
|
| 110 |
+
encoder and the RIFF muxer are inside the binary: a pipeline that ends in a
|
| 111 |
+
shell-out to a 20 MB dependency is not a pipeline you can ship. If you want an
|
| 112 |
+
mp4 for a browser, remux it yourself; the model never needs one.
|
| 113 |
+
|
| 114 |
+
**On a GPU.** The device path is opt-in for this model while its kernels earn
|
| 115 |
+
their keep (see below):
|
| 116 |
+
|
| 117 |
+
```bash
|
| 118 |
+
CMF_MMH3_GPU=1 cortiq animate mmh3-turbo-q4tp.cmf --prompt "β¦" --out corgi.avi
|
| 119 |
```
|
| 120 |
|
| 121 |
+
### Options that matter
|
| 122 |
+
|
| 123 |
+
| flag | default | what it does |
|
| 124 |
+
|---|---|---|
|
| 125 |
+
| `--width` / `--height` | 512 Γ 288 | multiples of 32. The trained short edge is 768; below ~256 the model drifts off-distribution |
|
| 126 |
+
| `--frames` | 39 | at 24 fps, snapped **up** to the model's 17k+5 grid: 5, 22, 39, 56, β¦ 124. 124 β 5 s, and 124β362 is the validated range |
|
| 127 |
+
| `--steps` | 4 | what the Turbo LoRA is trained for. More still helps a little |
|
| 128 |
+
| `--seed` | 42 | same seed, same prompt, same size β the same clip, byte for byte |
|
| 129 |
+
| `--quality` | 92 | JPEG quality of the AVI's frames |
|
| 130 |
+
| `--stock-sampler` | off | integrate the audio on the video's clock, as a single-schedule sampler does. Wrong at 4 steps β it is here to hear how wrong |
|
| 131 |
|
| 132 |
+
| environment | what it does |
|
| 133 |
+
|---|---|
|
| 134 |
+
| `CMF_MMH3_GPU=1` | opt into the device path |
|
| 135 |
+
| `CMF_THREADS=n` | cap the worker pool (defaults to the machine's cores) |
|
| 136 |
+
| `CMF_ANIM_PROF=1` | per-step rms of both latent streams and both velocities |
|
| 137 |
+
|
| 138 |
+
### What it needs
|
| 139 |
+
|
| 140 |
+
RAM at least the file's size β 24 GB β or every step faults on non-resident
|
| 141 |
+
pages; the weights are memory-mapped, not read. Disk: 24 GB. A GPU is optional
|
| 142 |
+
and wants ~14 GB of VRAM for the DiT's planes. No network access at run time.
|
| 143 |
|
| 144 |
## What it costs to run
|
| 145 |
|
| 146 |
The file is memory-mapped, so plan on RAM at least its size or every step
|
| 147 |
touches non-resident pages.
|
| 148 |
|
| 149 |
+
512Γ288, 39 frames, 4 steps, one machine β 48 CPU cores and one RTX PRO 6000
|
| 150 |
+
Blackwell:
|
| 151 |
|
| 152 |
+
| | denoise | decode | total |
|
| 153 |
+
|---|---|---|---|
|
| 154 |
+
| host | 198.2 s | 147.8 s | **346.5 s** |
|
| 155 |
+
| `CMF_MMH3_GPU=1` | 96.8 s | 74.7 s | **172.0 s** |
|
| 156 |
|
| 157 |
+
and the smaller size, on the device: 256Γ160 over 22 frames in **29.2 s**.
|
| 158 |
|
| 159 |
+
Nearly all of the decode is the video VAE β the vocoder is 4 s of it. The
|
| 160 |
+
packed sequence is `[text | audio | video]` and everything attends to
|
| 161 |
+
everything, so cost grows with the token count and then with its square: a
|
| 162 |
512Γ288 second is five times the tokens of a 256Γ160 one.
|
| 163 |
|
| 164 |
+
**A free 2Γ on the decoder, if you want it.** The video VAE decodes in
|
| 165 |
+
256-pixel tiles, always, and grows the OVERLAP rather than the tile count β
|
| 166 |
+
so a 288-pixel edge is covered by two 256-pixel tiles overlapping by 224, and
|
| 167 |
+
you pay for 512 rows to get 288. An edge of exactly 256 is one tile. 512Γ256
|
| 168 |
+
therefore decodes three tiles where 512Γ288 decodes six, for 89% of the
|
| 169 |
+
pixels. The schedule is the reference's and this port reproduces it exactly;
|
| 170 |
+
picking an edge that lands on it is free.
|
| 171 |
+
|
| 172 |
+
**Host and device do not agree to the last bit, and neither is wrong.** The
|
| 173 |
+
host arm quantizes activations to int8 (`CMF_SDOT`) where the device
|
| 174 |
+
dequantizes to f32, so the two renders differ by a few per cent in latent rms
|
| 175 |
+
and visibly in fine texture. Set `CMF_SDOT=0` on both sides to compare
|
| 176 |
+
arithmetic instead of that approximation.
|
| 177 |
+
|
| 178 |
+
**Why the device is opt-in.** Getting it right took three fixes, and one
|
| 179 |
+
thing is still held back.
|
| 180 |
+
|
| 181 |
+
The engine's blocked f32 GEMM cached its weight-side device buffer **by
|
| 182 |
+
pointer address**. Every batched attention allocates one k/v scratch pair per
|
| 183 |
+
call and refills it per head β same address, different matrix β so head 0's
|
| 184 |
+
keys came back for every head, on the GPU only, silently. It is keyed on a
|
| 185 |
+
content fingerprint now. The same GEMM also took every job over 4 M MACs on
|
| 186 |
+
sight with no CPU arm to lose to, which on this model's decoder was three
|
| 187 |
+
times *slower* than the host it displaced; it goes through the same
|
| 188 |
+
measure-don't-assume probe as every other op class now, and on this stack the
|
| 189 |
+
probe hands that work back (0.24 ms device against 0.13 host) while sending
|
| 190 |
+
the weight GEMMs to the card (25.8 ms against 92.0).
|
| 191 |
+
|
| 192 |
+
Still held: **the cooperative-matrix kernel runs this model out of f16 range.**
|
| 193 |
+
At 256Γ160 the render is correct; at 512Γ288 the audio stream goes NaN on the
|
| 194 |
+
second sampling step and the video follows. Bisected β `CMF_BAKE_GPU=0` does
|
| 195 |
+
not help, `CMF_COOP=0` does β so `cortiq animate` pins `CMF_COOP=0`. Tensor
|
| 196 |
+
cores are worth having here; the kernel needs to carry a scale before it can
|
| 197 |
+
carry these activations. That is the next real speedup in this model.
|
| 198 |
|
| 199 |
## What the conversion did
|
| 200 |
|