Instructions to use infosave/MiniMax-H3-Turbo-cmf with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- cortiq
How to use infosave/MiniMax-H3-Turbo-cmf with cortiq:
# one Rust binary, no additional dependencies cargo install cortiq-cli # or a prebuilt binary from github.com/infosave2007/cmf/releases hf download infosave/MiniMax-H3-Turbo-cmf --include "*.cmf" --local-dir . ls *.cmf # some repos ship more than one quantization
cortiq animate FILE.cmf --prompt "a corgi in a chef hat flipping a pancake" --out clip.avi
- Notebooks
- Google Colab
- Kaggle
File size: 9,587 Bytes
30cd106 5fb07cf 39c267f 5fb07cf 39c267f 5fb07cf 30cd106 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 | ---
license: apache-2.0
base_model:
- Comfy-Org/MiniMax-H3
- larryvrh/MiniMax-H3-Turbo-Lora
base_model_relation: quantized
pipeline_tag: text-to-video
tags:
- cmf
- cortiq
- video
- audio
- 4-bit
---
# MiniMax-H3 Turbo β one 23.5 GB file, no Python
[MiniMax-H3](https://huggingface.co/Comfy-Org/MiniMax-H3) renders video and
synchronized stereo audio from one prompt, in one transformer, on two flow
schedules. [larryvrh's Turbo LoRA](https://huggingface.co/larryvrh/MiniMax-H3-Turbo-Lora)
brings it to four sampling steps. This is both of them in the
[CMF container](https://github.com/infosave2007/cmf) β the DiT, the Qwen3-VL
prompt encoder, the video VAE decoder and the audio vocoder in a single
memory-mapped file β running on `cortiq`, a Rust binary with no ML framework
underneath.
| | reference checkout | here |
|---|---|---|
| diffusion model | 66.3 GB (bf16) | β |
| prompt encoder | 51.5 GB (bf16) | β |
| video + audio VAE | 5.8 GB | β |
| Turbo LoRA | 0.8 GB | β |
| **total** | **124.4 GB, four files + a ComfyUI checkout** | **23.5 GB, one file** |
47.83 B parameters, 2 361 tensors, `cortiq verify` clean.
## What comes out

*"A corgi in a chef hat flipping a pancake, sizzling sounds and a cheerful bark."*
β 512Γ288, 39 frames at 24 fps, seed 42, **four steps**, nothing but the prompt.
The GIF is silent; the audio is the point, so take the
**[mp4](https://huggingface.co/infosave/MiniMax-H3-Turbo-cmf/resolve/main/samples/corgi_512x288_4step.mp4)**.
It is not a second model: the same transformer denoises both streams in one
packed sequence, on two different flow schedules.
[`samples/`](https://huggingface.co/infosave/MiniMax-H3-Turbo-cmf/tree/main/samples)
also holds the AVI `cortiq animate` actually wrote and its `.wav` β the mp4 and
the GIF are remuxes for the browser, and the runtime itself never touches
ffmpeg.
The LoRA is not a separate download: it is merged into the weights, so the file
IS the 4-step model.
**Text-to-video only.** The release also takes first/last keyframes (`fl2va`)
and reference images, videos and audio (`ref2va`); those paths are not ported
and the vision tower is not packed. What is here is `t2va`: prompt in, video
and audio out.
## Getting a file
```bash
huggingface-cli download infosave/MiniMax-H3-Turbo-cmf \
--include 'parts-q4tp/part_*' --local-dir .
cat parts-q4tp/part_* > mmh3-turbo-q4tp.cmf
cortiq animate mmh3-turbo-q4tp.cmf \
--prompt "A corgi in a chef hat flipping a pancake, sizzling sounds and a cheerful bark." \
--width 512 --height 288 --frames 39 --out corgi.avi
```
`cortiq animate` writes an MJPEG+PCM AVI and the stereo track beside it as
`.wav`. There is no ffmpeg in the loop and nothing to install: the JPEG encoder
and the RIFF muxer are part of the binary, because a pipeline that ends in a
shell-out to a 20 MB dependency is not a pipeline you can ship.
Width and height are multiples of 32. Frame counts snap **up** to the model's
17k+5 grid (5, 22, 39, 56, β¦ 124), which is the grid the temporal VAE and the
DiT's frame-span pattern agree on. 4 steps is what the LoRA is trained for;
more still helps a little.
## What it costs to run
The file is memory-mapped, so plan on RAM at least its size or every step
touches non-resident pages.
Measured on 48 CPU cores, 4 steps:
| | tokens in the pack | denoise (4 steps) | decode | total |
|---|---|---|---|---|
| 256Γ160, 22 frames (0.9 s) | 375 | 45.0 s | 10.7 s | **55.7 s** |
| 512Γ288, 39 frames (1.6 s) | 1 879 | 220.8 s | 150.0 s | **371.3 s** |
Nearly all of the decode is the video VAE β the vocoder is 4 s of the 150.
The packed sequence is `[text | audio | video]` and everything attends to
everything, so cost grows with the token count and then with its square. A
512Γ288 second is five times the tokens of a 256Γ160 one.
**The GPU path is off, on purpose.** `cortiq`'s wide-GEMM arm on wgpu is
measured three times faster than the host here and it is also wrong: on an
RTX PRO 6000 Blackwell the DiT's first step disagrees with the CPU (video
velocity rms 1.32 against 1.73, audio 0.16 against 1.01) and the second step
returns NaN β a flat grey frame and a clipped waveform. The op probe cannot
notice: it only times the two arms, and it runs the bad one for real while it
is deciding. So `cortiq animate` pins `CMF_GPU=0` before the backend comes up,
which is early enough to be sure. `CMF_MMH3_GPU=1` opts back in if you want to
watch it fail. Fixing that kernel is the next thing worth doing to this model β
it is the whole difference between minutes and hours.
## What the conversion did
**The adaLN collapse.** Forty per cent of the released DiT is one matrix per
block: `adaln_proj.linear` is `[96768, 2688]`, 520 MB at bf16, **13 B of the
model's 33 B parameters** β for a map whose input is one number, the timestep.
Its output over the whole schedule is a one-dimensional curve in R^96768, and
Comfy-Org's `pruned` checkpoints already ship it as one: an `adaln_t_table` of
`[1025, 8]` shared by every block and per-block weights of `[96768, 8]`.
Measured against the full matrix on block 0 (`tools/mmh3_fetch.py check`, which
range-reads 520 MB out of the 66 GB file rather than downloading it):
```
adaln max|Ξ| 8.0e-4 rms 8.7e-5 against a signal of rms 0.464
time-curve singular values 1..12, relative:
1.00e0 2.96e-1 1.05e-1 6.63e-2 6.60e-3 2.11e-3 5.61e-4 2.92e-4
3.67e-5 2.73e-5 1.32e-5 1.34e-6
```
The ninth singular value is already 3.7e-5 of the first. Rank eight is not an
approximation anyone should feel nervous about; the 26 GB is redundant.
The Turbo LoRA is written against the FULL matrix (`lora_A` is `[16, 2688]`),
which is why the ComfyUI node re-injects the time conditioning at run time when
the base is pruned. `cortiq animate-pack` does it once, at conversion:
```
adaln(t) = W_p Β· u(t) + b + B Β· (A Β· silu(e(t)))
= [W_p | B] Β· [u(t) ; A Β· silu(e(t))]
```
β a rank-24 curve, driven by a `[1025, 24]` table per block. 4.6 MB a block
instead of 520, with the LoRA already inside it.
**The rest.**
- **Backbone** β bf16 β `q4tp`, 4.16 bits a weight with a predicted per-row
scale ladder. The LoRA's rank-64 update is merged before quantizing.
- **Prompt encoder** β Qwen3-VL-32B truncated to 50 layers, 51.5 GB β 12.2 GB.
It is the largest single component of the file and it runs once per
generation.
- **Video VAE β decoder only.** It is a ViT3D, not a conv stack: 36 transformer
blocks over the latent grid and one linear that expands each cell into a
4Γ16Γ16 block of pixels. The 3-D causal CNN encoder is a third of the
checkpoint and text-to-video never runs it.
- **Audio VAE β decoder only**, f16. Quantizing a vocoder buys 45 MB and costs
audible hiss. Its 254 kaiser-sinc resampling filters are read from the
checkpoint rather than re-derived β the design formula is in the code as a
fallback, but a filter you compute is a filter that can drift from the one
the weights were trained against.
- Integrity: 47.83 B parameters over 2 361 tensors; `cortiq verify` checks
every one against the directory's hashes.
## On parity
Established, not assumed, and separately for each of the four stacks. The
reference is ComfyUI's own module, run on a toy checkpoint carrying the
release's real tensor names and the release's real schedules β `tools/`
builds them, `tools/mmh3_toy_gate.sh` runs the diff. The packs are exact f32
on purpose: `q4tp`'s noise floor sits an order of magnitude above the
arithmetic difference these are looking for, so quantizing here would pass a
broken port.
| stack | worst | rms | signal rms |
|---|---|---|---|
| DiT β video velocity | 8.8e-5 | 2.1e-5 | 0.515 |
| DiT β audio velocity | 5.2e-5 | 2.5e-5 | 0.409 |
| DiT β token refiner | 8.3e-7 | 2.6e-7 | 1.003 |
| Qwen3-VL encoder | 1.1e-6 | 3.3e-7 | 0.812 |
| video VAE decoder | 4.2e-7 | 4.0e-8 | 0.470 |
| audio VAE decoder | 1.7e-9 | 3.5e-10 | 8.9e-4 |
A dozen conventions in this model pass at one token and fail differently at a
hundred, which is why the toys are not one-vector unit tests: the packed
layout's cursor, the video time axis's 1,4,4,4,4 span pattern, which 96 of 128
head dimensions rotate, the adaLN row order (timestep-major, modality-minor),
the video VAE's 256-pixel tiling β global attention makes a tile a different
computation from a whole frame, so the tiling is part of the output, not a
memory strategy β and the audio stream's separate clock.
## Two clocks
The video and audio latents ride different flow schedules (shift 12 and 3).
The sampler walks the video grid, which at four steps is
`1, 0.973, 0.923, 0.8, 0`, and integrates the audio on its own remap of it.
Stepping both on the video grid is what a stock sampler does; it is fine at
twenty steps and wrong at four, because over the last interval ΞΟ_a and ΞΟ_v
differ by a factor of three and no per-step slope correction survives a step
that large. `--stock-sampler` reproduces the broken behaviour if you want to
hear it.
## Provenance
Weights derive from MiniMax's H3 release as repackaged by Comfy-Org, and from
larryvrh's Turbo LoRA; both remain under their own licences. The Turbo LoRA is
a **preview** β its own card notes plastic-looking skin and over-sharp grain at
`ckpt850`, and nothing here changes that. The CMF container and the cortiq
runtime are Apache-2.0 (see the repository's LICENSE and PATENTS.md).
|