Instructions to use infosave/MiniMax-H3-Turbo-cmf with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- cortiq
How to use infosave/MiniMax-H3-Turbo-cmf with cortiq:
# one Rust binary, no additional dependencies cargo install cortiq-cli # or a prebuilt binary from github.com/infosave2007/cmf/releases hf download infosave/MiniMax-H3-Turbo-cmf --include "*.cmf" --local-dir . ls *.cmf # some repos ship more than one quantization
cortiq animate FILE.cmf --prompt "a corgi in a chef hat flipping a pancake" --out clip.avi
- Notebooks
- Google Colab
- Kaggle
Upload README.md with huggingface_hub
Browse files
README.md
ADDED
|
@@ -0,0 +1,193 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: apache-2.0
|
| 3 |
+
base_model:
|
| 4 |
+
- Comfy-Org/MiniMax-H3
|
| 5 |
+
- larryvrh/MiniMax-H3-Turbo-Lora
|
| 6 |
+
base_model_relation: quantized
|
| 7 |
+
pipeline_tag: text-to-video
|
| 8 |
+
tags:
|
| 9 |
+
- cmf
|
| 10 |
+
- cortiq
|
| 11 |
+
- video
|
| 12 |
+
- audio
|
| 13 |
+
- 4-bit
|
| 14 |
+
---
|
| 15 |
+
|
| 16 |
+
# MiniMax-H3 Turbo β one 23.5 GB file, no Python
|
| 17 |
+
|
| 18 |
+
[MiniMax-H3](https://huggingface.co/Comfy-Org/MiniMax-H3) renders video and
|
| 19 |
+
synchronized stereo audio from one prompt, in one transformer, on two flow
|
| 20 |
+
schedules. [larryvrh's Turbo LoRA](https://huggingface.co/larryvrh/MiniMax-H3-Turbo-Lora)
|
| 21 |
+
brings it to four sampling steps. This is both of them in the
|
| 22 |
+
[CMF container](https://github.com/infosave2007/cmf) β the DiT, the Qwen3-VL
|
| 23 |
+
prompt encoder, the video VAE decoder and the audio vocoder in a single
|
| 24 |
+
memory-mapped file β running on `cortiq`, a Rust binary with no ML framework
|
| 25 |
+
underneath.
|
| 26 |
+
|
| 27 |
+
| | reference checkout | here |
|
| 28 |
+
|---|---|---|
|
| 29 |
+
| diffusion model | 66.3 GB (bf16) | β |
|
| 30 |
+
| prompt encoder | 51.5 GB (bf16) | β |
|
| 31 |
+
| video + audio VAE | 5.8 GB | β |
|
| 32 |
+
| Turbo LoRA | 0.8 GB | β |
|
| 33 |
+
| **total** | **124.4 GB, four files + a ComfyUI checkout** | **23.5 GB, one file** |
|
| 34 |
+
|
| 35 |
+
47.83 B parameters, 2 361 tensors, `cortiq verify` clean.
|
| 36 |
+
|
| 37 |
+
The LoRA is not a separate download: it is merged into the weights, so the file
|
| 38 |
+
IS the 4-step model.
|
| 39 |
+
|
| 40 |
+
**Text-to-video only.** The release also takes first/last keyframes (`fl2va`)
|
| 41 |
+
and reference images, videos and audio (`ref2va`); those paths are not ported
|
| 42 |
+
and the vision tower is not packed. What is here is `t2va`: prompt in, video
|
| 43 |
+
and audio out.
|
| 44 |
+
|
| 45 |
+
## Getting a file
|
| 46 |
+
|
| 47 |
+
```bash
|
| 48 |
+
huggingface-cli download infosave/MiniMax-H3-Turbo-cmf \
|
| 49 |
+
--include 'parts-q4tp/part_*' --local-dir .
|
| 50 |
+
cat parts-q4tp/part_* > mmh3-turbo-q4tp.cmf
|
| 51 |
+
|
| 52 |
+
cortiq animate mmh3-turbo-q4tp.cmf \
|
| 53 |
+
--prompt "A corgi in a chef hat flipping a pancake, sizzling sounds and a cheerful bark." \
|
| 54 |
+
--width 512 --height 288 --frames 39 --out corgi.avi
|
| 55 |
+
```
|
| 56 |
+
|
| 57 |
+
`cortiq animate` writes an MJPEG+PCM AVI and the stereo track beside it as
|
| 58 |
+
`.wav`. There is no ffmpeg in the loop and nothing to install: the JPEG encoder
|
| 59 |
+
and the RIFF muxer are part of the binary, because a pipeline that ends in a
|
| 60 |
+
shell-out to a 20 MB dependency is not a pipeline you can ship.
|
| 61 |
+
|
| 62 |
+
Width and height are multiples of 32. Frame counts snap **up** to the model's
|
| 63 |
+
17k+5 grid (5, 22, 39, 56, β¦ 124), which is the grid the temporal VAE and the
|
| 64 |
+
DiT's frame-span pattern agree on. 4 steps is what the LoRA is trained for;
|
| 65 |
+
more still helps a little.
|
| 66 |
+
|
| 67 |
+
## What it costs to run
|
| 68 |
+
|
| 69 |
+
The file is memory-mapped, so plan on RAM at least its size or every step
|
| 70 |
+
touches non-resident pages.
|
| 71 |
+
|
| 72 |
+
Measured on 48 CPU cores, 4 steps:
|
| 73 |
+
|
| 74 |
+
| | tokens in the pack | denoise (4 steps) | decode | total |
|
| 75 |
+
|---|---|---|---|---|
|
| 76 |
+
| 256Γ160, 22 frames (0.9 s) | 375 | 45.0 s | 10.7 s | **55.7 s** |
|
| 77 |
+
| 512Γ288, 39 frames (1.6 s) | 1 879 | 220.8 s | 150.0 s | **371.3 s** |
|
| 78 |
+
|
| 79 |
+
Nearly all of the decode is the video VAE β the vocoder is 4 s of the 150.
|
| 80 |
+
|
| 81 |
+
The packed sequence is `[text | audio | video]` and everything attends to
|
| 82 |
+
everything, so cost grows with the token count and then with its square. A
|
| 83 |
+
512Γ288 second is five times the tokens of a 256Γ160 one.
|
| 84 |
+
|
| 85 |
+
**The GPU path is off, on purpose.** `cortiq`'s wide-GEMM arm on wgpu is
|
| 86 |
+
measured three times faster than the host here and it is also wrong: on an
|
| 87 |
+
RTX PRO 6000 Blackwell the DiT's first step disagrees with the CPU (video
|
| 88 |
+
velocity rms 1.32 against 1.73, audio 0.16 against 1.01) and the second step
|
| 89 |
+
returns NaN β a flat grey frame and a clipped waveform. The op probe cannot
|
| 90 |
+
notice: it only times the two arms, and it runs the bad one for real while it
|
| 91 |
+
is deciding. So `cortiq animate` pins `CMF_GPU=0` before the backend comes up,
|
| 92 |
+
which is early enough to be sure. `CMF_MMH3_GPU=1` opts back in if you want to
|
| 93 |
+
watch it fail. Fixing that kernel is the next thing worth doing to this model β
|
| 94 |
+
it is the whole difference between minutes and hours.
|
| 95 |
+
|
| 96 |
+
## What the conversion did
|
| 97 |
+
|
| 98 |
+
**The adaLN collapse.** Forty per cent of the released DiT is one matrix per
|
| 99 |
+
block: `adaln_proj.linear` is `[96768, 2688]`, 520 MB at bf16, **13 B of the
|
| 100 |
+
model's 33 B parameters** β for a map whose input is one number, the timestep.
|
| 101 |
+
Its output over the whole schedule is a one-dimensional curve in R^96768, and
|
| 102 |
+
Comfy-Org's `pruned` checkpoints already ship it as one: an `adaln_t_table` of
|
| 103 |
+
`[1025, 8]` shared by every block and per-block weights of `[96768, 8]`.
|
| 104 |
+
|
| 105 |
+
Measured against the full matrix on block 0 (`tools/mmh3_fetch.py check`, which
|
| 106 |
+
range-reads 520 MB out of the 66 GB file rather than downloading it):
|
| 107 |
+
|
| 108 |
+
```
|
| 109 |
+
adaln max|Ξ| 8.0e-4 rms 8.7e-5 against a signal of rms 0.464
|
| 110 |
+
time-curve singular values 1..12, relative:
|
| 111 |
+
1.00e0 2.96e-1 1.05e-1 6.63e-2 6.60e-3 2.11e-3 5.61e-4 2.92e-4
|
| 112 |
+
3.67e-5 2.73e-5 1.32e-5 1.34e-6
|
| 113 |
+
```
|
| 114 |
+
|
| 115 |
+
The ninth singular value is already 3.7e-5 of the first. Rank eight is not an
|
| 116 |
+
approximation anyone should feel nervous about; the 26 GB is redundant.
|
| 117 |
+
|
| 118 |
+
The Turbo LoRA is written against the FULL matrix (`lora_A` is `[16, 2688]`),
|
| 119 |
+
which is why the ComfyUI node re-injects the time conditioning at run time when
|
| 120 |
+
the base is pruned. `cortiq animate-pack` does it once, at conversion:
|
| 121 |
+
|
| 122 |
+
```
|
| 123 |
+
adaln(t) = W_p Β· u(t) + b + B Β· (A Β· silu(e(t)))
|
| 124 |
+
= [W_p | B] Β· [u(t) ; A Β· silu(e(t))]
|
| 125 |
+
```
|
| 126 |
+
|
| 127 |
+
β a rank-24 curve, driven by a `[1025, 24]` table per block. 4.6 MB a block
|
| 128 |
+
instead of 520, with the LoRA already inside it.
|
| 129 |
+
|
| 130 |
+
**The rest.**
|
| 131 |
+
|
| 132 |
+
- **Backbone** β bf16 β `q4tp`, 4.16 bits a weight with a predicted per-row
|
| 133 |
+
scale ladder. The LoRA's rank-64 update is merged before quantizing.
|
| 134 |
+
- **Prompt encoder** β Qwen3-VL-32B truncated to 50 layers, 51.5 GB β 12.2 GB.
|
| 135 |
+
It is the largest single component of the file and it runs once per
|
| 136 |
+
generation.
|
| 137 |
+
- **Video VAE β decoder only.** It is a ViT3D, not a conv stack: 36 transformer
|
| 138 |
+
blocks over the latent grid and one linear that expands each cell into a
|
| 139 |
+
4Γ16Γ16 block of pixels. The 3-D causal CNN encoder is a third of the
|
| 140 |
+
checkpoint and text-to-video never runs it.
|
| 141 |
+
- **Audio VAE β decoder only**, f16. Quantizing a vocoder buys 45 MB and costs
|
| 142 |
+
audible hiss. Its 254 kaiser-sinc resampling filters are read from the
|
| 143 |
+
checkpoint rather than re-derived β the design formula is in the code as a
|
| 144 |
+
fallback, but a filter you compute is a filter that can drift from the one
|
| 145 |
+
the weights were trained against.
|
| 146 |
+
- Integrity: 47.83 B parameters over 2 361 tensors; `cortiq verify` checks
|
| 147 |
+
every one against the directory's hashes.
|
| 148 |
+
|
| 149 |
+
## On parity
|
| 150 |
+
|
| 151 |
+
Established, not assumed, and separately for each of the four stacks. The
|
| 152 |
+
reference is ComfyUI's own module, run on a toy checkpoint carrying the
|
| 153 |
+
release's real tensor names and the release's real schedules β `tools/`
|
| 154 |
+
builds them, `tools/mmh3_toy_gate.sh` runs the diff. The packs are exact f32
|
| 155 |
+
on purpose: `q4tp`'s noise floor sits an order of magnitude above the
|
| 156 |
+
arithmetic difference these are looking for, so quantizing here would pass a
|
| 157 |
+
broken port.
|
| 158 |
+
|
| 159 |
+
| stack | worst | rms | signal rms |
|
| 160 |
+
|---|---|---|---|
|
| 161 |
+
| DiT β video velocity | 8.8e-5 | 2.1e-5 | 0.515 |
|
| 162 |
+
| DiT β audio velocity | 5.2e-5 | 2.5e-5 | 0.409 |
|
| 163 |
+
| DiT β token refiner | 8.3e-7 | 2.6e-7 | 1.003 |
|
| 164 |
+
| Qwen3-VL encoder | 1.1e-6 | 3.3e-7 | 0.812 |
|
| 165 |
+
| video VAE decoder | 4.2e-7 | 4.0e-8 | 0.470 |
|
| 166 |
+
| audio VAE decoder | 1.7e-9 | 3.5e-10 | 8.9e-4 |
|
| 167 |
+
|
| 168 |
+
A dozen conventions in this model pass at one token and fail differently at a
|
| 169 |
+
hundred, which is why the toys are not one-vector unit tests: the packed
|
| 170 |
+
layout's cursor, the video time axis's 1,4,4,4,4 span pattern, which 96 of 128
|
| 171 |
+
head dimensions rotate, the adaLN row order (timestep-major, modality-minor),
|
| 172 |
+
the video VAE's 256-pixel tiling β global attention makes a tile a different
|
| 173 |
+
computation from a whole frame, so the tiling is part of the output, not a
|
| 174 |
+
memory strategy β and the audio stream's separate clock.
|
| 175 |
+
|
| 176 |
+
## Two clocks
|
| 177 |
+
|
| 178 |
+
The video and audio latents ride different flow schedules (shift 12 and 3).
|
| 179 |
+
The sampler walks the video grid, which at four steps is
|
| 180 |
+
`1, 0.973, 0.923, 0.8, 0`, and integrates the audio on its own remap of it.
|
| 181 |
+
Stepping both on the video grid is what a stock sampler does; it is fine at
|
| 182 |
+
twenty steps and wrong at four, because over the last interval ΞΟ_a and ΞΟ_v
|
| 183 |
+
differ by a factor of three and no per-step slope correction survives a step
|
| 184 |
+
that large. `--stock-sampler` reproduces the broken behaviour if you want to
|
| 185 |
+
hear it.
|
| 186 |
+
|
| 187 |
+
## Provenance
|
| 188 |
+
|
| 189 |
+
Weights derive from MiniMax's H3 release as repackaged by Comfy-Org, and from
|
| 190 |
+
larryvrh's Turbo LoRA; both remain under their own licences. The Turbo LoRA is
|
| 191 |
+
a **preview** β its own card notes plastic-looking skin and over-sharp grain at
|
| 192 |
+
`ckpt850`, and nothing here changes that. The CMF container and the cortiq
|
| 193 |
+
runtime are Apache-2.0 (see the repository's LICENSE and PATENTS.md).
|