Instructions to use infosave/MiniMax-H3-Turbo-cmf with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- cortiq
How to use infosave/MiniMax-H3-Turbo-cmf with cortiq:
# one Rust binary, no additional dependencies cargo install cortiq-cli # or a prebuilt binary from github.com/infosave2007/cmf/releases hf download infosave/MiniMax-H3-Turbo-cmf --include "*.cmf" --local-dir . ls *.cmf # some repos ship more than one quantization
cortiq animate FILE.cmf --prompt "a corgi in a chef hat flipping a pancake" --out clip.avi
- Notebooks
- Google Colab
- Kaggle
Download README.md from infosave/MiniMax-H3-Turbo-cmf: direct link, hf CLI and curl.
- Browser
- Download file 8.71 kB
-
https://huggingface.co/infosave/MiniMax-H3-Turbo-cmf/resolve/7e089f91cf94420fd0e913efe4c78142f36bb7ef/README.md
- Command line
-
hf download hf://infosave/MiniMax-H3-Turbo-cmf@7e089f91cf94420fd0e913efe4c78142f36bb7ef/README.md
-
curl -L -o README.md https://huggingface.co/infosave/MiniMax-H3-Turbo-cmf/resolve/7e089f91cf94420fd0e913efe4c78142f36bb7ef/README.md
license: apache-2.0
base_model:
- Comfy-Org/MiniMax-H3
- larryvrh/MiniMax-H3-Turbo-Lora
base_model_relation: quantized
pipeline_tag: text-to-video
tags:
- cmf
- cortiq
- video
- audio
- 4-bit
MiniMax-H3 Turbo β one 23.5 GB file, no Python
MiniMax-H3 renders video and
synchronized stereo audio from one prompt, in one transformer, on two flow
schedules. larryvrh's Turbo LoRA
brings it to four sampling steps. This is both of them in the
CMF container β the DiT, the Qwen3-VL
prompt encoder, the video VAE decoder and the audio vocoder in a single
memory-mapped file β running on cortiq, a Rust binary with no ML framework
underneath.
| reference checkout | here | |
|---|---|---|
| diffusion model | 66.3 GB (bf16) | β |
| prompt encoder | 51.5 GB (bf16) | β |
| video + audio VAE | 5.8 GB | β |
| Turbo LoRA | 0.8 GB | β |
| total | 124.4 GB, four files + a ComfyUI checkout | 23.5 GB, one file |
47.83 B parameters, 2 361 tensors, cortiq verify clean.
The LoRA is not a separate download: it is merged into the weights, so the file IS the 4-step model.
Text-to-video only. The release also takes first/last keyframes (fl2va)
and reference images, videos and audio (ref2va); those paths are not ported
and the vision tower is not packed. What is here is t2va: prompt in, video
and audio out.
Getting a file
huggingface-cli download infosave/MiniMax-H3-Turbo-cmf \
--include 'parts-q4tp/part_*' --local-dir .
cat parts-q4tp/part_* > mmh3-turbo-q4tp.cmf
cortiq animate mmh3-turbo-q4tp.cmf \
--prompt "A corgi in a chef hat flipping a pancake, sizzling sounds and a cheerful bark." \
--width 512 --height 288 --frames 39 --out corgi.avi
cortiq animate writes an MJPEG+PCM AVI and the stereo track beside it as
.wav. There is no ffmpeg in the loop and nothing to install: the JPEG encoder
and the RIFF muxer are part of the binary, because a pipeline that ends in a
shell-out to a 20 MB dependency is not a pipeline you can ship.
Width and height are multiples of 32. Frame counts snap up to the model's 17k+5 grid (5, 22, 39, 56, β¦ 124), which is the grid the temporal VAE and the DiT's frame-span pattern agree on. 4 steps is what the LoRA is trained for; more still helps a little.
What it costs to run
The file is memory-mapped, so plan on RAM at least its size or every step touches non-resident pages.
Measured on 48 CPU cores, 4 steps:
| tokens in the pack | denoise (4 steps) | decode | total | |
|---|---|---|---|---|
| 256Γ160, 22 frames (0.9 s) | 375 | 45.0 s | 10.7 s | 55.7 s |
| 512Γ288, 39 frames (1.6 s) | 1 879 | 220.8 s | 150.0 s | 371.3 s |
Nearly all of the decode is the video VAE β the vocoder is 4 s of the 150.
The packed sequence is [text | audio | video] and everything attends to
everything, so cost grows with the token count and then with its square. A
512Γ288 second is five times the tokens of a 256Γ160 one.
The GPU path is off, on purpose. cortiq's wide-GEMM arm on wgpu is
measured three times faster than the host here and it is also wrong: on an
RTX PRO 6000 Blackwell the DiT's first step disagrees with the CPU (video
velocity rms 1.32 against 1.73, audio 0.16 against 1.01) and the second step
returns NaN β a flat grey frame and a clipped waveform. The op probe cannot
notice: it only times the two arms, and it runs the bad one for real while it
is deciding. So cortiq animate pins CMF_GPU=0 before the backend comes up,
which is early enough to be sure. CMF_MMH3_GPU=1 opts back in if you want to
watch it fail. Fixing that kernel is the next thing worth doing to this model β
it is the whole difference between minutes and hours.
What the conversion did
The adaLN collapse. Forty per cent of the released DiT is one matrix per
block: adaln_proj.linear is [96768, 2688], 520 MB at bf16, 13 B of the
model's 33 B parameters β for a map whose input is one number, the timestep.
Its output over the whole schedule is a one-dimensional curve in R^96768, and
Comfy-Org's pruned checkpoints already ship it as one: an adaln_t_table of
[1025, 8] shared by every block and per-block weights of [96768, 8].
Measured against the full matrix on block 0 (tools/mmh3_fetch.py check, which
range-reads 520 MB out of the 66 GB file rather than downloading it):
adaln max|Ξ| 8.0e-4 rms 8.7e-5 against a signal of rms 0.464
time-curve singular values 1..12, relative:
1.00e0 2.96e-1 1.05e-1 6.63e-2 6.60e-3 2.11e-3 5.61e-4 2.92e-4
3.67e-5 2.73e-5 1.32e-5 1.34e-6
The ninth singular value is already 3.7e-5 of the first. Rank eight is not an approximation anyone should feel nervous about; the 26 GB is redundant.
The Turbo LoRA is written against the FULL matrix (lora_A is [16, 2688]),
which is why the ComfyUI node re-injects the time conditioning at run time when
the base is pruned. cortiq animate-pack does it once, at conversion:
adaln(t) = W_p Β· u(t) + b + B Β· (A Β· silu(e(t)))
= [W_p | B] Β· [u(t) ; A Β· silu(e(t))]
β a rank-24 curve, driven by a [1025, 24] table per block. 4.6 MB a block
instead of 520, with the LoRA already inside it.
The rest.
- Backbone β bf16 β
q4tp, 4.16 bits a weight with a predicted per-row scale ladder. The LoRA's rank-64 update is merged before quantizing. - Prompt encoder β Qwen3-VL-32B truncated to 50 layers, 51.5 GB β 12.2 GB. It is the largest single component of the file and it runs once per generation.
- Video VAE β decoder only. It is a ViT3D, not a conv stack: 36 transformer blocks over the latent grid and one linear that expands each cell into a 4Γ16Γ16 block of pixels. The 3-D causal CNN encoder is a third of the checkpoint and text-to-video never runs it.
- Audio VAE β decoder only, f16. Quantizing a vocoder buys 45 MB and costs audible hiss. Its 254 kaiser-sinc resampling filters are read from the checkpoint rather than re-derived β the design formula is in the code as a fallback, but a filter you compute is a filter that can drift from the one the weights were trained against.
- Integrity: 47.83 B parameters over 2 361 tensors;
cortiq verifychecks every one against the directory's hashes.
On parity
Established, not assumed, and separately for each of the four stacks. The
reference is ComfyUI's own module, run on a toy checkpoint carrying the
release's real tensor names and the release's real schedules β tools/
builds them, tools/mmh3_toy_gate.sh runs the diff. The packs are exact f32
on purpose: q4tp's noise floor sits an order of magnitude above the
arithmetic difference these are looking for, so quantizing here would pass a
broken port.
| stack | worst | rms | signal rms |
|---|---|---|---|
| DiT β video velocity | 8.8e-5 | 2.1e-5 | 0.515 |
| DiT β audio velocity | 5.2e-5 | 2.5e-5 | 0.409 |
| DiT β token refiner | 8.3e-7 | 2.6e-7 | 1.003 |
| Qwen3-VL encoder | 1.1e-6 | 3.3e-7 | 0.812 |
| video VAE decoder | 4.2e-7 | 4.0e-8 | 0.470 |
| audio VAE decoder | 1.7e-9 | 3.5e-10 | 8.9e-4 |
A dozen conventions in this model pass at one token and fail differently at a hundred, which is why the toys are not one-vector unit tests: the packed layout's cursor, the video time axis's 1,4,4,4,4 span pattern, which 96 of 128 head dimensions rotate, the adaLN row order (timestep-major, modality-minor), the video VAE's 256-pixel tiling β global attention makes a tile a different computation from a whole frame, so the tiling is part of the output, not a memory strategy β and the audio stream's separate clock.
Two clocks
The video and audio latents ride different flow schedules (shift 12 and 3).
The sampler walks the video grid, which at four steps is
1, 0.973, 0.923, 0.8, 0, and integrates the audio on its own remap of it.
Stepping both on the video grid is what a stock sampler does; it is fine at
twenty steps and wrong at four, because over the last interval ΞΟ_a and ΞΟ_v
differ by a factor of three and no per-step slope correction survives a step
that large. --stock-sampler reproduces the broken behaviour if you want to
hear it.
Provenance
Weights derive from MiniMax's H3 release as repackaged by Comfy-Org, and from
larryvrh's Turbo LoRA; both remain under their own licences. The Turbo LoRA is
a preview β its own card notes plastic-looking skin and over-sharp grain at
ckpt850, and nothing here changes that. The CMF container and the cortiq
runtime are Apache-2.0 (see the repository's LICENSE and PATENTS.md).