--- license: apache-2.0 base_model: - Comfy-Org/MiniMax-H3 - larryvrh/MiniMax-H3-Turbo-Lora base_model_relation: quantized pipeline_tag: text-to-video tags: - cmf - cortiq - video - audio - 4-bit --- # MiniMax-H3 Turbo — one 23.5 GB file, no Python [MiniMax-H3](https://huggingface.co/Comfy-Org/MiniMax-H3) renders video and synchronized stereo audio from one prompt, in one transformer, on two flow schedules. [larryvrh's Turbo LoRA](https://huggingface.co/larryvrh/MiniMax-H3-Turbo-Lora) brings it to four sampling steps. This is both of them in the [CMF container](https://github.com/infosave2007/cmf) — the DiT, the Qwen3-VL prompt encoder, the video VAE decoder and the audio vocoder in a single memory-mapped file — running on `cortiq`, a Rust binary with no ML framework underneath. | | reference checkout | here | |---|---|---| | diffusion model | 66.3 GB (bf16) | — | | prompt encoder | 51.5 GB (bf16) | — | | video + audio VAE | 5.8 GB | — | | Turbo LoRA | 0.8 GB | — | | **total** | **124.4 GB, four files + a ComfyUI checkout** | **23.5 GB, one file** | 47.83 B parameters, 2 361 tensors, `cortiq verify` clean. The LoRA is not a separate download: it is merged into the weights, so the file IS the 4-step model. **Text-to-video only.** The release also takes first/last keyframes (`fl2va`) and reference images, videos and audio (`ref2va`); those paths are not ported and the vision tower is not packed. What is here is `t2va`: prompt in, video and audio out. ## Getting a file ```bash huggingface-cli download infosave/MiniMax-H3-Turbo-cmf \ --include 'parts-q4tp/part_*' --local-dir . cat parts-q4tp/part_* > mmh3-turbo-q4tp.cmf cortiq animate mmh3-turbo-q4tp.cmf \ --prompt "A corgi in a chef hat flipping a pancake, sizzling sounds and a cheerful bark." \ --width 512 --height 288 --frames 39 --out corgi.avi ``` `cortiq animate` writes an MJPEG+PCM AVI and the stereo track beside it as `.wav`. There is no ffmpeg in the loop and nothing to install: the JPEG encoder and the RIFF muxer are part of the binary, because a pipeline that ends in a shell-out to a 20 MB dependency is not a pipeline you can ship. Width and height are multiples of 32. Frame counts snap **up** to the model's 17k+5 grid (5, 22, 39, 56, … 124), which is the grid the temporal VAE and the DiT's frame-span pattern agree on. 4 steps is what the LoRA is trained for; more still helps a little. ## What it costs to run The file is memory-mapped, so plan on RAM at least its size or every step touches non-resident pages. Measured on 48 CPU cores, 4 steps: | | tokens in the pack | denoise (4 steps) | decode | total | |---|---|---|---|---| | 256×160, 22 frames (0.9 s) | 375 | 45.0 s | 10.7 s | **55.7 s** | | 512×288, 39 frames (1.6 s) | 1 879 | 220.8 s | 150.0 s | **371.3 s** | Nearly all of the decode is the video VAE — the vocoder is 4 s of the 150. The packed sequence is `[text | audio | video]` and everything attends to everything, so cost grows with the token count and then with its square. A 512×288 second is five times the tokens of a 256×160 one. **The GPU path is off, on purpose.** `cortiq`'s wide-GEMM arm on wgpu is measured three times faster than the host here and it is also wrong: on an RTX PRO 6000 Blackwell the DiT's first step disagrees with the CPU (video velocity rms 1.32 against 1.73, audio 0.16 against 1.01) and the second step returns NaN — a flat grey frame and a clipped waveform. The op probe cannot notice: it only times the two arms, and it runs the bad one for real while it is deciding. So `cortiq animate` pins `CMF_GPU=0` before the backend comes up, which is early enough to be sure. `CMF_MMH3_GPU=1` opts back in if you want to watch it fail. Fixing that kernel is the next thing worth doing to this model — it is the whole difference between minutes and hours. ## What the conversion did **The adaLN collapse.** Forty per cent of the released DiT is one matrix per block: `adaln_proj.linear` is `[96768, 2688]`, 520 MB at bf16, **13 B of the model's 33 B parameters** — for a map whose input is one number, the timestep. Its output over the whole schedule is a one-dimensional curve in R^96768, and Comfy-Org's `pruned` checkpoints already ship it as one: an `adaln_t_table` of `[1025, 8]` shared by every block and per-block weights of `[96768, 8]`. Measured against the full matrix on block 0 (`tools/mmh3_fetch.py check`, which range-reads 520 MB out of the 66 GB file rather than downloading it): ``` adaln max|Δ| 8.0e-4 rms 8.7e-5 against a signal of rms 0.464 time-curve singular values 1..12, relative: 1.00e0 2.96e-1 1.05e-1 6.63e-2 6.60e-3 2.11e-3 5.61e-4 2.92e-4 3.67e-5 2.73e-5 1.32e-5 1.34e-6 ``` The ninth singular value is already 3.7e-5 of the first. Rank eight is not an approximation anyone should feel nervous about; the 26 GB is redundant. The Turbo LoRA is written against the FULL matrix (`lora_A` is `[16, 2688]`), which is why the ComfyUI node re-injects the time conditioning at run time when the base is pruned. `cortiq animate-pack` does it once, at conversion: ``` adaln(t) = W_p · u(t) + b + B · (A · silu(e(t))) = [W_p | B] · [u(t) ; A · silu(e(t))] ``` — a rank-24 curve, driven by a `[1025, 24]` table per block. 4.6 MB a block instead of 520, with the LoRA already inside it. **The rest.** - **Backbone** — bf16 → `q4tp`, 4.16 bits a weight with a predicted per-row scale ladder. The LoRA's rank-64 update is merged before quantizing. - **Prompt encoder** — Qwen3-VL-32B truncated to 50 layers, 51.5 GB → 12.2 GB. It is the largest single component of the file and it runs once per generation. - **Video VAE — decoder only.** It is a ViT3D, not a conv stack: 36 transformer blocks over the latent grid and one linear that expands each cell into a 4×16×16 block of pixels. The 3-D causal CNN encoder is a third of the checkpoint and text-to-video never runs it. - **Audio VAE — decoder only**, f16. Quantizing a vocoder buys 45 MB and costs audible hiss. Its 254 kaiser-sinc resampling filters are read from the checkpoint rather than re-derived — the design formula is in the code as a fallback, but a filter you compute is a filter that can drift from the one the weights were trained against. - Integrity: 47.83 B parameters over 2 361 tensors; `cortiq verify` checks every one against the directory's hashes. ## On parity Established, not assumed, and separately for each of the four stacks. The reference is ComfyUI's own module, run on a toy checkpoint carrying the release's real tensor names and the release's real schedules — `tools/` builds them, `tools/mmh3_toy_gate.sh` runs the diff. The packs are exact f32 on purpose: `q4tp`'s noise floor sits an order of magnitude above the arithmetic difference these are looking for, so quantizing here would pass a broken port. | stack | worst | rms | signal rms | |---|---|---|---| | DiT — video velocity | 8.8e-5 | 2.1e-5 | 0.515 | | DiT — audio velocity | 5.2e-5 | 2.5e-5 | 0.409 | | DiT — token refiner | 8.3e-7 | 2.6e-7 | 1.003 | | Qwen3-VL encoder | 1.1e-6 | 3.3e-7 | 0.812 | | video VAE decoder | 4.2e-7 | 4.0e-8 | 0.470 | | audio VAE decoder | 1.7e-9 | 3.5e-10 | 8.9e-4 | A dozen conventions in this model pass at one token and fail differently at a hundred, which is why the toys are not one-vector unit tests: the packed layout's cursor, the video time axis's 1,4,4,4,4 span pattern, which 96 of 128 head dimensions rotate, the adaLN row order (timestep-major, modality-minor), the video VAE's 256-pixel tiling — global attention makes a tile a different computation from a whole frame, so the tiling is part of the output, not a memory strategy — and the audio stream's separate clock. ## Two clocks The video and audio latents ride different flow schedules (shift 12 and 3). The sampler walks the video grid, which at four steps is `1, 0.973, 0.923, 0.8, 0`, and integrates the audio on its own remap of it. Stepping both on the video grid is what a stock sampler does; it is fine at twenty steps and wrong at four, because over the last interval Δσ_a and Δσ_v differ by a factor of three and no per-step slope correction survives a step that large. `--stock-sampler` reproduces the broken behaviour if you want to hear it. ## Provenance Weights derive from MiniMax's H3 release as repackaged by Comfy-Org, and from larryvrh's Turbo LoRA; both remain under their own licences. The Turbo LoRA is a **preview** — its own card notes plastic-looking skin and over-sharp grain at `ckpt850`, and nothing here changes that. The CMF container and the cortiq runtime are Apache-2.0 (see the repository's LICENSE and PATENTS.md).