File size: 9,587 Bytes
30cd106
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
5fb07cf
 
39c267f
5fb07cf
 
 
39c267f
 
 
 
 
 
 
 
 
5fb07cf
30cd106
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
---
license: apache-2.0
base_model:
- Comfy-Org/MiniMax-H3
- larryvrh/MiniMax-H3-Turbo-Lora
base_model_relation: quantized
pipeline_tag: text-to-video
tags:
- cmf
- cortiq
- video
- audio
- 4-bit
---

# MiniMax-H3 Turbo β€” one 23.5 GB file, no Python

[MiniMax-H3](https://huggingface.co/Comfy-Org/MiniMax-H3) renders video and
synchronized stereo audio from one prompt, in one transformer, on two flow
schedules. [larryvrh's Turbo LoRA](https://huggingface.co/larryvrh/MiniMax-H3-Turbo-Lora)
brings it to four sampling steps. This is both of them in the
[CMF container](https://github.com/infosave2007/cmf) β€” the DiT, the Qwen3-VL
prompt encoder, the video VAE decoder and the audio vocoder in a single
memory-mapped file β€” running on `cortiq`, a Rust binary with no ML framework
underneath.

| | reference checkout | here |
|---|---|---|
| diffusion model | 66.3 GB (bf16) | β€” |
| prompt encoder | 51.5 GB (bf16) | β€” |
| video + audio VAE | 5.8 GB | β€” |
| Turbo LoRA | 0.8 GB | β€” |
| **total** | **124.4 GB, four files + a ComfyUI checkout** | **23.5 GB, one file** |

47.83 B parameters, 2 361 tensors, `cortiq verify` clean.

## What comes out

![A corgi in a chef hat over a pan, four-step render](https://huggingface.co/infosave/MiniMax-H3-Turbo-cmf/resolve/main/samples/corgi_512x288_4step.gif)

*"A corgi in a chef hat flipping a pancake, sizzling sounds and a cheerful bark."*
β€” 512Γ—288, 39 frames at 24 fps, seed 42, **four steps**, nothing but the prompt.

The GIF is silent; the audio is the point, so take the
**[mp4](https://huggingface.co/infosave/MiniMax-H3-Turbo-cmf/resolve/main/samples/corgi_512x288_4step.mp4)**.
It is not a second model: the same transformer denoises both streams in one
packed sequence, on two different flow schedules.
[`samples/`](https://huggingface.co/infosave/MiniMax-H3-Turbo-cmf/tree/main/samples)
also holds the AVI `cortiq animate` actually wrote and its `.wav` β€” the mp4 and
the GIF are remuxes for the browser, and the runtime itself never touches
ffmpeg.

The LoRA is not a separate download: it is merged into the weights, so the file
IS the 4-step model.

**Text-to-video only.** The release also takes first/last keyframes (`fl2va`)
and reference images, videos and audio (`ref2va`); those paths are not ported
and the vision tower is not packed. What is here is `t2va`: prompt in, video
and audio out.

## Getting a file

```bash
huggingface-cli download infosave/MiniMax-H3-Turbo-cmf \
  --include 'parts-q4tp/part_*' --local-dir .
cat parts-q4tp/part_* > mmh3-turbo-q4tp.cmf

cortiq animate mmh3-turbo-q4tp.cmf \
  --prompt "A corgi in a chef hat flipping a pancake, sizzling sounds and a cheerful bark." \
  --width 512 --height 288 --frames 39 --out corgi.avi
```

`cortiq animate` writes an MJPEG+PCM AVI and the stereo track beside it as
`.wav`. There is no ffmpeg in the loop and nothing to install: the JPEG encoder
and the RIFF muxer are part of the binary, because a pipeline that ends in a
shell-out to a 20 MB dependency is not a pipeline you can ship.

Width and height are multiples of 32. Frame counts snap **up** to the model's
17k+5 grid (5, 22, 39, 56, … 124), which is the grid the temporal VAE and the
DiT's frame-span pattern agree on. 4 steps is what the LoRA is trained for;
more still helps a little.

## What it costs to run

The file is memory-mapped, so plan on RAM at least its size or every step
touches non-resident pages.

Measured on 48 CPU cores, 4 steps:

| | tokens in the pack | denoise (4 steps) | decode | total |
|---|---|---|---|---|
| 256Γ—160, 22 frames (0.9 s) | 375 | 45.0 s | 10.7 s | **55.7 s** |
| 512Γ—288, 39 frames (1.6 s) | 1 879 | 220.8 s | 150.0 s | **371.3 s** |

Nearly all of the decode is the video VAE β€” the vocoder is 4 s of the 150.

The packed sequence is `[text | audio | video]` and everything attends to
everything, so cost grows with the token count and then with its square. A
512Γ—288 second is five times the tokens of a 256Γ—160 one.

**The GPU path is off, on purpose.** `cortiq`'s wide-GEMM arm on wgpu is
measured three times faster than the host here and it is also wrong: on an
RTX PRO 6000 Blackwell the DiT's first step disagrees with the CPU (video
velocity rms 1.32 against 1.73, audio 0.16 against 1.01) and the second step
returns NaN β€” a flat grey frame and a clipped waveform. The op probe cannot
notice: it only times the two arms, and it runs the bad one for real while it
is deciding. So `cortiq animate` pins `CMF_GPU=0` before the backend comes up,
which is early enough to be sure. `CMF_MMH3_GPU=1` opts back in if you want to
watch it fail. Fixing that kernel is the next thing worth doing to this model β€”
it is the whole difference between minutes and hours.

## What the conversion did

**The adaLN collapse.** Forty per cent of the released DiT is one matrix per
block: `adaln_proj.linear` is `[96768, 2688]`, 520 MB at bf16, **13 B of the
model's 33 B parameters** β€” for a map whose input is one number, the timestep.
Its output over the whole schedule is a one-dimensional curve in R^96768, and
Comfy-Org's `pruned` checkpoints already ship it as one: an `adaln_t_table` of
`[1025, 8]` shared by every block and per-block weights of `[96768, 8]`.

Measured against the full matrix on block 0 (`tools/mmh3_fetch.py check`, which
range-reads 520 MB out of the 66 GB file rather than downloading it):

```
adaln  max|Ξ”| 8.0e-4   rms 8.7e-5   against a signal of rms 0.464
time-curve singular values 1..12, relative:
  1.00e0 2.96e-1 1.05e-1 6.63e-2 6.60e-3 2.11e-3 5.61e-4 2.92e-4
  3.67e-5 2.73e-5 1.32e-5 1.34e-6
```

The ninth singular value is already 3.7e-5 of the first. Rank eight is not an
approximation anyone should feel nervous about; the 26 GB is redundant.

The Turbo LoRA is written against the FULL matrix (`lora_A` is `[16, 2688]`),
which is why the ComfyUI node re-injects the time conditioning at run time when
the base is pruned. `cortiq animate-pack` does it once, at conversion:

```
adaln(t) = W_p Β· u(t) + b + B Β· (A Β· silu(e(t)))
         = [W_p | B] Β· [u(t) ; A Β· silu(e(t))]
```

β€” a rank-24 curve, driven by a `[1025, 24]` table per block. 4.6 MB a block
instead of 520, with the LoRA already inside it.

**The rest.**

- **Backbone** β€” bf16 β†’ `q4tp`, 4.16 bits a weight with a predicted per-row
  scale ladder. The LoRA's rank-64 update is merged before quantizing.
- **Prompt encoder** β€” Qwen3-VL-32B truncated to 50 layers, 51.5 GB β†’ 12.2 GB.
  It is the largest single component of the file and it runs once per
  generation.
- **Video VAE β€” decoder only.** It is a ViT3D, not a conv stack: 36 transformer
  blocks over the latent grid and one linear that expands each cell into a
  4Γ—16Γ—16 block of pixels. The 3-D causal CNN encoder is a third of the
  checkpoint and text-to-video never runs it.
- **Audio VAE β€” decoder only**, f16. Quantizing a vocoder buys 45 MB and costs
  audible hiss. Its 254 kaiser-sinc resampling filters are read from the
  checkpoint rather than re-derived β€” the design formula is in the code as a
  fallback, but a filter you compute is a filter that can drift from the one
  the weights were trained against.
- Integrity: 47.83 B parameters over 2 361 tensors; `cortiq verify` checks
  every one against the directory's hashes.

## On parity

Established, not assumed, and separately for each of the four stacks. The
reference is ComfyUI's own module, run on a toy checkpoint carrying the
release's real tensor names and the release's real schedules β€” `tools/`
builds them, `tools/mmh3_toy_gate.sh` runs the diff. The packs are exact f32
on purpose: `q4tp`'s noise floor sits an order of magnitude above the
arithmetic difference these are looking for, so quantizing here would pass a
broken port.

| stack | worst | rms | signal rms |
|---|---|---|---|
| DiT β€” video velocity | 8.8e-5 | 2.1e-5 | 0.515 |
| DiT β€” audio velocity | 5.2e-5 | 2.5e-5 | 0.409 |
| DiT β€” token refiner | 8.3e-7 | 2.6e-7 | 1.003 |
| Qwen3-VL encoder | 1.1e-6 | 3.3e-7 | 0.812 |
| video VAE decoder | 4.2e-7 | 4.0e-8 | 0.470 |
| audio VAE decoder | 1.7e-9 | 3.5e-10 | 8.9e-4 |

A dozen conventions in this model pass at one token and fail differently at a
hundred, which is why the toys are not one-vector unit tests: the packed
layout's cursor, the video time axis's 1,4,4,4,4 span pattern, which 96 of 128
head dimensions rotate, the adaLN row order (timestep-major, modality-minor),
the video VAE's 256-pixel tiling β€” global attention makes a tile a different
computation from a whole frame, so the tiling is part of the output, not a
memory strategy β€” and the audio stream's separate clock.

## Two clocks

The video and audio latents ride different flow schedules (shift 12 and 3).
The sampler walks the video grid, which at four steps is
`1, 0.973, 0.923, 0.8, 0`, and integrates the audio on its own remap of it.
Stepping both on the video grid is what a stock sampler does; it is fine at
twenty steps and wrong at four, because over the last interval Δσ_a and Δσ_v
differ by a factor of three and no per-step slope correction survives a step
that large. `--stock-sampler` reproduces the broken behaviour if you want to
hear it.

## Provenance

Weights derive from MiniMax's H3 release as repackaged by Comfy-Org, and from
larryvrh's Turbo LoRA; both remain under their own licences. The Turbo LoRA is
a **preview** β€” its own card notes plastic-looking skin and over-sharp grain at
`ckpt850`, and nothing here changes that. The CMF container and the cortiq
runtime are Apache-2.0 (see the repository's LICENSE and PATENTS.md).