infosave commited on
Commit
30cd106
Β·
verified Β·
1 Parent(s): 8b8d895

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +193 -0
README.md ADDED
@@ -0,0 +1,193 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ base_model:
4
+ - Comfy-Org/MiniMax-H3
5
+ - larryvrh/MiniMax-H3-Turbo-Lora
6
+ base_model_relation: quantized
7
+ pipeline_tag: text-to-video
8
+ tags:
9
+ - cmf
10
+ - cortiq
11
+ - video
12
+ - audio
13
+ - 4-bit
14
+ ---
15
+
16
+ # MiniMax-H3 Turbo β€” one 23.5 GB file, no Python
17
+
18
+ [MiniMax-H3](https://huggingface.co/Comfy-Org/MiniMax-H3) renders video and
19
+ synchronized stereo audio from one prompt, in one transformer, on two flow
20
+ schedules. [larryvrh's Turbo LoRA](https://huggingface.co/larryvrh/MiniMax-H3-Turbo-Lora)
21
+ brings it to four sampling steps. This is both of them in the
22
+ [CMF container](https://github.com/infosave2007/cmf) β€” the DiT, the Qwen3-VL
23
+ prompt encoder, the video VAE decoder and the audio vocoder in a single
24
+ memory-mapped file β€” running on `cortiq`, a Rust binary with no ML framework
25
+ underneath.
26
+
27
+ | | reference checkout | here |
28
+ |---|---|---|
29
+ | diffusion model | 66.3 GB (bf16) | β€” |
30
+ | prompt encoder | 51.5 GB (bf16) | β€” |
31
+ | video + audio VAE | 5.8 GB | β€” |
32
+ | Turbo LoRA | 0.8 GB | β€” |
33
+ | **total** | **124.4 GB, four files + a ComfyUI checkout** | **23.5 GB, one file** |
34
+
35
+ 47.83 B parameters, 2 361 tensors, `cortiq verify` clean.
36
+
37
+ The LoRA is not a separate download: it is merged into the weights, so the file
38
+ IS the 4-step model.
39
+
40
+ **Text-to-video only.** The release also takes first/last keyframes (`fl2va`)
41
+ and reference images, videos and audio (`ref2va`); those paths are not ported
42
+ and the vision tower is not packed. What is here is `t2va`: prompt in, video
43
+ and audio out.
44
+
45
+ ## Getting a file
46
+
47
+ ```bash
48
+ huggingface-cli download infosave/MiniMax-H3-Turbo-cmf \
49
+ --include 'parts-q4tp/part_*' --local-dir .
50
+ cat parts-q4tp/part_* > mmh3-turbo-q4tp.cmf
51
+
52
+ cortiq animate mmh3-turbo-q4tp.cmf \
53
+ --prompt "A corgi in a chef hat flipping a pancake, sizzling sounds and a cheerful bark." \
54
+ --width 512 --height 288 --frames 39 --out corgi.avi
55
+ ```
56
+
57
+ `cortiq animate` writes an MJPEG+PCM AVI and the stereo track beside it as
58
+ `.wav`. There is no ffmpeg in the loop and nothing to install: the JPEG encoder
59
+ and the RIFF muxer are part of the binary, because a pipeline that ends in a
60
+ shell-out to a 20 MB dependency is not a pipeline you can ship.
61
+
62
+ Width and height are multiples of 32. Frame counts snap **up** to the model's
63
+ 17k+5 grid (5, 22, 39, 56, … 124), which is the grid the temporal VAE and the
64
+ DiT's frame-span pattern agree on. 4 steps is what the LoRA is trained for;
65
+ more still helps a little.
66
+
67
+ ## What it costs to run
68
+
69
+ The file is memory-mapped, so plan on RAM at least its size or every step
70
+ touches non-resident pages.
71
+
72
+ Measured on 48 CPU cores, 4 steps:
73
+
74
+ | | tokens in the pack | denoise (4 steps) | decode | total |
75
+ |---|---|---|---|---|
76
+ | 256Γ—160, 22 frames (0.9 s) | 375 | 45.0 s | 10.7 s | **55.7 s** |
77
+ | 512Γ—288, 39 frames (1.6 s) | 1 879 | 220.8 s | 150.0 s | **371.3 s** |
78
+
79
+ Nearly all of the decode is the video VAE β€” the vocoder is 4 s of the 150.
80
+
81
+ The packed sequence is `[text | audio | video]` and everything attends to
82
+ everything, so cost grows with the token count and then with its square. A
83
+ 512Γ—288 second is five times the tokens of a 256Γ—160 one.
84
+
85
+ **The GPU path is off, on purpose.** `cortiq`'s wide-GEMM arm on wgpu is
86
+ measured three times faster than the host here and it is also wrong: on an
87
+ RTX PRO 6000 Blackwell the DiT's first step disagrees with the CPU (video
88
+ velocity rms 1.32 against 1.73, audio 0.16 against 1.01) and the second step
89
+ returns NaN β€” a flat grey frame and a clipped waveform. The op probe cannot
90
+ notice: it only times the two arms, and it runs the bad one for real while it
91
+ is deciding. So `cortiq animate` pins `CMF_GPU=0` before the backend comes up,
92
+ which is early enough to be sure. `CMF_MMH3_GPU=1` opts back in if you want to
93
+ watch it fail. Fixing that kernel is the next thing worth doing to this model β€”
94
+ it is the whole difference between minutes and hours.
95
+
96
+ ## What the conversion did
97
+
98
+ **The adaLN collapse.** Forty per cent of the released DiT is one matrix per
99
+ block: `adaln_proj.linear` is `[96768, 2688]`, 520 MB at bf16, **13 B of the
100
+ model's 33 B parameters** β€” for a map whose input is one number, the timestep.
101
+ Its output over the whole schedule is a one-dimensional curve in R^96768, and
102
+ Comfy-Org's `pruned` checkpoints already ship it as one: an `adaln_t_table` of
103
+ `[1025, 8]` shared by every block and per-block weights of `[96768, 8]`.
104
+
105
+ Measured against the full matrix on block 0 (`tools/mmh3_fetch.py check`, which
106
+ range-reads 520 MB out of the 66 GB file rather than downloading it):
107
+
108
+ ```
109
+ adaln max|Ξ”| 8.0e-4 rms 8.7e-5 against a signal of rms 0.464
110
+ time-curve singular values 1..12, relative:
111
+ 1.00e0 2.96e-1 1.05e-1 6.63e-2 6.60e-3 2.11e-3 5.61e-4 2.92e-4
112
+ 3.67e-5 2.73e-5 1.32e-5 1.34e-6
113
+ ```
114
+
115
+ The ninth singular value is already 3.7e-5 of the first. Rank eight is not an
116
+ approximation anyone should feel nervous about; the 26 GB is redundant.
117
+
118
+ The Turbo LoRA is written against the FULL matrix (`lora_A` is `[16, 2688]`),
119
+ which is why the ComfyUI node re-injects the time conditioning at run time when
120
+ the base is pruned. `cortiq animate-pack` does it once, at conversion:
121
+
122
+ ```
123
+ adaln(t) = W_p Β· u(t) + b + B Β· (A Β· silu(e(t)))
124
+ = [W_p | B] Β· [u(t) ; A Β· silu(e(t))]
125
+ ```
126
+
127
+ β€” a rank-24 curve, driven by a `[1025, 24]` table per block. 4.6 MB a block
128
+ instead of 520, with the LoRA already inside it.
129
+
130
+ **The rest.**
131
+
132
+ - **Backbone** β€” bf16 β†’ `q4tp`, 4.16 bits a weight with a predicted per-row
133
+ scale ladder. The LoRA's rank-64 update is merged before quantizing.
134
+ - **Prompt encoder** β€” Qwen3-VL-32B truncated to 50 layers, 51.5 GB β†’ 12.2 GB.
135
+ It is the largest single component of the file and it runs once per
136
+ generation.
137
+ - **Video VAE β€” decoder only.** It is a ViT3D, not a conv stack: 36 transformer
138
+ blocks over the latent grid and one linear that expands each cell into a
139
+ 4Γ—16Γ—16 block of pixels. The 3-D causal CNN encoder is a third of the
140
+ checkpoint and text-to-video never runs it.
141
+ - **Audio VAE β€” decoder only**, f16. Quantizing a vocoder buys 45 MB and costs
142
+ audible hiss. Its 254 kaiser-sinc resampling filters are read from the
143
+ checkpoint rather than re-derived β€” the design formula is in the code as a
144
+ fallback, but a filter you compute is a filter that can drift from the one
145
+ the weights were trained against.
146
+ - Integrity: 47.83 B parameters over 2 361 tensors; `cortiq verify` checks
147
+ every one against the directory's hashes.
148
+
149
+ ## On parity
150
+
151
+ Established, not assumed, and separately for each of the four stacks. The
152
+ reference is ComfyUI's own module, run on a toy checkpoint carrying the
153
+ release's real tensor names and the release's real schedules β€” `tools/`
154
+ builds them, `tools/mmh3_toy_gate.sh` runs the diff. The packs are exact f32
155
+ on purpose: `q4tp`'s noise floor sits an order of magnitude above the
156
+ arithmetic difference these are looking for, so quantizing here would pass a
157
+ broken port.
158
+
159
+ | stack | worst | rms | signal rms |
160
+ |---|---|---|---|
161
+ | DiT β€” video velocity | 8.8e-5 | 2.1e-5 | 0.515 |
162
+ | DiT β€” audio velocity | 5.2e-5 | 2.5e-5 | 0.409 |
163
+ | DiT β€” token refiner | 8.3e-7 | 2.6e-7 | 1.003 |
164
+ | Qwen3-VL encoder | 1.1e-6 | 3.3e-7 | 0.812 |
165
+ | video VAE decoder | 4.2e-7 | 4.0e-8 | 0.470 |
166
+ | audio VAE decoder | 1.7e-9 | 3.5e-10 | 8.9e-4 |
167
+
168
+ A dozen conventions in this model pass at one token and fail differently at a
169
+ hundred, which is why the toys are not one-vector unit tests: the packed
170
+ layout's cursor, the video time axis's 1,4,4,4,4 span pattern, which 96 of 128
171
+ head dimensions rotate, the adaLN row order (timestep-major, modality-minor),
172
+ the video VAE's 256-pixel tiling β€” global attention makes a tile a different
173
+ computation from a whole frame, so the tiling is part of the output, not a
174
+ memory strategy β€” and the audio stream's separate clock.
175
+
176
+ ## Two clocks
177
+
178
+ The video and audio latents ride different flow schedules (shift 12 and 3).
179
+ The sampler walks the video grid, which at four steps is
180
+ `1, 0.973, 0.923, 0.8, 0`, and integrates the audio on its own remap of it.
181
+ Stepping both on the video grid is what a stock sampler does; it is fine at
182
+ twenty steps and wrong at four, because over the last interval Δσ_a and Δσ_v
183
+ differ by a factor of three and no per-step slope correction survives a step
184
+ that large. `--stock-sampler` reproduces the broken behaviour if you want to
185
+ hear it.
186
+
187
+ ## Provenance
188
+
189
+ Weights derive from MiniMax's H3 release as repackaged by Comfy-Org, and from
190
+ larryvrh's Turbo LoRA; both remain under their own licences. The Turbo LoRA is
191
+ a **preview** β€” its own card notes plastic-looking skin and over-sharp grain at
192
+ `ckpt850`, and nothing here changes that. The CMF container and the cortiq
193
+ runtime are Apache-2.0 (see the repository's LICENSE and PATENTS.md).