Instructions to use infosave/MiniMax-H3-Turbo-cmf with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- cortiq
How to use infosave/MiniMax-H3-Turbo-cmf with cortiq:
# one Rust binary, no additional dependencies cargo install cortiq-cli # or a prebuilt binary from github.com/infosave2007/cmf/releases hf download infosave/MiniMax-H3-Turbo-cmf --include "*.cmf" --local-dir . ls *.cmf # some repos ship more than one quantization
cortiq animate FILE.cmf --prompt "a corgi in a chef hat flipping a pancake" --out clip.avi
- Notebooks
- Google Colab
- Kaggle
File size: 41,539 Bytes
30cd106 701638b 30cd106 0914add 30cd106 e8a0878 5fb07cf e8a0878 5fb07cf e8a0878 39c267f e8a0878 5fb07cf e8a0878 30cd106 e8a0878 30eec76 0914add e8a0878 77201ae 24d7fb2 e8a0878 0914add e8a0878 77201ae e8a0878 30eec76 30cd106 37d210a 0914add 37d210a 30cd106 37d210a 30cd106 37d210a 30cd106 37d210a 76c9afc 37d210a 76c9afc 30cd106 60fd8b2 76c9afc 60fd8b2 37d210a 76c9afc 37d210a 76c9afc 37d210a 41b5099 37d210a 30cd106 37d210a 30cd106 7097d8d 41b5099 72a3ac3 41b5099 7c4e218 0914add e8a0878 0914add 4fa8278 60fd8b2 e14c9c9 4fa8278 e14c9c9 4fa8278 60fd8b2 4fa8278 e14c9c9 4fa8278 0914add 7097d8d 0914add 7097d8d 0914add 7097d8d 79fdefb 7097d8d 79fdefb 7097d8d 0914add 7097d8d 0914add 7097d8d 79fdefb 7097d8d 0914add 7097d8d 30cd106 60fd8b2 30cd106 60fd8b2 94c6382 30cd106 94c6382 60fd8b2 30cd106 60fd8b2 37d210a 30cd106 37d210a 60fd8b2 37d210a 0914add 37d210a ce3aa90 30cd106 72a3ac3 7419377 24d7fb2 7419377 1669c83 72a3ac3 be7e6c4 72a3ac3 30cd106 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 465 466 467 468 469 470 471 472 473 474 475 476 477 478 479 480 481 482 483 484 485 486 487 488 489 490 491 492 493 494 495 496 497 498 499 500 501 502 503 504 505 506 507 508 509 510 511 512 513 514 515 516 517 518 519 520 521 522 523 524 525 526 527 528 529 530 531 532 533 534 535 536 537 538 539 540 541 542 543 544 545 546 547 548 549 550 551 552 553 554 555 556 557 558 559 560 561 562 563 564 565 566 567 568 569 570 571 572 573 574 575 576 577 578 579 580 581 582 583 584 585 586 587 588 589 590 591 592 593 594 595 596 597 598 599 600 601 602 603 604 605 606 607 608 609 610 611 612 613 614 615 616 617 618 619 620 621 622 623 624 625 626 627 628 629 630 631 632 633 634 635 636 637 638 639 640 641 642 643 644 645 646 647 648 649 650 651 652 653 654 655 656 657 658 659 660 661 662 663 664 665 666 667 668 669 670 671 672 673 674 675 676 677 678 679 680 681 682 683 684 685 686 687 688 689 690 691 692 693 694 695 696 697 698 699 700 701 702 703 704 705 706 707 708 709 710 711 712 713 714 715 716 717 718 719 720 721 722 723 724 725 726 727 728 729 730 731 732 733 734 735 736 737 738 739 740 741 742 743 744 745 746 747 748 749 750 751 752 753 754 755 756 757 758 759 760 761 762 763 764 765 766 767 768 769 770 771 772 773 774 775 776 777 778 779 780 781 782 783 784 785 786 787 788 789 790 791 792 793 | ---
library_name: cortiq
license: apache-2.0
base_model:
- Comfy-Org/MiniMax-H3
- larryvrh/MiniMax-H3-Turbo-Lora
base_model_relation: quantized
pipeline_tag: text-to-video
tags:
- cmf
- cortiq
- video
- audio
- 4-bit
---
# MiniMax-H3 Turbo β one file, no Python
[MiniMax-H3](https://huggingface.co/Comfy-Org/MiniMax-H3) renders video and
synchronized stereo audio from one prompt, in one transformer, on two flow
schedules. [larryvrh's Turbo LoRA](https://huggingface.co/larryvrh/MiniMax-H3-Turbo-Lora)
brings it to four sampling steps. This is both of them in the
[CMF container](https://github.com/infosave2007/cmf) β the DiT, the Qwen3-VL
prompt encoder, the video VAE decoder and the audio vocoder in a single
memory-mapped file β running on `cortiq`, a Rust binary with no ML framework
underneath.
The reference checkout is 124.4 GB across four files plus a ComfyUI install.
Here it is **one file between 13.2 and 23.9 GB**, and which one you take is
decided by your VRAM.
## Same prompt, three files
*"A corgi in a chef hat flipping a pancake, sizzling sounds and a cheerful bark."*
β 512Γ288, 39 frames, seed 42, four steps, nothing but the prompt.
|  |  |  |
|---|---|---|
| **23.9 GB** β full 32B prompt encoder | **13.2 GB** β 4B encoder + ClipProj | **18.7 GB** β two-bit encoder |
| the reference | same scene, plainer set | a different animal, no pan |
The 13.2 GB file is the one to take if you have 16β20 GB of VRAM: it holds the
whole run resident instead of paging, and it is still four bits everywhere.
Two bits is published as a dead end, not a choice β
[why, with numbers](#making-it-smaller).
The GIFs are silent and the audio is half the model, so take an
**[mp4](https://huggingface.co/infosave/MiniMax-H3-Turbo-cmf/resolve/main/samples/ab_clipproj.mp4)**.
One transformer denoises picture and sound together in one packed sequence.
### Which file
| file | size | VRAM | keyframes | |
|---|---|---|---|---|
| `mmh3-turbo-fl2va-q4tp.cmf` | 23.94 GB | 24 GB+ | yes | **the default** |
| `mmh3-turbo-q4tp.cmf` | 23.47 GB | 24 GB+ | no | same, without the vision tower |
| **`mmh3-turbo-clipproj4b-q4tp.cmf`** | **13.16 GB** | **16β20 GB** | no | **the small one**, peaks at 15.1 GB |
| **`mmh3-turbo-clipproj4b-fl2va-q4tp.cmf`** | **14.48 GB** | **16β24 GB** | **yes** | the small one WITH start/end frames β the 4B vision tower and the VAE encoder join the compact build |
| **`mmh3-turbo-clipproj4b-fl2va-v2-q8_2f.cmf`** | **26.90 GB** | **32 GB+** | **yes** | **eight bits**: the two-field int8, `w = qΒ·row[o]Β·col[i]`. More weight fidelity, and it runs on the card as of 0.5.94 β still without the fused kernels q4tp has |
| `mmh3-turbo-fl2va-q2tp.cmf` | 18.74 GB | β | yes | don't render with this |
Text-to-video everywhere; the `fl2va` files also take a first and/or last frame
(`--first-frame`/`--last-frame`, binary P6 PPM). The compact
`clipproj4b-fl2va` file was verified frame-for-frame: the conditioning
image comes back as frame 0 of the render. One honest caveat: its
ClipProj projection was fitted on text-only encoder activations, so with
a picture in the prompt the conditioning quality on complex scenes is
still being compared against the full-encoder fl2va files β report what
you see in the discussions.
([keyframes](#keyframe-to-video)). The release's third path, `ref2va`, is not
ported. The Turbo LoRA is merged into the weights, so the file IS the 4-step
model β nothing else to download. 47.83 B parameters, 2 361 tensors,
`cortiq verify` clean.
## Keyframe to video

```bash
cortiq animate mmh3-turbo-fl2va-q4tp.cmf \
--prompt "the corgi lifts the pan and flips the pancake high, sizzling" \
--first-frame keyframe.ppm --out flip.avi
```
One picture conditions the run twice, and both halves matter. Its VAE latent
becomes a row the DiT holds at a timestep of its own near 1 β a condition, not
noise being removed β and never denoises. The picture ITSELF goes to the prompt
encoder through Qwen3-VL's vision tower, as `"<Picture 1>: "` and a vision
block: at 512Γ288 that is 144 tokens of the 168 the prompt above carries.
Leave one out and the model is conditioned on something the reference never
conditions on.
`--last-frame` anchors the other end. The first frame is a geometry anchor and
is stretched to the canvas; the last one follows and is cover-cropped, which is
what the reference does with each. Frames come in as binary P6 PPM.
## Running it
### 1. Get the runtime
`cortiq` is one Rust binary. Either install it β
```bash
cargo install cortiq-cli # needs Rust 1.85+; brings the GPU backend
```
β or take a prebuilt archive from the
[latest release](https://github.com/infosave2007/cmf/releases/latest)
(Linux x86-64, macOS on Apple Silicon and Intel, Windows x86-64 and ARM64;
each ships a `.sha256`). Nothing else is required: no Python, no PyTorch, no
CUDA toolkit, no ffmpeg.
Check it took:
```bash
cortiq --version
```
### 2. Get the weights
One file. Pick it from [Which file](#which-file) above β this is the
text-to-video default; swap the name for `mmh3-turbo-clipproj4b-q4tp.cmf` if
you are on a 20 GB card.
```bash
pip install -U "huggingface_hub[cli]" # only to fetch the file
hf download infosave/MiniMax-H3-Turbo-cmf mmh3-turbo-q4tp.cmf --local-dir .
```
Confirm it arrived whole β the container carries a hash per tensor:
```bash
cortiq verify mmh3-turbo-q4tp.cmf # β β all tensor hashes match
cortiq info mmh3-turbo-q4tp.cmf # β arch, layers, 47.83B params
```
### 3. Render
```bash
cortiq animate mmh3-turbo-q4tp.cmf \
--prompt "A corgi in a chef hat flipping a pancake, sizzling sounds and a cheerful bark." \
--width 512 --height 288 --frames 39 --steps 4 --seed 42 \
--out corgi.avi
```
That writes `corgi.avi` β MJPEG video with PCM stereo, playable in VLC, mpv,
QuickTime and Windows Media Player β and `corgi.wav` beside it. The JPEG
encoder and the RIFF muxer are inside the binary: a pipeline that ends in a
shell-out to a 20 MB dependency is not a pipeline you can ship. If you want an
mp4 for a browser, remux it yourself; the model never needs one.
**On a GPU.** Nothing to opt into any more: `cortiq` probes this file's own
first qkv weight against the host on startup and takes the device arm only if
they agree, so the arm that renders is the arm that was checked.
```bash
cortiq animate mmh3-turbo-q4tp.cmf --prompt "β¦" --out corgi.avi
```
On one RTX 5090, 512Γ288, 39 frames: **60.2 s at the default four steps**
(91.6 s at eight, 42.8 at two). `RUST_LOG=info` prints a per-stage breakdown
for every run, and the full table is under *What it costs to run* below. The
whole pipeline stays on the card β
the DiT block, both VAE decoders, and the vocoder's dilated convolutions β so
nothing but the finished frames crosses the bus. `CMF_MMH3_GPU=0` forces the
host path if you want to compare.
**Two cards.** A render does not split across them, and should not: the DiT
block and both decoders are already resident, so a second card would only add
a bus crossing to a pipeline that no longer has one. Two cards double your
*clips*, not your clip β run two processes, one pinned to each:
```bash
CMF_GPU_ADAPTER=0 cortiq animate model.cmf --prompt "β¦" --seed 1 --out a.avi &
CMF_GPU_ADAPTER=1 cortiq animate model.cmf --prompt "β¦" --seed 2 --out b.avi &
wait
```
`cortiq gpu` lists the cards and their indices. For text models the same
binary both splits and replicates across cards β see
[docs/MULTI_GPU.md](https://github.com/infosave2007/cmf/blob/master/docs/MULTI_GPU.md).
### Options that matter
| flag | default | what it does |
|---|---|---|
| `--width` / `--height` | 512 Γ 288 | multiples of 32. The trained short edge is 768; below ~256 the model drifts off-distribution |
| `--frames` | 39 | at 24 fps, snapped **up** to the model's 17k+5 grid: 5, 22, 39, 56, β¦ 124. 124 β 5 s, and 124β362 is the validated range |
| `--steps` | 4 | what the Turbo LoRA is trained for. More still helps a little |
| `--seed` | 42 | same seed, same prompt, same size β the same clip, byte for byte. This is now true on the GPU too: the op arbitration used to alternate arms on real data while it made up its mind, and two runs of one binary could differ |
| `--quality` | 92 | JPEG quality of the AVI's frames |
| `--stock-sampler` | off | integrate the audio on the video's clock, as a single-schedule sampler does. Wrong at 4 steps β it is here to hear how wrong |
| environment | what it does |
|---|---|
| `CMF_MMH3_GPU=1` / `=0` | force the device or the host path instead of letting the parity probe choose |
| `CMF_GPU_PROBE=0` | pin the op arbitration (already the default for `animate`, so a seed reproduces) |
| `CMF_THREADS=n` | cap the worker pool (defaults to the machine's cores) |
| `CMF_ANIM_PROF=1` | per-step rms of both latent streams and both velocities |
| `CMF_MM_KILL=0` | never fall back to the host on a slow device op. The engine treats three consecutive over-budget GEMMs as "another process owns the card" and finishes the run on the CPU; a weight paging in from disk is exempt, but on a machine where the file exceeds RAM the first steps can be slow for reasons that are not contention (see the field notes) |
### What it needs
RAM at least the file's size β 24 GB β or every step faults on non-resident
pages; the weights are memory-mapped, not read. Disk: 24 GB. A GPU is optional
and wants ~14 GB of VRAM for the DiT's planes. No network access at run time.
`mmh3-turbo-clipproj4b-q4tp.cmf` asks for 14 GB of RAM and disk instead, and
peaks at 15.1 GB of VRAM β a whole-run maximum, polled, not a snapshot. That is
what makes it the one to reach for on a 20 GB card: below that the prompt
encoder pages, and on this workload paging costs more than any kernel.
### Field notes: 24 GB Macs and 20 GB cards
Two users measured what this card could not, and the engine changed on
their reports (HF discussions #1, #2 and #4). The numbers, as they sent them:
| machine | file | render | result |
|---|---|---|---|
| RTX 3080 20 GB, Windows | `mmh3-turbo-q4tp` (25.2 GB) | 512Γ288Γ39, 4 steps | **1218 s** β the encoder did not fit the 19 GB budget and paged every step |
| RTX 3080 20 GB, Windows | `mmh3-turbo-clipproj4b-q4tp` (13.2 GB) | same | **105 s** β resident; denoise 62.8 s, video VAE 30.6 s |
| Mac mini M4 24 GB | `mmh3-turbo-clipproj4b-q4tp` | 512Γ256Γ39, 4 steps | **174 s** β denoise 122 s, video VAE 41 s |
| Mac mini M4 24 GB | `mmh3-turbo-clipproj4b-fl2va-q4tp` (14.5 GB) | 22 frames from a keyframe | **92 s**, no swap |
| Mac mini M4 24 GB | `mmh3-turbo-fl2va-q4tp` (25.7 GB), 0.5.79 | 10 / 40 frames | **48 s / 140 s per denoise step** on the GPU, 20 GB resident, 0.3 GB swap |
| Mac mini M4 24 GB | same, 150 frames | | 8 GB of swap and a sawtooth β the activation cache no longer fits |
| Mac mini M4 24 GB | `mmh3-turbo-clipproj4b-fl2va-v2-q4tp` (14.5 GB) | 448Γ768, 50 frames | **136 s per denoise step**, faces hold through the clip; 90 frames at that size is the paging threshold |
| Mac mini M4 24 GB | `mmh3-turbo-fl2va-q4tp` (25.7 GB) | 512Γ256 | safe to ~70β90 frames; **768Γ448 caps at 40β50** (90 frames: 22.2 GB RAM + 1.1 GB swap, the GPU stalls) |
| Mac mini M4 24 GB | any file, `--first-frame` | | the video-VAE **encoder** of the keyframe ran ~100β140 s on the CPU (`encode 0/1 (140.8s)`); text-to-video skips it (0.1 s) |
What follows from them, if you are on such a machine:
- **On a 24 GB Mac the 25.7 GB file works after 0.5.79, for clips of
β€ 40 frames.** The prompt encoder's pages are released after the text
encode; what remains β DiT, VAE decoders and the activation cache β fits
up to ~40 frames. Past that the cache spills to swap. For longer clips
chain 40-frame chunks (`--last-frame` of one render becomes
`--first-frame` of the next) or use the `clipproj4b` files, which leave
the room.
- **The Metal driver budgets by buffer size, not resident pages.** The
weight arena's overlapping windows read as ~27 GB to the driver even
after the encoder is released, so the first ops after the encode page
from the SSD and can take seconds. Before 0.5.80 a single such op tripped
the contention kill and the rest of the run walked the CPU (>60 s a
step); the kill now needs three consecutive strikes, exempts weights
that were not resident, and `CMF_MM_KILL=0` turns it off. If a run still
says `device contended, CPU for the rest of the process` on a machine
nobody else is using, that variable is the answer.
- **The kill is disarmed during the prompt encode (0.5.80).** The
encoder is a one-shot pass over 12 GB of weights; on a machine the file
does not fit it streams from disk, and its GEMMs run over any budget
for reasons that are not contention β the report that had to gut
`mm_kill` in the source was on exactly that. The kill now arms only
when the denoise loop starts, and three strikes are still needed.
- **The keyframe encode is CPU work today.** With `--first-frame` the
video-VAE encoder runs the reference picture on the host (~100β140 s
on an M4 for a 448Γ768 frame; 0.1 s without a keyframe). 0.5.80 puts
the pool's workers on the performance cores (they were landing on the
E-cores with the P-cores asleep β that report's `asitop`), which
shortens it; a device path for the encoder's 3-D convolutions is the
real fix and is on the list.
- **Frame budgets on 24 GB, from that user's sweep:** with the 25.7 GB
file 512Γ256 is safe to ~70β90 frames and 768Γ448 caps at 40β50; with
the 14.5 GB v2 file 448Γ768 at 50 frames is the sweet spot (136 s per
step) and 90 frames is where paging starts. Past those the machine
swaps and the GPU stalls to nothing β better to chain clips than to
push the frame count.
- **The draft β final workflow.** Block the shot on `clipproj4b-fl2va-v2`
(90 s on the Mac; v2 holds faces through a clip β "25 GB-level face
consistency at 14 GB speed", that user's words), then pay the full
encoder once for the final take. Identity of *specific* real people is
still the full encoder's territory.
- **`--height 256` instead of 288** halves the video-VAE decode on every
machine (three 256-pixel tiles instead of six).
- **Voices are prompt space *here*, not in the release.** Speaker identity,
timbre, pace and emotion respond to stage directions in the prompt, and a
fixed seed keeps the same actor across takes β which is how to work with
this file today. But the earlier claim on this card that "H3 has no
reference-audio input" was **wrong**, and a user was right to push back
(discussion #3): the release is tagged `audio-to-audio-video` and
`video-to-audio-video`, and the DiT takes both as conditioning rows. What
is missing is on our side and it is not training β see
[What this port does not do yet](#what-this-port-does-not-do-yet).
## Draft small, refine big (two-stage with LTX-2.5)
Since 0.5.100 the Turbo draft can hand its frames to
[LTX-2.5](https://huggingface.co/infosave/LTX-2.5-cmf) for a short
refinement pass at full resolution β the draft supplies composition and
motion, LTX reinvents the detail:
```bash
cortiq animate mmh3-turbo-clipproj4b-fl2va-v2-q4tp.cmf \
--prompt "β¦" --width 512 --height 288 --frames 39 --frames-dir draft/
# upscale the frames 2x (any bicubic) and trim to 33 β
# H3 renders 17k+5 frames, the LTX VAE takes 8k+1
cortiq ltx-video --model ltx25-q4tp.cmf --prompt "β¦same promptβ¦" \
--video draft2x/ --video-strength 0.45 --width 1024 --height 576 --frames 33
```
On a 5090: draft 40 s, refine 499 s at 1024Γ576. A/B of the same seed is
in `samples/refine_ab_*.mp4`. The two containers don't want RAM at the
same time β run the stages sequentially. `--video-strength` is the dial:
0.45 keeps the draft's motion, higher repaints more.
## Making it smaller
Half this file is the PROMPT ENCODER β 12.2 GB of Qwen3-VL against the DiT that
actually draws. Two ways to act on that were tried, and the three clips at the
top of this card are the result: **squeezing the encoder does not work,
replacing it does.** Here is why, in both directions.
### Two bits: smaller, faster, and answering a different question
The obvious next cut is the DeepSeek-V4 policy: gate/up at two bits,
everything else at four. It builds β 23.94 GB down to **18.74**, and
with a device kernel of its own it renders *faster* than the four-bit
file β 217.3 s against 258.4 when the pair was measured, because there is
less weight to move. (Both numbers are from the build of that day; the
four-bit file renders the same clip in 60.2 s now. The ratio is what
carries over, not the seconds.)
`cortiq verify` passes. The file is here.
It also stops following the prompt, which is why it is not the one to
reach for.
| | |
|---|---|
|  |  |
| `q4tp` β 23.94 GB | `q2tp` β 18.74 GB |
Same prompt, same seed, same four steps. On the left the corgi is behind
a pan with batter in it, drawn flat and clean, which is what was asked
for. On the right it is a different animal in a different style with
**no pan and no pancake at all**, over a washed-out ground with visible
texture noise. That is not a quantizer trading detail for size; that is
a model answering a different question.
The likely culprit is where the two bits landed. Half this file is the
PROMPT ENCODER, and the policy put two bits on its gate/up planes along
with the DiT's β so the loss falls on the part that decides what the
clip is about, not on the part that draws it. A two-bit build confined
to the DiT would save ~2.9 GB instead of 5.2 and is the version worth
measuring next β the packer's policy is one predicate,
`is_wide_plane`, if you want to try it. The file above is published so
that experiment starts from something rather than nothing; four bits is
what to render with.
### A smaller encoder: replace it, don't squeeze it
The two-bit section above ends on a diagnosis: half this file is the PROMPT
ENCODER, and squeezing it is what broke the prompt. There is a second way to act
on that diagnosis β **don't compress the encoder, replace it.**
[ClipProj](https://github.com/nicolab28/ComfyUI-ClipProj) fits a ridge
regression from a small Qwen3-VL's hidden state into the space the DiT was
conditioned on. Same tokenizer, same family, one affine map:
```text
cond = ((h - mean_in) / std_in) @ W [+ GELU residual] * std_out + mean_out
```
`mmh3-turbo-clipproj4b-q4tp.cmf` is the text-to-video file with Qwen3-VL-4B
tapped at layer 24 standing in for the 32B tapped at 50, plus the 304 MB
projection kept **exact** β f32, no quantization on the piece that carries the
whole substitution. Everything else is copied through byte for byte: DiT and
both VAE decoders, still `q4tp`. Nothing anywhere is two bits.
**23.47 GB β 13.16 GB.** The encoder went from 552 tensors to 277.
#### Does it still mean the same thing?
That is measurable without rendering a frame: dump what the DiT actually
receives (`CMF_TE_DUMP=<path>`) from both files on one prompt, and take the
cosine per token.
| | mean cosine to the 32B | worst token |
|---|---|---|
| 4B tapped at **24** | **0.9198** | 0.7982 |
| 4B tapped at 25 | 0.8452 | 0.5910 |
| *floor β two different tokens of the 32B itself* | *0.5914* | |
The tap index is 0-based, and off-by-one is a real failure mode rather than a
theoretical one: at 25 the worst token sits **on the floor**, uncorrelated.
Token 0 is not projected at all but replaced by a stored `sink_out` β the
attention sink is an outlier no regression fits β and it lands at cosine
**0.9999** against the teacher's, which is the sharpest single confirmation
that the mechanism is wired right.
#### What it buys
The one number that transfers between machines is the **encode stage**, because
that is the only stage ClipProj replaced. Same prompt, seed and four steps,
512x288x39, measured inside one run:
| | prompt encode |
|---|---|
| `mmh3-turbo-q4tp` β 32B tapped at 50 | 49.2 s |
| `mmh3-turbo-clipproj4b-q4tp` β 4B tapped at 24 | **1.9 s** |
Everything after that is residency, and residency is a property of YOUR card,
not of this file. On the 24 GB RTX 3090 these were taken on, the whole render
came out roughly half the time of the 23.47 GB file and the video VAE decode
more than halved β but totals on that box drifted between 140 s and 210 s for
one unchanged configuration as the card warmed, so treat the ratio as a
direction and measure your own.
The point is the 20 GB card this variant exists for, where the difference is
not a ratio but a cliff: a reader of this repo measured **20 minutes** for one
512x288 clip on an RTX 3080 20 GB, paging the 23.47 GB file through a ~19 GB
weight budget. At 13.16 GB, with a polled whole-run peak of 15.1 GB of VRAM,
there is nothing left to page.
The parity probe takes the device arm on both files at the same `rel rms
4.65e-3`, which is its own small proof that the DiT came through untouched.
#### What it costs in quality
Judge it on the clip at the top of this section, not on a frame: a still
catches the pancake at rest and reads as a loss that is not there.
Across the 39 frames the 4B does the whole job: corgi, chef hat, flat clean
style, and the pancake **leaves the surface and comes back**, the dish empty
underneath at the top of the toss. The verb survives β that is what `q2tp`
lost, along with the animal and the pan.
What drifts is set dressing. The 32B renders the cooking surface as a griddle
and puts patterned wallpaper behind; the 4B gives a rimmed white plate and a
plain ground. A 0.92 cosine is close, not equal, and where it is not equal is
the scene's furniture rather than its subject or its action.
So the ordering is not "smaller is worse". The two-bit file answers a different
question; this one answers the right question with a plainer set, at 56% of
the size. If you have the VRAM for the 32B encoder, use it β
its framing is richer. If you are paging, this is the better trade, and unlike
the two-bit build it is a trade rather than a loss.
Sound is untouched either way: the audio branch never sees the prompt encoder.
[The clip with its audio.](https://huggingface.co/infosave/MiniMax-H3-Turbo-cmf/resolve/main/samples/ab_clipproj.mp4)
#### Building one
```bash
cortiq animate-pack \
--in mmh3-turbo-q4tp.cmf \
--te qwen3vl_4b_bf16.safetensors --te-layers 24 \
--clip-proj mmh3-4b-ClipProj-celeb-mlp.safetensors \
--quant q4tp --out mmh3-turbo-clipproj4b-q4tp.cmf
```
`--te-layers` is not a size knob, it is the tap: layers above it never execute,
so packing them is pure file. `--clip-proj` carries the projection in exact and
replaces the encoder **as a whole component** β matching tensor names alone
would leave `te.layers.24..49` of the 32B behind, six gigabytes that
`num_hidden_layers` then excludes from the forward while disk and VRAM budget
still pay for them.
The encoder is [`Comfy-Org/Qwen3-VL`](https://huggingface.co/Comfy-Org/Qwen3-VL)
`text_encoders/qwen3vl_4b_bf16.safetensors`, and the projection is
[`NicoLab28/ClipProj-MiniMax-H3`](https://huggingface.co/NicoLab28/ClipProj-MiniMax-H3)
`mmh3-4b-ClipProj-celeb-mlp.safetensors` β use the file the projection was
fitted on, since the map is only valid for those exact weights. An 8B
projection is published too and scores higher on its author's corpus (0.8037
against 0.7930); it costs about 4 GB more.
## What it costs to run
The file is memory-mapped, so plan on RAM at least its size or every step
touches non-resident pages.
One RTX 5090, 4 steps, 39 frames. `RUST_LOG=info` prints this breakdown for
every run:
| | text | denoise | video VAE | audio VAE | total |
|---|---|---|---|---|---|
| 512Γ288 | 2.5 s | 34.8 s | 16.6 s | 4.3 s | **60.2 s** |
| 512Γ256 | 2.3 s | 31.4 s | 8.4 s | 4.3 s | **48.5 s** |
| 256Γ160, 22 frames | | | | | **15.9 s** |
| 512Γ288, host only (`CMF_MMH3_GPU=0`) | 2.6 s | 363.5 s | 271.5 s | 7.1 s | **646.1 s** |
The card is **10.7Γ the host path** on the same machine, same seed. At 8
steps a 512Γ288 clip is **91.6 s**; at 2 it is 42.8. The same 4-step
render took 172 s when this card was written β the pipeline has since moved
onto the card end to end, both VAE decoders with it.
Nearly all of the decode is the video VAE β the vocoder is 4.3 s of it. The
packed sequence is `[text | audio | video]` and everything attends to
everything, so cost grows with the token count and then with its square: a
512Γ288 second is five times the tokens of a 256Γ160 one.
**A free 2Γ on the decoder, if you want it.** The video VAE decodes in
256-pixel tiles, always, and grows the OVERLAP rather than the tile count β
so a 288-pixel edge is covered by two 256-pixel tiles overlapping by 224, and
you pay for 512 rows to get 288. An edge of exactly 256 is one tile. 512Γ256
therefore decodes three tiles where 512Γ288 decodes six, for 89% of the
pixels. Measured on the current build: the video VAE goes 16.6 s β 8.4 s and
the whole render 60.2 s β 48.5 s. The schedule is the reference's and this
port reproduces it exactly; picking an edge that lands on it is free.
**Host and device do not agree to the last bit, and neither is wrong.** The
host arm quantizes activations to int8 (`CMF_SDOT`) where the device
dequantizes to f32, so the two renders differ by a few per cent in latent rms
and visibly in fine texture. Set `CMF_SDOT=0` on both sides to compare
arithmetic instead of that approximation.
**Why the device took a while to trust.** It is no longer opt-in β the parity
probe decides per file β but getting there took three fixes, and one thing is
still held back.
The engine's blocked f32 GEMM cached its weight-side device buffer **by
pointer address**. Every batched attention allocates one k/v scratch pair per
call and refills it per head β same address, different matrix β so head 0's
keys came back for every head, on the GPU only, silently. It is keyed on a
content fingerprint now. The same GEMM also took every job over 4 M MACs on
sight with no CPU arm to lose to, which on this model's decoder was three
times *slower* than the host it displaced; it goes through the same
measure-don't-assume probe as every other op class now, and on this stack the
probe hands that work back (0.24 ms device against 0.13 host) while sending
the weight GEMMs to the card (25.8 ms against 92.0).
Still held: **the cooperative-matrix kernel runs this model out of f16 range.**
At 256Γ160 the render is correct; at 512Γ288 the audio stream goes NaN on the
second sampling step and the video follows. Bisected β `CMF_BAKE_GPU=0` does
not help, `CMF_COOP=0` does β so `cortiq animate` pins `CMF_COOP=0`.
That hold is specific to this model, not a verdict on the kernel: the image
model on the same card and the same kernel renders 20.5 s without it against
14.8 with, and the two agree to 42.6 dB β the price of f16 operands, which
the kernel documents, not a fault. MiniMax-H3's activations are simply larger.
Giving that kernel a scale is the next real speedup here.
## What the conversion did
**The adaLN collapse.** Forty per cent of the released DiT is one matrix per
block: `adaln_proj.linear` is `[96768, 2688]`, 520 MB at bf16, **13 B of the
model's 33 B parameters** β for a map whose input is one number, the timestep.
Its output over the whole schedule is a one-dimensional curve in R^96768, and
Comfy-Org's `pruned` checkpoints already ship it as one: an `adaln_t_table` of
`[1025, 8]` shared by every block and per-block weights of `[96768, 8]`.
Measured against the full matrix on block 0 (`tools/mmh3_fetch.py check`, which
range-reads 520 MB out of the 66 GB file rather than downloading it):
```
adaln max|Ξ| 8.0e-4 rms 8.7e-5 against a signal of rms 0.464
time-curve singular values 1..12, relative:
1.00e0 2.96e-1 1.05e-1 6.63e-2 6.60e-3 2.11e-3 5.61e-4 2.92e-4
3.67e-5 2.73e-5 1.32e-5 1.34e-6
```
The ninth singular value is already 3.7e-5 of the first. Rank eight is not an
approximation anyone should feel nervous about; the 26 GB is redundant.
The Turbo LoRA is written against the FULL matrix (`lora_A` is `[16, 2688]`),
which is why the ComfyUI node re-injects the time conditioning at run time when
the base is pruned. `cortiq animate-pack` does it once, at conversion:
```
adaln(t) = W_p Β· u(t) + b + B Β· (A Β· silu(e(t)))
= [W_p | B] Β· [u(t) ; A Β· silu(e(t))]
```
β a rank-24 curve, driven by a `[1025, 24]` table per block. 4.6 MB a block
instead of 520, with the LoRA already inside it.
**The rest.**
- **Backbone** β bf16 β `q4tp`, 4.16 bits a weight with a predicted per-row
scale ladder. The LoRA's rank-64 update is merged before quantizing.
- **Prompt encoder** β Qwen3-VL-32B truncated to 50 layers, 51.5 GB β 12.2 GB.
It is the largest single component of the file and it runs once per
generation.
- **Video VAE β decoder only.** It is a ViT3D, not a conv stack: 36 transformer
blocks over the latent grid and one linear that expands each cell into a
4Γ16Γ16 block of pixels. The 3-D causal CNN encoder is a third of the
checkpoint and text-to-video never runs it.
- **Audio VAE β decoder only**, f16. Quantizing a vocoder buys 45 MB and costs
audible hiss. Its 254 kaiser-sinc resampling filters are read from the
checkpoint rather than re-derived β the design formula is in the code as a
fallback, but a filter you compute is a filter that can drift from the one
the weights were trained against.
- Integrity: 47.83 B parameters over 2 361 tensors; `cortiq verify` checks
every one against the directory's hashes.
## On parity
Established, not assumed, and separately for each of the four stacks. The
reference is ComfyUI's own module, run on a toy checkpoint carrying the
release's real tensor names and the release's real schedules β `tools/`
builds them, `tools/mmh3_toy_gate.sh` runs the diff. The packs are exact f32
on purpose: `q4tp`'s noise floor sits an order of magnitude above the
arithmetic difference these are looking for, so quantizing here would pass a
broken port.
| stack | worst | rms | signal rms |
|---|---|---|---|
| DiT β video velocity | 8.8e-5 | 2.1e-5 | 0.515 |
| DiT β audio velocity | 5.2e-5 | 2.5e-5 | 0.409 |
| DiT β token refiner | 8.3e-7 | 2.6e-7 | 1.003 |
| Qwen3-VL encoder | 1.1e-6 | 3.3e-7 | 0.812 |
| video VAE decoder | 4.2e-7 | 4.0e-8 | 0.470 |
| audio VAE decoder | 1.7e-9 | 3.5e-10 | 8.9e-4 |
A dozen conventions in this model pass at one token and fail differently at a
hundred, which is why the toys are not one-vector unit tests: the packed
layout's cursor, the video time axis's 1,4,4,4,4 span pattern, which 96 of 128
head dimensions rotate, the adaLN row order (timestep-major, modality-minor),
the video VAE's 256-pixel tiling β global attention makes a tile a different
computation from a whole frame, so the tiling is part of the output, not a
memory strategy β and the audio stream's separate clock.
## Two clocks
The video and audio latents ride different flow schedules (shift 12 and 3).
The sampler walks the video grid, which at four steps is
`1, 0.973, 0.923, 0.8, 0`, and integrates the audio on its own remap of it.
Stepping both on the video grid is what a stock sampler does; it is fine at
twenty steps and wrong at four, because over the last interval ΞΟ_a and ΞΟ_v
differ by a factor of three and no per-step slope correction survives a step
that large. `--stock-sampler` reproduces the broken behaviour if you want to
hear it.
## What this port does not do yet
The release is tagged for six conditioning paths. This container packs one of
them β `fl2va`, text and/or keyframes β video + audio. What the others need is
listed here so nobody has to guess whether it is a missing feature or a
missing possibility. **None of them needs training.**
| the release's path | what it takes | what is missing here |
|---|---|---|
| `audio-to-audio-video` | a reference soundtrack | the audio VAE's **encoder half**. The container carries the decoder only β `pack_audio_vae` skips the encoder, `pre_block` and the mean/logs heads as unused. The DiT side already exists: the packed layout has a reference-audio segment kind and its own condition timestep |
| `video-to-audio-video` | a reference clip | the sampler plumbing. The video VAE **encoder is already packed** in the `fl2va` files β it is what encodes `--first-frame` β so this is a layout and CLI change, not a weight change |
| `ref2va` (subject / character references) | 1βN reference images | a **different DiT checkpoint** (`minimax_h3_ref2va_*`) with its own turbo LoRA. It would be a second container, not a flag on this one |
Two more, on the runtime side: the latent upscaler published for H3
(a 345 M-parameter 3-D conv net) is not ported, and there is no fused
device kernel for adapter branches on this model β see below.
## The eight-bit build, and what it costs today
`mmh3-turbo-clipproj4b-fl2va-v2-q8_2f.cmf` (26.90 GB, 26.21 B parameters,
`cortiq verify` clean) packs the DiT as **`q8_2f`** β the two-field int8,
`w = qΒ·row[o]Β·col[i]`: eight bits with a second scale field along the *input*
axis, which is where an activation-outlier channel shows up from the weight
side. As a codec it is strictly more faithful than the four-bit ladder.
**It renders on the card β after two gates were fixed, and the story is
worth reading if you pack your own containers.**
The first run of this file left the RTX 5090 idle at 2 MiB and took
**357.3 s**. Neither cause was a missing kernel:
1. **The startup probe matched on `Q4TiledP` by name and dtype.** Finding no
four-bit qkv weight it declared the host path for the whole render β for a
codec that has a device GEMM of its own. The two-field int8 folds its
column field into the activation and what is left is the per-row int8
kernel both backends already ship. The probe now asks the tensor which
entry point its codec has (`QTensor::device_matmat`).
2. **The weight-residency budget was the whole heap minus a gigabyte.** That
survived only because every container published before this one had
weights far under it; 24 GB of eight-bit weights took the card and the
first scratch allocation died with `wgpu error: Out of Memory`. A quarter
of the heap is held back now β 32 GB card β 24 GB of weights.
With both: **171.5 s** on the same card and clip (103.3 s denoise,
62.2 s video VAE), against 357.3 s on the host, and the probe agrees with
the CPU arm to 5.77e-3.
It is still slower than the four-bit file's 60.2 s, and that part *is* about
kernels: `q4tp` has the fused qkv β attention β output submission and the
packed FFN, `q8_2f` goes through the generic per-op GEMM. Fusing those for the
two-field codec is the next piece of work, and it is worth roughly what the
fusion is worth on q4tp.
## The latent upscaler
Render small, resize the **latent**, decode once:
```bash
cortiq animate mmh3-turbo-clipproj4b-fl2va-v2-q4tp.cmf \
--prompt "β¦" --width 512 --height 288 --frames 39 \
--upscale minimax_h3_latent_upscaler_3d_fp16.safetensors --upscale-by 2.0 \
--out big.avi
```
The net is [LBH-123-AI's](https://huggingface.co/LBH-123-AI/Minimax_h3_latent_Upscaler)
β 345 M parameters, twelve residual blocks and six temporal convolutions on
each side of a trilinear resize, with a scalar scale embedding modulating
every block. It is ported into the engine and gated against the node's own
torch module on the same weights: **worst 6.68e-6, relative rms 3.92e-7**.
|  |  |
|---|---|
| the 512Γ288 render, resampled to 1024Γ576 | the same latent through the net, decoded at 1024Γ576 |
Same seed, same prompt, same denoise. On the left a resampler has nothing to
add and softens hair, coat seams and the wet road; on the right the detail is
generated. **Measured on an M4 (24 GB), 22 frames: denoise 76.9 s, upscale
13.8 s, VAE 98.2 s β 193 s for a 1024Γ576 clip** whose denoise only ever paid
for 512Γ288. The upscale itself is the cheapest stage in that list.
Why it matters on a 24 GB machine: the alternative is decoding through a
5 B-parameter VAE, resizing pixels and encoding again β the expensive path β
or interpolating the latent, which is the cheap path that ghosts. This is
neither. You keep the frame budget of the small render and get the detail of
the large one, and the file is loaded beside the container like an adapter,
so it stays under its own licence.
## Reference clips
```bash
cortiq animate model.cmf --prompt "β¦" --video ref_frames/ --video-stride 8
```
Every eighth `.ppm` of the directory becomes a condition pinned to its own
moment of the render, and the source clip is mapped onto the render's length
by position β a 100-frame reference conditions a 39-frame clip at the same
*moments*, not the same indices. This is the `fl2va` keyframe path with more
than two frames. It is **not** a port of the release's own `v2v` node, which
conditions differently; what it gives you is composition and motion carried
across a chain of shots, which is what the keyframe hack in discussion #6 was
reaching for.
## Adapters at runtime
Community LoRAs for H3 run against this container as they ship:
```bash
cortiq animate mmh3-turbo-clipproj4b-fl2va-v2-q4tp.cmf \
--prompt "r34l1sm a woman in a red raincoat on a neon street, close-up" \
--lora h3-realism-people.safetensors --lora-strength 0.8 \
--out take.avi
```
`--lora` reads a `.safetensors` in any of the three conventions in the wild
(`diffusion_model.β¦`, `base_model.model.dit.β¦`, or the bare module path), at
F32/F16/BF16, with either `lora_A`/`lora_B` or `lora_down`/`lora_up` naming.
It binds `attn.qkv_proj`, `attn.out_proj`, `mlp.fc1`, `mlp.fc2` on all fifty
blocks and on the two token-refiner blocks, and it prints what it bound:
```
lora: rank 32, 104/104 branches bound
```
Branches it cannot bind are named rather than dropped in silence. The one real
gap is `adaln_proj.linear`: this container carries the modulation as a rank-24
curve over the timestep (that collapse is 40% of the released model's weights
and most of why the file is 14 GB), and folding an adaLN update into it needs
the time embedding, which only the packer has. Adapters that touch adaLN β
the spatial-physics and streaming ones β therefore land partially, and say so.
Bake them in instead with `animate-pack --lora β¦ --time-embedder β¦`, which is
exactly how the Turbo LoRA got into these files.
**What it costs.** On an M4 (24 GB), 512Γ288, 22 frames, four steps, with
`fal/MiniMax-H3-Realism-People-LoRA` (rank 32, 104 branches, all of which
bind), the two runs taken back to back with memory free: **denoise 133.6 s
with the adapter against 126.1 s without β 1.06Γ.** The video VAE, which no
adapter touches, moved 53.4 β 56.9 s between those same runs, so on this
machine the adapter costs at or below the noise. The branch itself is half a
per cent of the projection's arithmetic and rides inside the base GEMM's Metal
submission; what it can cost is the *attention* fusion standing down, because
a branch on `qkv_proj` or `out_proj` needs the panels that fusion keeps on the
card β visible where the DiT is fully resident, hidden where weights stream.
`--lora-strength 0` reproduces the base render byte for byte.
**Which parts of an adapter matter.** `CMF_LORA_PROBE=1` prints every branch by
its measured contribution `βsΒ·ΞYβ/βYβ`, and `CMF_LORA_ROUTE=<r>` switches off
the ones below `r`. For the Realism adapter the loudest branch is 150Γ the
quietest, blocks 0β2 contribute nothing it would miss, and 41 of its 104
branches carry the look:
```
lora branches by contribution βsΞYβ/βYβ (41 of 104 live):
0.1633 on blocks.13.attn.qkv_proj
0.0946 on blocks.9.attn.qkv_proj
β¦
0.0011 off blocks.1.attn.qkv_proj
```
Rendered on those 41, the clip still looks like the adapter and not like the
base (13.34 dB PSNR against the base, where the full adapter is 13.76, and
16.75 dB against the full adapter). It did not make the step shorter on Metal
β the reason, and where routing does pay, is in
[docs/LORA.md](https://github.com/infosave2007/cmf/blob/master/docs/LORA.md).
## Provenance
Weights derive from MiniMax's H3 release as repackaged by Comfy-Org, and from
larryvrh's Turbo LoRA; both remain under their own licences. The Turbo LoRA is
a **preview** β its own card notes plastic-looking skin and over-sharp grain at
`ckpt850`, and nothing here changes that. The CMF container and the cortiq
runtime are Apache-2.0 (see the repository's LICENSE and PATENTS.md).
|