infosave's picture
Upload README.md with huggingface_hub
30cd106 verified
|
Raw History Blame
8.71 kB
metadata
license: apache-2.0
base_model:
  - Comfy-Org/MiniMax-H3
  - larryvrh/MiniMax-H3-Turbo-Lora
base_model_relation: quantized
pipeline_tag: text-to-video
tags:
  - cmf
  - cortiq
  - video
  - audio
  - 4-bit

MiniMax-H3 Turbo β€” one 23.5 GB file, no Python

MiniMax-H3 renders video and synchronized stereo audio from one prompt, in one transformer, on two flow schedules. larryvrh's Turbo LoRA brings it to four sampling steps. This is both of them in the CMF container β€” the DiT, the Qwen3-VL prompt encoder, the video VAE decoder and the audio vocoder in a single memory-mapped file β€” running on cortiq, a Rust binary with no ML framework underneath.

reference checkout here
diffusion model 66.3 GB (bf16) β€”
prompt encoder 51.5 GB (bf16) β€”
video + audio VAE 5.8 GB β€”
Turbo LoRA 0.8 GB β€”
total 124.4 GB, four files + a ComfyUI checkout 23.5 GB, one file

47.83 B parameters, 2 361 tensors, cortiq verify clean.

The LoRA is not a separate download: it is merged into the weights, so the file IS the 4-step model.

Text-to-video only. The release also takes first/last keyframes (fl2va) and reference images, videos and audio (ref2va); those paths are not ported and the vision tower is not packed. What is here is t2va: prompt in, video and audio out.

Getting a file

huggingface-cli download infosave/MiniMax-H3-Turbo-cmf \
  --include 'parts-q4tp/part_*' --local-dir .
cat parts-q4tp/part_* > mmh3-turbo-q4tp.cmf

cortiq animate mmh3-turbo-q4tp.cmf \
  --prompt "A corgi in a chef hat flipping a pancake, sizzling sounds and a cheerful bark." \
  --width 512 --height 288 --frames 39 --out corgi.avi

cortiq animate writes an MJPEG+PCM AVI and the stereo track beside it as .wav. There is no ffmpeg in the loop and nothing to install: the JPEG encoder and the RIFF muxer are part of the binary, because a pipeline that ends in a shell-out to a 20 MB dependency is not a pipeline you can ship.

Width and height are multiples of 32. Frame counts snap up to the model's 17k+5 grid (5, 22, 39, 56, … 124), which is the grid the temporal VAE and the DiT's frame-span pattern agree on. 4 steps is what the LoRA is trained for; more still helps a little.

What it costs to run

The file is memory-mapped, so plan on RAM at least its size or every step touches non-resident pages.

Measured on 48 CPU cores, 4 steps:

tokens in the pack denoise (4 steps) decode total
256Γ—160, 22 frames (0.9 s) 375 45.0 s 10.7 s 55.7 s
512Γ—288, 39 frames (1.6 s) 1 879 220.8 s 150.0 s 371.3 s

Nearly all of the decode is the video VAE β€” the vocoder is 4 s of the 150.

The packed sequence is [text | audio | video] and everything attends to everything, so cost grows with the token count and then with its square. A 512Γ—288 second is five times the tokens of a 256Γ—160 one.

The GPU path is off, on purpose. cortiq's wide-GEMM arm on wgpu is measured three times faster than the host here and it is also wrong: on an RTX PRO 6000 Blackwell the DiT's first step disagrees with the CPU (video velocity rms 1.32 against 1.73, audio 0.16 against 1.01) and the second step returns NaN β€” a flat grey frame and a clipped waveform. The op probe cannot notice: it only times the two arms, and it runs the bad one for real while it is deciding. So cortiq animate pins CMF_GPU=0 before the backend comes up, which is early enough to be sure. CMF_MMH3_GPU=1 opts back in if you want to watch it fail. Fixing that kernel is the next thing worth doing to this model β€” it is the whole difference between minutes and hours.

What the conversion did

The adaLN collapse. Forty per cent of the released DiT is one matrix per block: adaln_proj.linear is [96768, 2688], 520 MB at bf16, 13 B of the model's 33 B parameters β€” for a map whose input is one number, the timestep. Its output over the whole schedule is a one-dimensional curve in R^96768, and Comfy-Org's pruned checkpoints already ship it as one: an adaln_t_table of [1025, 8] shared by every block and per-block weights of [96768, 8].

Measured against the full matrix on block 0 (tools/mmh3_fetch.py check, which range-reads 520 MB out of the 66 GB file rather than downloading it):

adaln  max|Ξ”| 8.0e-4   rms 8.7e-5   against a signal of rms 0.464
time-curve singular values 1..12, relative:
  1.00e0 2.96e-1 1.05e-1 6.63e-2 6.60e-3 2.11e-3 5.61e-4 2.92e-4
  3.67e-5 2.73e-5 1.32e-5 1.34e-6

The ninth singular value is already 3.7e-5 of the first. Rank eight is not an approximation anyone should feel nervous about; the 26 GB is redundant.

The Turbo LoRA is written against the FULL matrix (lora_A is [16, 2688]), which is why the ComfyUI node re-injects the time conditioning at run time when the base is pruned. cortiq animate-pack does it once, at conversion:

adaln(t) = W_p Β· u(t) + b + B Β· (A Β· silu(e(t)))
         = [W_p | B] Β· [u(t) ; A Β· silu(e(t))]

β€” a rank-24 curve, driven by a [1025, 24] table per block. 4.6 MB a block instead of 520, with the LoRA already inside it.

The rest.

  • Backbone β€” bf16 β†’ q4tp, 4.16 bits a weight with a predicted per-row scale ladder. The LoRA's rank-64 update is merged before quantizing.
  • Prompt encoder β€” Qwen3-VL-32B truncated to 50 layers, 51.5 GB β†’ 12.2 GB. It is the largest single component of the file and it runs once per generation.
  • Video VAE β€” decoder only. It is a ViT3D, not a conv stack: 36 transformer blocks over the latent grid and one linear that expands each cell into a 4Γ—16Γ—16 block of pixels. The 3-D causal CNN encoder is a third of the checkpoint and text-to-video never runs it.
  • Audio VAE β€” decoder only, f16. Quantizing a vocoder buys 45 MB and costs audible hiss. Its 254 kaiser-sinc resampling filters are read from the checkpoint rather than re-derived β€” the design formula is in the code as a fallback, but a filter you compute is a filter that can drift from the one the weights were trained against.
  • Integrity: 47.83 B parameters over 2 361 tensors; cortiq verify checks every one against the directory's hashes.

On parity

Established, not assumed, and separately for each of the four stacks. The reference is ComfyUI's own module, run on a toy checkpoint carrying the release's real tensor names and the release's real schedules β€” tools/ builds them, tools/mmh3_toy_gate.sh runs the diff. The packs are exact f32 on purpose: q4tp's noise floor sits an order of magnitude above the arithmetic difference these are looking for, so quantizing here would pass a broken port.

stack worst rms signal rms
DiT β€” video velocity 8.8e-5 2.1e-5 0.515
DiT β€” audio velocity 5.2e-5 2.5e-5 0.409
DiT β€” token refiner 8.3e-7 2.6e-7 1.003
Qwen3-VL encoder 1.1e-6 3.3e-7 0.812
video VAE decoder 4.2e-7 4.0e-8 0.470
audio VAE decoder 1.7e-9 3.5e-10 8.9e-4

A dozen conventions in this model pass at one token and fail differently at a hundred, which is why the toys are not one-vector unit tests: the packed layout's cursor, the video time axis's 1,4,4,4,4 span pattern, which 96 of 128 head dimensions rotate, the adaLN row order (timestep-major, modality-minor), the video VAE's 256-pixel tiling β€” global attention makes a tile a different computation from a whole frame, so the tiling is part of the output, not a memory strategy β€” and the audio stream's separate clock.

Two clocks

The video and audio latents ride different flow schedules (shift 12 and 3). The sampler walks the video grid, which at four steps is 1, 0.973, 0.923, 0.8, 0, and integrates the audio on its own remap of it. Stepping both on the video grid is what a stock sampler does; it is fine at twenty steps and wrong at four, because over the last interval Δσ_a and Δσ_v differ by a factor of three and no per-step slope correction survives a step that large. --stock-sampler reproduces the broken behaviour if you want to hear it.

Provenance

Weights derive from MiniMax's H3 release as repackaged by Comfy-Org, and from larryvrh's Turbo LoRA; both remain under their own licences. The Turbo LoRA is a preview β€” its own card notes plastic-looking skin and over-sharp grain at ckpt850, and nothing here changes that. The CMF container and the cortiq runtime are Apache-2.0 (see the repository's LICENSE and PATENTS.md).