Instructions to use infosave/MiniMax-H3-Turbo-cmf with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- cortiq
How to use infosave/MiniMax-H3-Turbo-cmf with cortiq:
# one Rust binary, no additional dependencies cargo install cortiq-cli # or a prebuilt binary from github.com/infosave2007/cmf/releases hf download infosave/MiniMax-H3-Turbo-cmf --include "*.cmf" --local-dir . ls *.cmf # some repos ship more than one quantization
cortiq animate FILE.cmf --prompt "a corgi in a chef hat flipping a pancake" --out clip.avi
- Notebooks
- Google Colab
- Kaggle
0.5.92: runtime adapters, the branch router, and what this port does not do yet
Browse files
README.md
CHANGED
|
@@ -270,10 +270,15 @@ What follows from them, if you are on such a machine:
|
|
| 270 |
still the full encoder's territory.
|
| 271 |
- **`--height 256` instead of 288** halves the video-VAE decode on every
|
| 272 |
machine (three 256-pixel tiles instead of six).
|
| 273 |
-
- **Voices are prompt space
|
| 274 |
-
|
| 275 |
-
|
| 276 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 277 |
|
| 278 |
## Making it smaller
|
| 279 |
|
|
@@ -592,6 +597,85 @@ differ by a factor of three and no per-step slope correction survives a step
|
|
| 592 |
that large. `--stock-sampler` reproduces the broken behaviour if you want to
|
| 593 |
hear it.
|
| 594 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 595 |
## Provenance
|
| 596 |
|
| 597 |
Weights derive from MiniMax's H3 release as repackaged by Comfy-Org, and from
|
|
|
|
| 270 |
still the full encoder's territory.
|
| 271 |
- **`--height 256` instead of 288** halves the video-VAE decode on every
|
| 272 |
machine (three 256-pixel tiles instead of six).
|
| 273 |
+
- **Voices are prompt space *here*, not in the release.** Speaker identity,
|
| 274 |
+
timbre, pace and emotion respond to stage directions in the prompt, and a
|
| 275 |
+
fixed seed keeps the same actor across takes β which is how to work with
|
| 276 |
+
this file today. But the earlier claim on this card that "H3 has no
|
| 277 |
+
reference-audio input" was **wrong**, and a user was right to push back
|
| 278 |
+
(discussion #3): the release is tagged `audio-to-audio-video` and
|
| 279 |
+
`video-to-audio-video`, and the DiT takes both as conditioning rows. What
|
| 280 |
+
is missing is on our side and it is not training β see
|
| 281 |
+
[What this port does not do yet](#what-this-port-does-not-do-yet).
|
| 282 |
|
| 283 |
## Making it smaller
|
| 284 |
|
|
|
|
| 597 |
that large. `--stock-sampler` reproduces the broken behaviour if you want to
|
| 598 |
hear it.
|
| 599 |
|
| 600 |
+
## What this port does not do yet
|
| 601 |
+
|
| 602 |
+
The release is tagged for six conditioning paths. This container packs one of
|
| 603 |
+
them β `fl2va`, text and/or keyframes β video + audio. What the others need is
|
| 604 |
+
listed here so nobody has to guess whether it is a missing feature or a
|
| 605 |
+
missing possibility. **None of them needs training.**
|
| 606 |
+
|
| 607 |
+
| the release's path | what it takes | what is missing here |
|
| 608 |
+
|---|---|---|
|
| 609 |
+
| `audio-to-audio-video` | a reference soundtrack | the audio VAE's **encoder half**. The container carries the decoder only β `pack_audio_vae` skips the encoder, `pre_block` and the mean/logs heads as unused. The DiT side already exists: the packed layout has a reference-audio segment kind and its own condition timestep |
|
| 610 |
+
| `video-to-audio-video` | a reference clip | the sampler plumbing. The video VAE **encoder is already packed** in the `fl2va` files β it is what encodes `--first-frame` β so this is a layout and CLI change, not a weight change |
|
| 611 |
+
| `ref2va` (subject / character references) | 1βN reference images | a **different DiT checkpoint** (`minimax_h3_ref2va_*`) with its own turbo LoRA. It would be a second container, not a flag on this one |
|
| 612 |
+
|
| 613 |
+
Two more, on the runtime side: the latent upscaler published for H3
|
| 614 |
+
(a 345 M-parameter 3-D conv net) is not ported, and there is no fused
|
| 615 |
+
device kernel for adapter branches on this model β see below.
|
| 616 |
+
|
| 617 |
+
## Adapters at runtime
|
| 618 |
+
|
| 619 |
+
Community LoRAs for H3 run against this container as they ship:
|
| 620 |
+
|
| 621 |
+
```bash
|
| 622 |
+
cortiq animate mmh3-turbo-clipproj4b-fl2va-v2-q4tp.cmf \
|
| 623 |
+
--prompt "r34l1sm a woman in a red raincoat on a neon street, close-up" \
|
| 624 |
+
--lora h3-realism-people.safetensors --lora-strength 0.8 \
|
| 625 |
+
--out take.avi
|
| 626 |
+
```
|
| 627 |
+
|
| 628 |
+
`--lora` reads a `.safetensors` in any of the three conventions in the wild
|
| 629 |
+
(`diffusion_model.β¦`, `base_model.model.dit.β¦`, or the bare module path), at
|
| 630 |
+
F32/F16/BF16, with either `lora_A`/`lora_B` or `lora_down`/`lora_up` naming.
|
| 631 |
+
It binds `attn.qkv_proj`, `attn.out_proj`, `mlp.fc1`, `mlp.fc2` on all fifty
|
| 632 |
+
blocks and on the two token-refiner blocks, and it prints what it bound:
|
| 633 |
+
|
| 634 |
+
```
|
| 635 |
+
lora: rank 32, 104/104 branches bound
|
| 636 |
+
```
|
| 637 |
+
|
| 638 |
+
Branches it cannot bind are named rather than dropped in silence. The one real
|
| 639 |
+
gap is `adaln_proj.linear`: this container carries the modulation as a rank-24
|
| 640 |
+
curve over the timestep (that collapse is 40% of the released model's weights
|
| 641 |
+
and most of why the file is 14 GB), and folding an adaLN update into it needs
|
| 642 |
+
the time embedding, which only the packer has. Adapters that touch adaLN β
|
| 643 |
+
the spatial-physics and streaming ones β therefore land partially, and say so.
|
| 644 |
+
Bake them in instead with `animate-pack --lora β¦ --time-embedder β¦`, which is
|
| 645 |
+
exactly how the Turbo LoRA got into these files.
|
| 646 |
+
|
| 647 |
+
**What it costs.** Measured on a Mac mini M4 (24 GB), 512Γ288, 22 frames, four
|
| 648 |
+
steps, with `fal/MiniMax-H3-Realism-People-LoRA` (rank 32, 104 branches, all
|
| 649 |
+
of which bind): **33.7 s a step with the adapter against 29.8 s without**, the
|
| 650 |
+
two runs taken back to back. The branch itself is half a per cent of the
|
| 651 |
+
projection's arithmetic and rides inside the base GEMM's Metal submission; what
|
| 652 |
+
it costs is the *attention* fusion standing down, because a branch on
|
| 653 |
+
`qkv_proj` or `out_proj` needs the panels that fusion keeps on the card. Two
|
| 654 |
+
honest notes: `--lora-strength 0` reproduces the base render byte for byte, and
|
| 655 |
+
the same base render drifted from 24.6 to 29.8 s a step over fifteen minutes on
|
| 656 |
+
that machine, so read the adapter as costing roughly a tenth of a step, not a
|
| 657 |
+
doubling.
|
| 658 |
+
|
| 659 |
+
**Which parts of an adapter matter.** `CMF_LORA_PROBE=1` prints every branch by
|
| 660 |
+
its measured contribution `βsΒ·ΞYβ/βYβ`, and `CMF_LORA_ROUTE=<r>` switches off
|
| 661 |
+
the ones below `r`. For the Realism adapter the loudest branch is 150Γ the
|
| 662 |
+
quietest, blocks 0β2 contribute nothing it would miss, and 41 of its 104
|
| 663 |
+
branches carry the look:
|
| 664 |
+
|
| 665 |
+
```
|
| 666 |
+
lora branches by contribution βsΞYβ/βYβ (41 of 104 live):
|
| 667 |
+
0.1633 on blocks.13.attn.qkv_proj
|
| 668 |
+
0.0946 on blocks.9.attn.qkv_proj
|
| 669 |
+
β¦
|
| 670 |
+
0.0011 off blocks.1.attn.qkv_proj
|
| 671 |
+
```
|
| 672 |
+
|
| 673 |
+
Rendered on those 41, the clip still looks like the adapter and not like the
|
| 674 |
+
base (13.34 dB PSNR against the base, where the full adapter is 13.76, and
|
| 675 |
+
16.75 dB against the full adapter). It did not make the step shorter on Metal
|
| 676 |
+
β the reason, and where routing does pay, is in
|
| 677 |
+
[docs/LORA.md](https://github.com/infosave2007/cmf/blob/master/docs/LORA.md).
|
| 678 |
+
|
| 679 |
## Provenance
|
| 680 |
|
| 681 |
Weights derive from MiniMax's H3 release as repackaged by Comfy-Org, and from
|