infosave commited on
Commit
72a3ac3
Β·
verified Β·
1 Parent(s): 41b5099

0.5.92: runtime adapters, the branch router, and what this port does not do yet

Browse files
Files changed (1) hide show
  1. README.md +88 -4
README.md CHANGED
@@ -270,10 +270,15 @@ What follows from them, if you are on such a machine:
270
  still the full encoder's territory.
271
  - **`--height 256` instead of 288** halves the video-VAE decode on every
272
  machine (three 256-pixel tiles instead of six).
273
- - **Voices are prompt space.** There is no reference-audio input β€” H3
274
- conditions on text and keyframes only β€” but speaker identity, timbre,
275
- pace and emotion respond to stage directions in the prompt, and a fixed
276
- seed keeps the same actor across takes.
 
 
 
 
 
277
 
278
  ## Making it smaller
279
 
@@ -592,6 +597,85 @@ differ by a factor of three and no per-step slope correction survives a step
592
  that large. `--stock-sampler` reproduces the broken behaviour if you want to
593
  hear it.
594
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
595
  ## Provenance
596
 
597
  Weights derive from MiniMax's H3 release as repackaged by Comfy-Org, and from
 
270
  still the full encoder's territory.
271
  - **`--height 256` instead of 288** halves the video-VAE decode on every
272
  machine (three 256-pixel tiles instead of six).
273
+ - **Voices are prompt space *here*, not in the release.** Speaker identity,
274
+ timbre, pace and emotion respond to stage directions in the prompt, and a
275
+ fixed seed keeps the same actor across takes β€” which is how to work with
276
+ this file today. But the earlier claim on this card that "H3 has no
277
+ reference-audio input" was **wrong**, and a user was right to push back
278
+ (discussion #3): the release is tagged `audio-to-audio-video` and
279
+ `video-to-audio-video`, and the DiT takes both as conditioning rows. What
280
+ is missing is on our side and it is not training β€” see
281
+ [What this port does not do yet](#what-this-port-does-not-do-yet).
282
 
283
  ## Making it smaller
284
 
 
597
  that large. `--stock-sampler` reproduces the broken behaviour if you want to
598
  hear it.
599
 
600
+ ## What this port does not do yet
601
+
602
+ The release is tagged for six conditioning paths. This container packs one of
603
+ them β€” `fl2va`, text and/or keyframes β†’ video + audio. What the others need is
604
+ listed here so nobody has to guess whether it is a missing feature or a
605
+ missing possibility. **None of them needs training.**
606
+
607
+ | the release's path | what it takes | what is missing here |
608
+ |---|---|---|
609
+ | `audio-to-audio-video` | a reference soundtrack | the audio VAE's **encoder half**. The container carries the decoder only β€” `pack_audio_vae` skips the encoder, `pre_block` and the mean/logs heads as unused. The DiT side already exists: the packed layout has a reference-audio segment kind and its own condition timestep |
610
+ | `video-to-audio-video` | a reference clip | the sampler plumbing. The video VAE **encoder is already packed** in the `fl2va` files β€” it is what encodes `--first-frame` β€” so this is a layout and CLI change, not a weight change |
611
+ | `ref2va` (subject / character references) | 1–N reference images | a **different DiT checkpoint** (`minimax_h3_ref2va_*`) with its own turbo LoRA. It would be a second container, not a flag on this one |
612
+
613
+ Two more, on the runtime side: the latent upscaler published for H3
614
+ (a 345 M-parameter 3-D conv net) is not ported, and there is no fused
615
+ device kernel for adapter branches on this model β€” see below.
616
+
617
+ ## Adapters at runtime
618
+
619
+ Community LoRAs for H3 run against this container as they ship:
620
+
621
+ ```bash
622
+ cortiq animate mmh3-turbo-clipproj4b-fl2va-v2-q4tp.cmf \
623
+ --prompt "r34l1sm a woman in a red raincoat on a neon street, close-up" \
624
+ --lora h3-realism-people.safetensors --lora-strength 0.8 \
625
+ --out take.avi
626
+ ```
627
+
628
+ `--lora` reads a `.safetensors` in any of the three conventions in the wild
629
+ (`diffusion_model.…`, `base_model.model.dit.…`, or the bare module path), at
630
+ F32/F16/BF16, with either `lora_A`/`lora_B` or `lora_down`/`lora_up` naming.
631
+ It binds `attn.qkv_proj`, `attn.out_proj`, `mlp.fc1`, `mlp.fc2` on all fifty
632
+ blocks and on the two token-refiner blocks, and it prints what it bound:
633
+
634
+ ```
635
+ lora: rank 32, 104/104 branches bound
636
+ ```
637
+
638
+ Branches it cannot bind are named rather than dropped in silence. The one real
639
+ gap is `adaln_proj.linear`: this container carries the modulation as a rank-24
640
+ curve over the timestep (that collapse is 40% of the released model's weights
641
+ and most of why the file is 14 GB), and folding an adaLN update into it needs
642
+ the time embedding, which only the packer has. Adapters that touch adaLN β€”
643
+ the spatial-physics and streaming ones β€” therefore land partially, and say so.
644
+ Bake them in instead with `animate-pack --lora … --time-embedder …`, which is
645
+ exactly how the Turbo LoRA got into these files.
646
+
647
+ **What it costs.** Measured on a Mac mini M4 (24 GB), 512Γ—288, 22 frames, four
648
+ steps, with `fal/MiniMax-H3-Realism-People-LoRA` (rank 32, 104 branches, all
649
+ of which bind): **33.7 s a step with the adapter against 29.8 s without**, the
650
+ two runs taken back to back. The branch itself is half a per cent of the
651
+ projection's arithmetic and rides inside the base GEMM's Metal submission; what
652
+ it costs is the *attention* fusion standing down, because a branch on
653
+ `qkv_proj` or `out_proj` needs the panels that fusion keeps on the card. Two
654
+ honest notes: `--lora-strength 0` reproduces the base render byte for byte, and
655
+ the same base render drifted from 24.6 to 29.8 s a step over fifteen minutes on
656
+ that machine, so read the adapter as costing roughly a tenth of a step, not a
657
+ doubling.
658
+
659
+ **Which parts of an adapter matter.** `CMF_LORA_PROBE=1` prints every branch by
660
+ its measured contribution `β€–sΒ·Ξ”Yβ€–/β€–Yβ€–`, and `CMF_LORA_ROUTE=<r>` switches off
661
+ the ones below `r`. For the Realism adapter the loudest branch is 150Γ— the
662
+ quietest, blocks 0–2 contribute nothing it would miss, and 41 of its 104
663
+ branches carry the look:
664
+
665
+ ```
666
+ lora branches by contribution β€–sΞ”Yβ€–/β€–Yβ€– (41 of 104 live):
667
+ 0.1633 on blocks.13.attn.qkv_proj
668
+ 0.0946 on blocks.9.attn.qkv_proj
669
+ …
670
+ 0.0011 off blocks.1.attn.qkv_proj
671
+ ```
672
+
673
+ Rendered on those 41, the clip still looks like the adapter and not like the
674
+ base (13.34 dB PSNR against the base, where the full adapter is 13.76, and
675
+ 16.75 dB against the full adapter). It did not make the step shorter on Metal
676
+ β€” the reason, and where routing does pay, is in
677
+ [docs/LORA.md](https://github.com/infosave2007/cmf/blob/master/docs/LORA.md).
678
+
679
  ## Provenance
680
 
681
  Weights derive from MiniMax's H3 release as repackaged by Comfy-Org, and from