Spaces:
Running on Zero
Running on Zero
MiniMax-H3 character-swap LoRA demo (ref2va, split conditioner)
Browse files- .gitattributes +5 -0
- README.md +142 -6
- app.py +785 -0
- examples/cafe_scene.mp4 +3 -0
- examples/character_anime.png +3 -0
- examples/character_hiker.png +3 -0
- examples/character_sheet_orin.png +3 -0
- examples/workshop_scene.mp4 +3 -0
- h3_fbc.py +357 -0
- h3_split_blocks.py +147 -0
- ncii_guard.py +90 -0
- packages.txt +1 -0
- requirements.txt +33 -0
.gitattributes
CHANGED
|
@@ -33,3 +33,8 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
|
|
| 33 |
*.zip filter=lfs diff=lfs merge=lfs -text
|
| 34 |
*.zst filter=lfs diff=lfs merge=lfs -text
|
| 35 |
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 33 |
*.zip filter=lfs diff=lfs merge=lfs -text
|
| 34 |
*.zst filter=lfs diff=lfs merge=lfs -text
|
| 35 |
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
| 36 |
+
examples/cafe_scene.mp4 filter=lfs diff=lfs merge=lfs -text
|
| 37 |
+
examples/character_anime.png filter=lfs diff=lfs merge=lfs -text
|
| 38 |
+
examples/character_hiker.png filter=lfs diff=lfs merge=lfs -text
|
| 39 |
+
examples/character_sheet_orin.png filter=lfs diff=lfs merge=lfs -text
|
| 40 |
+
examples/workshop_scene.mp4 filter=lfs diff=lfs merge=lfs -text
|
README.md
CHANGED
|
@@ -1,13 +1,149 @@
|
|
| 1 |
---
|
| 2 |
-
title:
|
| 3 |
-
emoji:
|
| 4 |
-
colorFrom:
|
| 5 |
-
colorTo:
|
| 6 |
sdk: gradio
|
| 7 |
sdk_version: 6.28.0
|
| 8 |
-
python_version: '3.12'
|
| 9 |
app_file: app.py
|
| 10 |
pinned: false
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 11 |
---
|
| 12 |
|
| 13 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
---
|
| 2 |
+
title: MiniMax-H3 Character Swap LoRA
|
| 3 |
+
emoji: 🎭
|
| 4 |
+
colorFrom: red
|
| 5 |
+
colorTo: yellow
|
| 6 |
sdk: gradio
|
| 7 |
sdk_version: 6.28.0
|
|
|
|
| 8 |
app_file: app.py
|
| 9 |
pinned: false
|
| 10 |
+
short_description: Swap one character in a clip for a reference character
|
| 11 |
+
python_version: "3.12"
|
| 12 |
+
startup_duration_timeout: 1h
|
| 13 |
+
models:
|
| 14 |
+
- akatz-ai/MiniMax-H3-Character-Swap-LoRA
|
| 15 |
+
- multimodalart/MiniMax-H3-Pruned
|
| 16 |
+
- MiniMaxAI/MiniMax-H3
|
| 17 |
+
datasets:
|
| 18 |
+
- akatz-ai/H3-Character-Swap-v1
|
| 19 |
---
|
| 20 |
|
| 21 |
+
# MiniMax-H3 Character Swap LoRA
|
| 22 |
+
|
| 23 |
+
A demo of [`akatz-ai/MiniMax-H3-Character-Swap-LoRA`](https://huggingface.co/akatz-ai/MiniMax-H3-Character-Swap-LoRA),
|
| 24 |
+
Akatz Labs' experimental character-replacement adapter for
|
| 25 |
+
[MiniMax-H3](https://huggingface.co/MiniMaxAI/MiniMax-H3)'s `ref2va` partition. Hand it a scene clip and a character
|
| 26 |
+
reference, name who to replace, and it puts the reference character into the shot — identity, outfit and art style
|
| 27 |
+
carried over — while the background, camera, lighting and everyone else stay where they were. Video and its
|
| 28 |
+
synchronized soundtrack come out of a single denoising pass.
|
| 29 |
+
|
| 30 |
+
## The request
|
| 31 |
+
|
| 32 |
+
Two references, **in the order the model reads them**, and a short targeting instruction:
|
| 33 |
+
|
| 34 |
+
| Slot | Label in the prompt | What it is |
|
| 35 |
+
|---|---|---|
|
| 36 |
+
| reference 1 | `<Picture 1>` | the replacement character — a portrait or a full character sheet |
|
| 37 |
+
| reference 2 | `<Video 1>` | the clip to edit, 5 frames to 15 s |
|
| 38 |
+
|
| 39 |
+
> Swap the man in the purple shirt in \<Video 1\> with the character in \<Picture 1\>.
|
| 40 |
+
|
| 41 |
+
That order is the one thing about this request worth being careful with, and it is **not** the order the prompt
|
| 42 |
+
names them in. MiniMax-H3 numbers a reference's label *per modality*, so the first image is `<Picture 1>` and the
|
| 43 |
+
first video is `<Video 1>` whichever way round they are packed — but the packed order still fixes the shared
|
| 44 |
+
audio/video rotary clock, so the same two references swapped round are a different request. The LoRA was trained
|
| 45 |
+
with the character sheet first: `control_path: [character_references, scene_videos]` in
|
| 46 |
+
`configs/trained-run-1000.json` of [`akatz-ai/H3-Character-Swap-v1`](https://huggingface.co/datasets/akatz-ai/H3-Character-Swap-v1).
|
| 47 |
+
`collect()` builds the list that way and nothing else in the app reorders it.
|
| 48 |
+
|
| 49 |
+
No trigger word was trained. Strength **1.0** is the card's recommendation; **0** is the base `ref2va` model, which
|
| 50 |
+
is the comparison the adapter was judged against, so the slider doubles as an A/B.
|
| 51 |
+
|
| 52 |
+
## What it is good at, and what it is not
|
| 53 |
+
|
| 54 |
+
The card is candid, and this demo does not oversell it. It is a **1,000-update experimental** adapter:
|
| 55 |
+
|
| 56 |
+
* background and scene preservation improved over the base model in the author's local comparisons — qualitative,
|
| 57 |
+
not a benchmark,
|
| 58 |
+
* motion timing, facial expressions and hard cuts remain unreliable; a hard cut can become a zoom or a gradual
|
| 59 |
+
reposition,
|
| 60 |
+
* long windows drift in framing and placement. **Short continuous shots of roughly 3–5 s** are the promising range,
|
| 61 |
+
* two-character inference was tested but multi-character replacement was never supervised,
|
| 62 |
+
* the soundtrack is generated, not carried over. Audio preservation in the author's later review was a remux, which
|
| 63 |
+
does not repair lip-sync drift.
|
| 64 |
+
|
| 65 |
+
Its 94 training edits targeted **single still frames**, with five-frame static clips standing in for `<Video 1>` —
|
| 66 |
+
which is exactly what the examples below are. Real moving footage goes in the same slot and is what the 40
|
| 67 |
+
preservation clips regularized, but it is the harder case.
|
| 68 |
+
|
| 69 |
+
## Defaults, and where they come from
|
| 70 |
+
|
| 71 |
+
`1344x768` at 24 fps, 73 frames (3.04 s), 28 steps, seed 904231 — the `sample` block of the LoRA's own
|
| 72 |
+
`configs/trained-run-1000.json`. The canvas dropdown follows the scene clip's aspect ratio on upload, because a
|
| 73 |
+
character swap is asked to keep the source framing and a portrait clip generated on a landscape canvas is a
|
| 74 |
+
recomposed shot before the model has done anything. "Match the scene clip's length" generates for as long as the
|
| 75 |
+
clip runs whenever that is a length MiniMax-H3 generates (2–14 s); the five-frame example clips are not, so they
|
| 76 |
+
fall through to the slider.
|
| 77 |
+
|
| 78 |
+
## How it is deployed
|
| 79 |
+
|
| 80 |
+
MiniMax-H3 is 195.9 GiB in bfloat16 and a ZeroGPU Space is evicted at 150 GB of storage, so `MiniMaxH3Blocks` is
|
| 81 |
+
cut at its `text_encoder` step. This Space is the **denoising half** of `ref2va` — the `transformer_ref` partition
|
| 82 |
+
and the two autoencoders — and the 62 GiB Qwen3-VL conditioner runs in
|
| 83 |
+
[`multimodalart/qwen3vl-conditioner`](https://huggingface.co/spaces/multimodalart/qwen3vl-conditioner), which this
|
| 84 |
+
Space calls over the gradio API for every request. `prompt_embeds` + `text_token_tags` is the whole wire format;
|
| 85 |
+
`reference_encoder` stays here, next to the autoencoders it runs. `h3_split_blocks.py` is the subclass that removes
|
| 86 |
+
the step.
|
| 87 |
+
|
| 88 |
+
The DiT is [`multimodalart/MiniMax-H3-Pruned`](https://huggingface.co/multimodalart/MiniMax-H3-Pruned)'s
|
| 89 |
+
`transformer_ref`: the released partition with its AdaLN input projections folded onto their reachable rank, 37.5
|
| 90 |
+
GiB instead of 61.7. That is the **same checkpoint family the LoRA was trained against** — ai-toolkit trained it on
|
| 91 |
+
Comfy-Org's `minimax_h3_ref2va_pruned_int8_convrot`, an int8 ConvRot quantization of these weights — and everything
|
| 92 |
+
the adapter touches is identical between the pruned and the released partition. The run's
|
| 93 |
+
`network_kwargs.ignore_if_contains = ["adaln_proj"]` kept it off the timestep path, which is the only place the two
|
| 94 |
+
differ.
|
| 95 |
+
|
| 96 |
+
**GPU time is priced per request, not per Space.** MiniMax-H3 attends over one packed sequence, and on this half the
|
| 97 |
+
references dominate its length: a 2048-short-edge character sheet is thousands of conditioning rows on top of the
|
| 98 |
+
generated ones. `get_duration` evaluates a fitted cost model over the sequence it is about to denoise — the
|
| 99 |
+
conditioner's exact token count plus the references measured from metadata — instead of reserving a flat ceiling for
|
| 100 |
+
everything, because the pool reserves whatever number it is given.
|
| 101 |
+
|
| 102 |
+
A first-block cache (`h3_fbc.py`, ported from `duckyshell/ComfyUI-MiniMaxH3-FirstBlockCache` with an audio
|
| 103 |
+
exemption) skips blocks 1–49 on steps whose block-0 residual has barely moved. `H3_FBC=0` restores the uncached
|
| 104 |
+
trajectory exactly.
|
| 105 |
+
|
| 106 |
+
There is no AoTI on this Space, unlike its siblings: a compiled block package binds the base module's weights by
|
| 107 |
+
fully qualified name and would run straight past the PEFT branch the adapter lives in.
|
| 108 |
+
|
| 109 |
+
## Space variables
|
| 110 |
+
|
| 111 |
+
| Variable | Default | Meaning |
|
| 112 |
+
|---|---|---|
|
| 113 |
+
| `H3_LORA_SCALE` | `1.0` | Default adapter strength. |
|
| 114 |
+
| `H3_CONDITIONER` | `multimodalart/qwen3vl-conditioner` | The public Space this one asks for embeddings; the client passes no token, so the call runs on the caller's own quota. |
|
| 115 |
+
| `H3_MODEL_REPO` | `multimodalart/MiniMax-H3-Pruned` | The diffusers-layout DiT. |
|
| 116 |
+
| `H3_ATTENTION` | `_native_cudnn` | cuDNN's fused kernel, 10–20% faster than the SDPA default. The two float32 VAEs are pinned to torch SDPA, which cuDNN has no kernel for. |
|
| 117 |
+
| `H3_FBC` / `H3_FBC_THRESHOLD` | `1` / `0.05` | First-block cache and its relative-L1 gate. |
|
| 118 |
+
| `H3_GPU_SIZE` | `xlarge` | ZeroGPU allocation size. `large` does not fit. |
|
| 119 |
+
| `H3_PLACEMENT` | `lazy` | Moves the partition onto the card on the first GPU call and leaves it there. |
|
| 120 |
+
|
| 121 |
+
## Safety
|
| 122 |
+
|
| 123 |
+
Every request here carries an image and a video reference — the edit-on-a-real-photo case — so
|
| 124 |
+
[`hfmlsoc/ncii-light-guard-v01`](https://huggingface.co/hfmlsoc/ncii-light-guard-v01) screens the prompt on all of
|
| 125 |
+
them, before the conditioner call and before any GPU is booked. It runs in its own subprocess
|
| 126 |
+
(`ncii_guard.py`): loaded in the main process, its torch activity poisons every later ZeroGPU fork.
|
| 127 |
+
|
| 128 |
+
## Example assets
|
| 129 |
+
|
| 130 |
+
All five files in `examples/` are from
|
| 131 |
+
[`akatz-ai/H3-Character-Swap-v1`](https://huggingface.co/datasets/akatz-ai/H3-Character-Swap-v1), the LoRA's own
|
| 132 |
+
training set — Apache-2.0 for Akatz Labs' synthetic contributions. They are the dataset's `CS001`, `CS051` and
|
| 133 |
+
`CS090` edits, with the instructions their own captions carry:
|
| 134 |
+
|
| 135 |
+
| File | Dataset path |
|
| 136 |
+
|---|---|
|
| 137 |
+
| `cafe_scene.mp4` | `checks/smoke-data/edits/train/scene_videos/CS001.mp4` (shared with `CS051`) |
|
| 138 |
+
| `character_hiker.png` | `.../character_references/CS001.png` |
|
| 139 |
+
| `character_anime.png` | `.../character_references/CS051.png` — a cross-style swap |
|
| 140 |
+
| `workshop_scene.mp4` | `.../scene_videos/CS090.mp4` |
|
| 141 |
+
| `character_sheet_orin.png` | `.../character_references/CS090.png` — a multi-view character sheet |
|
| 142 |
+
|
| 143 |
+
## License
|
| 144 |
+
|
| 145 |
+
The adapter is distributed under the
|
| 146 |
+
[MiniMax H3 Community License Agreement](https://huggingface.co/akatz-ai/MiniMax-H3-Character-Swap-LoRA/blob/main/LICENSE),
|
| 147 |
+
**not** Apache-2.0, and that agreement excludes the US, EU, UK and Republic of Korea from its standard territorial
|
| 148 |
+
grant. Read the upstream terms; nothing here extends them. The dataset's own Apache-2.0 covers the example assets
|
| 149 |
+
only and does not replace the model's terms.
|
app.py
ADDED
|
@@ -0,0 +1,785 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""MiniMax-H3 Character Swap LoRA — replace one character in a shot with a reference character.
|
| 2 |
+
|
| 3 |
+
`akatz-ai/MiniMax-H3-Character-Swap-LoRA` is a rank-16 ai-toolkit adapter for MiniMax-H3's **`ref2va`** partition.
|
| 4 |
+
A request is two references, in the order the LoRA was trained with — the character sheet as `<Picture 1>` and the
|
| 5 |
+
scene clip as `<Video 1>` — plus a short targeting instruction naming who to replace.
|
| 6 |
+
|
| 7 |
+
This Space is the denoising half of `ref2va`: the `transformer_ref` partition and the two autoencoders. The Qwen3-VL
|
| 8 |
+
conditioner runs in [`multimodalart/qwen3vl-conditioner`](https://huggingface.co/spaces/multimodalart/qwen3vl-conditioner),
|
| 9 |
+
which this Space calls over the gradio API for every request — MiniMax-H3 is 195.9 GiB in bfloat16 and a ZeroGPU
|
| 10 |
+
Space is evicted at 150 GB of storage, so the two halves cannot live together.
|
| 11 |
+
"""
|
| 12 |
+
|
| 13 |
+
from __future__ import annotations
|
| 14 |
+
|
| 15 |
+
import os
|
| 16 |
+
import tempfile
|
| 17 |
+
import time
|
| 18 |
+
import traceback
|
| 19 |
+
from functools import cache
|
| 20 |
+
|
| 21 |
+
# Before anything that could initialize CUDA: `import spaces` patches `torch.cuda` so the DiT load can happen at
|
| 22 |
+
# startup rather than on GPU time.
|
| 23 |
+
import spaces
|
| 24 |
+
import gradio as gr
|
| 25 |
+
import torch
|
| 26 |
+
|
| 27 |
+
VERSION = "h3-character-swap/1"
|
| 28 |
+
|
| 29 |
+
# The AdaLN-pruned `ref2va` partition in diffusers layout: 37.5 GiB against the released partition's 61.7 GiB, and
|
| 30 |
+
# the **same** checkpoint family this LoRA was trained against — ai-toolkit trained it on Comfy-Org's
|
| 31 |
+
# `minimax_h3_ref2va_pruned_int8_convrot`, i.e. an int8 ConvRot quantization of these very weights. Everything the
|
| 32 |
+
# LoRA touches (the block stack's attention and feed-forward projections, and the token refiner's) is identical
|
| 33 |
+
# between the pruned and the released partition; only the AdaLN input width differs, and the LoRA excludes
|
| 34 |
+
# `adaln_proj` by construction (`network_kwargs.ignore_if_contains`).
|
| 35 |
+
MODEL_REPO = os.environ.get("H3_MODEL_REPO", "multimodalart/MiniMax-H3-Pruned")
|
| 36 |
+
LORA_REPO = os.environ.get("H3_LORA_REPO", "akatz-ai/MiniMax-H3-Character-Swap-LoRA")
|
| 37 |
+
LORA_FILE = os.environ.get("H3_LORA_FILE", "h3_character_swap_pro4500_1000.safetensors")
|
| 38 |
+
ADAPTER = "character_swap"
|
| 39 |
+
# The model card's own recommendation: "Start with the final checkpoint at strength 1.0."
|
| 40 |
+
DEFAULT_LORA_SCALE = float(os.environ.get("H3_LORA_SCALE", "1.0"))
|
| 41 |
+
|
| 42 |
+
CONDITIONER_SPACE = os.environ.get("H3_CONDITIONER", "multimodalart/qwen3vl-conditioner")
|
| 43 |
+
# `lazy` moves the whole partition onto the card on the first GPU call and leaves it there. Startup placement is not
|
| 44 |
+
# an option: `spaces`' startup `torch.pack()` writes every startup-resident CUDA tensor to a second copy on disk, and
|
| 45 |
+
# the weights plus their pack do not fit the storage quota next to the VAEs.
|
| 46 |
+
PLACEMENT = os.environ.get("H3_PLACEMENT", "lazy").lower()
|
| 47 |
+
# cuDNN's fused attention, 10-20% faster than the SDPA default on this pool and needs nothing installed.
|
| 48 |
+
# flash-attention 3 is sm90-only and this card is sm120 (the `zero-a10g` flavour name is legacy).
|
| 49 |
+
ATTENTION = os.environ.get("H3_ATTENTION", "_native_cudnn").lower()
|
| 50 |
+
GPU_SIZE = os.environ.get("H3_GPU_SIZE", "xlarge")
|
| 51 |
+
MIN_GPU_DURATION = int(os.environ.get("H3_GPU_DURATION_MIN", "120"))
|
| 52 |
+
MAX_GPU_DURATION = int(os.environ.get("H3_GPU_DURATION_MAX", "1500"))
|
| 53 |
+
|
| 54 |
+
# Must stay identical to the conditioner's table: the *label* goes over the wire, so a canvas that half does not
|
| 55 |
+
# know is rejected there and surfaces as a failure here.
|
| 56 |
+
CANVASES = {
|
| 57 |
+
# 16:9
|
| 58 |
+
"960x544 · 16:9 fast": (544, 960),
|
| 59 |
+
"1024x576 · 16:9 fast": (576, 1024),
|
| 60 |
+
"1152x640 · 16:9": (640, 1152),
|
| 61 |
+
"1280x704 · 16:9": (704, 1280),
|
| 62 |
+
"1344x768 · 16:9 full": (768, 1344),
|
| 63 |
+
# 9:16
|
| 64 |
+
"544x960 · 9:16 fast": (960, 544),
|
| 65 |
+
"640x1152 · 9:16": (1152, 640),
|
| 66 |
+
"768x1344 · 9:16 full": (1344, 768),
|
| 67 |
+
# 1:1
|
| 68 |
+
"544x544 · 1:1 fast": (544, 544),
|
| 69 |
+
"768x768 · 1:1 full": (768, 768),
|
| 70 |
+
"1024x1024 · 1:1 max": (1024, 1024),
|
| 71 |
+
# 4:3 / 3:4
|
| 72 |
+
"768x576 · 4:3 fast": (576, 768),
|
| 73 |
+
"1024x768 · 4:3 full": (768, 1024),
|
| 74 |
+
"576x768 · 3:4 fast": (768, 576),
|
| 75 |
+
"768x1024 · 3:4 full": (1024, 768),
|
| 76 |
+
# 21:9
|
| 77 |
+
"1152x512 · 21:9 fast": (512, 1152),
|
| 78 |
+
"1536x672 · 21:9 full": (672, 1536),
|
| 79 |
+
}
|
| 80 |
+
# The LoRA's own training bucket: 1344x768 at 24 fps, which is also the resolution its author sampled at.
|
| 81 |
+
DEFAULT_CANVAS = "1344x768 · 16:9 full"
|
| 82 |
+
# The canvases `pick_canvas` may snap a scene clip to — one per aspect family, at the LoRA's own 768 short edge.
|
| 83 |
+
AUTO_CANVASES = (
|
| 84 |
+
"1344x768 · 16:9 full",
|
| 85 |
+
"768x1344 · 9:16 full",
|
| 86 |
+
"768x768 · 1:1 full",
|
| 87 |
+
"1024x768 · 4:3 full",
|
| 88 |
+
"768x1024 · 3:4 full",
|
| 89 |
+
"1536x672 · 21:9 full",
|
| 90 |
+
)
|
| 91 |
+
|
| 92 |
+
FPS, FRAMES_PER_CHUNK, LATENTS_PER_CHUNK = 24, 17, 5
|
| 93 |
+
# It is the *snapped* frame count the ceiling has to hold for: 15 s is 360 frames, which rounds up to 362, i.e.
|
| 94 |
+
# 15.083 s, and is refused. 14 is the last whole second that survives the snap.
|
| 95 |
+
MAX_UI_DURATION, MIN_DURATION = 14, 2
|
| 96 |
+
# 73 frames, 3.04 s — the `num_frames` the LoRA's own training run sampled at, and the length of its regularization
|
| 97 |
+
# clips. Short, continuous shots are what its card reports as the promising range.
|
| 98 |
+
DEFAULT_DURATION = 3
|
| 99 |
+
# The scene clip is a reference video. The floor is five frames rather than the two seconds a motion reference wants,
|
| 100 |
+
# because a five-frame static clip is exactly what every character-swap edit in the training set used for `<Video 1>`.
|
| 101 |
+
MIN_REFERENCE_VIDEO, MAX_REFERENCE_VIDEO = 5 / FPS, 15.0
|
| 102 |
+
DEFAULT_STEPS = 28
|
| 103 |
+
|
| 104 |
+
# Seconds of GPU one request needs, from the packed sequence it is about to denoise: linear in the rows for the
|
| 105 |
+
# matmuls, quadratic for the attention. Fitted on the `t2va` half and checked against live `ref2va` requests.
|
| 106 |
+
STEP_LINEAR, STEP_QUADRATIC, SAFETY = 1.1745e-4, 3.8396e-9, 1.3
|
| 107 |
+
# The lazy `PIPE.to("cuda")` a cold worker pays inside its first GPU call; every request carries it, because nothing
|
| 108 |
+
# here knows whether the worker it lands on is cold.
|
| 109 |
+
PLACEMENT_ALLOWANCE = int(os.environ.get("H3_PLACEMENT_ALLOWANCE", "90"))
|
| 110 |
+
AUDIO_LATENTS_PER_SECOND, AUDIO_CHANNELS = 40, 2
|
| 111 |
+
REFERENCE_IMAGE_SHORT_EDGE, CANVAS_MULTIPLE = 2048, 32
|
| 112 |
+
DECODE_BASE, DECODE_PER_DEFAULT_CANVAS, DEFAULT_CANVAS_PIXELS = 15, 25, 960 * 544 * 124
|
| 113 |
+
|
| 114 |
+
|
| 115 |
+
def snap_frames(seconds: float) -> int:
|
| 116 |
+
"""The frame count MiniMax-H3's video VAE can decode: the next `17 * n + 5` at 24 fps."""
|
| 117 |
+
frames = max(1, round(float(seconds) * FPS))
|
| 118 |
+
while frames % FRAMES_PER_CHUNK != LATENTS_PER_CHUNK:
|
| 119 |
+
frames += 1
|
| 120 |
+
return frames
|
| 121 |
+
|
| 122 |
+
|
| 123 |
+
def lower_duration_floor(seconds: float = MIN_DURATION) -> None:
|
| 124 |
+
"""Let the pipeline generate below its 5 s floor. 56 frames (2.33 s) is fine on the released checkpoint."""
|
| 125 |
+
from diffusers.modular_pipelines.minimax_h3.modular_pipeline import MiniMaxH3ModularPipeline
|
| 126 |
+
|
| 127 |
+
MiniMaxH3ModularPipeline.min_duration = property(lambda self: float(seconds))
|
| 128 |
+
|
| 129 |
+
|
| 130 |
+
def video_latent_frames(num_frames: int) -> int:
|
| 131 |
+
"""`17 * n + 5` frames become `5 * n + 2` video latents."""
|
| 132 |
+
return 5 * ((num_frames - LATENTS_PER_CHUNK) // FRAMES_PER_CHUNK) + 2
|
| 133 |
+
|
| 134 |
+
|
| 135 |
+
def target_rows(height: int, width: int, num_frames: int) -> int:
|
| 136 |
+
"""The generated rows of the packed sequence: video patched `(1, 2, 2)`, plus two audio rows per latent."""
|
| 137 |
+
video = video_latent_frames(num_frames) * (height // CANVAS_MULTIPLE) * (width // CANVAS_MULTIPLE)
|
| 138 |
+
return video + round(num_frames / FPS * AUDIO_LATENTS_PER_SECOND) * AUDIO_CHANNELS
|
| 139 |
+
|
| 140 |
+
|
| 141 |
+
def reference_rows(references: list[tuple[str, str]], num_frames: int) -> int:
|
| 142 |
+
"""The rows the two reference blocks add, from metadata alone — no decode."""
|
| 143 |
+
from PIL import Image
|
| 144 |
+
|
| 145 |
+
from diffusers.modular_pipelines.minimax_h3.modular_pipeline import resolve_canvas_size
|
| 146 |
+
|
| 147 |
+
rows = 0
|
| 148 |
+
for kind, path in references:
|
| 149 |
+
if kind == "image":
|
| 150 |
+
width, height = Image.open(path).size
|
| 151 |
+
scale = REFERENCE_IMAGE_SHORT_EDGE / min(width, height)
|
| 152 |
+
resolved = [
|
| 153 |
+
max(CANVAS_MULTIPLE, round(edge * scale / CANVAS_MULTIPLE) * CANVAS_MULTIPLE)
|
| 154 |
+
for edge in (height, width)
|
| 155 |
+
]
|
| 156 |
+
rows += (resolved[0] // CANVAS_MULTIPLE) * (resolved[1] // CANVAS_MULTIPLE)
|
| 157 |
+
continue
|
| 158 |
+
|
| 159 |
+
video_seconds, audio_seconds = probe(path)
|
| 160 |
+
if kind == "video" and video_seconds is not None:
|
| 161 |
+
import av
|
| 162 |
+
|
| 163 |
+
with av.open(path) as container:
|
| 164 |
+
stream = container.streams.video[0]
|
| 165 |
+
source_height, source_width = stream.height, stream.width
|
| 166 |
+
canvas_height, canvas_width = resolve_canvas_size(source_width, source_height, CANVAS_MULTIPLE)
|
| 167 |
+
frames = min(round(video_seconds * FPS), num_frames)
|
| 168 |
+
snapped = max(1, (frames - LATENTS_PER_CHUNK) // FRAMES_PER_CHUNK) * FRAMES_PER_CHUNK + LATENTS_PER_CHUNK
|
| 169 |
+
rows += (
|
| 170 |
+
video_latent_frames(snapped)
|
| 171 |
+
* (canvas_height // CANVAS_MULTIPLE)
|
| 172 |
+
* (canvas_width // CANVAS_MULTIPLE)
|
| 173 |
+
)
|
| 174 |
+
if audio_seconds is not None:
|
| 175 |
+
seconds = min(audio_seconds, num_frames / FPS)
|
| 176 |
+
rows += round(seconds * AUDIO_LATENTS_PER_SECOND) * AUDIO_CHANNELS
|
| 177 |
+
return rows
|
| 178 |
+
|
| 179 |
+
|
| 180 |
+
def get_duration(
|
| 181 |
+
prompt_embeds, text_token_tags, references, height, width, num_frames, steps, seed, lora_scale, *_, **__
|
| 182 |
+
):
|
| 183 |
+
"""Seconds of GPU to reserve for one request. Takes the arguments of the `@spaces.GPU` function it decorates, and
|
| 184 |
+
tolerates the `gr.Progress` `spaces` injects."""
|
| 185 |
+
sequence = (
|
| 186 |
+
int(text_token_tags.shape[0])
|
| 187 |
+
+ reference_rows(references, num_frames)
|
| 188 |
+
+ target_rows(height, width, num_frames)
|
| 189 |
+
)
|
| 190 |
+
denoise = int(steps) * (STEP_LINEAR * sequence + STEP_QUADRATIC * sequence**2) * SAFETY
|
| 191 |
+
# The two reference encoders ahead of the loop, and the two decoders plus the mux after it. Both scale with what
|
| 192 |
+
# they are handed rather than with the step count.
|
| 193 |
+
encode = 5 + reference_rows(references, num_frames) * 1e-3
|
| 194 |
+
decode = DECODE_BASE + DECODE_PER_DEFAULT_CANVAS * (height * width * num_frames) / DEFAULT_CANVAS_PIXELS
|
| 195 |
+
total = PLACEMENT_ALLOWANCE + encode + denoise + decode + 10
|
| 196 |
+
duration = max(MIN_GPU_DURATION, min(MAX_GPU_DURATION, int(total)))
|
| 197 |
+
print(f"[{VERSION}] S={sequence} -> reserving {duration}s ({denoise:.0f}s of denoise at {steps} steps)", flush=True)
|
| 198 |
+
return duration
|
| 199 |
+
|
| 200 |
+
|
| 201 |
+
PIPE = None
|
| 202 |
+
MANAGER = None
|
| 203 |
+
LOAD_ERROR: str | None = None
|
| 204 |
+
LORA_STATUS: str = ""
|
| 205 |
+
|
| 206 |
+
|
| 207 |
+
# ── LoRA loading (runtime PEFT adapter) ──────────────────────────────────────
|
| 208 |
+
#
|
| 209 |
+
# The adapter targets the checkpoint's own module names (`diffusion_model.blocks.N.attn.qkv_proj`, `mlp.fc1`/`fc2`,
|
| 210 |
+
# `attn.out_proj`, and the same four under `token_refiner.blocks.N` — 208 modules, 416 tensors). diffusers serves
|
| 211 |
+
# the converted port, so every name — and, for two of them, the row layout of `lora_B` — has to be pushed through
|
| 212 |
+
# the same transforms `scripts/convert_minimax_h3_to_diffusers.py` applied to the base weights:
|
| 213 |
+
#
|
| 214 |
+
# * `blocks.` -> `transformer_blocks.`, `token_refiner.blocks.` -> `token_refiner.refiner_blocks.`,
|
| 215 |
+
# `attn.out_proj` -> `attn.to_out.0`, `mlp.fc2` -> `ff.net.2` (pure renames),
|
| 216 |
+
# * `mlp.fc1` -> `ff.net.0.proj` with its two fused halves swapped, because diffusers' SwiGLU reads
|
| 217 |
+
# `[value; gate]` where the checkpoint stores `[gate; value]`,
|
| 218 |
+
# * `attn.qkv_proj` -> `to_q` / `to_k` / `to_v`: split `lora_B`'s 21504 rows into **contiguous** thirds of 7168
|
| 219 |
+
# (`num_attention_heads * attention_head_dim`).
|
| 220 |
+
#
|
| 221 |
+
# That last split is where MiniMax-H3 LoRAs diverge. diffusers' own `_convert_non_diffusers_minimax_h3_lora_to_
|
| 222 |
+
# diffusers` documents two fused-QKV layouts: DiffSynth-Studio exports (identifiable by the peft `.default.` infix)
|
| 223 |
+
# are *per-head interleaved* and must be de-interleaved before the split, whereas ai-toolkit exports under the
|
| 224 |
+
# `diffusion_model.` prefix are already `[q_all; k_all; v_all]` and must **not** be. This file's keys are
|
| 225 |
+
# `diffusion_model.…lora_A.weight` with `ai-toolkit 0.13.21` in its metadata, so it is the second kind: contiguous
|
| 226 |
+
# thirds, no reorder. De-interleaving it anyway would scatter each head's q/k/v across all three projections and
|
| 227 |
+
# turn the adapter into structured noise on all 52 attention blocks.
|
| 228 |
+
#
|
| 229 |
+
# Row transforms only ever touch `lora_B`, so the three projections share one `lora_A`: the rows of `B @ A` are just
|
| 230 |
+
# the rows of `B`, which makes the split exact rather than an approximation.
|
| 231 |
+
#
|
| 232 |
+
# There is no AdaLN branch to reconcile: the run's `network_kwargs.ignore_if_contains = ["adaln_proj"]` kept the
|
| 233 |
+
# adapter off the timestep path entirely, which is also why the pruned partition takes it verbatim.
|
| 234 |
+
|
| 235 |
+
|
| 236 |
+
def _lora_target_name(source_name: str) -> str:
|
| 237 |
+
"""Map an original-checkpoint module path to the diffusers one (no `.lora_A/B.*` suffix)."""
|
| 238 |
+
if source_name.startswith("token_refiner.blocks."):
|
| 239 |
+
return source_name.replace("token_refiner.blocks.", "token_refiner.refiner_blocks.", 1)
|
| 240 |
+
if source_name.startswith("blocks."):
|
| 241 |
+
return source_name.replace("blocks.", "transformer_blocks.", 1)
|
| 242 |
+
return source_name
|
| 243 |
+
|
| 244 |
+
|
| 245 |
+
def _lora_modules(name: str, a_weight, b_weight, inner_dim: int):
|
| 246 |
+
"""Yield `(diffusers_module_path, lora_A, lora_B)` for one original module path."""
|
| 247 |
+
target = _lora_target_name(name)
|
| 248 |
+
|
| 249 |
+
if target.endswith(".attn.qkv_proj"):
|
| 250 |
+
prefix = target.removesuffix("qkv_proj")
|
| 251 |
+
for kind, part in zip(("q", "k", "v"), b_weight.split(inner_dim, dim=0)):
|
| 252 |
+
yield f"{prefix}to_{kind}", a_weight, part.contiguous()
|
| 253 |
+
elif target.endswith(".mlp.fc1"):
|
| 254 |
+
# SwiGLU gate/value swap: the checkpoint stores [gate, value]; diffusers stores [value, gate].
|
| 255 |
+
gate, value = b_weight.chunk(2, dim=0)
|
| 256 |
+
yield target.replace(".mlp.fc1", ".ff.net.0.proj"), a_weight, torch.cat([value, gate]).contiguous()
|
| 257 |
+
elif target.endswith(".mlp.fc2"):
|
| 258 |
+
yield target.replace(".mlp.fc2", ".ff.net.2"), a_weight, b_weight
|
| 259 |
+
elif target.endswith(".attn.out_proj"):
|
| 260 |
+
yield target.replace(".attn.out_proj", ".attn.to_out.0"), a_weight, b_weight
|
| 261 |
+
else:
|
| 262 |
+
yield target, a_weight, b_weight
|
| 263 |
+
|
| 264 |
+
|
| 265 |
+
def build_lora_state_dict(transformer) -> tuple[dict, int, int]:
|
| 266 |
+
"""Download the LoRA and remap it into a PEFT-format state dict for the diffusers transformer.
|
| 267 |
+
|
| 268 |
+
Every target is validated against the transformer's own parameter shapes, and an unresolved one is fatal: a
|
| 269 |
+
silently dropped target means the name mapping is wrong and the Space would serve a half-applied adapter that
|
| 270 |
+
still *looks* like it worked.
|
| 271 |
+
"""
|
| 272 |
+
from huggingface_hub import hf_hub_download
|
| 273 |
+
from safetensors.torch import load_file
|
| 274 |
+
|
| 275 |
+
raw = load_file(hf_hub_download(LORA_REPO, LORA_FILE))
|
| 276 |
+
|
| 277 |
+
prefix, suffix_a, suffix_b = "diffusion_model.", ".lora_A.weight", ".lora_B.weight"
|
| 278 |
+
unexpected = [k for k in raw if not (k.startswith(prefix) and k.endswith((suffix_a, suffix_b)))]
|
| 279 |
+
if unexpected:
|
| 280 |
+
raise ValueError(f"{LORA_FILE} holds {len(unexpected)} unexpected tensors, e.g. {unexpected[:5]}")
|
| 281 |
+
bases = sorted({k[len(prefix) : -len(suffix_a)] for k in raw if k.endswith(suffix_a)})
|
| 282 |
+
if not bases:
|
| 283 |
+
raise ValueError(f"No `{prefix}*{suffix_a}` / `{suffix_b}` pairs found in {LORA_FILE}")
|
| 284 |
+
|
| 285 |
+
ranks = set()
|
| 286 |
+
for name in bases:
|
| 287 |
+
if f"{prefix}{name}{suffix_b}" not in raw:
|
| 288 |
+
raise ValueError(f"LoRA is missing the lora_B twin of {prefix}{name}{suffix_a}")
|
| 289 |
+
ranks.add(raw[f"{prefix}{name}{suffix_a}"].shape[0])
|
| 290 |
+
if len(ranks) != 1:
|
| 291 |
+
raise ValueError(f"LoRA mixes ranks {sorted(ranks)}; this loader assumes a single rank")
|
| 292 |
+
rank = ranks.pop()
|
| 293 |
+
|
| 294 |
+
inner_dim = transformer.config.num_attention_heads * transformer.config.attention_head_dim
|
| 295 |
+
base_shapes = {key: tuple(value.shape) for key, value in transformer.state_dict().items()}
|
| 296 |
+
|
| 297 |
+
state_dict: dict[str, torch.Tensor] = {}
|
| 298 |
+
missed: list[str] = []
|
| 299 |
+
for name in bases:
|
| 300 |
+
a_weight = raw[f"{prefix}{name}{suffix_a}"]
|
| 301 |
+
b_weight = raw[f"{prefix}{name}{suffix_b}"]
|
| 302 |
+
for module, a_part, b_part in _lora_modules(name, a_weight, b_weight, inner_dim):
|
| 303 |
+
base = base_shapes.get(f"{module}.weight")
|
| 304 |
+
if base is None:
|
| 305 |
+
missed.append(module)
|
| 306 |
+
continue
|
| 307 |
+
# `W` is [out, in]; the adapter must be `lora_B` [out, r] @ `lora_A` [r, in].
|
| 308 |
+
if (b_part.shape[0], a_part.shape[1]) != base:
|
| 309 |
+
raise ValueError(
|
| 310 |
+
f"LoRA delta for `{module}` would be {(b_part.shape[0], a_part.shape[1])}, "
|
| 311 |
+
f"base weight is {base}"
|
| 312 |
+
)
|
| 313 |
+
state_dict[f"{module}.lora_A.weight"] = a_part
|
| 314 |
+
state_dict[f"{module}.lora_B.weight"] = b_part
|
| 315 |
+
if missed:
|
| 316 |
+
raise ValueError(
|
| 317 |
+
f"{len(missed)} LoRA targets matched no transformer weight, e.g. {missed[:5]}. "
|
| 318 |
+
"The LoRA and the diffusers transformer disagree on module naming."
|
| 319 |
+
)
|
| 320 |
+
|
| 321 |
+
return state_dict, rank, len(bases)
|
| 322 |
+
|
| 323 |
+
|
| 324 |
+
def load_and_apply_lora(transformer) -> str:
|
| 325 |
+
"""Attach the LoRA as a PEFT adapter so its strength stays a per-request knob."""
|
| 326 |
+
state_dict, rank, targets = build_lora_state_dict(transformer)
|
| 327 |
+
# `prefix=None` because the keys are already the transformer's own module paths, and no `network_alphas` because
|
| 328 |
+
# the file carries no `.alpha` tensors — PEFT then sets alpha == rank, which is also what the run trained with
|
| 329 |
+
# (`linear: 16`, `linear_alpha: 16`), so the adapter's intrinsic scale is 1.0 and the slider *is* the strength.
|
| 330 |
+
transformer.load_lora_adapter(state_dict, prefix=None, adapter_name=ADAPTER)
|
| 331 |
+
transformer.set_adapters([ADAPTER], [DEFAULT_LORA_SCALE])
|
| 332 |
+
return (
|
| 333 |
+
f"LoRA attached · {targets} checkpoint targets -> {len(state_dict) // 2} diffusers modules, "
|
| 334 |
+
f"rank {rank}, alpha == rank · default strength {DEFAULT_LORA_SCALE:g} · {LORA_FILE}"
|
| 335 |
+
)
|
| 336 |
+
|
| 337 |
+
|
| 338 |
+
def check_prompt(prompt: str) -> None:
|
| 339 |
+
"""The NCII guard. Every request here carries an image and a video reference — the edit-on-a-real-photo case — so
|
| 340 |
+
it runs on all of them, before the conditioner call and the denoise booking, and a refused prompt costs no GPU
|
| 341 |
+
time on either half. The classifier lives in `ncii_guard`'s spawned subprocess; in the main process it kills
|
| 342 |
+
every GPU worker."""
|
| 343 |
+
import ncii_guard
|
| 344 |
+
|
| 345 |
+
flag = ncii_guard.classify(prompt)
|
| 346 |
+
if flag["label"] == "ncii":
|
| 347 |
+
print(f"[guard] prompt refused (ncii {flag['score']:.2f}): {prompt!r}", flush=True)
|
| 348 |
+
raise gr.Error("This prompt was flagged by a content filter and wasn't run.")
|
| 349 |
+
|
| 350 |
+
|
| 351 |
+
def load_models() -> str | None:
|
| 352 |
+
"""Load the denoising half at startup, but *not* onto the card.
|
| 353 |
+
|
| 354 |
+
`MiniMaxH3Ref2VAGeneratorBlocks` declares `transformer_ref`, `vae`, `audio_vae`, the two schedulers and
|
| 355 |
+
`video_processor`, so `load_components` fetches exactly those subfolders — `text_encoder/` and the `transformer/`
|
| 356 |
+
partition are never touched. Both autoencoders carry `_keep_in_fp32_modules` over every module and stay float32:
|
| 357 |
+
a bfloat16 audio VAE decodes the soundtrack roughly 20 dB too quiet.
|
| 358 |
+
"""
|
| 359 |
+
global PIPE, MANAGER, LOAD_ERROR, LORA_STATUS
|
| 360 |
+
|
| 361 |
+
if PIPE is not None or LOAD_ERROR is not None:
|
| 362 |
+
return LOAD_ERROR
|
| 363 |
+
|
| 364 |
+
started = time.time()
|
| 365 |
+
try:
|
| 366 |
+
from diffusers import ComponentsManager
|
| 367 |
+
|
| 368 |
+
from h3_split_blocks import MiniMaxH3Ref2VAGeneratorBlocks
|
| 369 |
+
|
| 370 |
+
lower_duration_floor()
|
| 371 |
+
manager = ComponentsManager()
|
| 372 |
+
blocks = MiniMaxH3Ref2VAGeneratorBlocks()
|
| 373 |
+
print(f"[{VERSION}] loading {[c.name for c in blocks.expected_components]} from {MODEL_REPO} ...", flush=True)
|
| 374 |
+
pipe = blocks.init_pipeline(MODEL_REPO, components_manager=manager, collection="h3")
|
| 375 |
+
# The pruned DiT is served as remote code (`transformer_ref/modeling_minimax_h3_pruned.py`, reached through
|
| 376 |
+
# the `AutoModel` type hint in `modular_model_index.json`). `load_components` forwards `trust_remote_code`
|
| 377 |
+
# only to components that live in the pipeline's own repo, so the two VAEs and the schedulers — which
|
| 378 |
+
# `modular_model_index.json` still points at `MiniMaxAI/MiniMax-H3` — never see it.
|
| 379 |
+
pipe.load_components(dtype=torch.bfloat16, trust_remote_code=True)
|
| 380 |
+
|
| 381 |
+
# Both VAEs first, and explicitly. `set_attention_backend` also sets the registry's *global* backend, which
|
| 382 |
+
# every processor that was not stamped falls through to, and the float32 audio VAE has no cuDNN kernel:
|
| 383 |
+
# `RuntimeError: No available kernel. Aborting execution.` in its causal encoder attention.
|
| 384 |
+
pipe.vae.set_attention_backend("native")
|
| 385 |
+
pipe.audio_vae.set_attention_backend("native")
|
| 386 |
+
pipe.transformer_ref.set_attention_backend(ATTENTION)
|
| 387 |
+
|
| 388 |
+
# Before any GPU placement, so the adapter's parameters travel with the transformer. No AoTI on this Space
|
| 389 |
+
# for the same reason: a compiled block package binds the *base* module's weights by fully qualified name
|
| 390 |
+
# and would run straight past the PEFT branch.
|
| 391 |
+
LORA_STATUS = load_and_apply_lora(pipe.transformer_ref)
|
| 392 |
+
print(f"[{VERSION}] {LORA_STATUS}", flush=True)
|
| 393 |
+
|
| 394 |
+
# Installed per request rather than here — a cached residual only means something within the schedule it was
|
| 395 |
+
# measured on, so the state cannot outlive one generation.
|
| 396 |
+
import h3_fbc
|
| 397 |
+
|
| 398 |
+
print(f"[h3-fbc] {h3_fbc.status()}", flush=True)
|
| 399 |
+
|
| 400 |
+
if PLACEMENT == "offload":
|
| 401 |
+
manager.enable_auto_cpu_offload(device="cuda")
|
| 402 |
+
_arm_decode_hooks(pipe)
|
| 403 |
+
|
| 404 |
+
PIPE, MANAGER = pipe, manager
|
| 405 |
+
print(f"[{VERSION}] ready in {time.time() - started:.0f}s", flush=True)
|
| 406 |
+
except Exception as error:
|
| 407 |
+
traceback.print_exc()
|
| 408 |
+
LOAD_ERROR = (
|
| 409 |
+
f"**Loading `{MODEL_REPO}` + `{LORA_REPO}` failed** after {time.time() - started:.0f}s: "
|
| 410 |
+
f"`{type(error).__name__}: {error}`"
|
| 411 |
+
)
|
| 412 |
+
return LOAD_ERROR
|
| 413 |
+
|
| 414 |
+
|
| 415 |
+
def _arm_decode_hooks(pipe):
|
| 416 |
+
"""Make the offload hooks fire for the two VAEs.
|
| 417 |
+
|
| 418 |
+
`enable_auto_cpu_offload` wraps `forward`, and the reference-encoder and decode blocks call `vae.encode/decode(...)`
|
| 419 |
+
directly, so the hook never runs and the VAE is still on the host when the latents arrive on the card.
|
| 420 |
+
"""
|
| 421 |
+
for name in ("vae", "audio_vae"):
|
| 422 |
+
module = getattr(pipe, name)
|
| 423 |
+
for method in ("encode", "decode"):
|
| 424 |
+
inner = getattr(module, method)
|
| 425 |
+
|
| 426 |
+
def armed(*args, _module=module, _inner=inner, **kwargs):
|
| 427 |
+
hook = getattr(_module, "_hf_hook", None)
|
| 428 |
+
if hook is not None:
|
| 429 |
+
hook.pre_forward(_module)
|
| 430 |
+
return _inner(*args, **kwargs)
|
| 431 |
+
|
| 432 |
+
setattr(module, method, armed)
|
| 433 |
+
|
| 434 |
+
|
| 435 |
+
@cache
|
| 436 |
+
def conditioner():
|
| 437 |
+
"""The other half, over the gradio API. `gradio_client` attaches the caller's own ZeroGPU token per call, so the
|
| 438 |
+
conditioner's booking is billed to whoever asked for the video."""
|
| 439 |
+
from gradio_client import Client
|
| 440 |
+
|
| 441 |
+
return Client(CONDITIONER_SPACE)
|
| 442 |
+
|
| 443 |
+
|
| 444 |
+
def probe(path: str) -> tuple[float | None, float | None]:
|
| 445 |
+
"""`(video seconds, audio seconds)` of a media file, either being `None` when the stream is absent."""
|
| 446 |
+
import av
|
| 447 |
+
|
| 448 |
+
def seconds(stream, container):
|
| 449 |
+
if stream.duration is not None and stream.time_base is not None:
|
| 450 |
+
return float(stream.duration * stream.time_base)
|
| 451 |
+
return None if container.duration is None else container.duration / av.time_base
|
| 452 |
+
|
| 453 |
+
with av.open(path) as container:
|
| 454 |
+
video = seconds(container.streams.video[0], container) if container.streams.video else None
|
| 455 |
+
audio = seconds(container.streams.audio[0], container) if container.streams.audio else None
|
| 456 |
+
return video, audio
|
| 457 |
+
|
| 458 |
+
|
| 459 |
+
def collect(character_path: str, scene_path: str) -> list[tuple[str, str]]:
|
| 460 |
+
"""The `(kind, path)` references of a request, **in the order the model reads them**.
|
| 461 |
+
|
| 462 |
+
The character sheet first, the scene clip second. That order is not cosmetic and it is not the order the prompt
|
| 463 |
+
names them in: it numbers the labels of MiniMax-H3's prompt presentation and it advances the shared audio/video
|
| 464 |
+
rotary clock, so the same two references swapped round are a different request. It is the order the LoRA was
|
| 465 |
+
trained with — `control_path: [character_references, scene_videos]` in `configs/trained-run-1000.json` of
|
| 466 |
+
`akatz-ai/H3-Character-Swap-v1` — which is why the captions read `<Video 1>` before `<Picture 1>` while the
|
| 467 |
+
references go the other way: the labels are numbered per modality, not by position.
|
| 468 |
+
"""
|
| 469 |
+
return [("image", character_path), ("video", scene_path)]
|
| 470 |
+
|
| 471 |
+
|
| 472 |
+
def build_references(references: list[tuple[str, str]]):
|
| 473 |
+
"""The `(kind, path)` references of a request as decoded reference dataclasses, in packed order. `from_file`
|
| 474 |
+
brings the rates along: a video its own frame rate and soundtrack."""
|
| 475 |
+
from diffusers.modular_pipelines.minimax_h3 import (
|
| 476 |
+
MiniMaxH3ImageReference,
|
| 477 |
+
MiniMaxH3VideoReference,
|
| 478 |
+
)
|
| 479 |
+
|
| 480 |
+
classes = {"image": MiniMaxH3ImageReference, "video": MiniMaxH3VideoReference}
|
| 481 |
+
return [classes[kind].from_file(path) for kind, path in references]
|
| 482 |
+
|
| 483 |
+
|
| 484 |
+
def pick_canvas(scene_path: str | None) -> str:
|
| 485 |
+
"""The generated canvas whose aspect ratio is closest to the scene clip's, at the LoRA's own 768 short edge.
|
| 486 |
+
|
| 487 |
+
A character swap is asked to keep the source framing, so a portrait clip generated on a landscape canvas is a
|
| 488 |
+
recomposed shot before the model has done anything. Wired to the clip's `change`, so the dropdown shows what a
|
| 489 |
+
request will actually use and stays overridable.
|
| 490 |
+
"""
|
| 491 |
+
if not scene_path:
|
| 492 |
+
return gr.update()
|
| 493 |
+
try:
|
| 494 |
+
import av
|
| 495 |
+
|
| 496 |
+
with av.open(scene_path) as container:
|
| 497 |
+
stream = container.streams.video[0]
|
| 498 |
+
ratio = stream.width / stream.height
|
| 499 |
+
except Exception:
|
| 500 |
+
return gr.update()
|
| 501 |
+
best = min(AUTO_CANVASES, key=lambda label: abs(CANVASES[label][1] / CANVASES[label][0] - ratio))
|
| 502 |
+
return gr.update(value=best)
|
| 503 |
+
|
| 504 |
+
|
| 505 |
+
def check(prompt: str, character_path: str | None, scene_path: str | None) -> None:
|
| 506 |
+
"""The request's own rules, before anything is uploaded or a card is allocated."""
|
| 507 |
+
if not prompt or not prompt.strip():
|
| 508 |
+
raise gr.Error("Name who to replace — MiniMax-H3 always takes a prompt.")
|
| 509 |
+
if not scene_path:
|
| 510 |
+
raise gr.Error("Add the scene clip to edit; it is the request's `<Video 1>`.")
|
| 511 |
+
if not character_path:
|
| 512 |
+
raise gr.Error("Add the character reference to swap in; it is the request's `<Picture 1>`.")
|
| 513 |
+
video_seconds, _ = probe(scene_path)
|
| 514 |
+
if video_seconds is None:
|
| 515 |
+
raise gr.Error("That scene clip has no video stream.")
|
| 516 |
+
if not MIN_REFERENCE_VIDEO <= video_seconds <= MAX_REFERENCE_VIDEO:
|
| 517 |
+
raise gr.Error(
|
| 518 |
+
f"The scene clip is {video_seconds:.2f} s. Use one between {MIN_REFERENCE_VIDEO:.2f} and "
|
| 519 |
+
f"{MAX_REFERENCE_VIDEO:g} seconds."
|
| 520 |
+
)
|
| 521 |
+
|
| 522 |
+
|
| 523 |
+
def encode_remote(prompt, references, canvas, num_frames):
|
| 524 |
+
"""`/encode_ref2va` on the conditioner Space: a safetensors file holding `prompt_embeds` + `text_token_tags`,
|
| 525 |
+
with the resolved `height` / `width` / `num_frames` in its metadata, plus the plan.
|
| 526 |
+
|
| 527 |
+
`canvas` is the label. `media` and `kinds` are parallel and ordered, and the references go over because `ref2va`'s
|
| 528 |
+
presentation puts a vision block in front of the prompt for every image and every merged video frame pair.
|
| 529 |
+
"""
|
| 530 |
+
from gradio_client import handle_file
|
| 531 |
+
from safetensors import safe_open
|
| 532 |
+
|
| 533 |
+
path, plan = conditioner().predict(
|
| 534 |
+
prompt=prompt,
|
| 535 |
+
media=[handle_file(path) for _, path in references],
|
| 536 |
+
kinds=",".join(kind for kind, _ in references),
|
| 537 |
+
canvas=canvas,
|
| 538 |
+
num_frames=num_frames,
|
| 539 |
+
rewrite_prompt=False,
|
| 540 |
+
api_name="/encode_ref2va",
|
| 541 |
+
)
|
| 542 |
+
with safe_open(path, framework="pt") as handle:
|
| 543 |
+
return handle.get_tensor("prompt_embeds"), handle.get_tensor("text_token_tags"), handle.metadata(), plan
|
| 544 |
+
|
| 545 |
+
|
| 546 |
+
@spaces.GPU(duration=get_duration, size=GPU_SIZE)
|
| 547 |
+
def _generate(prompt_embeds, text_token_tags, references, height, width, num_frames, steps, seed, lora_scale):
|
| 548 |
+
"""The only thing on GPU time: the two reference encoders, the packed-sequence denoise loop and the decoders.
|
| 549 |
+
|
| 550 |
+
References cross as paths and are decoded here; only the three generated outputs come back. A `@spaces.GPU`
|
| 551 |
+
argument crosses a process boundary by pickling, a 5 s 1344x768 reference video is 370 MB of expanded frames, and
|
| 552 |
+
the full `PipelineState` still holds the packed latents and the rotary grid on the card.
|
| 553 |
+
"""
|
| 554 |
+
import h3_fbc
|
| 555 |
+
|
| 556 |
+
if PLACEMENT == "lazy":
|
| 557 |
+
PIPE.to("cuda")
|
| 558 |
+
|
| 559 |
+
# Per request, inside the worker: the adapter's scale is a python attribute on its PEFT layers, so setting it on
|
| 560 |
+
# this fork's copy leaves every concurrent request with its own strength.
|
| 561 |
+
PIPE.transformer_ref.set_adapters([ADAPTER], [float(lora_scale)])
|
| 562 |
+
|
| 563 |
+
with h3_fbc.enabled(PIPE.transformer_ref, steps=int(steps)):
|
| 564 |
+
state = PIPE(
|
| 565 |
+
prompt_embeds=prompt_embeds.to("cuda"),
|
| 566 |
+
text_token_tags=text_token_tags,
|
| 567 |
+
references=build_references(references),
|
| 568 |
+
height=height,
|
| 569 |
+
width=width,
|
| 570 |
+
num_frames=num_frames,
|
| 571 |
+
num_inference_steps=int(steps),
|
| 572 |
+
generator=torch.Generator("cpu").manual_seed(int(seed)),
|
| 573 |
+
)
|
| 574 |
+
return state.get("videos")[0], state.get("audio")[0].cpu(), state.get("sampling_rate")
|
| 575 |
+
|
| 576 |
+
|
| 577 |
+
def generate(
|
| 578 |
+
# The first three are the columns `gr.Examples` varies, and they lead the signature for that reason: an example
|
| 579 |
+
# row is applied to `inputs` positionally. Every other parameter carries the same default as its UI component,
|
| 580 |
+
# so clicking an example and pressing Generate behave identically.
|
| 581 |
+
prompt: str,
|
| 582 |
+
scene_video: str | None = None,
|
| 583 |
+
character_image: str | None = None,
|
| 584 |
+
canvas: str = DEFAULT_CANVAS,
|
| 585 |
+
match_length: bool = True,
|
| 586 |
+
duration: float = DEFAULT_DURATION,
|
| 587 |
+
steps: int = DEFAULT_STEPS,
|
| 588 |
+
seed: int = 904231,
|
| 589 |
+
lora_scale: float = DEFAULT_LORA_SCALE,
|
| 590 |
+
progress=gr.Progress(track_tqdm=True),
|
| 591 |
+
):
|
| 592 |
+
"""Replace one character in a clip with the character in a reference image, keeping the rest of the shot.
|
| 593 |
+
|
| 594 |
+
Args:
|
| 595 |
+
prompt: who to replace and with what, naming the references as `<Video 1>` and `<Picture 1>`, e.g.
|
| 596 |
+
"Swap the man in the purple shirt in <Video 1> with the character in <Picture 1>."
|
| 597 |
+
scene_video: path to the clip to edit — the request's `<Video 1>`, 5 frames to 15 seconds.
|
| 598 |
+
character_image: path to the replacement character's portrait or character sheet — `<Picture 1>`.
|
| 599 |
+
canvas: the generated canvas, as one of this Space's labels.
|
| 600 |
+
match_length: generate for as long as the scene clip runs, when that length is one MiniMax-H3 generates.
|
| 601 |
+
duration: seconds to generate when `match_length` is off, or the clip is too short to set it.
|
| 602 |
+
steps: denoising steps. The LoRA's own training run sampled at 28.
|
| 603 |
+
seed: RNG seed.
|
| 604 |
+
lora_scale: character-swap adapter strength. 1.0 is the card's recommendation; 0 is the base model.
|
| 605 |
+
|
| 606 |
+
Returns:
|
| 607 |
+
Path to the generated MP4, video and its jointly generated soundtrack.
|
| 608 |
+
"""
|
| 609 |
+
if LOAD_ERROR:
|
| 610 |
+
raise gr.Error(LOAD_ERROR)
|
| 611 |
+
if PIPE is None:
|
| 612 |
+
raise gr.Error("The denoiser is still loading.")
|
| 613 |
+
|
| 614 |
+
from diffusers.utils import encode_video
|
| 615 |
+
|
| 616 |
+
check(prompt, character_image, scene_video)
|
| 617 |
+
check_prompt(prompt)
|
| 618 |
+
references = collect(character_image, scene_video)
|
| 619 |
+
|
| 620 |
+
scene_seconds, _ = probe(scene_video)
|
| 621 |
+
requested_seconds = float(duration)
|
| 622 |
+
if match_length and scene_seconds is not None and MIN_DURATION <= scene_seconds <= MAX_UI_DURATION:
|
| 623 |
+
requested_seconds = scene_seconds
|
| 624 |
+
num_frames = snap_frames(requested_seconds)
|
| 625 |
+
|
| 626 |
+
progress(0.0, desc="Reading the prompt and the two references ...")
|
| 627 |
+
conditioned = time.time()
|
| 628 |
+
try:
|
| 629 |
+
prompt_embeds, text_token_tags, metadata, plan = encode_remote(prompt, references, canvas, num_frames)
|
| 630 |
+
except gr.Error:
|
| 631 |
+
raise
|
| 632 |
+
except Exception as error:
|
| 633 |
+
# gradio only puts the exception *type* on the wire, so the useful half of a conditioner-side failure is in
|
| 634 |
+
# that Space's logs.
|
| 635 |
+
traceback.print_exc()
|
| 636 |
+
raise gr.Error(
|
| 637 |
+
f"The conditioner ({CONDITIONER_SPACE}) failed with `{type(error).__name__}: {error}`. "
|
| 638 |
+
"Its logs carry the full traceback."
|
| 639 |
+
) from error
|
| 640 |
+
condition_seconds = time.time() - conditioned
|
| 641 |
+
height, width, num_frames = (int(metadata[key]) for key in ("height", "width", "num_frames"))
|
| 642 |
+
|
| 643 |
+
progress(0.1, desc=f"Swapping over {num_frames / FPS:.1f} s at {width}x{height} ...")
|
| 644 |
+
started = time.time()
|
| 645 |
+
frames, audio, sampling_rate = _generate(
|
| 646 |
+
prompt_embeds, text_token_tags, references, height, width, num_frames, steps, seed, lora_scale
|
| 647 |
+
)
|
| 648 |
+
generate_seconds = time.time() - started
|
| 649 |
+
|
| 650 |
+
directory = os.path.join(tempfile.gettempdir(), "h3-outputs")
|
| 651 |
+
os.makedirs(directory, exist_ok=True)
|
| 652 |
+
path = os.path.join(directory, f"h3-character-swap-{int(time.time() * 1000)}.mp4")
|
| 653 |
+
encode_video(frames, fps=FPS, output_path=path, audio=audio, audio_sample_rate=sampling_rate)
|
| 654 |
+
|
| 655 |
+
print(
|
| 656 |
+
f"[{VERSION}] `{width}x{height}`, {num_frames} frames ({num_frames / FPS:.3f} s), {int(steps)} steps, "
|
| 657 |
+
f"strength {float(lora_scale):g} · conditioner {condition_seconds:.0f}s "
|
| 658 |
+
f"({plan['num_text_tokens']} tokens) · denoise + decode {generate_seconds:.0f}s "
|
| 659 |
+
f"({generate_seconds / max(1, int(steps)):.1f} s/step) · seed {int(seed)}",
|
| 660 |
+
flush=True,
|
| 661 |
+
)
|
| 662 |
+
return path
|
| 663 |
+
|
| 664 |
+
|
| 665 |
+
import ncii_guard
|
| 666 |
+
|
| 667 |
+
ncii_guard.start()
|
| 668 |
+
load_models()
|
| 669 |
+
|
| 670 |
+
INTRO = """# MiniMax-H3 Character Swap LoRA
|
| 671 |
+
|
| 672 |
+
<div align="center">
|
| 673 |
+
<a href="https://huggingface.co/akatz-ai/MiniMax-H3-Character-Swap-LoRA" target="_blank" rel="noopener"><strong>[ LoRA ]</strong></a>
|
| 674 |
+
<a href="https://huggingface.co/datasets/akatz-ai/H3-Character-Swap-v1" target="_blank" rel="noopener"><strong>[ dataset ]</strong></a>
|
| 675 |
+
<a href="https://huggingface.co/MiniMaxAI/MiniMax-H3" target="_blank" rel="noopener"><strong>[ base model ]</strong></a>
|
| 676 |
+
</div>
|
| 677 |
+
|
| 678 |
+
Give it a **scene clip** and a **character reference**, then name who to replace. Akatz Labs' experimental
|
| 679 |
+
character-swap adapter for MiniMax-H3 `ref2va` puts the reference character into the shot — identity, outfit and
|
| 680 |
+
art style carried over — while the background, the camera and everyone else stay where they were. Video and its
|
| 681 |
+
soundtrack come out of one denoising pass.
|
| 682 |
+
|
| 683 |
+
Refer to the two references the way the LoRA was trained to read them: the clip is **`<Video 1>`** and the
|
| 684 |
+
character is **`<Picture 1>`**.
|
| 685 |
+
"""
|
| 686 |
+
|
| 687 |
+
NOTES = """It is an experimental 1,000-update adapter and its own card is candid about the limits: motion timing,
|
| 688 |
+
facial expressions and hard cuts are unreliable, long windows drift in framing, and close-up expressions may not
|
| 689 |
+
match the source performance. Short continuous shots of about 3–5 seconds are the promising range.
|
| 690 |
+
|
| 691 |
+
Its training targets were **single still frames** with five-frame static scene clips, which is what the examples
|
| 692 |
+
below are — a still scene, handed over as the short clip `<Video 1>` wants. Real moving footage works the same way
|
| 693 |
+
and is what the 40 preservation clips in the dataset regularized, but it is the harder case.
|
| 694 |
+
|
| 695 |
+
Strength 1.0 is the author's recommendation; 0 is the base `ref2va` model, which is the comparison the adapter was
|
| 696 |
+
judged against.
|
| 697 |
+
"""
|
| 698 |
+
|
| 699 |
+
CSS = """
|
| 700 |
+
.main.fillable { max-width: 1250px !important; }
|
| 701 |
+
.dark .gradio-container { color: var(--body-text-color); }
|
| 702 |
+
"""
|
| 703 |
+
|
| 704 |
+
with gr.Blocks(title="MiniMax-H3 Character Swap LoRA", theme=gr.themes.Citrus(), css=CSS) as demo:
|
| 705 |
+
gr.Markdown(INTRO)
|
| 706 |
+
|
| 707 |
+
with gr.Row():
|
| 708 |
+
with gr.Column():
|
| 709 |
+
prompt = gr.Textbox(
|
| 710 |
+
label="Prompt — name the person to replace",
|
| 711 |
+
lines=3,
|
| 712 |
+
value="Swap the man in the purple shirt in <Video 1> with the character in <Picture 1>.",
|
| 713 |
+
placeholder="Swap the woman on the left in <Video 1> with the character in <Picture 1>.",
|
| 714 |
+
)
|
| 715 |
+
with gr.Row():
|
| 716 |
+
scene = gr.Video(label="Scene clip · <Video 1>", height=260)
|
| 717 |
+
character = gr.Image(label="Character reference · <Picture 1>", type="filepath", height=260)
|
| 718 |
+
run = gr.Button("Swap the character", variant="primary", size="lg")
|
| 719 |
+
with gr.Accordion("Advanced options", open=False):
|
| 720 |
+
lora_scale = gr.Slider(
|
| 721 |
+
label="Character-swap LoRA strength",
|
| 722 |
+
minimum=0.0,
|
| 723 |
+
maximum=1.5,
|
| 724 |
+
step=0.05,
|
| 725 |
+
value=DEFAULT_LORA_SCALE,
|
| 726 |
+
)
|
| 727 |
+
canvas = gr.Dropdown(
|
| 728 |
+
label="Canvas (follows the scene clip's aspect ratio)",
|
| 729 |
+
choices=list(CANVASES),
|
| 730 |
+
value=DEFAULT_CANVAS,
|
| 731 |
+
)
|
| 732 |
+
match_length = gr.Checkbox(label="Match the scene clip's length", value=True)
|
| 733 |
+
duration = gr.Slider(
|
| 734 |
+
label="Duration (s) — used when the clip is shorter than the model generates",
|
| 735 |
+
minimum=MIN_DURATION,
|
| 736 |
+
maximum=MAX_UI_DURATION,
|
| 737 |
+
step=1,
|
| 738 |
+
value=DEFAULT_DURATION,
|
| 739 |
+
)
|
| 740 |
+
steps = gr.Slider(label="Steps", minimum=10, maximum=40, step=1, value=DEFAULT_STEPS)
|
| 741 |
+
seed = gr.Number(label="Seed", value=904231, precision=0)
|
| 742 |
+
|
| 743 |
+
with gr.Column():
|
| 744 |
+
result = gr.Video(label="Swapped clip + soundtrack")
|
| 745 |
+
gr.Markdown(NOTES)
|
| 746 |
+
|
| 747 |
+
scene.change(pick_canvas, scene, canvas, show_progress="hidden", api_name=False)
|
| 748 |
+
|
| 749 |
+
request = [prompt, scene, character, canvas, match_length, duration, steps, seed, lora_scale]
|
| 750 |
+
|
| 751 |
+
gr.Examples(
|
| 752 |
+
# The dataset's own edits, with the instructions its captions carry: a photoreal swap, a cross-style swap
|
| 753 |
+
# onto an illustrated character, and a multi-view character sheet.
|
| 754 |
+
examples=[
|
| 755 |
+
[
|
| 756 |
+
"Swap the man in the purple shirt in <Video 1> with the character in <Picture 1>.",
|
| 757 |
+
"examples/cafe_scene.mp4",
|
| 758 |
+
"examples/character_hiker.png",
|
| 759 |
+
],
|
| 760 |
+
[
|
| 761 |
+
"Swap the man in the purple shirt in <Video 1> with the character in <Picture 1>. Keep the "
|
| 762 |
+
"replacement character's identity, outfit and art style from <Picture 1>. Preserve the source "
|
| 763 |
+
"video's camera, background, lighting, objects and all other people.",
|
| 764 |
+
"examples/cafe_scene.mp4",
|
| 765 |
+
"examples/character_anime.png",
|
| 766 |
+
],
|
| 767 |
+
[
|
| 768 |
+
"Swap the person in <Video 1> with the character in <Picture 1>. Do not show the reference sheet "
|
| 769 |
+
"or its background.",
|
| 770 |
+
"examples/workshop_scene.mp4",
|
| 771 |
+
"examples/character_sheet_orin.png",
|
| 772 |
+
],
|
| 773 |
+
],
|
| 774 |
+
inputs=[prompt, scene, character],
|
| 775 |
+
outputs=result,
|
| 776 |
+
fn=generate,
|
| 777 |
+
cache_examples=True,
|
| 778 |
+
cache_mode="lazy",
|
| 779 |
+
)
|
| 780 |
+
|
| 781 |
+
run.click(generate, request, result, api_name="generate")
|
| 782 |
+
|
| 783 |
+
|
| 784 |
+
if __name__ == "__main__":
|
| 785 |
+
demo.launch(show_error=True, max_threads=1000)
|
examples/cafe_scene.mp4
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:3d1dbc1b503df47708aa2b1a49d2987078e1edb2954d0f28ee30650823ce0488
|
| 3 |
+
size 923206
|
examples/character_anime.png
ADDED
|
Git LFS Details
|
examples/character_hiker.png
ADDED
|
Git LFS Details
|
examples/character_sheet_orin.png
ADDED
|
Git LFS Details
|
examples/workshop_scene.mp4
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:fa1b341d2bb03a786b95457a1705d6fda46e7ebaa0baf205bf8a36c70662ee4a
|
| 3 |
+
size 766362
|
h3_fbc.py
ADDED
|
@@ -0,0 +1,357 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""First-block cache for MiniMax-H3: skip the trunk on steps whose block-0 residual barely moved.
|
| 2 |
+
|
| 3 |
+
Block 0 and the final AdaLN head run at the true timestep on **every** step; only blocks 1..49 are skipped, and only
|
| 4 |
+
while the residual they would have been handed looks like the one from the last step that actually ran. The decision
|
| 5 |
+
signal is the relative L1 between this step's block-0 residual (`block0_out - block0_in`) and the residual of the last
|
| 6 |
+
*computed* step, over the whole packed sequence. On a skip the trunk's contribution is replayed as a cached residual,
|
| 7 |
+
`(final_trunk_out - block0_out)` of that computed step.
|
| 8 |
+
|
| 9 |
+
Ported from `duckyshell/ComfyUI-MiniMaxH3-FirstBlockCache` (`nodes.py` @ 725973c) — same signal, same protected window
|
| 10 |
+
(10%-95% of the schedule, converted through the video shift of 12.0) and the same cap of two consecutive skips, which
|
| 11 |
+
is their "H3 Safe" preset at threshold 0.08.
|
| 12 |
+
|
| 13 |
+
**The audio exemption is not theirs.** duckyshell has no audio term anywhere: the decision is ~98% video by row count
|
| 14 |
+
and every audio row rides the stale trunk residual, which is what costs the soundtrack its energy. Audio runs its own
|
| 15 |
+
schedule (shift 3 against video's 12, diverging by up to 11x in rate across 20 steps), so on a skip the audio rows of
|
| 16 |
+
the trunk output are instead a 2-point **linear** extrapolation of the last two actually-computed audio features, in
|
| 17 |
+
the audio sigma coordinate — `xmarre/ComfyUI-Spectrum-MiniMax-H3`'s `audio_blend_weight=0.0` path, applied where
|
| 18 |
+
Spectrum applies it (the post-trunk hidden feature, ahead of the head that still runs at the true timestep) and fixed
|
| 19 |
+
to the right coordinate. Following ComfyUI PR #15390, no carried audio tensor is ever mutated: every write is a fresh
|
| 20 |
+
tensor out of `index_copy`.
|
| 21 |
+
|
| 22 |
+
It earns its place. Running this Space's own request at threshold 0.08 with `H3_FBC_AUDIO_EXEMPT=0` — the duckyshell
|
| 23 |
+
mechanism verbatim, same 13 skipped forwards, same video to within 0.2 dB — costs the soundtrack **16.7% of its RMS**
|
| 24 |
+
(0.0483 against an uncached 0.0580), while the exemption holds it at 1.04x. That is the same direction the offline
|
| 25 |
+
study measured on a near-silent clip, three times the size on one with real audio energy.
|
| 26 |
+
|
| 27 |
+
**A threshold does not travel across step counts.** The signature shrinks as the schedule is subdivided, so the same
|
| 28 |
+
number gates far more loosely at more steps. Measured on this Space — 960x544, 124 frames, one image reference,
|
| 29 |
+
seed 42, AoTI blocks, at its **default 28 steps** (27 forwards) — against the same request with `H3_FBC=0`:
|
| 30 |
+
|
| 31 |
+
| threshold | skipped | denoise loop | end to end | audio RMS vs uncached |
|
| 32 |
+
|---|---|---|---|---|
|
| 33 |
+
| 0.03 | 0 / 27 | 1.02x | 1.06x | 1.000 (bitwise-identical audio) |
|
| 34 |
+
| 0.05 | 9 / 27 | 1.46x | 1.36x | 0.956 |
|
| 35 |
+
| 0.08 | 13 / 27 | 2.19x | 1.82x | 1.040 |
|
| 36 |
+
|
| 37 |
+
The offline study calibrated 0.08 over a 20-step schedule, where it skipped 7 of 19 forwards. At 28 steps that same
|
| 38 |
+
0.08 skips 13 of 27 — nearly half — and the sampled video visibly re-rolls its background detail. **0.05 is the
|
| 39 |
+
default** because it reproduces the skip fraction that study validated (33% here against 37% there); 0.08 is the
|
| 40 |
+
aggressive setting. Below the signature floor — 0.03 skipped nothing at all here — a threshold buys nothing and still
|
| 41 |
+
pays for the signature, so lower is not safer, it is just slower.
|
| 42 |
+
|
| 43 |
+
A cached request is not the uncached one: the trajectory moves, so the video is a different sample of the same prompt
|
| 44 |
+
(same shot, same subject, same quality — different signage and background detail). `H3_FBC=0` restores today's output
|
| 45 |
+
exactly, and is worth reaching for when a request has to reproduce a specific earlier result.
|
| 46 |
+
|
| 47 |
+
This composes with `h3_aoti`: that module patches each of the 50 blocks' own `forward` and they stay a real
|
| 48 |
+
`ModuleList`, so skipping the trunk simply does not call blocks 1..49 that step. `LazyAOTIModel` rebinds its constants
|
| 49 |
+
whenever the weights dict it is handed changes identity, and the first forward of a request never skips, so every
|
| 50 |
+
block has bound its own weights before any step is cached.
|
| 51 |
+
"""
|
| 52 |
+
|
| 53 |
+
from __future__ import annotations
|
| 54 |
+
|
| 55 |
+
import contextlib
|
| 56 |
+
import os
|
| 57 |
+
import types
|
| 58 |
+
|
| 59 |
+
ENABLED = os.environ.get("H3_FBC", "1") == "1"
|
| 60 |
+
# Relative-L1 gate on the block-0 residual, calibrated at this Space's default 28 steps — see the table above.
|
| 61 |
+
THRESHOLD = float(os.environ.get("H3_FBC_THRESHOLD", "0.05"))
|
| 62 |
+
# duckyshell's cap. Without it the gate compares against an ever-older computed step and drifts away unbounded.
|
| 63 |
+
MAX_CONSECUTIVE_HITS = int(os.environ.get("H3_FBC_MAX_CONSECUTIVE", "2"))
|
| 64 |
+
# The protected head and tail of the schedule, as fractions, converted to sigma through the video shift below.
|
| 65 |
+
START_PERCENT = float(os.environ.get("H3_FBC_START_PERCENT", "0.10"))
|
| 66 |
+
END_PERCENT = float(os.environ.get("H3_FBC_END_PERCENT", "0.95"))
|
| 67 |
+
# `MiniMaxH3SetTimestepsStep` builds the video schedule at shift 12.0 and the audio one at 3.0.
|
| 68 |
+
VIDEO_SHIFT = float(os.environ.get("H3_FBC_VIDEO_SHIFT", "12.0"))
|
| 69 |
+
AUDIO_EXEMPT = os.environ.get("H3_FBC_AUDIO_EXEMPT", "1") == "1"
|
| 70 |
+
|
| 71 |
+
# Every keyword `MiniMaxH3LoopDenoiser` passes. It filters the packed-sequence layout through
|
| 72 |
+
# `inspect.signature(transformer.forward).parameters`, so a replacement forward that drops a name silently stops
|
| 73 |
+
# receiving it; `install` refuses rather than let that happen quietly.
|
| 74 |
+
FORWARD_PARAMETERS = (
|
| 75 |
+
"hidden_states",
|
| 76 |
+
"audio_hidden_states",
|
| 77 |
+
"encoder_hidden_states",
|
| 78 |
+
"timestep",
|
| 79 |
+
"timestep_indices",
|
| 80 |
+
"token_tags",
|
| 81 |
+
"position_ids",
|
| 82 |
+
"video_indices",
|
| 83 |
+
"audio_indices",
|
| 84 |
+
"text_indices",
|
| 85 |
+
"attention_kwargs",
|
| 86 |
+
"return_dict",
|
| 87 |
+
)
|
| 88 |
+
|
| 89 |
+
|
| 90 |
+
def status() -> str:
|
| 91 |
+
return (
|
| 92 |
+
f"first-block cache **on** · threshold `{THRESHOLD}` · audio exemption "
|
| 93 |
+
f"{'on' if AUDIO_EXEMPT else 'off'}"
|
| 94 |
+
if ENABLED
|
| 95 |
+
else "first-block cache **off** (`H3_FBC=1` to skip the trunk on steady steps)"
|
| 96 |
+
)
|
| 97 |
+
|
| 98 |
+
|
| 99 |
+
def _shifted_sigma(u: float, shift: float) -> float:
|
| 100 |
+
return shift * u / (1.0 + (shift - 1.0) * u)
|
| 101 |
+
|
| 102 |
+
|
| 103 |
+
def _rel_l1(current, previous) -> float:
|
| 104 |
+
numerator = (current.float() - previous.float()).abs().mean()
|
| 105 |
+
denominator = previous.float().abs().mean().clamp(min=1e-8)
|
| 106 |
+
return float((numerator / denominator).item())
|
| 107 |
+
|
| 108 |
+
|
| 109 |
+
class _State:
|
| 110 |
+
def __init__(self, threshold: float, steps: int, audio_exempt: bool):
|
| 111 |
+
self.threshold = threshold
|
| 112 |
+
self.steps = steps
|
| 113 |
+
self.audio_exempt = audio_exempt
|
| 114 |
+
# duckyshell reads the window as sigma bounds: a flow model's sigma at `u = 1 - percent`, shifted.
|
| 115 |
+
self.start_sigma = _shifted_sigma(1.0 - START_PERCENT, VIDEO_SHIFT)
|
| 116 |
+
self.end_sigma = _shifted_sigma(1.0 - END_PERCENT, VIDEO_SHIFT)
|
| 117 |
+
self.original = None
|
| 118 |
+
self.failed = False
|
| 119 |
+
self.consecutive_hits = 0
|
| 120 |
+
self.prev_first_residual = None
|
| 121 |
+
self.tail_residual = None
|
| 122 |
+
self.audio_history = [] # [(sigma_audio, audio rows of the trunk output)], newest last, at most two
|
| 123 |
+
self.computed = 0
|
| 124 |
+
self.skipped = 0
|
| 125 |
+
|
| 126 |
+
|
| 127 |
+
def _cached_forward(
|
| 128 |
+
self,
|
| 129 |
+
state,
|
| 130 |
+
hidden_states,
|
| 131 |
+
audio_hidden_states,
|
| 132 |
+
encoder_hidden_states,
|
| 133 |
+
timestep,
|
| 134 |
+
timestep_indices,
|
| 135 |
+
token_tags,
|
| 136 |
+
position_ids,
|
| 137 |
+
video_indices,
|
| 138 |
+
audio_indices,
|
| 139 |
+
text_indices,
|
| 140 |
+
return_dict,
|
| 141 |
+
):
|
| 142 |
+
"""`MiniMaxH3Transformer3DModel.forward` with the block loop split at block 0.
|
| 143 |
+
|
| 144 |
+
Everything outside the loop is that method verbatim, at the `diffusers` commit `requirements.txt` pins; keep the
|
| 145 |
+
two in step when the pin moves.
|
| 146 |
+
"""
|
| 147 |
+
import torch
|
| 148 |
+
|
| 149 |
+
from diffusers.models.transformers.transformer_minimax_h3 import (
|
| 150 |
+
MINIMAX_H3_MODALITY_NUM,
|
| 151 |
+
MiniMaxH3TransformerOutput,
|
| 152 |
+
)
|
| 153 |
+
|
| 154 |
+
sequence_length = position_ids.shape[0]
|
| 155 |
+
rotary_emb = self.rope(position_ids)
|
| 156 |
+
|
| 157 |
+
video_embeds = self.proj_in(hidden_states.to(self.proj_in.weight.dtype))
|
| 158 |
+
audio_embeds = self.audio_proj_in(audio_hidden_states.to(self.audio_proj_in.weight.dtype))
|
| 159 |
+
text_embeds = self.context_embedder(encoder_hidden_states.to(self.context_embedder.weight.dtype))
|
| 160 |
+
text_embeds = self.token_refiner(text_embeds)
|
| 161 |
+
|
| 162 |
+
packed = text_embeds.new_zeros((text_embeds.shape[0], sequence_length, text_embeds.shape[-1]))
|
| 163 |
+
packed = packed.index_copy(1, text_indices, text_embeds)
|
| 164 |
+
packed = packed.index_copy(1, video_indices, video_embeds.to(text_embeds.dtype))
|
| 165 |
+
packed = packed.index_copy(1, audio_indices, audio_embeds.to(text_embeds.dtype))
|
| 166 |
+
|
| 167 |
+
temb = self.time_proj(timestep)
|
| 168 |
+
temb = self.time_embedder(temb.to(self.time_embedder.linear_1.weight.dtype))
|
| 169 |
+
adaln_indices = timestep_indices * MINIMAX_H3_MODALITY_NUM + token_tags.clamp(min=0)
|
| 170 |
+
|
| 171 |
+
attention_mask = None
|
| 172 |
+
is_pad = token_tags < 0
|
| 173 |
+
if bool(is_pad.any()):
|
| 174 |
+
attention_mask = is_pad[None, :] == is_pad[:, None]
|
| 175 |
+
|
| 176 |
+
blocks = self.transformer_blocks
|
| 177 |
+
block0_out = blocks[0](packed, temb, adaln_indices, rotary_emb, attention_mask)
|
| 178 |
+
first_residual = block0_out - packed
|
| 179 |
+
|
| 180 |
+
# The generated rows trail their modality's index list, so the last row of each carries that stream's live noise
|
| 181 |
+
# level. The scheduler exposes `timesteps = 1 - sigmas[:-1]`.
|
| 182 |
+
sigma_video = 1.0 - float(timestep[timestep_indices[video_indices[-1]]].item())
|
| 183 |
+
sigma_audio = 1.0 - float(timestep[timestep_indices[audio_indices[-1]]].item())
|
| 184 |
+
|
| 185 |
+
use_cache = False
|
| 186 |
+
if (
|
| 187 |
+
state.prev_first_residual is not None
|
| 188 |
+
and state.tail_residual is not None
|
| 189 |
+
and state.prev_first_residual.shape == first_residual.shape
|
| 190 |
+
and state.consecutive_hits < MAX_CONSECUTIVE_HITS
|
| 191 |
+
and state.end_sigma <= sigma_video <= state.start_sigma
|
| 192 |
+
):
|
| 193 |
+
use_cache = _rel_l1(first_residual, state.prev_first_residual) <= state.threshold
|
| 194 |
+
|
| 195 |
+
if use_cache:
|
| 196 |
+
state.consecutive_hits += 1
|
| 197 |
+
state.skipped += 1
|
| 198 |
+
trunk_out = block0_out + state.tail_residual
|
| 199 |
+
if state.audio_exempt and state.audio_history:
|
| 200 |
+
sigma_prev, feature_prev = state.audio_history[-1]
|
| 201 |
+
if len(state.audio_history) == 2 and abs(sigma_prev - state.audio_history[-2][0]) > 1e-8:
|
| 202 |
+
sigma_prev2, feature_prev2 = state.audio_history[-2]
|
| 203 |
+
ratio = (sigma_audio - sigma_prev) / (sigma_prev - sigma_prev2)
|
| 204 |
+
audio_feature = feature_prev + (feature_prev - feature_prev2) * ratio
|
| 205 |
+
else:
|
| 206 |
+
audio_feature = feature_prev
|
| 207 |
+
trunk_out = trunk_out.index_copy(1, audio_indices, audio_feature.to(trunk_out.dtype))
|
| 208 |
+
else:
|
| 209 |
+
state.consecutive_hits = 0
|
| 210 |
+
state.computed += 1
|
| 211 |
+
trunk_out = block0_out
|
| 212 |
+
for block in blocks[1:]:
|
| 213 |
+
trunk_out = block(trunk_out, temb, adaln_indices, rotary_emb, attention_mask)
|
| 214 |
+
state.tail_residual = (trunk_out - block0_out).detach()
|
| 215 |
+
state.prev_first_residual = first_residual.detach()
|
| 216 |
+
if state.audio_exempt:
|
| 217 |
+
state.audio_history.append((sigma_audio, trunk_out.index_select(1, audio_indices).detach().float()))
|
| 218 |
+
state.audio_history = state.audio_history[-2:]
|
| 219 |
+
|
| 220 |
+
out = self.norm_out(trunk_out, temb, timestep_indices).to(self.proj_out.weight.dtype)
|
| 221 |
+
video_output = self.proj_out(out).index_select(1, video_indices)
|
| 222 |
+
audio_output = self.audio_proj_out(out).index_select(1, audio_indices)
|
| 223 |
+
|
| 224 |
+
if not return_dict:
|
| 225 |
+
return (video_output, audio_output)
|
| 226 |
+
return MiniMaxH3TransformerOutput(sample=video_output, audio_sample=audio_output)
|
| 227 |
+
|
| 228 |
+
|
| 229 |
+
def _forward(
|
| 230 |
+
self,
|
| 231 |
+
hidden_states,
|
| 232 |
+
audio_hidden_states,
|
| 233 |
+
encoder_hidden_states,
|
| 234 |
+
timestep,
|
| 235 |
+
timestep_indices,
|
| 236 |
+
token_tags,
|
| 237 |
+
position_ids,
|
| 238 |
+
video_indices,
|
| 239 |
+
audio_indices,
|
| 240 |
+
text_indices,
|
| 241 |
+
attention_kwargs=None,
|
| 242 |
+
return_dict: bool = True,
|
| 243 |
+
):
|
| 244 |
+
"""The installed forward. Anything it cannot serve — a LoRA scale, an unexpected layout, a bug — is handed to the
|
| 245 |
+
original forward instead, for this call and every later one, so a cached request can degrade to an uncached one but
|
| 246 |
+
never to a failed one."""
|
| 247 |
+
state = self._h3_fbc
|
| 248 |
+
original = dict(
|
| 249 |
+
hidden_states=hidden_states,
|
| 250 |
+
audio_hidden_states=audio_hidden_states,
|
| 251 |
+
encoder_hidden_states=encoder_hidden_states,
|
| 252 |
+
timestep=timestep,
|
| 253 |
+
timestep_indices=timestep_indices,
|
| 254 |
+
token_tags=token_tags,
|
| 255 |
+
position_ids=position_ids,
|
| 256 |
+
video_indices=video_indices,
|
| 257 |
+
audio_indices=audio_indices,
|
| 258 |
+
text_indices=text_indices,
|
| 259 |
+
attention_kwargs=attention_kwargs,
|
| 260 |
+
return_dict=return_dict,
|
| 261 |
+
)
|
| 262 |
+
# `apply_lora_scale` decorates the real forward and this one is not it, so a request that actually scales a LoRA
|
| 263 |
+
# goes down the original path rather than silently losing its scale.
|
| 264 |
+
if state.failed or (attention_kwargs or {}).get("scale") is not None:
|
| 265 |
+
return state.original(**original)
|
| 266 |
+
try:
|
| 267 |
+
return _cached_forward(
|
| 268 |
+
self,
|
| 269 |
+
state,
|
| 270 |
+
hidden_states,
|
| 271 |
+
audio_hidden_states,
|
| 272 |
+
encoder_hidden_states,
|
| 273 |
+
timestep,
|
| 274 |
+
timestep_indices,
|
| 275 |
+
token_tags,
|
| 276 |
+
position_ids,
|
| 277 |
+
video_indices,
|
| 278 |
+
audio_indices,
|
| 279 |
+
text_indices,
|
| 280 |
+
return_dict,
|
| 281 |
+
)
|
| 282 |
+
except Exception as error:
|
| 283 |
+
state.failed = True
|
| 284 |
+
print(f"[h3-fbc] disabled for this request ({type(error).__name__}: {error}); running uncached", flush=True)
|
| 285 |
+
return state.original(**original)
|
| 286 |
+
|
| 287 |
+
|
| 288 |
+
def install(transformer, steps: int = 0, threshold: float = THRESHOLD, audio_exempt: bool = AUDIO_EXEMPT) -> bool:
|
| 289 |
+
"""Bind the caching forward onto `transformer`. Returns whether it went on.
|
| 290 |
+
|
| 291 |
+
`accelerate`'s `add_hook_to_module` — what `ComponentsManager.enable_auto_cpu_offload` installs — moves the real
|
| 292 |
+
forward to `_old_forward` and puts its own onload wrapper in `forward`. Replacing `forward` there would step over
|
| 293 |
+
the wrapper and run the block stack against weights still on the host, so the replacement goes into `_old_forward`
|
| 294 |
+
whenever the hook is present.
|
| 295 |
+
"""
|
| 296 |
+
import inspect
|
| 297 |
+
|
| 298 |
+
if getattr(transformer, "_h3_fbc", None) is not None:
|
| 299 |
+
return True
|
| 300 |
+
|
| 301 |
+
hooked = hasattr(transformer, "_hf_hook") and hasattr(transformer, "_old_forward")
|
| 302 |
+
current = transformer._old_forward if hooked else transformer.forward
|
| 303 |
+
missing = [name for name in FORWARD_PARAMETERS if name not in inspect.signature(current).parameters]
|
| 304 |
+
if missing:
|
| 305 |
+
print(f"[h3-fbc] this transformer's forward has no {missing}; running uncached", flush=True)
|
| 306 |
+
return False
|
| 307 |
+
if not hasattr(transformer, "transformer_blocks") or len(transformer.transformer_blocks) < 2:
|
| 308 |
+
print("[h3-fbc] no block stack to skip; running uncached", flush=True)
|
| 309 |
+
return False
|
| 310 |
+
|
| 311 |
+
state = _State(threshold, steps, audio_exempt)
|
| 312 |
+
state.original = current
|
| 313 |
+
transformer._h3_fbc = state
|
| 314 |
+
bound = types.MethodType(_forward, transformer)
|
| 315 |
+
if hooked:
|
| 316 |
+
transformer._old_forward = bound
|
| 317 |
+
else:
|
| 318 |
+
transformer.forward = bound
|
| 319 |
+
return True
|
| 320 |
+
|
| 321 |
+
|
| 322 |
+
def uninstall(transformer) -> None:
|
| 323 |
+
state = getattr(transformer, "_h3_fbc", None)
|
| 324 |
+
if state is None:
|
| 325 |
+
return
|
| 326 |
+
if hasattr(transformer, "_hf_hook") and hasattr(transformer, "_old_forward"):
|
| 327 |
+
transformer._old_forward = state.original
|
| 328 |
+
else:
|
| 329 |
+
transformer.__dict__.pop("forward", None)
|
| 330 |
+
del transformer._h3_fbc
|
| 331 |
+
total = state.computed + state.skipped
|
| 332 |
+
if total:
|
| 333 |
+
print(
|
| 334 |
+
f"[h3-fbc] {state.skipped}/{total} forwards served from cache "
|
| 335 |
+
f"(threshold {state.threshold}, audio exemption {'on' if state.audio_exempt else 'off'})",
|
| 336 |
+
flush=True,
|
| 337 |
+
)
|
| 338 |
+
|
| 339 |
+
|
| 340 |
+
@contextlib.contextmanager
|
| 341 |
+
def enabled(transformer, steps: int = 0):
|
| 342 |
+
"""Cache the trunk for the duration of one request. The state is per-request by construction — a residual only ever
|
| 343 |
+
means something within the schedule it was measured on — and nothing in here can raise into the request."""
|
| 344 |
+
installed = False
|
| 345 |
+
if ENABLED:
|
| 346 |
+
try:
|
| 347 |
+
installed = install(transformer, steps=steps)
|
| 348 |
+
except Exception as error:
|
| 349 |
+
print(f"[h3-fbc] install failed ({type(error).__name__}: {error}); running uncached", flush=True)
|
| 350 |
+
try:
|
| 351 |
+
yield installed
|
| 352 |
+
finally:
|
| 353 |
+
if installed:
|
| 354 |
+
try:
|
| 355 |
+
uninstall(transformer)
|
| 356 |
+
except Exception as error:
|
| 357 |
+
print(f"[h3-fbc] uninstall failed ({type(error).__name__}: {error})", flush=True)
|
h3_split_blocks.py
ADDED
|
@@ -0,0 +1,147 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""The halves of a **split** MiniMax-H3 deployment, for both of its checkpoint partitions.
|
| 2 |
+
|
| 3 |
+
MiniMax-H3 is 195.9 GiB in bfloat16 and a ZeroGPU Space is evicted at 150 GB of storage, so `MiniMaxH3Blocks` is cut
|
| 4 |
+
at its `text_encoder` step: the 62.14 GiB Qwen3-VL runs in the conditioner Space, everything else in a generator
|
| 5 |
+
Space, and `prompt_embeds` + `text_token_tags` is the whole wire format between them.
|
| 6 |
+
|
| 7 |
+
`resize` / `setup` run on **both** sides: they own no pretrained component, and each half needs the canvas and the
|
| 8 |
+
prepared keyframes or normalized references. Both conditioner halves also return the resolved `height` / `width` /
|
| 9 |
+
`num_frames`, which the generating half pins rather than re-deriving.
|
| 10 |
+
|
| 11 |
+
Two things the blocks leave to the caller: a keyframe reaches them EXIF-transposed and in RGB, and the `t2va` / `fl2va`
|
| 12 |
+
frame count is aligned to `17 * n + 5` before the call, since that arithmetic lives on the denoising side of the cut.
|
| 13 |
+
"""
|
| 14 |
+
|
| 15 |
+
from diffusers.modular_pipelines.minimax_h3.before_encoder import MiniMaxH3Ref2VASetupStep
|
| 16 |
+
from diffusers.modular_pipelines.minimax_h3.decoders import MiniMaxH3AfterDenoiseStep
|
| 17 |
+
from diffusers.modular_pipelines.minimax_h3.encoders import (
|
| 18 |
+
MiniMaxH3Ref2VAReferenceEncoderStep,
|
| 19 |
+
MiniMaxH3Ref2VATextEncoderStep,
|
| 20 |
+
MiniMaxH3TextEncoderStep,
|
| 21 |
+
)
|
| 22 |
+
from diffusers.modular_pipelines.minimax_h3.modular_blocks_minimax_h3 import (
|
| 23 |
+
MiniMaxH3AutoKeyframeVaeEncoderStep,
|
| 24 |
+
MiniMaxH3AutoResizeStep,
|
| 25 |
+
MiniMaxH3CoreDenoiseStep,
|
| 26 |
+
MiniMaxH3DecodeStep,
|
| 27 |
+
MiniMaxH3Ref2VACoreDenoiseStep,
|
| 28 |
+
_generation_outputs,
|
| 29 |
+
)
|
| 30 |
+
from diffusers.modular_pipelines.modular_pipeline import SequentialPipelineBlocks
|
| 31 |
+
from diffusers.modular_pipelines.modular_pipeline_utils import OutputParam
|
| 32 |
+
|
| 33 |
+
|
| 34 |
+
def _wire_outputs(num_frames: bool = True) -> list[OutputParam]:
|
| 35 |
+
"""The wire format of the split. `num_frames` is declared by the `ref2va` half alone, whose setup resolves one."""
|
| 36 |
+
return [
|
| 37 |
+
OutputParam.template("prompt_embeds"),
|
| 38 |
+
OutputParam("text_token_tags", description="The per-row modality tag of every row of `prompt_embeds`."),
|
| 39 |
+
OutputParam("height", type_hint=int, description="Resolved height of the generated video in pixels."),
|
| 40 |
+
OutputParam("width", type_hint=int, description="Resolved width of the generated video in pixels."),
|
| 41 |
+
*(
|
| 42 |
+
[OutputParam("num_frames", type_hint=int, description="Resolved number of frames, of the form 17 * n + 5.")]
|
| 43 |
+
if num_frames
|
| 44 |
+
else []
|
| 45 |
+
),
|
| 46 |
+
]
|
| 47 |
+
|
| 48 |
+
|
| 49 |
+
class MiniMaxH3ConditionerBlocks(SequentialPipelineBlocks):
|
| 50 |
+
"""The conditioner half of a split MiniMax-H3: the keyframes on the canvas plus the Qwen3-VL read at layer 50."""
|
| 51 |
+
|
| 52 |
+
model_name = "minimax-h3"
|
| 53 |
+
block_classes = [MiniMaxH3AutoResizeStep, MiniMaxH3TextEncoderStep]
|
| 54 |
+
block_names = ["resize", "text_encoder"]
|
| 55 |
+
|
| 56 |
+
@property
|
| 57 |
+
def description(self):
|
| 58 |
+
return (
|
| 59 |
+
"The conditioner half of a split MiniMax-H3 deployment: puts the keyframes onto the target canvas and "
|
| 60 |
+
"encodes MiniMax-H3's presentation of the request into the `prompt_embeds` / `text_token_tags` pair the "
|
| 61 |
+
"denoising half consumes. The frame count is the caller's to align."
|
| 62 |
+
)
|
| 63 |
+
|
| 64 |
+
@property
|
| 65 |
+
def outputs(self):
|
| 66 |
+
return _wire_outputs(num_frames=False)
|
| 67 |
+
|
| 68 |
+
|
| 69 |
+
class MiniMaxH3GeneratorBlocks(SequentialPipelineBlocks):
|
| 70 |
+
"""The denoising half of a split MiniMax-H3: `MiniMaxH3Blocks` with its `text_encoder` step removed."""
|
| 71 |
+
|
| 72 |
+
model_name = "minimax-h3"
|
| 73 |
+
block_classes = [
|
| 74 |
+
MiniMaxH3AutoResizeStep,
|
| 75 |
+
MiniMaxH3AutoKeyframeVaeEncoderStep,
|
| 76 |
+
MiniMaxH3CoreDenoiseStep,
|
| 77 |
+
MiniMaxH3AfterDenoiseStep,
|
| 78 |
+
MiniMaxH3DecodeStep,
|
| 79 |
+
]
|
| 80 |
+
block_names = ["resize", "vae_encoder", "denoise", "after_denoise", "decode"]
|
| 81 |
+
|
| 82 |
+
@property
|
| 83 |
+
def description(self):
|
| 84 |
+
return (
|
| 85 |
+
"The denoising half of a split MiniMax-H3 deployment: the `t2va` / `fl2va` branch of `MiniMaxH3Blocks` "
|
| 86 |
+
"without its text-encoder step, so `prompt_embeds` and `text_token_tags` come in as inputs and the "
|
| 87 |
+
"62.14 GiB Qwen3-VL conditioner is never loaded here."
|
| 88 |
+
)
|
| 89 |
+
|
| 90 |
+
@property
|
| 91 |
+
def outputs(self):
|
| 92 |
+
return _generation_outputs()
|
| 93 |
+
|
| 94 |
+
|
| 95 |
+
class MiniMaxH3Ref2VAConditionerBlocks(SequentialPipelineBlocks):
|
| 96 |
+
"""The conditioner half of a split `ref2va`: the resolved plan plus the Qwen3-VL read at its 50th layer.
|
| 97 |
+
|
| 98 |
+
Component for component this is `MiniMaxH3ConditionerBlocks`, so one conditioner Space serves both partitions.
|
| 99 |
+
What differs is the presentation: `ref2va` prepends a label per reference and a vision block per image and per
|
| 100 |
+
merged video frame pair, so the references themselves have to reach this half.
|
| 101 |
+
"""
|
| 102 |
+
|
| 103 |
+
model_name = "minimax-h3"
|
| 104 |
+
block_classes = [MiniMaxH3Ref2VASetupStep, MiniMaxH3Ref2VATextEncoderStep]
|
| 105 |
+
block_names = ["setup", "text_encoder"]
|
| 106 |
+
|
| 107 |
+
@property
|
| 108 |
+
def description(self):
|
| 109 |
+
return (
|
| 110 |
+
"The conditioner half of a split MiniMax-H3 `ref2va` deployment: resolves the request plan (canvas, frame "
|
| 111 |
+
"count, references normalized onto MiniMax-H3's own rates and resolutions) and encodes MiniMax-H3's "
|
| 112 |
+
"presentation of it into the `prompt_embeds` / `text_token_tags` pair the denoising half consumes."
|
| 113 |
+
)
|
| 114 |
+
|
| 115 |
+
@property
|
| 116 |
+
def outputs(self):
|
| 117 |
+
return _wire_outputs()
|
| 118 |
+
|
| 119 |
+
|
| 120 |
+
class MiniMaxH3Ref2VAGeneratorBlocks(SequentialPipelineBlocks):
|
| 121 |
+
"""The denoising half of a split `ref2va`: the `ref2va` branch with its `text_encoder` step removed.
|
| 122 |
+
|
| 123 |
+
`reference_encoder` stays here, next to the two autoencoders it runs: its output shapes are where every reference
|
| 124 |
+
block's geometry in the packed layout comes from.
|
| 125 |
+
"""
|
| 126 |
+
|
| 127 |
+
model_name = "minimax-h3"
|
| 128 |
+
block_classes = [
|
| 129 |
+
MiniMaxH3Ref2VASetupStep,
|
| 130 |
+
MiniMaxH3Ref2VAReferenceEncoderStep,
|
| 131 |
+
MiniMaxH3Ref2VACoreDenoiseStep,
|
| 132 |
+
MiniMaxH3AfterDenoiseStep,
|
| 133 |
+
MiniMaxH3DecodeStep,
|
| 134 |
+
]
|
| 135 |
+
block_names = ["setup", "reference_encoder", "denoise", "after_denoise", "decode"]
|
| 136 |
+
|
| 137 |
+
@property
|
| 138 |
+
def description(self):
|
| 139 |
+
return (
|
| 140 |
+
"The denoising half of a split MiniMax-H3 `ref2va` deployment: the `ref2va` branch of `MiniMaxH3Blocks` "
|
| 141 |
+
"without its text-encoder step, so `prompt_embeds` and `text_token_tags` come in as inputs and the "
|
| 142 |
+
"62.14 GiB Qwen3-VL conditioner is never loaded here. The transformer is the `transformer_ref` partition."
|
| 143 |
+
)
|
| 144 |
+
|
| 145 |
+
@property
|
| 146 |
+
def outputs(self):
|
| 147 |
+
return _generation_outputs()
|
ncii_guard.py
ADDED
|
@@ -0,0 +1,90 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""The NCII prompt guard, in its own process.
|
| 2 |
+
|
| 3 |
+
[`hfmlsoc/ncii-light-guard-v01`](https://huggingface.co/hfmlsoc/ncii-light-guard-v01) is a 270M CPU text
|
| 4 |
+
classifier scoring the NCII risk of an edit prompt. It cannot live in the main process: with it loaded there,
|
| 5 |
+
every subsequent `@spaces.GPU` worker dies at `worker_init` with `RuntimeError: No CUDA GPUs are available` —
|
| 6 |
+
the fork inherits whatever CUDA driver state the classifier's torch activity left behind, and a factory reboot
|
| 7 |
+
does not clear it. It cannot be a `multiprocessing.spawn` child either: spawn re-imports the parent's main
|
| 8 |
+
module, and on a Space that main module is `app.py` — the child would re-run the whole startup, `start()`
|
| 9 |
+
included. So the classifier runs this file as a plain subprocess — a fresh interpreter that never sees `spaces`
|
| 10 |
+
— and answers over stdin/stdout, one JSON object per line.
|
| 11 |
+
"""
|
| 12 |
+
|
| 13 |
+
from __future__ import annotations
|
| 14 |
+
|
| 15 |
+
import json
|
| 16 |
+
import os
|
| 17 |
+
import select
|
| 18 |
+
import subprocess
|
| 19 |
+
import sys
|
| 20 |
+
import threading
|
| 21 |
+
|
| 22 |
+
GUARD_REPO = "hfmlsoc/ncii-light-guard-v01"
|
| 23 |
+
|
| 24 |
+
_lock = threading.Lock()
|
| 25 |
+
_process: subprocess.Popen | None = None
|
| 26 |
+
|
| 27 |
+
|
| 28 |
+
def _read(timeout: float) -> dict:
|
| 29 |
+
readable, _, _ = select.select([_process.stdout], [], [], timeout)
|
| 30 |
+
if not readable:
|
| 31 |
+
raise TimeoutError(f"the guard did not answer within {timeout}s")
|
| 32 |
+
line = _process.stdout.readline()
|
| 33 |
+
if not line:
|
| 34 |
+
raise EOFError("the guard process died")
|
| 35 |
+
return json.loads(line)
|
| 36 |
+
|
| 37 |
+
|
| 38 |
+
def _spawn() -> None:
|
| 39 |
+
global _process
|
| 40 |
+
_process = subprocess.Popen(
|
| 41 |
+
[sys.executable, os.path.abspath(__file__)],
|
| 42 |
+
stdin=subprocess.PIPE,
|
| 43 |
+
stdout=subprocess.PIPE,
|
| 44 |
+
text=True,
|
| 45 |
+
bufsize=1,
|
| 46 |
+
)
|
| 47 |
+
# Generous: a cold cache downloads the checkpoint first.
|
| 48 |
+
assert _read(300.0) == {"status": "ready"}
|
| 49 |
+
|
| 50 |
+
|
| 51 |
+
def start() -> None:
|
| 52 |
+
"""Launch the worker and block until its model is up. Called once at startup; `classify` revives it if it dies."""
|
| 53 |
+
with _lock:
|
| 54 |
+
_spawn()
|
| 55 |
+
|
| 56 |
+
|
| 57 |
+
def classify(prompt: str, timeout: float = 60.0) -> dict:
|
| 58 |
+
"""`{'label': 'safe' | 'ncii', 'score': ...}` for one prompt, replacing a dead or wedged worker once."""
|
| 59 |
+
with _lock:
|
| 60 |
+
for attempt in (0, 1):
|
| 61 |
+
try:
|
| 62 |
+
if _process is None or _process.poll() is not None:
|
| 63 |
+
_spawn()
|
| 64 |
+
_process.stdin.write(json.dumps({"prompt": prompt}) + "\n")
|
| 65 |
+
_process.stdin.flush()
|
| 66 |
+
return _read(timeout)
|
| 67 |
+
except Exception:
|
| 68 |
+
if attempt:
|
| 69 |
+
raise
|
| 70 |
+
if _process is not None and _process.poll() is None:
|
| 71 |
+
_process.kill()
|
| 72 |
+
|
| 73 |
+
|
| 74 |
+
def _serve() -> None:
|
| 75 |
+
"""The child: plain torch on CPU. The protocol keeps the real stdout to itself — everything else
|
| 76 |
+
(download progress, warnings) is pushed over to stderr so it cannot corrupt a reply."""
|
| 77 |
+
protocol = os.fdopen(os.dup(1), "w", buffering=1)
|
| 78 |
+
os.dup2(2, 1)
|
| 79 |
+
|
| 80 |
+
from transformers import pipeline
|
| 81 |
+
|
| 82 |
+
classifier = pipeline("text-classification", model=GUARD_REPO, device="cpu")
|
| 83 |
+
protocol.write(json.dumps({"status": "ready"}) + "\n")
|
| 84 |
+
for line in sys.stdin:
|
| 85 |
+
result = classifier(json.loads(line)["prompt"], truncation=True)[0]
|
| 86 |
+
protocol.write(json.dumps({"label": result["label"], "score": float(result["score"])}) + "\n")
|
| 87 |
+
|
| 88 |
+
|
| 89 |
+
if __name__ == "__main__":
|
| 90 |
+
_serve()
|
packages.txt
ADDED
|
@@ -0,0 +1 @@
|
|
|
|
|
|
|
| 1 |
+
ffmpeg
|
requirements.txt
ADDED
|
@@ -0,0 +1,33 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# `diffusers` is installed from the canonical MiniMax-H3 pull request,
|
| 2 |
+
# https://github.com/huggingface/diffusers/pull/14371 ("Minimax h3 follow up (review & refactor)"), pinned to a
|
| 3 |
+
# **commit** rather than to its `minimax-h3-refactor` branch: the PR is a WIP and its head moves, and this Space's
|
| 4 |
+
# `h3_split_blocks.py` subclasses its block classes. Re-pin — and re-check `h3_split_blocks.py` against the block
|
| 5 |
+
# names of the new head — whenever the PR updates. (diffusers `main` has already renamed several of them.) It is
|
| 6 |
+
# also the commit `multimodalart/qwen3vl-conditioner` serves, and the two halves have to agree on the wire format.
|
| 7 |
+
#
|
| 8 |
+
# 665f578278365ea4a3318cb8c9b66ce6c01204b9 = refs/pull/14371/head at the time of this deploy
|
| 9 |
+
--extra-index-url https://download.pytorch.org/whl/cu130
|
| 10 |
+
diffusers @ git+https://github.com/huggingface/diffusers.git@665f578278365ea4a3318cb8c9b66ce6c01204b9
|
| 11 |
+
torch==2.11.0
|
| 12 |
+
torchvision==0.26.0
|
| 13 |
+
# A scene clip's soundtrack that is not already at the audio VAE's 32 kHz is resampled with `torchaudio`. Real
|
| 14 |
+
# camera and phone footage is 44.1/48 kHz, so this is hit by most uploads that are not silent.
|
| 15 |
+
torchaudio==2.11.0
|
| 16 |
+
# The Qwen3-VL processor decides the vision patch count, so a different minor changes the conditioning.
|
| 17 |
+
transformers==5.8.0
|
| 18 |
+
accelerate==1.14.0
|
| 19 |
+
# diffusers pins <2.
|
| 20 |
+
huggingface-hub==1.24.0
|
| 21 |
+
# The character-swap LoRA is attached as a runtime PEFT adapter (`load_lora_adapter` + `set_adapters`), not folded
|
| 22 |
+
# into the base weights, so its strength stays a per-request slider. Without `peft`, `load_lora_adapter` raises
|
| 23 |
+
# ModuleNotFoundError at startup and the Space would quietly serve the base `ref2va` model.
|
| 24 |
+
peft
|
| 25 |
+
# `gradio` and `spaces` are deliberately absent: the Space runtime installs both itself (gradio from `sdk_version`
|
| 26 |
+
# in README.md, `spaces` at whichever version it currently ships), so pinning either here is a resolver conflict
|
| 27 |
+
# rather than a version choice.
|
| 28 |
+
# PyAV decodes the scene clip and muxes the generated soundtrack onto the frames (`encode_video`).
|
| 29 |
+
av
|
| 30 |
+
pillow
|
| 31 |
+
numpy
|
| 32 |
+
requests
|
| 33 |
+
safetensors>=0.8.0
|