multimodalart HF Staff commited on
Commit
4b6691d
·
verified ·
1 Parent(s): baa3322

MiniMax-H3 character-swap LoRA demo (ref2va, split conditioner)

Browse files
.gitattributes CHANGED
@@ -33,3 +33,8 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
 
 
 
 
 
 
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
+ examples/cafe_scene.mp4 filter=lfs diff=lfs merge=lfs -text
37
+ examples/character_anime.png filter=lfs diff=lfs merge=lfs -text
38
+ examples/character_hiker.png filter=lfs diff=lfs merge=lfs -text
39
+ examples/character_sheet_orin.png filter=lfs diff=lfs merge=lfs -text
40
+ examples/workshop_scene.mp4 filter=lfs diff=lfs merge=lfs -text
README.md CHANGED
@@ -1,13 +1,149 @@
1
  ---
2
- title: Minimax H3 Character Swap Lora
3
- emoji: 📈
4
- colorFrom: purple
5
- colorTo: purple
6
  sdk: gradio
7
  sdk_version: 6.28.0
8
- python_version: '3.12'
9
  app_file: app.py
10
  pinned: false
 
 
 
 
 
 
 
 
 
11
  ---
12
 
13
- Check out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  ---
2
+ title: MiniMax-H3 Character Swap LoRA
3
+ emoji: 🎭
4
+ colorFrom: red
5
+ colorTo: yellow
6
  sdk: gradio
7
  sdk_version: 6.28.0
 
8
  app_file: app.py
9
  pinned: false
10
+ short_description: Swap one character in a clip for a reference character
11
+ python_version: "3.12"
12
+ startup_duration_timeout: 1h
13
+ models:
14
+ - akatz-ai/MiniMax-H3-Character-Swap-LoRA
15
+ - multimodalart/MiniMax-H3-Pruned
16
+ - MiniMaxAI/MiniMax-H3
17
+ datasets:
18
+ - akatz-ai/H3-Character-Swap-v1
19
  ---
20
 
21
+ # MiniMax-H3 Character Swap LoRA
22
+
23
+ A demo of [`akatz-ai/MiniMax-H3-Character-Swap-LoRA`](https://huggingface.co/akatz-ai/MiniMax-H3-Character-Swap-LoRA),
24
+ Akatz Labs' experimental character-replacement adapter for
25
+ [MiniMax-H3](https://huggingface.co/MiniMaxAI/MiniMax-H3)'s `ref2va` partition. Hand it a scene clip and a character
26
+ reference, name who to replace, and it puts the reference character into the shot — identity, outfit and art style
27
+ carried over — while the background, camera, lighting and everyone else stay where they were. Video and its
28
+ synchronized soundtrack come out of a single denoising pass.
29
+
30
+ ## The request
31
+
32
+ Two references, **in the order the model reads them**, and a short targeting instruction:
33
+
34
+ | Slot | Label in the prompt | What it is |
35
+ |---|---|---|
36
+ | reference 1 | `<Picture 1>` | the replacement character — a portrait or a full character sheet |
37
+ | reference 2 | `<Video 1>` | the clip to edit, 5 frames to 15 s |
38
+
39
+ > Swap the man in the purple shirt in \<Video 1\> with the character in \<Picture 1\>.
40
+
41
+ That order is the one thing about this request worth being careful with, and it is **not** the order the prompt
42
+ names them in. MiniMax-H3 numbers a reference's label *per modality*, so the first image is `<Picture 1>` and the
43
+ first video is `<Video 1>` whichever way round they are packed — but the packed order still fixes the shared
44
+ audio/video rotary clock, so the same two references swapped round are a different request. The LoRA was trained
45
+ with the character sheet first: `control_path: [character_references, scene_videos]` in
46
+ `configs/trained-run-1000.json` of [`akatz-ai/H3-Character-Swap-v1`](https://huggingface.co/datasets/akatz-ai/H3-Character-Swap-v1).
47
+ `collect()` builds the list that way and nothing else in the app reorders it.
48
+
49
+ No trigger word was trained. Strength **1.0** is the card's recommendation; **0** is the base `ref2va` model, which
50
+ is the comparison the adapter was judged against, so the slider doubles as an A/B.
51
+
52
+ ## What it is good at, and what it is not
53
+
54
+ The card is candid, and this demo does not oversell it. It is a **1,000-update experimental** adapter:
55
+
56
+ * background and scene preservation improved over the base model in the author's local comparisons — qualitative,
57
+ not a benchmark,
58
+ * motion timing, facial expressions and hard cuts remain unreliable; a hard cut can become a zoom or a gradual
59
+ reposition,
60
+ * long windows drift in framing and placement. **Short continuous shots of roughly 3–5 s** are the promising range,
61
+ * two-character inference was tested but multi-character replacement was never supervised,
62
+ * the soundtrack is generated, not carried over. Audio preservation in the author's later review was a remux, which
63
+ does not repair lip-sync drift.
64
+
65
+ Its 94 training edits targeted **single still frames**, with five-frame static clips standing in for `<Video 1>` —
66
+ which is exactly what the examples below are. Real moving footage goes in the same slot and is what the 40
67
+ preservation clips regularized, but it is the harder case.
68
+
69
+ ## Defaults, and where they come from
70
+
71
+ `1344x768` at 24 fps, 73 frames (3.04 s), 28 steps, seed 904231 — the `sample` block of the LoRA's own
72
+ `configs/trained-run-1000.json`. The canvas dropdown follows the scene clip's aspect ratio on upload, because a
73
+ character swap is asked to keep the source framing and a portrait clip generated on a landscape canvas is a
74
+ recomposed shot before the model has done anything. "Match the scene clip's length" generates for as long as the
75
+ clip runs whenever that is a length MiniMax-H3 generates (2–14 s); the five-frame example clips are not, so they
76
+ fall through to the slider.
77
+
78
+ ## How it is deployed
79
+
80
+ MiniMax-H3 is 195.9 GiB in bfloat16 and a ZeroGPU Space is evicted at 150 GB of storage, so `MiniMaxH3Blocks` is
81
+ cut at its `text_encoder` step. This Space is the **denoising half** of `ref2va` — the `transformer_ref` partition
82
+ and the two autoencoders — and the 62 GiB Qwen3-VL conditioner runs in
83
+ [`multimodalart/qwen3vl-conditioner`](https://huggingface.co/spaces/multimodalart/qwen3vl-conditioner), which this
84
+ Space calls over the gradio API for every request. `prompt_embeds` + `text_token_tags` is the whole wire format;
85
+ `reference_encoder` stays here, next to the autoencoders it runs. `h3_split_blocks.py` is the subclass that removes
86
+ the step.
87
+
88
+ The DiT is [`multimodalart/MiniMax-H3-Pruned`](https://huggingface.co/multimodalart/MiniMax-H3-Pruned)'s
89
+ `transformer_ref`: the released partition with its AdaLN input projections folded onto their reachable rank, 37.5
90
+ GiB instead of 61.7. That is the **same checkpoint family the LoRA was trained against** — ai-toolkit trained it on
91
+ Comfy-Org's `minimax_h3_ref2va_pruned_int8_convrot`, an int8 ConvRot quantization of these weights — and everything
92
+ the adapter touches is identical between the pruned and the released partition. The run's
93
+ `network_kwargs.ignore_if_contains = ["adaln_proj"]` kept it off the timestep path, which is the only place the two
94
+ differ.
95
+
96
+ **GPU time is priced per request, not per Space.** MiniMax-H3 attends over one packed sequence, and on this half the
97
+ references dominate its length: a 2048-short-edge character sheet is thousands of conditioning rows on top of the
98
+ generated ones. `get_duration` evaluates a fitted cost model over the sequence it is about to denoise — the
99
+ conditioner's exact token count plus the references measured from metadata — instead of reserving a flat ceiling for
100
+ everything, because the pool reserves whatever number it is given.
101
+
102
+ A first-block cache (`h3_fbc.py`, ported from `duckyshell/ComfyUI-MiniMaxH3-FirstBlockCache` with an audio
103
+ exemption) skips blocks 1–49 on steps whose block-0 residual has barely moved. `H3_FBC=0` restores the uncached
104
+ trajectory exactly.
105
+
106
+ There is no AoTI on this Space, unlike its siblings: a compiled block package binds the base module's weights by
107
+ fully qualified name and would run straight past the PEFT branch the adapter lives in.
108
+
109
+ ## Space variables
110
+
111
+ | Variable | Default | Meaning |
112
+ |---|---|---|
113
+ | `H3_LORA_SCALE` | `1.0` | Default adapter strength. |
114
+ | `H3_CONDITIONER` | `multimodalart/qwen3vl-conditioner` | The public Space this one asks for embeddings; the client passes no token, so the call runs on the caller's own quota. |
115
+ | `H3_MODEL_REPO` | `multimodalart/MiniMax-H3-Pruned` | The diffusers-layout DiT. |
116
+ | `H3_ATTENTION` | `_native_cudnn` | cuDNN's fused kernel, 10–20% faster than the SDPA default. The two float32 VAEs are pinned to torch SDPA, which cuDNN has no kernel for. |
117
+ | `H3_FBC` / `H3_FBC_THRESHOLD` | `1` / `0.05` | First-block cache and its relative-L1 gate. |
118
+ | `H3_GPU_SIZE` | `xlarge` | ZeroGPU allocation size. `large` does not fit. |
119
+ | `H3_PLACEMENT` | `lazy` | Moves the partition onto the card on the first GPU call and leaves it there. |
120
+
121
+ ## Safety
122
+
123
+ Every request here carries an image and a video reference — the edit-on-a-real-photo case — so
124
+ [`hfmlsoc/ncii-light-guard-v01`](https://huggingface.co/hfmlsoc/ncii-light-guard-v01) screens the prompt on all of
125
+ them, before the conditioner call and before any GPU is booked. It runs in its own subprocess
126
+ (`ncii_guard.py`): loaded in the main process, its torch activity poisons every later ZeroGPU fork.
127
+
128
+ ## Example assets
129
+
130
+ All five files in `examples/` are from
131
+ [`akatz-ai/H3-Character-Swap-v1`](https://huggingface.co/datasets/akatz-ai/H3-Character-Swap-v1), the LoRA's own
132
+ training set — Apache-2.0 for Akatz Labs' synthetic contributions. They are the dataset's `CS001`, `CS051` and
133
+ `CS090` edits, with the instructions their own captions carry:
134
+
135
+ | File | Dataset path |
136
+ |---|---|
137
+ | `cafe_scene.mp4` | `checks/smoke-data/edits/train/scene_videos/CS001.mp4` (shared with `CS051`) |
138
+ | `character_hiker.png` | `.../character_references/CS001.png` |
139
+ | `character_anime.png` | `.../character_references/CS051.png` — a cross-style swap |
140
+ | `workshop_scene.mp4` | `.../scene_videos/CS090.mp4` |
141
+ | `character_sheet_orin.png` | `.../character_references/CS090.png` — a multi-view character sheet |
142
+
143
+ ## License
144
+
145
+ The adapter is distributed under the
146
+ [MiniMax H3 Community License Agreement](https://huggingface.co/akatz-ai/MiniMax-H3-Character-Swap-LoRA/blob/main/LICENSE),
147
+ **not** Apache-2.0, and that agreement excludes the US, EU, UK and Republic of Korea from its standard territorial
148
+ grant. Read the upstream terms; nothing here extends them. The dataset's own Apache-2.0 covers the example assets
149
+ only and does not replace the model's terms.
app.py ADDED
@@ -0,0 +1,785 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """MiniMax-H3 Character Swap LoRA — replace one character in a shot with a reference character.
2
+
3
+ `akatz-ai/MiniMax-H3-Character-Swap-LoRA` is a rank-16 ai-toolkit adapter for MiniMax-H3's **`ref2va`** partition.
4
+ A request is two references, in the order the LoRA was trained with — the character sheet as `<Picture 1>` and the
5
+ scene clip as `<Video 1>` — plus a short targeting instruction naming who to replace.
6
+
7
+ This Space is the denoising half of `ref2va`: the `transformer_ref` partition and the two autoencoders. The Qwen3-VL
8
+ conditioner runs in [`multimodalart/qwen3vl-conditioner`](https://huggingface.co/spaces/multimodalart/qwen3vl-conditioner),
9
+ which this Space calls over the gradio API for every request — MiniMax-H3 is 195.9 GiB in bfloat16 and a ZeroGPU
10
+ Space is evicted at 150 GB of storage, so the two halves cannot live together.
11
+ """
12
+
13
+ from __future__ import annotations
14
+
15
+ import os
16
+ import tempfile
17
+ import time
18
+ import traceback
19
+ from functools import cache
20
+
21
+ # Before anything that could initialize CUDA: `import spaces` patches `torch.cuda` so the DiT load can happen at
22
+ # startup rather than on GPU time.
23
+ import spaces
24
+ import gradio as gr
25
+ import torch
26
+
27
+ VERSION = "h3-character-swap/1"
28
+
29
+ # The AdaLN-pruned `ref2va` partition in diffusers layout: 37.5 GiB against the released partition's 61.7 GiB, and
30
+ # the **same** checkpoint family this LoRA was trained against — ai-toolkit trained it on Comfy-Org's
31
+ # `minimax_h3_ref2va_pruned_int8_convrot`, i.e. an int8 ConvRot quantization of these very weights. Everything the
32
+ # LoRA touches (the block stack's attention and feed-forward projections, and the token refiner's) is identical
33
+ # between the pruned and the released partition; only the AdaLN input width differs, and the LoRA excludes
34
+ # `adaln_proj` by construction (`network_kwargs.ignore_if_contains`).
35
+ MODEL_REPO = os.environ.get("H3_MODEL_REPO", "multimodalart/MiniMax-H3-Pruned")
36
+ LORA_REPO = os.environ.get("H3_LORA_REPO", "akatz-ai/MiniMax-H3-Character-Swap-LoRA")
37
+ LORA_FILE = os.environ.get("H3_LORA_FILE", "h3_character_swap_pro4500_1000.safetensors")
38
+ ADAPTER = "character_swap"
39
+ # The model card's own recommendation: "Start with the final checkpoint at strength 1.0."
40
+ DEFAULT_LORA_SCALE = float(os.environ.get("H3_LORA_SCALE", "1.0"))
41
+
42
+ CONDITIONER_SPACE = os.environ.get("H3_CONDITIONER", "multimodalart/qwen3vl-conditioner")
43
+ # `lazy` moves the whole partition onto the card on the first GPU call and leaves it there. Startup placement is not
44
+ # an option: `spaces`' startup `torch.pack()` writes every startup-resident CUDA tensor to a second copy on disk, and
45
+ # the weights plus their pack do not fit the storage quota next to the VAEs.
46
+ PLACEMENT = os.environ.get("H3_PLACEMENT", "lazy").lower()
47
+ # cuDNN's fused attention, 10-20% faster than the SDPA default on this pool and needs nothing installed.
48
+ # flash-attention 3 is sm90-only and this card is sm120 (the `zero-a10g` flavour name is legacy).
49
+ ATTENTION = os.environ.get("H3_ATTENTION", "_native_cudnn").lower()
50
+ GPU_SIZE = os.environ.get("H3_GPU_SIZE", "xlarge")
51
+ MIN_GPU_DURATION = int(os.environ.get("H3_GPU_DURATION_MIN", "120"))
52
+ MAX_GPU_DURATION = int(os.environ.get("H3_GPU_DURATION_MAX", "1500"))
53
+
54
+ # Must stay identical to the conditioner's table: the *label* goes over the wire, so a canvas that half does not
55
+ # know is rejected there and surfaces as a failure here.
56
+ CANVASES = {
57
+ # 16:9
58
+ "960x544 · 16:9 fast": (544, 960),
59
+ "1024x576 · 16:9 fast": (576, 1024),
60
+ "1152x640 · 16:9": (640, 1152),
61
+ "1280x704 · 16:9": (704, 1280),
62
+ "1344x768 · 16:9 full": (768, 1344),
63
+ # 9:16
64
+ "544x960 · 9:16 fast": (960, 544),
65
+ "640x1152 · 9:16": (1152, 640),
66
+ "768x1344 · 9:16 full": (1344, 768),
67
+ # 1:1
68
+ "544x544 · 1:1 fast": (544, 544),
69
+ "768x768 · 1:1 full": (768, 768),
70
+ "1024x1024 · 1:1 max": (1024, 1024),
71
+ # 4:3 / 3:4
72
+ "768x576 · 4:3 fast": (576, 768),
73
+ "1024x768 · 4:3 full": (768, 1024),
74
+ "576x768 · 3:4 fast": (768, 576),
75
+ "768x1024 · 3:4 full": (1024, 768),
76
+ # 21:9
77
+ "1152x512 · 21:9 fast": (512, 1152),
78
+ "1536x672 · 21:9 full": (672, 1536),
79
+ }
80
+ # The LoRA's own training bucket: 1344x768 at 24 fps, which is also the resolution its author sampled at.
81
+ DEFAULT_CANVAS = "1344x768 · 16:9 full"
82
+ # The canvases `pick_canvas` may snap a scene clip to — one per aspect family, at the LoRA's own 768 short edge.
83
+ AUTO_CANVASES = (
84
+ "1344x768 · 16:9 full",
85
+ "768x1344 · 9:16 full",
86
+ "768x768 · 1:1 full",
87
+ "1024x768 · 4:3 full",
88
+ "768x1024 · 3:4 full",
89
+ "1536x672 · 21:9 full",
90
+ )
91
+
92
+ FPS, FRAMES_PER_CHUNK, LATENTS_PER_CHUNK = 24, 17, 5
93
+ # It is the *snapped* frame count the ceiling has to hold for: 15 s is 360 frames, which rounds up to 362, i.e.
94
+ # 15.083 s, and is refused. 14 is the last whole second that survives the snap.
95
+ MAX_UI_DURATION, MIN_DURATION = 14, 2
96
+ # 73 frames, 3.04 s — the `num_frames` the LoRA's own training run sampled at, and the length of its regularization
97
+ # clips. Short, continuous shots are what its card reports as the promising range.
98
+ DEFAULT_DURATION = 3
99
+ # The scene clip is a reference video. The floor is five frames rather than the two seconds a motion reference wants,
100
+ # because a five-frame static clip is exactly what every character-swap edit in the training set used for `<Video 1>`.
101
+ MIN_REFERENCE_VIDEO, MAX_REFERENCE_VIDEO = 5 / FPS, 15.0
102
+ DEFAULT_STEPS = 28
103
+
104
+ # Seconds of GPU one request needs, from the packed sequence it is about to denoise: linear in the rows for the
105
+ # matmuls, quadratic for the attention. Fitted on the `t2va` half and checked against live `ref2va` requests.
106
+ STEP_LINEAR, STEP_QUADRATIC, SAFETY = 1.1745e-4, 3.8396e-9, 1.3
107
+ # The lazy `PIPE.to("cuda")` a cold worker pays inside its first GPU call; every request carries it, because nothing
108
+ # here knows whether the worker it lands on is cold.
109
+ PLACEMENT_ALLOWANCE = int(os.environ.get("H3_PLACEMENT_ALLOWANCE", "90"))
110
+ AUDIO_LATENTS_PER_SECOND, AUDIO_CHANNELS = 40, 2
111
+ REFERENCE_IMAGE_SHORT_EDGE, CANVAS_MULTIPLE = 2048, 32
112
+ DECODE_BASE, DECODE_PER_DEFAULT_CANVAS, DEFAULT_CANVAS_PIXELS = 15, 25, 960 * 544 * 124
113
+
114
+
115
+ def snap_frames(seconds: float) -> int:
116
+ """The frame count MiniMax-H3's video VAE can decode: the next `17 * n + 5` at 24 fps."""
117
+ frames = max(1, round(float(seconds) * FPS))
118
+ while frames % FRAMES_PER_CHUNK != LATENTS_PER_CHUNK:
119
+ frames += 1
120
+ return frames
121
+
122
+
123
+ def lower_duration_floor(seconds: float = MIN_DURATION) -> None:
124
+ """Let the pipeline generate below its 5 s floor. 56 frames (2.33 s) is fine on the released checkpoint."""
125
+ from diffusers.modular_pipelines.minimax_h3.modular_pipeline import MiniMaxH3ModularPipeline
126
+
127
+ MiniMaxH3ModularPipeline.min_duration = property(lambda self: float(seconds))
128
+
129
+
130
+ def video_latent_frames(num_frames: int) -> int:
131
+ """`17 * n + 5` frames become `5 * n + 2` video latents."""
132
+ return 5 * ((num_frames - LATENTS_PER_CHUNK) // FRAMES_PER_CHUNK) + 2
133
+
134
+
135
+ def target_rows(height: int, width: int, num_frames: int) -> int:
136
+ """The generated rows of the packed sequence: video patched `(1, 2, 2)`, plus two audio rows per latent."""
137
+ video = video_latent_frames(num_frames) * (height // CANVAS_MULTIPLE) * (width // CANVAS_MULTIPLE)
138
+ return video + round(num_frames / FPS * AUDIO_LATENTS_PER_SECOND) * AUDIO_CHANNELS
139
+
140
+
141
+ def reference_rows(references: list[tuple[str, str]], num_frames: int) -> int:
142
+ """The rows the two reference blocks add, from metadata alone — no decode."""
143
+ from PIL import Image
144
+
145
+ from diffusers.modular_pipelines.minimax_h3.modular_pipeline import resolve_canvas_size
146
+
147
+ rows = 0
148
+ for kind, path in references:
149
+ if kind == "image":
150
+ width, height = Image.open(path).size
151
+ scale = REFERENCE_IMAGE_SHORT_EDGE / min(width, height)
152
+ resolved = [
153
+ max(CANVAS_MULTIPLE, round(edge * scale / CANVAS_MULTIPLE) * CANVAS_MULTIPLE)
154
+ for edge in (height, width)
155
+ ]
156
+ rows += (resolved[0] // CANVAS_MULTIPLE) * (resolved[1] // CANVAS_MULTIPLE)
157
+ continue
158
+
159
+ video_seconds, audio_seconds = probe(path)
160
+ if kind == "video" and video_seconds is not None:
161
+ import av
162
+
163
+ with av.open(path) as container:
164
+ stream = container.streams.video[0]
165
+ source_height, source_width = stream.height, stream.width
166
+ canvas_height, canvas_width = resolve_canvas_size(source_width, source_height, CANVAS_MULTIPLE)
167
+ frames = min(round(video_seconds * FPS), num_frames)
168
+ snapped = max(1, (frames - LATENTS_PER_CHUNK) // FRAMES_PER_CHUNK) * FRAMES_PER_CHUNK + LATENTS_PER_CHUNK
169
+ rows += (
170
+ video_latent_frames(snapped)
171
+ * (canvas_height // CANVAS_MULTIPLE)
172
+ * (canvas_width // CANVAS_MULTIPLE)
173
+ )
174
+ if audio_seconds is not None:
175
+ seconds = min(audio_seconds, num_frames / FPS)
176
+ rows += round(seconds * AUDIO_LATENTS_PER_SECOND) * AUDIO_CHANNELS
177
+ return rows
178
+
179
+
180
+ def get_duration(
181
+ prompt_embeds, text_token_tags, references, height, width, num_frames, steps, seed, lora_scale, *_, **__
182
+ ):
183
+ """Seconds of GPU to reserve for one request. Takes the arguments of the `@spaces.GPU` function it decorates, and
184
+ tolerates the `gr.Progress` `spaces` injects."""
185
+ sequence = (
186
+ int(text_token_tags.shape[0])
187
+ + reference_rows(references, num_frames)
188
+ + target_rows(height, width, num_frames)
189
+ )
190
+ denoise = int(steps) * (STEP_LINEAR * sequence + STEP_QUADRATIC * sequence**2) * SAFETY
191
+ # The two reference encoders ahead of the loop, and the two decoders plus the mux after it. Both scale with what
192
+ # they are handed rather than with the step count.
193
+ encode = 5 + reference_rows(references, num_frames) * 1e-3
194
+ decode = DECODE_BASE + DECODE_PER_DEFAULT_CANVAS * (height * width * num_frames) / DEFAULT_CANVAS_PIXELS
195
+ total = PLACEMENT_ALLOWANCE + encode + denoise + decode + 10
196
+ duration = max(MIN_GPU_DURATION, min(MAX_GPU_DURATION, int(total)))
197
+ print(f"[{VERSION}] S={sequence} -> reserving {duration}s ({denoise:.0f}s of denoise at {steps} steps)", flush=True)
198
+ return duration
199
+
200
+
201
+ PIPE = None
202
+ MANAGER = None
203
+ LOAD_ERROR: str | None = None
204
+ LORA_STATUS: str = ""
205
+
206
+
207
+ # ── LoRA loading (runtime PEFT adapter) ──────────────────────────────────────
208
+ #
209
+ # The adapter targets the checkpoint's own module names (`diffusion_model.blocks.N.attn.qkv_proj`, `mlp.fc1`/`fc2`,
210
+ # `attn.out_proj`, and the same four under `token_refiner.blocks.N` — 208 modules, 416 tensors). diffusers serves
211
+ # the converted port, so every name — and, for two of them, the row layout of `lora_B` — has to be pushed through
212
+ # the same transforms `scripts/convert_minimax_h3_to_diffusers.py` applied to the base weights:
213
+ #
214
+ # * `blocks.` -> `transformer_blocks.`, `token_refiner.blocks.` -> `token_refiner.refiner_blocks.`,
215
+ # `attn.out_proj` -> `attn.to_out.0`, `mlp.fc2` -> `ff.net.2` (pure renames),
216
+ # * `mlp.fc1` -> `ff.net.0.proj` with its two fused halves swapped, because diffusers' SwiGLU reads
217
+ # `[value; gate]` where the checkpoint stores `[gate; value]`,
218
+ # * `attn.qkv_proj` -> `to_q` / `to_k` / `to_v`: split `lora_B`'s 21504 rows into **contiguous** thirds of 7168
219
+ # (`num_attention_heads * attention_head_dim`).
220
+ #
221
+ # That last split is where MiniMax-H3 LoRAs diverge. diffusers' own `_convert_non_diffusers_minimax_h3_lora_to_
222
+ # diffusers` documents two fused-QKV layouts: DiffSynth-Studio exports (identifiable by the peft `.default.` infix)
223
+ # are *per-head interleaved* and must be de-interleaved before the split, whereas ai-toolkit exports under the
224
+ # `diffusion_model.` prefix are already `[q_all; k_all; v_all]` and must **not** be. This file's keys are
225
+ # `diffusion_model.…lora_A.weight` with `ai-toolkit 0.13.21` in its metadata, so it is the second kind: contiguous
226
+ # thirds, no reorder. De-interleaving it anyway would scatter each head's q/k/v across all three projections and
227
+ # turn the adapter into structured noise on all 52 attention blocks.
228
+ #
229
+ # Row transforms only ever touch `lora_B`, so the three projections share one `lora_A`: the rows of `B @ A` are just
230
+ # the rows of `B`, which makes the split exact rather than an approximation.
231
+ #
232
+ # There is no AdaLN branch to reconcile: the run's `network_kwargs.ignore_if_contains = ["adaln_proj"]` kept the
233
+ # adapter off the timestep path entirely, which is also why the pruned partition takes it verbatim.
234
+
235
+
236
+ def _lora_target_name(source_name: str) -> str:
237
+ """Map an original-checkpoint module path to the diffusers one (no `.lora_A/B.*` suffix)."""
238
+ if source_name.startswith("token_refiner.blocks."):
239
+ return source_name.replace("token_refiner.blocks.", "token_refiner.refiner_blocks.", 1)
240
+ if source_name.startswith("blocks."):
241
+ return source_name.replace("blocks.", "transformer_blocks.", 1)
242
+ return source_name
243
+
244
+
245
+ def _lora_modules(name: str, a_weight, b_weight, inner_dim: int):
246
+ """Yield `(diffusers_module_path, lora_A, lora_B)` for one original module path."""
247
+ target = _lora_target_name(name)
248
+
249
+ if target.endswith(".attn.qkv_proj"):
250
+ prefix = target.removesuffix("qkv_proj")
251
+ for kind, part in zip(("q", "k", "v"), b_weight.split(inner_dim, dim=0)):
252
+ yield f"{prefix}to_{kind}", a_weight, part.contiguous()
253
+ elif target.endswith(".mlp.fc1"):
254
+ # SwiGLU gate/value swap: the checkpoint stores [gate, value]; diffusers stores [value, gate].
255
+ gate, value = b_weight.chunk(2, dim=0)
256
+ yield target.replace(".mlp.fc1", ".ff.net.0.proj"), a_weight, torch.cat([value, gate]).contiguous()
257
+ elif target.endswith(".mlp.fc2"):
258
+ yield target.replace(".mlp.fc2", ".ff.net.2"), a_weight, b_weight
259
+ elif target.endswith(".attn.out_proj"):
260
+ yield target.replace(".attn.out_proj", ".attn.to_out.0"), a_weight, b_weight
261
+ else:
262
+ yield target, a_weight, b_weight
263
+
264
+
265
+ def build_lora_state_dict(transformer) -> tuple[dict, int, int]:
266
+ """Download the LoRA and remap it into a PEFT-format state dict for the diffusers transformer.
267
+
268
+ Every target is validated against the transformer's own parameter shapes, and an unresolved one is fatal: a
269
+ silently dropped target means the name mapping is wrong and the Space would serve a half-applied adapter that
270
+ still *looks* like it worked.
271
+ """
272
+ from huggingface_hub import hf_hub_download
273
+ from safetensors.torch import load_file
274
+
275
+ raw = load_file(hf_hub_download(LORA_REPO, LORA_FILE))
276
+
277
+ prefix, suffix_a, suffix_b = "diffusion_model.", ".lora_A.weight", ".lora_B.weight"
278
+ unexpected = [k for k in raw if not (k.startswith(prefix) and k.endswith((suffix_a, suffix_b)))]
279
+ if unexpected:
280
+ raise ValueError(f"{LORA_FILE} holds {len(unexpected)} unexpected tensors, e.g. {unexpected[:5]}")
281
+ bases = sorted({k[len(prefix) : -len(suffix_a)] for k in raw if k.endswith(suffix_a)})
282
+ if not bases:
283
+ raise ValueError(f"No `{prefix}*{suffix_a}` / `{suffix_b}` pairs found in {LORA_FILE}")
284
+
285
+ ranks = set()
286
+ for name in bases:
287
+ if f"{prefix}{name}{suffix_b}" not in raw:
288
+ raise ValueError(f"LoRA is missing the lora_B twin of {prefix}{name}{suffix_a}")
289
+ ranks.add(raw[f"{prefix}{name}{suffix_a}"].shape[0])
290
+ if len(ranks) != 1:
291
+ raise ValueError(f"LoRA mixes ranks {sorted(ranks)}; this loader assumes a single rank")
292
+ rank = ranks.pop()
293
+
294
+ inner_dim = transformer.config.num_attention_heads * transformer.config.attention_head_dim
295
+ base_shapes = {key: tuple(value.shape) for key, value in transformer.state_dict().items()}
296
+
297
+ state_dict: dict[str, torch.Tensor] = {}
298
+ missed: list[str] = []
299
+ for name in bases:
300
+ a_weight = raw[f"{prefix}{name}{suffix_a}"]
301
+ b_weight = raw[f"{prefix}{name}{suffix_b}"]
302
+ for module, a_part, b_part in _lora_modules(name, a_weight, b_weight, inner_dim):
303
+ base = base_shapes.get(f"{module}.weight")
304
+ if base is None:
305
+ missed.append(module)
306
+ continue
307
+ # `W` is [out, in]; the adapter must be `lora_B` [out, r] @ `lora_A` [r, in].
308
+ if (b_part.shape[0], a_part.shape[1]) != base:
309
+ raise ValueError(
310
+ f"LoRA delta for `{module}` would be {(b_part.shape[0], a_part.shape[1])}, "
311
+ f"base weight is {base}"
312
+ )
313
+ state_dict[f"{module}.lora_A.weight"] = a_part
314
+ state_dict[f"{module}.lora_B.weight"] = b_part
315
+ if missed:
316
+ raise ValueError(
317
+ f"{len(missed)} LoRA targets matched no transformer weight, e.g. {missed[:5]}. "
318
+ "The LoRA and the diffusers transformer disagree on module naming."
319
+ )
320
+
321
+ return state_dict, rank, len(bases)
322
+
323
+
324
+ def load_and_apply_lora(transformer) -> str:
325
+ """Attach the LoRA as a PEFT adapter so its strength stays a per-request knob."""
326
+ state_dict, rank, targets = build_lora_state_dict(transformer)
327
+ # `prefix=None` because the keys are already the transformer's own module paths, and no `network_alphas` because
328
+ # the file carries no `.alpha` tensors — PEFT then sets alpha == rank, which is also what the run trained with
329
+ # (`linear: 16`, `linear_alpha: 16`), so the adapter's intrinsic scale is 1.0 and the slider *is* the strength.
330
+ transformer.load_lora_adapter(state_dict, prefix=None, adapter_name=ADAPTER)
331
+ transformer.set_adapters([ADAPTER], [DEFAULT_LORA_SCALE])
332
+ return (
333
+ f"LoRA attached · {targets} checkpoint targets -> {len(state_dict) // 2} diffusers modules, "
334
+ f"rank {rank}, alpha == rank · default strength {DEFAULT_LORA_SCALE:g} · {LORA_FILE}"
335
+ )
336
+
337
+
338
+ def check_prompt(prompt: str) -> None:
339
+ """The NCII guard. Every request here carries an image and a video reference — the edit-on-a-real-photo case — so
340
+ it runs on all of them, before the conditioner call and the denoise booking, and a refused prompt costs no GPU
341
+ time on either half. The classifier lives in `ncii_guard`'s spawned subprocess; in the main process it kills
342
+ every GPU worker."""
343
+ import ncii_guard
344
+
345
+ flag = ncii_guard.classify(prompt)
346
+ if flag["label"] == "ncii":
347
+ print(f"[guard] prompt refused (ncii {flag['score']:.2f}): {prompt!r}", flush=True)
348
+ raise gr.Error("This prompt was flagged by a content filter and wasn't run.")
349
+
350
+
351
+ def load_models() -> str | None:
352
+ """Load the denoising half at startup, but *not* onto the card.
353
+
354
+ `MiniMaxH3Ref2VAGeneratorBlocks` declares `transformer_ref`, `vae`, `audio_vae`, the two schedulers and
355
+ `video_processor`, so `load_components` fetches exactly those subfolders — `text_encoder/` and the `transformer/`
356
+ partition are never touched. Both autoencoders carry `_keep_in_fp32_modules` over every module and stay float32:
357
+ a bfloat16 audio VAE decodes the soundtrack roughly 20 dB too quiet.
358
+ """
359
+ global PIPE, MANAGER, LOAD_ERROR, LORA_STATUS
360
+
361
+ if PIPE is not None or LOAD_ERROR is not None:
362
+ return LOAD_ERROR
363
+
364
+ started = time.time()
365
+ try:
366
+ from diffusers import ComponentsManager
367
+
368
+ from h3_split_blocks import MiniMaxH3Ref2VAGeneratorBlocks
369
+
370
+ lower_duration_floor()
371
+ manager = ComponentsManager()
372
+ blocks = MiniMaxH3Ref2VAGeneratorBlocks()
373
+ print(f"[{VERSION}] loading {[c.name for c in blocks.expected_components]} from {MODEL_REPO} ...", flush=True)
374
+ pipe = blocks.init_pipeline(MODEL_REPO, components_manager=manager, collection="h3")
375
+ # The pruned DiT is served as remote code (`transformer_ref/modeling_minimax_h3_pruned.py`, reached through
376
+ # the `AutoModel` type hint in `modular_model_index.json`). `load_components` forwards `trust_remote_code`
377
+ # only to components that live in the pipeline's own repo, so the two VAEs and the schedulers — which
378
+ # `modular_model_index.json` still points at `MiniMaxAI/MiniMax-H3` — never see it.
379
+ pipe.load_components(dtype=torch.bfloat16, trust_remote_code=True)
380
+
381
+ # Both VAEs first, and explicitly. `set_attention_backend` also sets the registry's *global* backend, which
382
+ # every processor that was not stamped falls through to, and the float32 audio VAE has no cuDNN kernel:
383
+ # `RuntimeError: No available kernel. Aborting execution.` in its causal encoder attention.
384
+ pipe.vae.set_attention_backend("native")
385
+ pipe.audio_vae.set_attention_backend("native")
386
+ pipe.transformer_ref.set_attention_backend(ATTENTION)
387
+
388
+ # Before any GPU placement, so the adapter's parameters travel with the transformer. No AoTI on this Space
389
+ # for the same reason: a compiled block package binds the *base* module's weights by fully qualified name
390
+ # and would run straight past the PEFT branch.
391
+ LORA_STATUS = load_and_apply_lora(pipe.transformer_ref)
392
+ print(f"[{VERSION}] {LORA_STATUS}", flush=True)
393
+
394
+ # Installed per request rather than here — a cached residual only means something within the schedule it was
395
+ # measured on, so the state cannot outlive one generation.
396
+ import h3_fbc
397
+
398
+ print(f"[h3-fbc] {h3_fbc.status()}", flush=True)
399
+
400
+ if PLACEMENT == "offload":
401
+ manager.enable_auto_cpu_offload(device="cuda")
402
+ _arm_decode_hooks(pipe)
403
+
404
+ PIPE, MANAGER = pipe, manager
405
+ print(f"[{VERSION}] ready in {time.time() - started:.0f}s", flush=True)
406
+ except Exception as error:
407
+ traceback.print_exc()
408
+ LOAD_ERROR = (
409
+ f"**Loading `{MODEL_REPO}` + `{LORA_REPO}` failed** after {time.time() - started:.0f}s: "
410
+ f"`{type(error).__name__}: {error}`"
411
+ )
412
+ return LOAD_ERROR
413
+
414
+
415
+ def _arm_decode_hooks(pipe):
416
+ """Make the offload hooks fire for the two VAEs.
417
+
418
+ `enable_auto_cpu_offload` wraps `forward`, and the reference-encoder and decode blocks call `vae.encode/decode(...)`
419
+ directly, so the hook never runs and the VAE is still on the host when the latents arrive on the card.
420
+ """
421
+ for name in ("vae", "audio_vae"):
422
+ module = getattr(pipe, name)
423
+ for method in ("encode", "decode"):
424
+ inner = getattr(module, method)
425
+
426
+ def armed(*args, _module=module, _inner=inner, **kwargs):
427
+ hook = getattr(_module, "_hf_hook", None)
428
+ if hook is not None:
429
+ hook.pre_forward(_module)
430
+ return _inner(*args, **kwargs)
431
+
432
+ setattr(module, method, armed)
433
+
434
+
435
+ @cache
436
+ def conditioner():
437
+ """The other half, over the gradio API. `gradio_client` attaches the caller's own ZeroGPU token per call, so the
438
+ conditioner's booking is billed to whoever asked for the video."""
439
+ from gradio_client import Client
440
+
441
+ return Client(CONDITIONER_SPACE)
442
+
443
+
444
+ def probe(path: str) -> tuple[float | None, float | None]:
445
+ """`(video seconds, audio seconds)` of a media file, either being `None` when the stream is absent."""
446
+ import av
447
+
448
+ def seconds(stream, container):
449
+ if stream.duration is not None and stream.time_base is not None:
450
+ return float(stream.duration * stream.time_base)
451
+ return None if container.duration is None else container.duration / av.time_base
452
+
453
+ with av.open(path) as container:
454
+ video = seconds(container.streams.video[0], container) if container.streams.video else None
455
+ audio = seconds(container.streams.audio[0], container) if container.streams.audio else None
456
+ return video, audio
457
+
458
+
459
+ def collect(character_path: str, scene_path: str) -> list[tuple[str, str]]:
460
+ """The `(kind, path)` references of a request, **in the order the model reads them**.
461
+
462
+ The character sheet first, the scene clip second. That order is not cosmetic and it is not the order the prompt
463
+ names them in: it numbers the labels of MiniMax-H3's prompt presentation and it advances the shared audio/video
464
+ rotary clock, so the same two references swapped round are a different request. It is the order the LoRA was
465
+ trained with — `control_path: [character_references, scene_videos]` in `configs/trained-run-1000.json` of
466
+ `akatz-ai/H3-Character-Swap-v1` — which is why the captions read `<Video 1>` before `<Picture 1>` while the
467
+ references go the other way: the labels are numbered per modality, not by position.
468
+ """
469
+ return [("image", character_path), ("video", scene_path)]
470
+
471
+
472
+ def build_references(references: list[tuple[str, str]]):
473
+ """The `(kind, path)` references of a request as decoded reference dataclasses, in packed order. `from_file`
474
+ brings the rates along: a video its own frame rate and soundtrack."""
475
+ from diffusers.modular_pipelines.minimax_h3 import (
476
+ MiniMaxH3ImageReference,
477
+ MiniMaxH3VideoReference,
478
+ )
479
+
480
+ classes = {"image": MiniMaxH3ImageReference, "video": MiniMaxH3VideoReference}
481
+ return [classes[kind].from_file(path) for kind, path in references]
482
+
483
+
484
+ def pick_canvas(scene_path: str | None) -> str:
485
+ """The generated canvas whose aspect ratio is closest to the scene clip's, at the LoRA's own 768 short edge.
486
+
487
+ A character swap is asked to keep the source framing, so a portrait clip generated on a landscape canvas is a
488
+ recomposed shot before the model has done anything. Wired to the clip's `change`, so the dropdown shows what a
489
+ request will actually use and stays overridable.
490
+ """
491
+ if not scene_path:
492
+ return gr.update()
493
+ try:
494
+ import av
495
+
496
+ with av.open(scene_path) as container:
497
+ stream = container.streams.video[0]
498
+ ratio = stream.width / stream.height
499
+ except Exception:
500
+ return gr.update()
501
+ best = min(AUTO_CANVASES, key=lambda label: abs(CANVASES[label][1] / CANVASES[label][0] - ratio))
502
+ return gr.update(value=best)
503
+
504
+
505
+ def check(prompt: str, character_path: str | None, scene_path: str | None) -> None:
506
+ """The request's own rules, before anything is uploaded or a card is allocated."""
507
+ if not prompt or not prompt.strip():
508
+ raise gr.Error("Name who to replace — MiniMax-H3 always takes a prompt.")
509
+ if not scene_path:
510
+ raise gr.Error("Add the scene clip to edit; it is the request's `<Video 1>`.")
511
+ if not character_path:
512
+ raise gr.Error("Add the character reference to swap in; it is the request's `<Picture 1>`.")
513
+ video_seconds, _ = probe(scene_path)
514
+ if video_seconds is None:
515
+ raise gr.Error("That scene clip has no video stream.")
516
+ if not MIN_REFERENCE_VIDEO <= video_seconds <= MAX_REFERENCE_VIDEO:
517
+ raise gr.Error(
518
+ f"The scene clip is {video_seconds:.2f} s. Use one between {MIN_REFERENCE_VIDEO:.2f} and "
519
+ f"{MAX_REFERENCE_VIDEO:g} seconds."
520
+ )
521
+
522
+
523
+ def encode_remote(prompt, references, canvas, num_frames):
524
+ """`/encode_ref2va` on the conditioner Space: a safetensors file holding `prompt_embeds` + `text_token_tags`,
525
+ with the resolved `height` / `width` / `num_frames` in its metadata, plus the plan.
526
+
527
+ `canvas` is the label. `media` and `kinds` are parallel and ordered, and the references go over because `ref2va`'s
528
+ presentation puts a vision block in front of the prompt for every image and every merged video frame pair.
529
+ """
530
+ from gradio_client import handle_file
531
+ from safetensors import safe_open
532
+
533
+ path, plan = conditioner().predict(
534
+ prompt=prompt,
535
+ media=[handle_file(path) for _, path in references],
536
+ kinds=",".join(kind for kind, _ in references),
537
+ canvas=canvas,
538
+ num_frames=num_frames,
539
+ rewrite_prompt=False,
540
+ api_name="/encode_ref2va",
541
+ )
542
+ with safe_open(path, framework="pt") as handle:
543
+ return handle.get_tensor("prompt_embeds"), handle.get_tensor("text_token_tags"), handle.metadata(), plan
544
+
545
+
546
+ @spaces.GPU(duration=get_duration, size=GPU_SIZE)
547
+ def _generate(prompt_embeds, text_token_tags, references, height, width, num_frames, steps, seed, lora_scale):
548
+ """The only thing on GPU time: the two reference encoders, the packed-sequence denoise loop and the decoders.
549
+
550
+ References cross as paths and are decoded here; only the three generated outputs come back. A `@spaces.GPU`
551
+ argument crosses a process boundary by pickling, a 5 s 1344x768 reference video is 370 MB of expanded frames, and
552
+ the full `PipelineState` still holds the packed latents and the rotary grid on the card.
553
+ """
554
+ import h3_fbc
555
+
556
+ if PLACEMENT == "lazy":
557
+ PIPE.to("cuda")
558
+
559
+ # Per request, inside the worker: the adapter's scale is a python attribute on its PEFT layers, so setting it on
560
+ # this fork's copy leaves every concurrent request with its own strength.
561
+ PIPE.transformer_ref.set_adapters([ADAPTER], [float(lora_scale)])
562
+
563
+ with h3_fbc.enabled(PIPE.transformer_ref, steps=int(steps)):
564
+ state = PIPE(
565
+ prompt_embeds=prompt_embeds.to("cuda"),
566
+ text_token_tags=text_token_tags,
567
+ references=build_references(references),
568
+ height=height,
569
+ width=width,
570
+ num_frames=num_frames,
571
+ num_inference_steps=int(steps),
572
+ generator=torch.Generator("cpu").manual_seed(int(seed)),
573
+ )
574
+ return state.get("videos")[0], state.get("audio")[0].cpu(), state.get("sampling_rate")
575
+
576
+
577
+ def generate(
578
+ # The first three are the columns `gr.Examples` varies, and they lead the signature for that reason: an example
579
+ # row is applied to `inputs` positionally. Every other parameter carries the same default as its UI component,
580
+ # so clicking an example and pressing Generate behave identically.
581
+ prompt: str,
582
+ scene_video: str | None = None,
583
+ character_image: str | None = None,
584
+ canvas: str = DEFAULT_CANVAS,
585
+ match_length: bool = True,
586
+ duration: float = DEFAULT_DURATION,
587
+ steps: int = DEFAULT_STEPS,
588
+ seed: int = 904231,
589
+ lora_scale: float = DEFAULT_LORA_SCALE,
590
+ progress=gr.Progress(track_tqdm=True),
591
+ ):
592
+ """Replace one character in a clip with the character in a reference image, keeping the rest of the shot.
593
+
594
+ Args:
595
+ prompt: who to replace and with what, naming the references as `<Video 1>` and `<Picture 1>`, e.g.
596
+ "Swap the man in the purple shirt in <Video 1> with the character in <Picture 1>."
597
+ scene_video: path to the clip to edit — the request's `<Video 1>`, 5 frames to 15 seconds.
598
+ character_image: path to the replacement character's portrait or character sheet — `<Picture 1>`.
599
+ canvas: the generated canvas, as one of this Space's labels.
600
+ match_length: generate for as long as the scene clip runs, when that length is one MiniMax-H3 generates.
601
+ duration: seconds to generate when `match_length` is off, or the clip is too short to set it.
602
+ steps: denoising steps. The LoRA's own training run sampled at 28.
603
+ seed: RNG seed.
604
+ lora_scale: character-swap adapter strength. 1.0 is the card's recommendation; 0 is the base model.
605
+
606
+ Returns:
607
+ Path to the generated MP4, video and its jointly generated soundtrack.
608
+ """
609
+ if LOAD_ERROR:
610
+ raise gr.Error(LOAD_ERROR)
611
+ if PIPE is None:
612
+ raise gr.Error("The denoiser is still loading.")
613
+
614
+ from diffusers.utils import encode_video
615
+
616
+ check(prompt, character_image, scene_video)
617
+ check_prompt(prompt)
618
+ references = collect(character_image, scene_video)
619
+
620
+ scene_seconds, _ = probe(scene_video)
621
+ requested_seconds = float(duration)
622
+ if match_length and scene_seconds is not None and MIN_DURATION <= scene_seconds <= MAX_UI_DURATION:
623
+ requested_seconds = scene_seconds
624
+ num_frames = snap_frames(requested_seconds)
625
+
626
+ progress(0.0, desc="Reading the prompt and the two references ...")
627
+ conditioned = time.time()
628
+ try:
629
+ prompt_embeds, text_token_tags, metadata, plan = encode_remote(prompt, references, canvas, num_frames)
630
+ except gr.Error:
631
+ raise
632
+ except Exception as error:
633
+ # gradio only puts the exception *type* on the wire, so the useful half of a conditioner-side failure is in
634
+ # that Space's logs.
635
+ traceback.print_exc()
636
+ raise gr.Error(
637
+ f"The conditioner ({CONDITIONER_SPACE}) failed with `{type(error).__name__}: {error}`. "
638
+ "Its logs carry the full traceback."
639
+ ) from error
640
+ condition_seconds = time.time() - conditioned
641
+ height, width, num_frames = (int(metadata[key]) for key in ("height", "width", "num_frames"))
642
+
643
+ progress(0.1, desc=f"Swapping over {num_frames / FPS:.1f} s at {width}x{height} ...")
644
+ started = time.time()
645
+ frames, audio, sampling_rate = _generate(
646
+ prompt_embeds, text_token_tags, references, height, width, num_frames, steps, seed, lora_scale
647
+ )
648
+ generate_seconds = time.time() - started
649
+
650
+ directory = os.path.join(tempfile.gettempdir(), "h3-outputs")
651
+ os.makedirs(directory, exist_ok=True)
652
+ path = os.path.join(directory, f"h3-character-swap-{int(time.time() * 1000)}.mp4")
653
+ encode_video(frames, fps=FPS, output_path=path, audio=audio, audio_sample_rate=sampling_rate)
654
+
655
+ print(
656
+ f"[{VERSION}] `{width}x{height}`, {num_frames} frames ({num_frames / FPS:.3f} s), {int(steps)} steps, "
657
+ f"strength {float(lora_scale):g} · conditioner {condition_seconds:.0f}s "
658
+ f"({plan['num_text_tokens']} tokens) · denoise + decode {generate_seconds:.0f}s "
659
+ f"({generate_seconds / max(1, int(steps)):.1f} s/step) · seed {int(seed)}",
660
+ flush=True,
661
+ )
662
+ return path
663
+
664
+
665
+ import ncii_guard
666
+
667
+ ncii_guard.start()
668
+ load_models()
669
+
670
+ INTRO = """# MiniMax-H3 Character Swap LoRA
671
+
672
+ <div align="center">
673
+ <a href="https://huggingface.co/akatz-ai/MiniMax-H3-Character-Swap-LoRA" target="_blank" rel="noopener"><strong>[ LoRA ]</strong></a> &nbsp;
674
+ <a href="https://huggingface.co/datasets/akatz-ai/H3-Character-Swap-v1" target="_blank" rel="noopener"><strong>[ dataset ]</strong></a> &nbsp;
675
+ <a href="https://huggingface.co/MiniMaxAI/MiniMax-H3" target="_blank" rel="noopener"><strong>[ base model ]</strong></a>
676
+ </div>
677
+
678
+ Give it a **scene clip** and a **character reference**, then name who to replace. Akatz Labs' experimental
679
+ character-swap adapter for MiniMax-H3 `ref2va` puts the reference character into the shot — identity, outfit and
680
+ art style carried over — while the background, the camera and everyone else stay where they were. Video and its
681
+ soundtrack come out of one denoising pass.
682
+
683
+ Refer to the two references the way the LoRA was trained to read them: the clip is **`<Video 1>`** and the
684
+ character is **`<Picture 1>`**.
685
+ """
686
+
687
+ NOTES = """It is an experimental 1,000-update adapter and its own card is candid about the limits: motion timing,
688
+ facial expressions and hard cuts are unreliable, long windows drift in framing, and close-up expressions may not
689
+ match the source performance. Short continuous shots of about 3–5 seconds are the promising range.
690
+
691
+ Its training targets were **single still frames** with five-frame static scene clips, which is what the examples
692
+ below are — a still scene, handed over as the short clip `<Video 1>` wants. Real moving footage works the same way
693
+ and is what the 40 preservation clips in the dataset regularized, but it is the harder case.
694
+
695
+ Strength 1.0 is the author's recommendation; 0 is the base `ref2va` model, which is the comparison the adapter was
696
+ judged against.
697
+ """
698
+
699
+ CSS = """
700
+ .main.fillable { max-width: 1250px !important; }
701
+ .dark .gradio-container { color: var(--body-text-color); }
702
+ """
703
+
704
+ with gr.Blocks(title="MiniMax-H3 Character Swap LoRA", theme=gr.themes.Citrus(), css=CSS) as demo:
705
+ gr.Markdown(INTRO)
706
+
707
+ with gr.Row():
708
+ with gr.Column():
709
+ prompt = gr.Textbox(
710
+ label="Prompt — name the person to replace",
711
+ lines=3,
712
+ value="Swap the man in the purple shirt in <Video 1> with the character in <Picture 1>.",
713
+ placeholder="Swap the woman on the left in <Video 1> with the character in <Picture 1>.",
714
+ )
715
+ with gr.Row():
716
+ scene = gr.Video(label="Scene clip · <Video 1>", height=260)
717
+ character = gr.Image(label="Character reference · <Picture 1>", type="filepath", height=260)
718
+ run = gr.Button("Swap the character", variant="primary", size="lg")
719
+ with gr.Accordion("Advanced options", open=False):
720
+ lora_scale = gr.Slider(
721
+ label="Character-swap LoRA strength",
722
+ minimum=0.0,
723
+ maximum=1.5,
724
+ step=0.05,
725
+ value=DEFAULT_LORA_SCALE,
726
+ )
727
+ canvas = gr.Dropdown(
728
+ label="Canvas (follows the scene clip's aspect ratio)",
729
+ choices=list(CANVASES),
730
+ value=DEFAULT_CANVAS,
731
+ )
732
+ match_length = gr.Checkbox(label="Match the scene clip's length", value=True)
733
+ duration = gr.Slider(
734
+ label="Duration (s) — used when the clip is shorter than the model generates",
735
+ minimum=MIN_DURATION,
736
+ maximum=MAX_UI_DURATION,
737
+ step=1,
738
+ value=DEFAULT_DURATION,
739
+ )
740
+ steps = gr.Slider(label="Steps", minimum=10, maximum=40, step=1, value=DEFAULT_STEPS)
741
+ seed = gr.Number(label="Seed", value=904231, precision=0)
742
+
743
+ with gr.Column():
744
+ result = gr.Video(label="Swapped clip + soundtrack")
745
+ gr.Markdown(NOTES)
746
+
747
+ scene.change(pick_canvas, scene, canvas, show_progress="hidden", api_name=False)
748
+
749
+ request = [prompt, scene, character, canvas, match_length, duration, steps, seed, lora_scale]
750
+
751
+ gr.Examples(
752
+ # The dataset's own edits, with the instructions its captions carry: a photoreal swap, a cross-style swap
753
+ # onto an illustrated character, and a multi-view character sheet.
754
+ examples=[
755
+ [
756
+ "Swap the man in the purple shirt in <Video 1> with the character in <Picture 1>.",
757
+ "examples/cafe_scene.mp4",
758
+ "examples/character_hiker.png",
759
+ ],
760
+ [
761
+ "Swap the man in the purple shirt in <Video 1> with the character in <Picture 1>. Keep the "
762
+ "replacement character's identity, outfit and art style from <Picture 1>. Preserve the source "
763
+ "video's camera, background, lighting, objects and all other people.",
764
+ "examples/cafe_scene.mp4",
765
+ "examples/character_anime.png",
766
+ ],
767
+ [
768
+ "Swap the person in <Video 1> with the character in <Picture 1>. Do not show the reference sheet "
769
+ "or its background.",
770
+ "examples/workshop_scene.mp4",
771
+ "examples/character_sheet_orin.png",
772
+ ],
773
+ ],
774
+ inputs=[prompt, scene, character],
775
+ outputs=result,
776
+ fn=generate,
777
+ cache_examples=True,
778
+ cache_mode="lazy",
779
+ )
780
+
781
+ run.click(generate, request, result, api_name="generate")
782
+
783
+
784
+ if __name__ == "__main__":
785
+ demo.launch(show_error=True, max_threads=1000)
examples/cafe_scene.mp4 ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:3d1dbc1b503df47708aa2b1a49d2987078e1edb2954d0f28ee30650823ce0488
3
+ size 923206
examples/character_anime.png ADDED

Git LFS Details

  • SHA256: 60c59df3023b638d049bf0c07b7d1186c7ac92730641c21d90fddb1c096ebda5
  • Pointer size: 131 Bytes
  • Size of remote file: 834 kB
examples/character_hiker.png ADDED

Git LFS Details

  • SHA256: 40f6e295b6859959709e4e9edb6d06953c2bfb53e9935a8fa472dbbd248c4e2a
  • Pointer size: 131 Bytes
  • Size of remote file: 862 kB
examples/character_sheet_orin.png ADDED

Git LFS Details

  • SHA256: 9c3d578ada04bd833893d2919f715e7419aa446a6c87b2d95d8eb80d7a711779
  • Pointer size: 132 Bytes
  • Size of remote file: 1.63 MB
examples/workshop_scene.mp4 ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:fa1b341d2bb03a786b95457a1705d6fda46e7ebaa0baf205bf8a36c70662ee4a
3
+ size 766362
h3_fbc.py ADDED
@@ -0,0 +1,357 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """First-block cache for MiniMax-H3: skip the trunk on steps whose block-0 residual barely moved.
2
+
3
+ Block 0 and the final AdaLN head run at the true timestep on **every** step; only blocks 1..49 are skipped, and only
4
+ while the residual they would have been handed looks like the one from the last step that actually ran. The decision
5
+ signal is the relative L1 between this step's block-0 residual (`block0_out - block0_in`) and the residual of the last
6
+ *computed* step, over the whole packed sequence. On a skip the trunk's contribution is replayed as a cached residual,
7
+ `(final_trunk_out - block0_out)` of that computed step.
8
+
9
+ Ported from `duckyshell/ComfyUI-MiniMaxH3-FirstBlockCache` (`nodes.py` @ 725973c) — same signal, same protected window
10
+ (10%-95% of the schedule, converted through the video shift of 12.0) and the same cap of two consecutive skips, which
11
+ is their "H3 Safe" preset at threshold 0.08.
12
+
13
+ **The audio exemption is not theirs.** duckyshell has no audio term anywhere: the decision is ~98% video by row count
14
+ and every audio row rides the stale trunk residual, which is what costs the soundtrack its energy. Audio runs its own
15
+ schedule (shift 3 against video's 12, diverging by up to 11x in rate across 20 steps), so on a skip the audio rows of
16
+ the trunk output are instead a 2-point **linear** extrapolation of the last two actually-computed audio features, in
17
+ the audio sigma coordinate — `xmarre/ComfyUI-Spectrum-MiniMax-H3`'s `audio_blend_weight=0.0` path, applied where
18
+ Spectrum applies it (the post-trunk hidden feature, ahead of the head that still runs at the true timestep) and fixed
19
+ to the right coordinate. Following ComfyUI PR #15390, no carried audio tensor is ever mutated: every write is a fresh
20
+ tensor out of `index_copy`.
21
+
22
+ It earns its place. Running this Space's own request at threshold 0.08 with `H3_FBC_AUDIO_EXEMPT=0` — the duckyshell
23
+ mechanism verbatim, same 13 skipped forwards, same video to within 0.2 dB — costs the soundtrack **16.7% of its RMS**
24
+ (0.0483 against an uncached 0.0580), while the exemption holds it at 1.04x. That is the same direction the offline
25
+ study measured on a near-silent clip, three times the size on one with real audio energy.
26
+
27
+ **A threshold does not travel across step counts.** The signature shrinks as the schedule is subdivided, so the same
28
+ number gates far more loosely at more steps. Measured on this Space — 960x544, 124 frames, one image reference,
29
+ seed 42, AoTI blocks, at its **default 28 steps** (27 forwards) — against the same request with `H3_FBC=0`:
30
+
31
+ | threshold | skipped | denoise loop | end to end | audio RMS vs uncached |
32
+ |---|---|---|---|---|
33
+ | 0.03 | 0 / 27 | 1.02x | 1.06x | 1.000 (bitwise-identical audio) |
34
+ | 0.05 | 9 / 27 | 1.46x | 1.36x | 0.956 |
35
+ | 0.08 | 13 / 27 | 2.19x | 1.82x | 1.040 |
36
+
37
+ The offline study calibrated 0.08 over a 20-step schedule, where it skipped 7 of 19 forwards. At 28 steps that same
38
+ 0.08 skips 13 of 27 — nearly half — and the sampled video visibly re-rolls its background detail. **0.05 is the
39
+ default** because it reproduces the skip fraction that study validated (33% here against 37% there); 0.08 is the
40
+ aggressive setting. Below the signature floor — 0.03 skipped nothing at all here — a threshold buys nothing and still
41
+ pays for the signature, so lower is not safer, it is just slower.
42
+
43
+ A cached request is not the uncached one: the trajectory moves, so the video is a different sample of the same prompt
44
+ (same shot, same subject, same quality — different signage and background detail). `H3_FBC=0` restores today's output
45
+ exactly, and is worth reaching for when a request has to reproduce a specific earlier result.
46
+
47
+ This composes with `h3_aoti`: that module patches each of the 50 blocks' own `forward` and they stay a real
48
+ `ModuleList`, so skipping the trunk simply does not call blocks 1..49 that step. `LazyAOTIModel` rebinds its constants
49
+ whenever the weights dict it is handed changes identity, and the first forward of a request never skips, so every
50
+ block has bound its own weights before any step is cached.
51
+ """
52
+
53
+ from __future__ import annotations
54
+
55
+ import contextlib
56
+ import os
57
+ import types
58
+
59
+ ENABLED = os.environ.get("H3_FBC", "1") == "1"
60
+ # Relative-L1 gate on the block-0 residual, calibrated at this Space's default 28 steps — see the table above.
61
+ THRESHOLD = float(os.environ.get("H3_FBC_THRESHOLD", "0.05"))
62
+ # duckyshell's cap. Without it the gate compares against an ever-older computed step and drifts away unbounded.
63
+ MAX_CONSECUTIVE_HITS = int(os.environ.get("H3_FBC_MAX_CONSECUTIVE", "2"))
64
+ # The protected head and tail of the schedule, as fractions, converted to sigma through the video shift below.
65
+ START_PERCENT = float(os.environ.get("H3_FBC_START_PERCENT", "0.10"))
66
+ END_PERCENT = float(os.environ.get("H3_FBC_END_PERCENT", "0.95"))
67
+ # `MiniMaxH3SetTimestepsStep` builds the video schedule at shift 12.0 and the audio one at 3.0.
68
+ VIDEO_SHIFT = float(os.environ.get("H3_FBC_VIDEO_SHIFT", "12.0"))
69
+ AUDIO_EXEMPT = os.environ.get("H3_FBC_AUDIO_EXEMPT", "1") == "1"
70
+
71
+ # Every keyword `MiniMaxH3LoopDenoiser` passes. It filters the packed-sequence layout through
72
+ # `inspect.signature(transformer.forward).parameters`, so a replacement forward that drops a name silently stops
73
+ # receiving it; `install` refuses rather than let that happen quietly.
74
+ FORWARD_PARAMETERS = (
75
+ "hidden_states",
76
+ "audio_hidden_states",
77
+ "encoder_hidden_states",
78
+ "timestep",
79
+ "timestep_indices",
80
+ "token_tags",
81
+ "position_ids",
82
+ "video_indices",
83
+ "audio_indices",
84
+ "text_indices",
85
+ "attention_kwargs",
86
+ "return_dict",
87
+ )
88
+
89
+
90
+ def status() -> str:
91
+ return (
92
+ f"first-block cache **on** · threshold `{THRESHOLD}` · audio exemption "
93
+ f"{'on' if AUDIO_EXEMPT else 'off'}"
94
+ if ENABLED
95
+ else "first-block cache **off** (`H3_FBC=1` to skip the trunk on steady steps)"
96
+ )
97
+
98
+
99
+ def _shifted_sigma(u: float, shift: float) -> float:
100
+ return shift * u / (1.0 + (shift - 1.0) * u)
101
+
102
+
103
+ def _rel_l1(current, previous) -> float:
104
+ numerator = (current.float() - previous.float()).abs().mean()
105
+ denominator = previous.float().abs().mean().clamp(min=1e-8)
106
+ return float((numerator / denominator).item())
107
+
108
+
109
+ class _State:
110
+ def __init__(self, threshold: float, steps: int, audio_exempt: bool):
111
+ self.threshold = threshold
112
+ self.steps = steps
113
+ self.audio_exempt = audio_exempt
114
+ # duckyshell reads the window as sigma bounds: a flow model's sigma at `u = 1 - percent`, shifted.
115
+ self.start_sigma = _shifted_sigma(1.0 - START_PERCENT, VIDEO_SHIFT)
116
+ self.end_sigma = _shifted_sigma(1.0 - END_PERCENT, VIDEO_SHIFT)
117
+ self.original = None
118
+ self.failed = False
119
+ self.consecutive_hits = 0
120
+ self.prev_first_residual = None
121
+ self.tail_residual = None
122
+ self.audio_history = [] # [(sigma_audio, audio rows of the trunk output)], newest last, at most two
123
+ self.computed = 0
124
+ self.skipped = 0
125
+
126
+
127
+ def _cached_forward(
128
+ self,
129
+ state,
130
+ hidden_states,
131
+ audio_hidden_states,
132
+ encoder_hidden_states,
133
+ timestep,
134
+ timestep_indices,
135
+ token_tags,
136
+ position_ids,
137
+ video_indices,
138
+ audio_indices,
139
+ text_indices,
140
+ return_dict,
141
+ ):
142
+ """`MiniMaxH3Transformer3DModel.forward` with the block loop split at block 0.
143
+
144
+ Everything outside the loop is that method verbatim, at the `diffusers` commit `requirements.txt` pins; keep the
145
+ two in step when the pin moves.
146
+ """
147
+ import torch
148
+
149
+ from diffusers.models.transformers.transformer_minimax_h3 import (
150
+ MINIMAX_H3_MODALITY_NUM,
151
+ MiniMaxH3TransformerOutput,
152
+ )
153
+
154
+ sequence_length = position_ids.shape[0]
155
+ rotary_emb = self.rope(position_ids)
156
+
157
+ video_embeds = self.proj_in(hidden_states.to(self.proj_in.weight.dtype))
158
+ audio_embeds = self.audio_proj_in(audio_hidden_states.to(self.audio_proj_in.weight.dtype))
159
+ text_embeds = self.context_embedder(encoder_hidden_states.to(self.context_embedder.weight.dtype))
160
+ text_embeds = self.token_refiner(text_embeds)
161
+
162
+ packed = text_embeds.new_zeros((text_embeds.shape[0], sequence_length, text_embeds.shape[-1]))
163
+ packed = packed.index_copy(1, text_indices, text_embeds)
164
+ packed = packed.index_copy(1, video_indices, video_embeds.to(text_embeds.dtype))
165
+ packed = packed.index_copy(1, audio_indices, audio_embeds.to(text_embeds.dtype))
166
+
167
+ temb = self.time_proj(timestep)
168
+ temb = self.time_embedder(temb.to(self.time_embedder.linear_1.weight.dtype))
169
+ adaln_indices = timestep_indices * MINIMAX_H3_MODALITY_NUM + token_tags.clamp(min=0)
170
+
171
+ attention_mask = None
172
+ is_pad = token_tags < 0
173
+ if bool(is_pad.any()):
174
+ attention_mask = is_pad[None, :] == is_pad[:, None]
175
+
176
+ blocks = self.transformer_blocks
177
+ block0_out = blocks[0](packed, temb, adaln_indices, rotary_emb, attention_mask)
178
+ first_residual = block0_out - packed
179
+
180
+ # The generated rows trail their modality's index list, so the last row of each carries that stream's live noise
181
+ # level. The scheduler exposes `timesteps = 1 - sigmas[:-1]`.
182
+ sigma_video = 1.0 - float(timestep[timestep_indices[video_indices[-1]]].item())
183
+ sigma_audio = 1.0 - float(timestep[timestep_indices[audio_indices[-1]]].item())
184
+
185
+ use_cache = False
186
+ if (
187
+ state.prev_first_residual is not None
188
+ and state.tail_residual is not None
189
+ and state.prev_first_residual.shape == first_residual.shape
190
+ and state.consecutive_hits < MAX_CONSECUTIVE_HITS
191
+ and state.end_sigma <= sigma_video <= state.start_sigma
192
+ ):
193
+ use_cache = _rel_l1(first_residual, state.prev_first_residual) <= state.threshold
194
+
195
+ if use_cache:
196
+ state.consecutive_hits += 1
197
+ state.skipped += 1
198
+ trunk_out = block0_out + state.tail_residual
199
+ if state.audio_exempt and state.audio_history:
200
+ sigma_prev, feature_prev = state.audio_history[-1]
201
+ if len(state.audio_history) == 2 and abs(sigma_prev - state.audio_history[-2][0]) > 1e-8:
202
+ sigma_prev2, feature_prev2 = state.audio_history[-2]
203
+ ratio = (sigma_audio - sigma_prev) / (sigma_prev - sigma_prev2)
204
+ audio_feature = feature_prev + (feature_prev - feature_prev2) * ratio
205
+ else:
206
+ audio_feature = feature_prev
207
+ trunk_out = trunk_out.index_copy(1, audio_indices, audio_feature.to(trunk_out.dtype))
208
+ else:
209
+ state.consecutive_hits = 0
210
+ state.computed += 1
211
+ trunk_out = block0_out
212
+ for block in blocks[1:]:
213
+ trunk_out = block(trunk_out, temb, adaln_indices, rotary_emb, attention_mask)
214
+ state.tail_residual = (trunk_out - block0_out).detach()
215
+ state.prev_first_residual = first_residual.detach()
216
+ if state.audio_exempt:
217
+ state.audio_history.append((sigma_audio, trunk_out.index_select(1, audio_indices).detach().float()))
218
+ state.audio_history = state.audio_history[-2:]
219
+
220
+ out = self.norm_out(trunk_out, temb, timestep_indices).to(self.proj_out.weight.dtype)
221
+ video_output = self.proj_out(out).index_select(1, video_indices)
222
+ audio_output = self.audio_proj_out(out).index_select(1, audio_indices)
223
+
224
+ if not return_dict:
225
+ return (video_output, audio_output)
226
+ return MiniMaxH3TransformerOutput(sample=video_output, audio_sample=audio_output)
227
+
228
+
229
+ def _forward(
230
+ self,
231
+ hidden_states,
232
+ audio_hidden_states,
233
+ encoder_hidden_states,
234
+ timestep,
235
+ timestep_indices,
236
+ token_tags,
237
+ position_ids,
238
+ video_indices,
239
+ audio_indices,
240
+ text_indices,
241
+ attention_kwargs=None,
242
+ return_dict: bool = True,
243
+ ):
244
+ """The installed forward. Anything it cannot serve — a LoRA scale, an unexpected layout, a bug — is handed to the
245
+ original forward instead, for this call and every later one, so a cached request can degrade to an uncached one but
246
+ never to a failed one."""
247
+ state = self._h3_fbc
248
+ original = dict(
249
+ hidden_states=hidden_states,
250
+ audio_hidden_states=audio_hidden_states,
251
+ encoder_hidden_states=encoder_hidden_states,
252
+ timestep=timestep,
253
+ timestep_indices=timestep_indices,
254
+ token_tags=token_tags,
255
+ position_ids=position_ids,
256
+ video_indices=video_indices,
257
+ audio_indices=audio_indices,
258
+ text_indices=text_indices,
259
+ attention_kwargs=attention_kwargs,
260
+ return_dict=return_dict,
261
+ )
262
+ # `apply_lora_scale` decorates the real forward and this one is not it, so a request that actually scales a LoRA
263
+ # goes down the original path rather than silently losing its scale.
264
+ if state.failed or (attention_kwargs or {}).get("scale") is not None:
265
+ return state.original(**original)
266
+ try:
267
+ return _cached_forward(
268
+ self,
269
+ state,
270
+ hidden_states,
271
+ audio_hidden_states,
272
+ encoder_hidden_states,
273
+ timestep,
274
+ timestep_indices,
275
+ token_tags,
276
+ position_ids,
277
+ video_indices,
278
+ audio_indices,
279
+ text_indices,
280
+ return_dict,
281
+ )
282
+ except Exception as error:
283
+ state.failed = True
284
+ print(f"[h3-fbc] disabled for this request ({type(error).__name__}: {error}); running uncached", flush=True)
285
+ return state.original(**original)
286
+
287
+
288
+ def install(transformer, steps: int = 0, threshold: float = THRESHOLD, audio_exempt: bool = AUDIO_EXEMPT) -> bool:
289
+ """Bind the caching forward onto `transformer`. Returns whether it went on.
290
+
291
+ `accelerate`'s `add_hook_to_module` — what `ComponentsManager.enable_auto_cpu_offload` installs — moves the real
292
+ forward to `_old_forward` and puts its own onload wrapper in `forward`. Replacing `forward` there would step over
293
+ the wrapper and run the block stack against weights still on the host, so the replacement goes into `_old_forward`
294
+ whenever the hook is present.
295
+ """
296
+ import inspect
297
+
298
+ if getattr(transformer, "_h3_fbc", None) is not None:
299
+ return True
300
+
301
+ hooked = hasattr(transformer, "_hf_hook") and hasattr(transformer, "_old_forward")
302
+ current = transformer._old_forward if hooked else transformer.forward
303
+ missing = [name for name in FORWARD_PARAMETERS if name not in inspect.signature(current).parameters]
304
+ if missing:
305
+ print(f"[h3-fbc] this transformer's forward has no {missing}; running uncached", flush=True)
306
+ return False
307
+ if not hasattr(transformer, "transformer_blocks") or len(transformer.transformer_blocks) < 2:
308
+ print("[h3-fbc] no block stack to skip; running uncached", flush=True)
309
+ return False
310
+
311
+ state = _State(threshold, steps, audio_exempt)
312
+ state.original = current
313
+ transformer._h3_fbc = state
314
+ bound = types.MethodType(_forward, transformer)
315
+ if hooked:
316
+ transformer._old_forward = bound
317
+ else:
318
+ transformer.forward = bound
319
+ return True
320
+
321
+
322
+ def uninstall(transformer) -> None:
323
+ state = getattr(transformer, "_h3_fbc", None)
324
+ if state is None:
325
+ return
326
+ if hasattr(transformer, "_hf_hook") and hasattr(transformer, "_old_forward"):
327
+ transformer._old_forward = state.original
328
+ else:
329
+ transformer.__dict__.pop("forward", None)
330
+ del transformer._h3_fbc
331
+ total = state.computed + state.skipped
332
+ if total:
333
+ print(
334
+ f"[h3-fbc] {state.skipped}/{total} forwards served from cache "
335
+ f"(threshold {state.threshold}, audio exemption {'on' if state.audio_exempt else 'off'})",
336
+ flush=True,
337
+ )
338
+
339
+
340
+ @contextlib.contextmanager
341
+ def enabled(transformer, steps: int = 0):
342
+ """Cache the trunk for the duration of one request. The state is per-request by construction — a residual only ever
343
+ means something within the schedule it was measured on — and nothing in here can raise into the request."""
344
+ installed = False
345
+ if ENABLED:
346
+ try:
347
+ installed = install(transformer, steps=steps)
348
+ except Exception as error:
349
+ print(f"[h3-fbc] install failed ({type(error).__name__}: {error}); running uncached", flush=True)
350
+ try:
351
+ yield installed
352
+ finally:
353
+ if installed:
354
+ try:
355
+ uninstall(transformer)
356
+ except Exception as error:
357
+ print(f"[h3-fbc] uninstall failed ({type(error).__name__}: {error})", flush=True)
h3_split_blocks.py ADDED
@@ -0,0 +1,147 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """The halves of a **split** MiniMax-H3 deployment, for both of its checkpoint partitions.
2
+
3
+ MiniMax-H3 is 195.9 GiB in bfloat16 and a ZeroGPU Space is evicted at 150 GB of storage, so `MiniMaxH3Blocks` is cut
4
+ at its `text_encoder` step: the 62.14 GiB Qwen3-VL runs in the conditioner Space, everything else in a generator
5
+ Space, and `prompt_embeds` + `text_token_tags` is the whole wire format between them.
6
+
7
+ `resize` / `setup` run on **both** sides: they own no pretrained component, and each half needs the canvas and the
8
+ prepared keyframes or normalized references. Both conditioner halves also return the resolved `height` / `width` /
9
+ `num_frames`, which the generating half pins rather than re-deriving.
10
+
11
+ Two things the blocks leave to the caller: a keyframe reaches them EXIF-transposed and in RGB, and the `t2va` / `fl2va`
12
+ frame count is aligned to `17 * n + 5` before the call, since that arithmetic lives on the denoising side of the cut.
13
+ """
14
+
15
+ from diffusers.modular_pipelines.minimax_h3.before_encoder import MiniMaxH3Ref2VASetupStep
16
+ from diffusers.modular_pipelines.minimax_h3.decoders import MiniMaxH3AfterDenoiseStep
17
+ from diffusers.modular_pipelines.minimax_h3.encoders import (
18
+ MiniMaxH3Ref2VAReferenceEncoderStep,
19
+ MiniMaxH3Ref2VATextEncoderStep,
20
+ MiniMaxH3TextEncoderStep,
21
+ )
22
+ from diffusers.modular_pipelines.minimax_h3.modular_blocks_minimax_h3 import (
23
+ MiniMaxH3AutoKeyframeVaeEncoderStep,
24
+ MiniMaxH3AutoResizeStep,
25
+ MiniMaxH3CoreDenoiseStep,
26
+ MiniMaxH3DecodeStep,
27
+ MiniMaxH3Ref2VACoreDenoiseStep,
28
+ _generation_outputs,
29
+ )
30
+ from diffusers.modular_pipelines.modular_pipeline import SequentialPipelineBlocks
31
+ from diffusers.modular_pipelines.modular_pipeline_utils import OutputParam
32
+
33
+
34
+ def _wire_outputs(num_frames: bool = True) -> list[OutputParam]:
35
+ """The wire format of the split. `num_frames` is declared by the `ref2va` half alone, whose setup resolves one."""
36
+ return [
37
+ OutputParam.template("prompt_embeds"),
38
+ OutputParam("text_token_tags", description="The per-row modality tag of every row of `prompt_embeds`."),
39
+ OutputParam("height", type_hint=int, description="Resolved height of the generated video in pixels."),
40
+ OutputParam("width", type_hint=int, description="Resolved width of the generated video in pixels."),
41
+ *(
42
+ [OutputParam("num_frames", type_hint=int, description="Resolved number of frames, of the form 17 * n + 5.")]
43
+ if num_frames
44
+ else []
45
+ ),
46
+ ]
47
+
48
+
49
+ class MiniMaxH3ConditionerBlocks(SequentialPipelineBlocks):
50
+ """The conditioner half of a split MiniMax-H3: the keyframes on the canvas plus the Qwen3-VL read at layer 50."""
51
+
52
+ model_name = "minimax-h3"
53
+ block_classes = [MiniMaxH3AutoResizeStep, MiniMaxH3TextEncoderStep]
54
+ block_names = ["resize", "text_encoder"]
55
+
56
+ @property
57
+ def description(self):
58
+ return (
59
+ "The conditioner half of a split MiniMax-H3 deployment: puts the keyframes onto the target canvas and "
60
+ "encodes MiniMax-H3's presentation of the request into the `prompt_embeds` / `text_token_tags` pair the "
61
+ "denoising half consumes. The frame count is the caller's to align."
62
+ )
63
+
64
+ @property
65
+ def outputs(self):
66
+ return _wire_outputs(num_frames=False)
67
+
68
+
69
+ class MiniMaxH3GeneratorBlocks(SequentialPipelineBlocks):
70
+ """The denoising half of a split MiniMax-H3: `MiniMaxH3Blocks` with its `text_encoder` step removed."""
71
+
72
+ model_name = "minimax-h3"
73
+ block_classes = [
74
+ MiniMaxH3AutoResizeStep,
75
+ MiniMaxH3AutoKeyframeVaeEncoderStep,
76
+ MiniMaxH3CoreDenoiseStep,
77
+ MiniMaxH3AfterDenoiseStep,
78
+ MiniMaxH3DecodeStep,
79
+ ]
80
+ block_names = ["resize", "vae_encoder", "denoise", "after_denoise", "decode"]
81
+
82
+ @property
83
+ def description(self):
84
+ return (
85
+ "The denoising half of a split MiniMax-H3 deployment: the `t2va` / `fl2va` branch of `MiniMaxH3Blocks` "
86
+ "without its text-encoder step, so `prompt_embeds` and `text_token_tags` come in as inputs and the "
87
+ "62.14 GiB Qwen3-VL conditioner is never loaded here."
88
+ )
89
+
90
+ @property
91
+ def outputs(self):
92
+ return _generation_outputs()
93
+
94
+
95
+ class MiniMaxH3Ref2VAConditionerBlocks(SequentialPipelineBlocks):
96
+ """The conditioner half of a split `ref2va`: the resolved plan plus the Qwen3-VL read at its 50th layer.
97
+
98
+ Component for component this is `MiniMaxH3ConditionerBlocks`, so one conditioner Space serves both partitions.
99
+ What differs is the presentation: `ref2va` prepends a label per reference and a vision block per image and per
100
+ merged video frame pair, so the references themselves have to reach this half.
101
+ """
102
+
103
+ model_name = "minimax-h3"
104
+ block_classes = [MiniMaxH3Ref2VASetupStep, MiniMaxH3Ref2VATextEncoderStep]
105
+ block_names = ["setup", "text_encoder"]
106
+
107
+ @property
108
+ def description(self):
109
+ return (
110
+ "The conditioner half of a split MiniMax-H3 `ref2va` deployment: resolves the request plan (canvas, frame "
111
+ "count, references normalized onto MiniMax-H3's own rates and resolutions) and encodes MiniMax-H3's "
112
+ "presentation of it into the `prompt_embeds` / `text_token_tags` pair the denoising half consumes."
113
+ )
114
+
115
+ @property
116
+ def outputs(self):
117
+ return _wire_outputs()
118
+
119
+
120
+ class MiniMaxH3Ref2VAGeneratorBlocks(SequentialPipelineBlocks):
121
+ """The denoising half of a split `ref2va`: the `ref2va` branch with its `text_encoder` step removed.
122
+
123
+ `reference_encoder` stays here, next to the two autoencoders it runs: its output shapes are where every reference
124
+ block's geometry in the packed layout comes from.
125
+ """
126
+
127
+ model_name = "minimax-h3"
128
+ block_classes = [
129
+ MiniMaxH3Ref2VASetupStep,
130
+ MiniMaxH3Ref2VAReferenceEncoderStep,
131
+ MiniMaxH3Ref2VACoreDenoiseStep,
132
+ MiniMaxH3AfterDenoiseStep,
133
+ MiniMaxH3DecodeStep,
134
+ ]
135
+ block_names = ["setup", "reference_encoder", "denoise", "after_denoise", "decode"]
136
+
137
+ @property
138
+ def description(self):
139
+ return (
140
+ "The denoising half of a split MiniMax-H3 `ref2va` deployment: the `ref2va` branch of `MiniMaxH3Blocks` "
141
+ "without its text-encoder step, so `prompt_embeds` and `text_token_tags` come in as inputs and the "
142
+ "62.14 GiB Qwen3-VL conditioner is never loaded here. The transformer is the `transformer_ref` partition."
143
+ )
144
+
145
+ @property
146
+ def outputs(self):
147
+ return _generation_outputs()
ncii_guard.py ADDED
@@ -0,0 +1,90 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """The NCII prompt guard, in its own process.
2
+
3
+ [`hfmlsoc/ncii-light-guard-v01`](https://huggingface.co/hfmlsoc/ncii-light-guard-v01) is a 270M CPU text
4
+ classifier scoring the NCII risk of an edit prompt. It cannot live in the main process: with it loaded there,
5
+ every subsequent `@spaces.GPU` worker dies at `worker_init` with `RuntimeError: No CUDA GPUs are available` —
6
+ the fork inherits whatever CUDA driver state the classifier's torch activity left behind, and a factory reboot
7
+ does not clear it. It cannot be a `multiprocessing.spawn` child either: spawn re-imports the parent's main
8
+ module, and on a Space that main module is `app.py` — the child would re-run the whole startup, `start()`
9
+ included. So the classifier runs this file as a plain subprocess — a fresh interpreter that never sees `spaces`
10
+ — and answers over stdin/stdout, one JSON object per line.
11
+ """
12
+
13
+ from __future__ import annotations
14
+
15
+ import json
16
+ import os
17
+ import select
18
+ import subprocess
19
+ import sys
20
+ import threading
21
+
22
+ GUARD_REPO = "hfmlsoc/ncii-light-guard-v01"
23
+
24
+ _lock = threading.Lock()
25
+ _process: subprocess.Popen | None = None
26
+
27
+
28
+ def _read(timeout: float) -> dict:
29
+ readable, _, _ = select.select([_process.stdout], [], [], timeout)
30
+ if not readable:
31
+ raise TimeoutError(f"the guard did not answer within {timeout}s")
32
+ line = _process.stdout.readline()
33
+ if not line:
34
+ raise EOFError("the guard process died")
35
+ return json.loads(line)
36
+
37
+
38
+ def _spawn() -> None:
39
+ global _process
40
+ _process = subprocess.Popen(
41
+ [sys.executable, os.path.abspath(__file__)],
42
+ stdin=subprocess.PIPE,
43
+ stdout=subprocess.PIPE,
44
+ text=True,
45
+ bufsize=1,
46
+ )
47
+ # Generous: a cold cache downloads the checkpoint first.
48
+ assert _read(300.0) == {"status": "ready"}
49
+
50
+
51
+ def start() -> None:
52
+ """Launch the worker and block until its model is up. Called once at startup; `classify` revives it if it dies."""
53
+ with _lock:
54
+ _spawn()
55
+
56
+
57
+ def classify(prompt: str, timeout: float = 60.0) -> dict:
58
+ """`{'label': 'safe' | 'ncii', 'score': ...}` for one prompt, replacing a dead or wedged worker once."""
59
+ with _lock:
60
+ for attempt in (0, 1):
61
+ try:
62
+ if _process is None or _process.poll() is not None:
63
+ _spawn()
64
+ _process.stdin.write(json.dumps({"prompt": prompt}) + "\n")
65
+ _process.stdin.flush()
66
+ return _read(timeout)
67
+ except Exception:
68
+ if attempt:
69
+ raise
70
+ if _process is not None and _process.poll() is None:
71
+ _process.kill()
72
+
73
+
74
+ def _serve() -> None:
75
+ """The child: plain torch on CPU. The protocol keeps the real stdout to itself — everything else
76
+ (download progress, warnings) is pushed over to stderr so it cannot corrupt a reply."""
77
+ protocol = os.fdopen(os.dup(1), "w", buffering=1)
78
+ os.dup2(2, 1)
79
+
80
+ from transformers import pipeline
81
+
82
+ classifier = pipeline("text-classification", model=GUARD_REPO, device="cpu")
83
+ protocol.write(json.dumps({"status": "ready"}) + "\n")
84
+ for line in sys.stdin:
85
+ result = classifier(json.loads(line)["prompt"], truncation=True)[0]
86
+ protocol.write(json.dumps({"label": result["label"], "score": float(result["score"])}) + "\n")
87
+
88
+
89
+ if __name__ == "__main__":
90
+ _serve()
packages.txt ADDED
@@ -0,0 +1 @@
 
 
1
+ ffmpeg
requirements.txt ADDED
@@ -0,0 +1,33 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # `diffusers` is installed from the canonical MiniMax-H3 pull request,
2
+ # https://github.com/huggingface/diffusers/pull/14371 ("Minimax h3 follow up (review & refactor)"), pinned to a
3
+ # **commit** rather than to its `minimax-h3-refactor` branch: the PR is a WIP and its head moves, and this Space's
4
+ # `h3_split_blocks.py` subclasses its block classes. Re-pin — and re-check `h3_split_blocks.py` against the block
5
+ # names of the new head — whenever the PR updates. (diffusers `main` has already renamed several of them.) It is
6
+ # also the commit `multimodalart/qwen3vl-conditioner` serves, and the two halves have to agree on the wire format.
7
+ #
8
+ # 665f578278365ea4a3318cb8c9b66ce6c01204b9 = refs/pull/14371/head at the time of this deploy
9
+ --extra-index-url https://download.pytorch.org/whl/cu130
10
+ diffusers @ git+https://github.com/huggingface/diffusers.git@665f578278365ea4a3318cb8c9b66ce6c01204b9
11
+ torch==2.11.0
12
+ torchvision==0.26.0
13
+ # A scene clip's soundtrack that is not already at the audio VAE's 32 kHz is resampled with `torchaudio`. Real
14
+ # camera and phone footage is 44.1/48 kHz, so this is hit by most uploads that are not silent.
15
+ torchaudio==2.11.0
16
+ # The Qwen3-VL processor decides the vision patch count, so a different minor changes the conditioning.
17
+ transformers==5.8.0
18
+ accelerate==1.14.0
19
+ # diffusers pins <2.
20
+ huggingface-hub==1.24.0
21
+ # The character-swap LoRA is attached as a runtime PEFT adapter (`load_lora_adapter` + `set_adapters`), not folded
22
+ # into the base weights, so its strength stays a per-request slider. Without `peft`, `load_lora_adapter` raises
23
+ # ModuleNotFoundError at startup and the Space would quietly serve the base `ref2va` model.
24
+ peft
25
+ # `gradio` and `spaces` are deliberately absent: the Space runtime installs both itself (gradio from `sdk_version`
26
+ # in README.md, `spaces` at whichever version it currently ships), so pinning either here is a resolver conflict
27
+ # rather than a version choice.
28
+ # PyAV decodes the scene clip and muxes the generated soundtrack onto the frames (`encode_video`).
29
+ av
30
+ pillow
31
+ numpy
32
+ requests
33
+ safetensors>=0.8.0