Spaces:
Running on Zero
Running on Zero
MiniMax-H3 Third-Person View LoRA demo
Browse files- .gitattributes +4 -0
- README.md +103 -6
- app.py +894 -0
- examples/doorway_character.jpg +3 -0
- examples/katana_character.jpg +3 -0
- examples/night_city.jpg +3 -0
- examples/ruined_village.jpg +0 -0
- examples/wyvern_boss.jpg +3 -0
- h3_fbc.py +357 -0
- h3_split_blocks.py +147 -0
- ncii_guard.py +90 -0
- packages.txt +1 -0
- requirements.txt +18 -0
.gitattributes
CHANGED
|
@@ -33,3 +33,7 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
|
|
| 33 |
*.zip filter=lfs diff=lfs merge=lfs -text
|
| 34 |
*.zst filter=lfs diff=lfs merge=lfs -text
|
| 35 |
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 33 |
*.zip filter=lfs diff=lfs merge=lfs -text
|
| 34 |
*.zst filter=lfs diff=lfs merge=lfs -text
|
| 35 |
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
| 36 |
+
examples/doorway_character.jpg filter=lfs diff=lfs merge=lfs -text
|
| 37 |
+
examples/katana_character.jpg filter=lfs diff=lfs merge=lfs -text
|
| 38 |
+
examples/night_city.jpg filter=lfs diff=lfs merge=lfs -text
|
| 39 |
+
examples/wyvern_boss.jpg filter=lfs diff=lfs merge=lfs -text
|
README.md
CHANGED
|
@@ -1,13 +1,110 @@
|
|
| 1 |
---
|
| 2 |
-
title:
|
| 3 |
-
emoji:
|
| 4 |
-
colorFrom:
|
| 5 |
-
colorTo:
|
| 6 |
sdk: gradio
|
| 7 |
sdk_version: 6.28.0
|
| 8 |
-
python_version: '3.13'
|
| 9 |
app_file: app.py
|
|
|
|
|
|
|
| 10 |
pinned: false
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 11 |
---
|
| 12 |
|
| 13 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
---
|
| 2 |
+
title: MiniMax-H3 Third-Person View LoRA
|
| 3 |
+
emoji: 🎮
|
| 4 |
+
colorFrom: yellow
|
| 5 |
+
colorTo: pink
|
| 6 |
sdk: gradio
|
| 7 |
sdk_version: 6.28.0
|
|
|
|
| 8 |
app_file: app.py
|
| 9 |
+
python_version: "3.12"
|
| 10 |
+
startup_duration_timeout: 1h
|
| 11 |
pinned: false
|
| 12 |
+
short_description: Design sheets to game cutscenes with HUD and audio
|
| 13 |
+
suggested_hardware: zero-a10g
|
| 14 |
+
models:
|
| 15 |
+
- MiniMaxAI/MiniMax-H3
|
| 16 |
+
- multimodalart/MiniMax-H3-Pruned
|
| 17 |
+
- WarmBloodAban/Minimax_H3_LoRAs
|
| 18 |
---
|
| 19 |
|
| 20 |
+
# MiniMax-H3 · Third-Person View LoRA
|
| 21 |
+
|
| 22 |
+
A demo of [`WarmBloodAban/Minimax_H3_LoRAs`](https://huggingface.co/WarmBloodAban/Minimax_H3_LoRAs)
|
| 23 |
+
(`Minimax-h3_Third_person_view.safetensors`) on [MiniMax-H3](https://huggingface.co/MiniMaxAI/MiniMax-H3),
|
| 24 |
+
MiniMax's 33B omni-modal model that generates video with natively synchronized audio.
|
| 25 |
+
|
| 26 |
+
The LoRA turns H3 into a **game-cutscene renderer**: over-the-shoulder and FPV camera language,
|
| 27 |
+
Unreal-Engine-looking lighting, and — the part it is actually distinctive for — **legible HUD furniture
|
| 28 |
+
drawn in-frame**: target-lock reticles, floating damage numbers, QTE prompts, translucent status bars.
|
| 29 |
+
The soundtrack (impacts, foley, score) is base H3's, generated in the same denoising pass.
|
| 30 |
+
|
| 31 |
+
## It is a *reference* LoRA, not a text-to-video one
|
| 32 |
+
|
| 33 |
+
The adapter's own metadata says `ss_base_model_version: minimax_h3_ref2va`, so it was trained on H3's
|
| 34 |
+
**ref2va** partition — the one that conditions on an ordered list of reference media, not the plain
|
| 35 |
+
text-to-video partition. So this Space is built around reference sheets: you upload a character sheet,
|
| 36 |
+
an environment plate, a boss design, and H3 keeps those identities consistent through the shot.
|
| 37 |
+
Running this adapter on the t2v partition would attach cleanly and quietly produce worse output; it
|
| 38 |
+
isn't offered.
|
| 39 |
+
|
| 40 |
+
## What the demo owes the model card
|
| 41 |
+
|
| 42 |
+
| The card says | Here |
|
| 43 |
+
|---|---|
|
| 44 |
+
| LoRA weight **0.6 – 0.85**, 0.7 recommended | A slider clamped to exactly that window, defaulting to 0.7. The adapter is attached with PEFT and re-scaled per request, so the window is live |
|
| 45 |
+
| A long structured prompt format — `subject_definitions:` / `summary:` / `retention_analysis:` / `detailed_description:` with `[Shot N]` beats / `overall_soundscape:` / `non_diegetic_music:`, cross-referencing `<Subject N>`, `<Environment N>`, `<Picture N>` | The UI **is** that format. Each reference slot you fill becomes a `<Subject N>`/`<Environment N>` bound to its `<Picture N>`; your one action line plus the camera preset become the `detailed_description:` beats; the `summary:` is prefixed `[reference generation]` as the card shows. The assembled document is shown after every run, and an override box takes a hand-written one |
|
| 46 |
+
| Trigger vocabulary in three groups — camera views (*third-person perspective*, *over-the-shoulder camera*, *first-person POV*, *FPV HUD*), game rendering (*rendered in Unreal Engine*, *gameplay sequence*, *combat stance*, *macro close-up*), UI elements (*transparent HUD elements*, *target-lock UI reticle*, *QTE UI prompt*, *floating damage text UI*) | The four camera presets each spend the camera-view + game-rendering tags their shot actually wants; the HUD checkbox spends the UI group. Nothing is sprinkled in decoratively |
|
| 47 |
+
| Base is the **INT8 Reference** architecture | The ref2va DiT, in bf16 on ZeroGPU's 96 GB tier — no quantization needed at this size |
|
| 48 |
+
|
| 49 |
+
## Architecture
|
| 50 |
+
|
| 51 |
+
MiniMax-H3 is 195.9 GiB in bf16 and a ZeroGPU Space is evicted at 150 GB of storage, so the pipeline is
|
| 52 |
+
cut at its `text_encoder` step — the split every MiniMax-H3 Space uses. The 62 GiB Qwen3-VL conditioner
|
| 53 |
+
runs in [`multimodalart/qwen3vl-conditioner`](https://huggingface.co/spaces/multimodalart/qwen3vl-conditioner)
|
| 54 |
+
and is called over the gradio API with each caller's own ZeroGPU token forwarded, so the booking is
|
| 55 |
+
billed to whoever asked for the video; `prompt_embeds` + `text_token_tags` is the whole wire format.
|
| 56 |
+
This Space holds the DiT and the two autoencoders. `h3_split_blocks.py` holds the denoising-half block
|
| 57 |
+
stack.
|
| 58 |
+
|
| 59 |
+
The DiT is [`multimodalart/MiniMax-H3-Pruned`](https://huggingface.co/multimodalart/MiniMax-H3-Pruned),
|
| 60 |
+
which folds the AdaLN input projections onto a rank-8 subspace: `transformer_ref` drops from 61.7 GiB
|
| 61 |
+
to 37.5 GiB with identical output. This LoRA touches no AdaLN module, so nothing needs projecting.
|
| 62 |
+
|
| 63 |
+
## How the LoRA is applied
|
| 64 |
+
|
| 65 |
+
Attached as a live **PEFT adapter** (`load_lora_adapter(..., adapter_name="third_person_view")`) and
|
| 66 |
+
re-scaled per request via `set_adapters`, rather than folded into the base weights — folding would be
|
| 67 |
+
faster to serve but would freeze the strength, and 0.6–0.85 is the knob this LoRA is actually steered
|
| 68 |
+
with. For the same reason the pipeline is **not** AoTI-compiled: a captured graph would route around
|
| 69 |
+
the PEFT-injected layers and silently drop the adapter.
|
| 70 |
+
|
| 71 |
+
The export is keyed on the *original* H3 partition (`diffusion_model.blocks.N.…`, 416 tensors, rank 64,
|
| 72 |
+
ai-toolkit), so applying it to the diffusers port means replaying the layout transforms
|
| 73 |
+
[`convert_minimax_h3_to_diffusers.py`](https://github.com/huggingface/diffusers/blob/main/scripts/convert_minimax_h3_to_diffusers.py)
|
| 74 |
+
applied to the base weights — a delta is only valid in the layout of the weight it is added to:
|
| 75 |
+
|
| 76 |
+
| Original LoRA target | diffusers parameter | Transform |
|
| 77 |
+
|---|---|---|
|
| 78 |
+
| `blocks.N.…` / `token_refiner.blocks.N.…` | `transformer_blocks.N.…` / `token_refiner.refiner_blocks.N.…` | rename |
|
| 79 |
+
| `attn.out_proj` | `attn.to_out.0` | rename |
|
| 80 |
+
| `attn.qkv_proj` | `attn.to_q` / `to_k` / `to_v` | split `lora_B`'s rows into contiguous thirds |
|
| 81 |
+
| `mlp.fc2` | `ff.net.2` | rename |
|
| 82 |
+
| `mlp.fc1` | `ff.net.0.proj` | **swap the two fused SwiGLU halves** (`[gate; value]` → `[value; gate]`) |
|
| 83 |
+
|
| 84 |
+
Both of the last two rows are easy to get wrong. The gate/value swap is a real difference between H3's
|
| 85 |
+
native SwiGLU packing and diffusers' `GEGLU`-style ordering — skip it and every MLP delta lands on the
|
| 86 |
+
wrong half of the activation. Conversely the contiguous QKV split is *only* correct for ai-toolkit
|
| 87 |
+
exports like this one; fal-trained H3 LoRAs store per-head-interleaved QKV rows and need de-interleaving
|
| 88 |
+
first. The two are distinguished by key naming, and this repo is unambiguously the former. Any target
|
| 89 |
+
that resolves to no transformer weight is a fatal error rather than a silent skip.
|
| 90 |
+
|
| 91 |
+
## Examples
|
| 92 |
+
|
| 93 |
+
Three ready-to-run briefs. The wyvern-boss and ruined-village reference sheets are frames lifted from
|
| 94 |
+
the LoRA author's own showcase clip in
|
| 95 |
+
[`WarmBloodAban/Minimax_H3_LoRAs`](https://huggingface.co/WarmBloodAban/Minimax_H3_LoRAs) (Apache-2.0) —
|
| 96 |
+
the closest thing to the inputs the adapter was demonstrated on. The character and city plates come from
|
| 97 |
+
[`linoyts/repo-to-space-example-inputs`](https://huggingface.co/datasets/linoyts/repo-to-space-example-inputs)
|
| 98 |
+
(CC0-1.0).
|
| 99 |
+
|
| 100 |
+
## Known limits
|
| 101 |
+
|
| 102 |
+
Five seconds of 24 fps video per run, at 960×544 or 1280×720. HUD text is *shaped* like HUD text and
|
| 103 |
+
usually isn't real words — that's a base-H3 limit the LoRA doesn't fix. Faces drift on fast whip-pans.
|
| 104 |
+
The first request after a cold start pays for weight download plus placement.
|
| 105 |
+
|
| 106 |
+
## License
|
| 107 |
+
|
| 108 |
+
MiniMax-H3 is covered by the **MiniMax H3 Community License Agreement** — note its attribution,
|
| 109 |
+
territorial-scope and revenue clauses. The LoRA is Apache-2.0. Read both before using any output of
|
| 110 |
+
this Space.
|
app.py
ADDED
|
@@ -0,0 +1,894 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""MiniMax-H3 · Third-Person View LoRA — game cutscenes from design reference sheets.
|
| 2 |
+
|
| 3 |
+
[`WarmBloodAban/Minimax_H3_LoRAs`](https://huggingface.co/WarmBloodAban/Minimax_H3_LoRAs) ships
|
| 4 |
+
`Minimax-h3_Third_person_view.safetensors`, a rank-64 ai-toolkit LoRA whose own metadata names its base as
|
| 5 |
+
`minimax_h3_ref2va` — the **reference** partition of MiniMax-H3, the one that conditions on an ordered list of
|
| 6 |
+
reference images rather than on a first frame. That is why this Space is a `ref2va` deployment: the adapter goes
|
| 7 |
+
onto `transformer_ref`, and the user's uploads arrive as `<Picture 1..3>` in the order they are given.
|
| 8 |
+
|
| 9 |
+
What the LoRA does, from its card: cinematic game cutscenes — third-person over-the-shoulder / spring-arm camera
|
| 10 |
+
tracking, first-person POV, whip-pan view transitions, and native game HUD overlays (target-lock reticles, Boss
|
| 11 |
+
health bars, QTE prompts, floating damage text). Its card is explicit that it wants MiniMax-H3's **structured**
|
| 12 |
+
prompt format (`subject_definitions` / `summary` / `retention_analysis` / `detailed_description` /
|
| 13 |
+
`overall_soundscape` / `non_diegetic_music`) with `<Subject N>` / `<Environment N>` / `<Picture N>` cross-references,
|
| 14 |
+
so this demo is a composer for exactly that document rather than a prompt box: you label each reference sheet, write
|
| 15 |
+
one line of action, pick a camera and whether the HUD is on, and the Space assembles the card's format around the
|
| 16 |
+
LoRA's own trigger vocabulary. The assembled document is shown next to the result, and an override box takes a
|
| 17 |
+
hand-written one.
|
| 18 |
+
|
| 19 |
+
Recipe, all from the card: LoRA weight **0.6–0.85, 0.7 recommended** (a live slider here, which is why the adapter
|
| 20 |
+
is attached with PEFT rather than folded into the weights).
|
| 21 |
+
|
| 22 |
+
Deployment follows the other MiniMax-H3 Spaces. H3 is 195.9 GiB in bfloat16 and a ZeroGPU Space is evicted at 150 GB
|
| 23 |
+
of storage, so `MiniMaxH3Ref2VAGeneratorBlocks` is cut at its `text_encoder` step: the 62 GiB Qwen3-VL conditioner
|
| 24 |
+
runs in [`multimodalart/qwen3vl-conditioner`](https://huggingface.co/spaces/multimodalart/qwen3vl-conditioner) and is
|
| 25 |
+
called over the gradio API, while this Space holds the DiT and the two autoencoders. The DiT is
|
| 26 |
+
[`multimodalart/MiniMax-H3-Pruned`](https://huggingface.co/multimodalart/MiniMax-H3-Pruned)'s `transformer_ref`
|
| 27 |
+
(37.5 GiB — the AdaLN input projections folded onto their reachable rank-8 subspace, 1.5e-5 relative, ~250x below one
|
| 28 |
+
bfloat16 step), which this LoRA does not touch: it adapts `attn.qkv_proj`, `attn.out_proj` and `mlp.fc1`/`fc2` only.
|
| 29 |
+
"""
|
| 30 |
+
|
| 31 |
+
from __future__ import annotations
|
| 32 |
+
|
| 33 |
+
import os
|
| 34 |
+
import tempfile
|
| 35 |
+
import time
|
| 36 |
+
import traceback
|
| 37 |
+
from functools import cache
|
| 38 |
+
|
| 39 |
+
# Before anything that could initialize CUDA: `import spaces` patches `torch.cuda` so the weights can be loaded at
|
| 40 |
+
# startup rather than on GPU time.
|
| 41 |
+
import spaces # noqa: F401
|
| 42 |
+
import gradio as gr
|
| 43 |
+
import torch
|
| 44 |
+
|
| 45 |
+
VERSION = "tpv-lora"
|
| 46 |
+
MODEL_REPO = os.environ.get("H3_MODEL_REPO", "multimodalart/MiniMax-H3-Pruned")
|
| 47 |
+
LORA_REPO = os.environ.get("H3_LORA_REPO", "WarmBloodAban/Minimax_H3_LoRAs")
|
| 48 |
+
LORA_FILE = os.environ.get("H3_LORA_FILE", "Minimax-h3_Third_person_view.safetensors")
|
| 49 |
+
ADAPTER = "third_person_view"
|
| 50 |
+
CONDITIONER_SPACE = os.environ.get("H3_CONDITIONER", "multimodalart/qwen3vl-conditioner")
|
| 51 |
+
# `lazy` moves the weights onto the card inside the first GPU call and leaves them there. Startup placement is not an
|
| 52 |
+
# option: `spaces`' startup `torch.pack()` writes a second on-disk copy of every startup-resident CUDA tensor, and
|
| 53 |
+
# 48 GB of weights plus its pack runs at the 150 GB Space storage quota.
|
| 54 |
+
PLACEMENT = os.environ.get("H3_PLACEMENT", "lazy").lower()
|
| 55 |
+
# cuDNN's fused attention is 10-20% faster than the SDPA default on this pool and needs nothing installed.
|
| 56 |
+
# flash-attention 3 is sm90-only and this card is sm120 (the `zero-a10g` flavour name is legacy).
|
| 57 |
+
ATTENTION = os.environ.get("H3_ATTENTION", "_native_cudnn").lower()
|
| 58 |
+
GPU_SIZE = os.environ.get("H3_GPU_SIZE", "xlarge")
|
| 59 |
+
MIN_GPU_DURATION = int(os.environ.get("H3_GPU_DURATION_MIN", "120"))
|
| 60 |
+
MAX_GPU_DURATION = int(os.environ.get("H3_GPU_DURATION_MAX", "1500"))
|
| 61 |
+
|
| 62 |
+
# The card's recommended weight window, verbatim: "LoRA Weight: 0.6 - 0.85 (0.7 is recommended as a starting point)".
|
| 63 |
+
WEIGHT_MIN, WEIGHT_MAX, DEFAULT_WEIGHT = 0.6, 0.85, 0.7
|
| 64 |
+
DEFAULT_STEPS = 20
|
| 65 |
+
DEFAULT_SEED = 42
|
| 66 |
+
|
| 67 |
+
# Must stay identical to the conditioner's table: the *label* goes over the wire, so a canvas that half does not know
|
| 68 |
+
# is rejected there and surfaces as a failure here.
|
| 69 |
+
CANVASES = {
|
| 70 |
+
# 16:9
|
| 71 |
+
"960x544 · 16:9 fast": (544, 960),
|
| 72 |
+
"1024x576 · 16:9 fast": (576, 1024),
|
| 73 |
+
"1152x640 · 16:9": (640, 1152),
|
| 74 |
+
"1280x704 · 16:9": (704, 1280),
|
| 75 |
+
"1344x768 · 16:9 full": (768, 1344),
|
| 76 |
+
# 21:9
|
| 77 |
+
"1152x512 · 21:9 fast": (512, 1152),
|
| 78 |
+
"1536x672 · 21:9 full": (672, 1536),
|
| 79 |
+
# 9:16
|
| 80 |
+
"544x960 · 9:16 fast": (960, 544),
|
| 81 |
+
"640x1152 · 9:16": (1152, 640),
|
| 82 |
+
# 4:3
|
| 83 |
+
"768x576 · 4:3 fast": (576, 768),
|
| 84 |
+
"1024x768 · 4:3 full": (768, 1024),
|
| 85 |
+
}
|
| 86 |
+
# Game cutscenes are widescreen; the cheapest 16:9 bucket keeps a default request inside a few minutes of GPU.
|
| 87 |
+
DEFAULT_CANVAS = "960x544 · 16:9 fast"
|
| 88 |
+
FPS, FRAMES_PER_CHUNK, LATENTS_PER_CHUNK = 24, 17, 5
|
| 89 |
+
MIN_DURATION, MAX_UI_DURATION = 2, 14
|
| 90 |
+
MAX_IMAGE_SLOTS, OPEN_IMAGE_SLOTS = 3, 2
|
| 91 |
+
|
| 92 |
+
# Seconds of GPU one request needs, from the packed sequence it is about to denoise: linear in the rows for the
|
| 93 |
+
# matmuls, quadratic for the attention. Fit shared with the other MiniMax-H3 Spaces on this pool.
|
| 94 |
+
STEP_LINEAR, STEP_QUADRATIC, SAFETY = 1.1745e-4, 3.8396e-9, 1.3
|
| 95 |
+
PLACEMENT_ALLOWANCE = int(os.environ.get("H3_PLACEMENT_ALLOWANCE", "90"))
|
| 96 |
+
AUDIO_LATENTS_PER_SECOND, AUDIO_CHANNELS = 40, 2
|
| 97 |
+
REFERENCE_IMAGE_SHORT_EDGE, CANVAS_MULTIPLE = 2048, 32
|
| 98 |
+
DECODE_BASE, DECODE_PER_DEFAULT_CANVAS, DEFAULT_CANVAS_PIXELS = 15, 25, 960 * 544 * 124
|
| 99 |
+
# A 2048-short-edge reference also becomes Qwen3-VL vision tokens in the text stream. Only the *estimate* label needs
|
| 100 |
+
# a number for them; the booking itself reads the tags the conditioner actually returned.
|
| 101 |
+
ESTIMATED_TEXT_ROWS, ESTIMATED_VISION_ROWS_PER_IMAGE = 900, 1800
|
| 102 |
+
|
| 103 |
+
# ── The LoRA's prompt contract ────────────────────────────────────────────────
|
| 104 |
+
# Roles, camera phrasings and the HUD clause are built out of the card's own "Key Trigger Words & Recommended Tags"
|
| 105 |
+
# (camera views: third-person perspective, over-the-shoulder camera, first-person POV, FPV HUD; game rendering:
|
| 106 |
+
# rendered in Unreal Engine, gameplay sequence, combat stance; UI elements: transparent HUD elements, target-lock UI
|
| 107 |
+
# reticle, QTE UI prompt, floating damage text UI) and out of the structured example prompt it publishes.
|
| 108 |
+
|
| 109 |
+
ROLES = {
|
| 110 |
+
"Character": ("subject", "character design reference sheet"),
|
| 111 |
+
"Creature / Boss": ("subject", "monster design reference sheet"),
|
| 112 |
+
"Prop / Weapon": ("subject", "prop design reference sheet"),
|
| 113 |
+
"Environment": ("environment", "environment design reference sheet"),
|
| 114 |
+
}
|
| 115 |
+
|
| 116 |
+
CAMERAS = {
|
| 117 |
+
"Third-person · over-the-shoulder": {
|
| 118 |
+
"genre": "third-person",
|
| 119 |
+
"summary": "tracked continuously by an over-the-shoulder spring-arm camera",
|
| 120 |
+
"aesthetic": "third-person over-the-shoulder spring-arm camera tracking",
|
| 121 |
+
"opening": "an over-the-shoulder camera locked 3.0 meters directly behind",
|
| 122 |
+
"continues": (
|
| 123 |
+
"The camera holds its over-the-back angle and adjusts its spring-arm distance with every movement, "
|
| 124 |
+
"with hit-impulse micro-shakes on each impact."
|
| 125 |
+
),
|
| 126 |
+
},
|
| 127 |
+
"Third-person · orbiting combat camera": {
|
| 128 |
+
"genre": "third-person",
|
| 129 |
+
"summary": "tracked by an orbiting third-person combat camera",
|
| 130 |
+
"aesthetic": "orbiting third-person combat camera with a dynamic spring-arm distance",
|
| 131 |
+
"opening": "a third-person combat camera orbiting 4.0 meters around",
|
| 132 |
+
"continues": (
|
| 133 |
+
"The camera orbits around the action and tightens its distance on every impact, keeping the combat "
|
| 134 |
+
"stance centred in frame."
|
| 135 |
+
),
|
| 136 |
+
},
|
| 137 |
+
"First-person POV (FPV)": {
|
| 138 |
+
"genre": "first-person",
|
| 139 |
+
"summary": "shown entirely in first-person POV with FPV framing",
|
| 140 |
+
"aesthetic": "first-person POV (FPV) camera with weapon-in-hand framing",
|
| 141 |
+
"opening": "a first-person POV camera looking out through the eyes of",
|
| 142 |
+
"continues": (
|
| 143 |
+
"The camera stays in first-person POV throughout, with FPV head-bob, fast view whip-pans and "
|
| 144 |
+
"weapon-in-hand framing at the bottom of frame."
|
| 145 |
+
),
|
| 146 |
+
},
|
| 147 |
+
"Whip-pan: third-person → first-person": {
|
| 148 |
+
"genre": "third-person",
|
| 149 |
+
"summary": "starting over-the-shoulder and whip-panning into first-person POV",
|
| 150 |
+
"aesthetic": "third-person to first-person perspective transition driven by a fast camera whip-pan",
|
| 151 |
+
"opening": "an over-the-shoulder camera locked 3.0 meters behind",
|
| 152 |
+
"continues": (
|
| 153 |
+
"Mid-sequence the camera whip-pans forward into a first-person POV view and holds it to the end of "
|
| 154 |
+
"the shot."
|
| 155 |
+
),
|
| 156 |
+
},
|
| 157 |
+
}
|
| 158 |
+
DEFAULT_CAMERA = "Third-person · over-the-shoulder"
|
| 159 |
+
|
| 160 |
+
HUD_CLAUSE = (
|
| 161 |
+
", and transparent combat HUD elements in the top-right corner: a target-lock UI reticle, a Boss health bar, "
|
| 162 |
+
"floating damage text UI and QTE UI prompts"
|
| 163 |
+
)
|
| 164 |
+
NO_HUD_CLAUSE = ", and a clean cinematic frame with no UI overlays"
|
| 165 |
+
DEFAULT_SOUNDSCAPE = (
|
| 166 |
+
"Footsteps on wet stone, weapon impacts and metallic blade clashes, monstrous roars, thruster and dash bursts, "
|
| 167 |
+
"and crisp action RPG combat UI audio effects."
|
| 168 |
+
)
|
| 169 |
+
DEFAULT_MUSIC = (
|
| 170 |
+
"An intense, high-tempo epic battle track, orchestral-electronic, with heavy industrial percussion and a low "
|
| 171 |
+
"pulsing synth bass that swells through the sequence."
|
| 172 |
+
)
|
| 173 |
+
|
| 174 |
+
|
| 175 |
+
def snap_frames(seconds: float) -> int:
|
| 176 |
+
"""The frame count MiniMax-H3's video VAE can decode: the next `17 * n + 5` at 24 fps."""
|
| 177 |
+
frames = max(1, round(float(seconds) * FPS))
|
| 178 |
+
while frames % FRAMES_PER_CHUNK != LATENTS_PER_CHUNK:
|
| 179 |
+
frames += 1
|
| 180 |
+
return frames
|
| 181 |
+
|
| 182 |
+
|
| 183 |
+
def lower_duration_floor(seconds: float = MIN_DURATION) -> None:
|
| 184 |
+
"""Let the pipeline generate below its 5 s floor. 56 frames (2.33 s) is fine on this checkpoint."""
|
| 185 |
+
from diffusers.modular_pipelines.minimax_h3.modular_pipeline import MiniMaxH3ModularPipeline
|
| 186 |
+
|
| 187 |
+
MiniMaxH3ModularPipeline.min_duration = property(lambda self: float(seconds))
|
| 188 |
+
|
| 189 |
+
|
| 190 |
+
def video_latent_frames(num_frames: int) -> int:
|
| 191 |
+
"""`17 * n + 5` frames become `5 * n + 2` video latents."""
|
| 192 |
+
return 5 * ((num_frames - LATENTS_PER_CHUNK) // FRAMES_PER_CHUNK) + 2
|
| 193 |
+
|
| 194 |
+
|
| 195 |
+
def target_rows(height: int, width: int, num_frames: int) -> int:
|
| 196 |
+
"""The generated rows of the packed sequence: video patched `(1, 2, 2)`, plus two audio rows per latent."""
|
| 197 |
+
video = video_latent_frames(num_frames) * (height // CANVAS_MULTIPLE) * (width // CANVAS_MULTIPLE)
|
| 198 |
+
return video + round(num_frames / FPS * AUDIO_LATENTS_PER_SECOND) * AUDIO_CHANNELS
|
| 199 |
+
|
| 200 |
+
|
| 201 |
+
def reference_rows(image_paths: list[str]) -> int:
|
| 202 |
+
"""The rows the image reference blocks add, from metadata alone — no decode.
|
| 203 |
+
|
| 204 |
+
A reference image is resized to a 2048-pixel short edge and encoded as a single frame, so a squarer reference is
|
| 205 |
+
a cheaper one.
|
| 206 |
+
"""
|
| 207 |
+
from PIL import Image
|
| 208 |
+
|
| 209 |
+
rows = 0
|
| 210 |
+
for path in image_paths:
|
| 211 |
+
if not path:
|
| 212 |
+
continue
|
| 213 |
+
try:
|
| 214 |
+
width, height = Image.open(path).size
|
| 215 |
+
except Exception:
|
| 216 |
+
width, height = 1024, 1024
|
| 217 |
+
scale = REFERENCE_IMAGE_SHORT_EDGE / min(width, height)
|
| 218 |
+
resolved = [
|
| 219 |
+
max(CANVAS_MULTIPLE, round(edge * scale / CANVAS_MULTIPLE) * CANVAS_MULTIPLE) for edge in (height, width)
|
| 220 |
+
]
|
| 221 |
+
rows += (resolved[0] // CANVAS_MULTIPLE) * (resolved[1] // CANVAS_MULTIPLE)
|
| 222 |
+
return rows
|
| 223 |
+
|
| 224 |
+
|
| 225 |
+
def denoise_seconds(sequence: int, steps: int) -> float:
|
| 226 |
+
return int(steps) * (STEP_LINEAR * sequence + STEP_QUADRATIC * sequence**2) * SAFETY
|
| 227 |
+
|
| 228 |
+
|
| 229 |
+
def get_duration(prompt_embeds, text_token_tags, image_paths, height, width, num_frames, steps, weight, seed, **_):
|
| 230 |
+
"""Seconds of GPU to reserve for one request, from the packed sequence it is about to denoise."""
|
| 231 |
+
sequence = int(text_token_tags.shape[0]) + reference_rows(image_paths) + target_rows(height, width, num_frames)
|
| 232 |
+
denoise = denoise_seconds(sequence, steps)
|
| 233 |
+
encode = 5 + reference_rows(image_paths) * 1e-3
|
| 234 |
+
decode = DECODE_BASE + DECODE_PER_DEFAULT_CANVAS * (height * width * num_frames) / DEFAULT_CANVAS_PIXELS
|
| 235 |
+
total = PLACEMENT_ALLOWANCE + encode + denoise + decode + 10
|
| 236 |
+
duration = max(MIN_GPU_DURATION, min(MAX_GPU_DURATION, int(total)))
|
| 237 |
+
print(f"[{VERSION}] S={sequence} -> reserving {duration}s ({denoise:.0f}s of denoise at {steps} steps)", flush=True)
|
| 238 |
+
return duration
|
| 239 |
+
|
| 240 |
+
|
| 241 |
+
def estimate_label(canvas, duration, steps, *image_paths) -> str:
|
| 242 |
+
"""The same estimate, phrased for the UI, so the cost of a canvas / duration / steps choice is visible up front."""
|
| 243 |
+
height, width = CANVASES.get(canvas, CANVASES[DEFAULT_CANVAS])
|
| 244 |
+
num_frames = snap_frames(duration)
|
| 245 |
+
paths = [path for path in image_paths if path]
|
| 246 |
+
refs = reference_rows(paths)
|
| 247 |
+
sequence = ESTIMATED_TEXT_ROWS + ESTIMATED_VISION_ROWS_PER_IMAGE * len(paths) + refs
|
| 248 |
+
sequence += target_rows(height, width, num_frames)
|
| 249 |
+
seconds = int(PLACEMENT_ALLOWANCE + 5 + denoise_seconds(sequence, steps) + DECODE_BASE)
|
| 250 |
+
return (
|
| 251 |
+
f"{width}x{height} · {num_frames} frames ({num_frames / FPS:.2f} s) · {int(steps)} steps · "
|
| 252 |
+
f"{len(paths)} reference{'' if len(paths) == 1 else 's'} → roughly "
|
| 253 |
+
f"**{seconds // 60}m {seconds % 60:02d}s** of GPU time"
|
| 254 |
+
)
|
| 255 |
+
|
| 256 |
+
|
| 257 |
+
# ── The LoRA, onto the diffusers port ────────────────────────────────────────
|
| 258 |
+
#
|
| 259 |
+
# `Minimax-h3_Third_person_view.safetensors` is an ai-toolkit export: 416 tensors, rank 64, no `.alpha`, keys
|
| 260 |
+
# `diffusion_model.blocks.N.{attn.qkv_proj,attn.out_proj,mlp.fc1,mlp.fc2}.lora_{A,B}.weight` plus the same four under
|
| 261 |
+
# `token_refiner.blocks.N`, and `ss_base_model_version: minimax_h3_ref2va` in its metadata. diffusers serves the
|
| 262 |
+
# *converted* port, so every name — and, for two of them, the row layout of `lora_B` — has to be pushed through the
|
| 263 |
+
# same transforms `scripts/convert_minimax_h3_to_diffusers.py` applied to the base weights. This reproduces
|
| 264 |
+
# `_convert_non_diffusers_minimax_h3_lora_to_diffusers`:
|
| 265 |
+
#
|
| 266 |
+
# * `blocks.` -> `transformer_blocks.`, `token_refiner.blocks.` -> `token_refiner.refiner_blocks.`,
|
| 267 |
+
# `attn.out_proj` -> `attn.to_out.0`, `mlp.fc2` -> `ff.net.2` (pure renames),
|
| 268 |
+
# * `mlp.fc1` -> `ff.net.0.proj` with its two fused halves swapped: the reference computes `fc2(silu(gate) * value)`
|
| 269 |
+
# from a fused `[gate; value]` while diffusers' `SwiGLU` computes `value * silu(gate)` from a fused
|
| 270 |
+
# `[value; gate]`, so the halves trade places,
|
| 271 |
+
# * `attn.qkv_proj` -> `to_q` / `to_k` / `to_v`: split `lora_B`'s 21504 rows into **contiguous** thirds of 7168
|
| 272 |
+
# (`num_attention_heads * attention_head_dim`).
|
| 273 |
+
#
|
| 274 |
+
# That last split is where MiniMax-H3 LoRAs diverge. DiffSynth-Studio exports run the raw checkpoint's *per-head
|
| 275 |
+
# interleaved* fused QKV and have to be de-interleaved before the split; they are identified by peft's `.default.`
|
| 276 |
+
# infix over these names. ai-toolkit exports under `diffusion_model.` are already `[q_all; k_all; v_all]` and must
|
| 277 |
+
# **not** be reordered — de-interleaving this file anyway would scatter each head's q/k/v across all three
|
| 278 |
+
# projections and turn the adapter into structured noise on all 50 blocks.
|
| 279 |
+
#
|
| 280 |
+
# Row transforms only ever touch `lora_B`, so the three attention projections share one `lora_A`: the rows of `B @ A`
|
| 281 |
+
# are the rows of `B`, which makes the split exact rather than an approximation.
|
| 282 |
+
|
| 283 |
+
|
| 284 |
+
def _lora_target_name(source_name: str) -> str:
|
| 285 |
+
"""Map an original-checkpoint module path to the diffusers one (no `.lora_A/B.*` suffix)."""
|
| 286 |
+
if source_name.startswith("token_refiner.blocks."):
|
| 287 |
+
return source_name.replace("token_refiner.blocks.", "token_refiner.refiner_blocks.", 1)
|
| 288 |
+
if source_name.startswith("blocks."):
|
| 289 |
+
return source_name.replace("blocks.", "transformer_blocks.", 1)
|
| 290 |
+
return source_name
|
| 291 |
+
|
| 292 |
+
|
| 293 |
+
def _lora_modules(name: str, a_weight, b_weight, inner_dim: int):
|
| 294 |
+
"""Yield `(diffusers_module_path, lora_A, lora_B)` for one original module path."""
|
| 295 |
+
target = _lora_target_name(name)
|
| 296 |
+
|
| 297 |
+
if target.endswith(".attn.qkv_proj"):
|
| 298 |
+
prefix = target.removesuffix("qkv_proj")
|
| 299 |
+
for kind, part in zip(("q", "k", "v"), b_weight.split(inner_dim, dim=0)):
|
| 300 |
+
yield f"{prefix}to_{kind}", a_weight, part.contiguous()
|
| 301 |
+
elif target.endswith(".mlp.fc1"):
|
| 302 |
+
gate, value = b_weight.chunk(2, dim=0)
|
| 303 |
+
yield target.replace(".mlp.fc1", ".ff.net.0.proj"), a_weight, torch.cat([value, gate]).contiguous()
|
| 304 |
+
elif target.endswith(".mlp.fc2"):
|
| 305 |
+
yield target.replace(".mlp.fc2", ".ff.net.2"), a_weight, b_weight
|
| 306 |
+
elif target.endswith(".attn.out_proj"):
|
| 307 |
+
yield target.replace(".attn.out_proj", ".attn.to_out.0"), a_weight, b_weight
|
| 308 |
+
else:
|
| 309 |
+
raise ValueError(
|
| 310 |
+
f"unexpected LoRA target `{name}`: this adapter is documented as attention + feed-forward only, so a "
|
| 311 |
+
"new module type means the checkpoint changed"
|
| 312 |
+
)
|
| 313 |
+
|
| 314 |
+
|
| 315 |
+
def build_lora_state_dict(transformer) -> tuple[dict, int, int]:
|
| 316 |
+
"""Download the LoRA and remap it into a PEFT-format state dict for the diffusers transformer.
|
| 317 |
+
|
| 318 |
+
Every target is validated against the transformer's own parameter shapes, and an unresolved one is fatal: a
|
| 319 |
+
silently dropped target means the name mapping is wrong and the Space would serve a half-applied adapter that
|
| 320 |
+
still *looks* like it worked.
|
| 321 |
+
"""
|
| 322 |
+
from huggingface_hub import hf_hub_download
|
| 323 |
+
from safetensors.torch import load_file
|
| 324 |
+
|
| 325 |
+
raw = load_file(hf_hub_download(LORA_REPO, LORA_FILE))
|
| 326 |
+
|
| 327 |
+
prefix, suffix_a, suffix_b = "diffusion_model.", ".lora_A.weight", ".lora_B.weight"
|
| 328 |
+
unexpected = [key for key in raw if not (key.startswith(prefix) and key.endswith((suffix_a, suffix_b)))]
|
| 329 |
+
if unexpected:
|
| 330 |
+
raise ValueError(f"{LORA_FILE} holds {len(unexpected)} unexpected tensors, e.g. {unexpected[:5]}")
|
| 331 |
+
bases = sorted({key[len(prefix) : -len(suffix_a)] for key in raw if key.endswith(suffix_a)})
|
| 332 |
+
if not bases:
|
| 333 |
+
raise ValueError(f"No `{prefix}*{suffix_a}` / `{suffix_b}` pairs found in {LORA_FILE}")
|
| 334 |
+
|
| 335 |
+
ranks = set()
|
| 336 |
+
for name in bases:
|
| 337 |
+
if f"{prefix}{name}{suffix_b}" not in raw:
|
| 338 |
+
raise ValueError(f"LoRA is missing the lora_B twin of {prefix}{name}{suffix_a}")
|
| 339 |
+
ranks.add(raw[f"{prefix}{name}{suffix_a}"].shape[0])
|
| 340 |
+
if len(ranks) != 1:
|
| 341 |
+
raise ValueError(f"LoRA mixes ranks {sorted(ranks)}; this loader assumes a single rank")
|
| 342 |
+
rank = ranks.pop()
|
| 343 |
+
|
| 344 |
+
inner_dim = transformer.config.num_attention_heads * transformer.config.attention_head_dim
|
| 345 |
+
base_shapes = {key: tuple(value.shape) for key, value in transformer.state_dict().items()}
|
| 346 |
+
|
| 347 |
+
state_dict: dict[str, torch.Tensor] = {}
|
| 348 |
+
missed: list[str] = []
|
| 349 |
+
for name in bases:
|
| 350 |
+
a_weight = raw[f"{prefix}{name}{suffix_a}"]
|
| 351 |
+
b_weight = raw[f"{prefix}{name}{suffix_b}"]
|
| 352 |
+
for module, a_part, b_part in _lora_modules(name, a_weight, b_weight, inner_dim):
|
| 353 |
+
base = base_shapes.get(f"{module}.weight")
|
| 354 |
+
if base is None:
|
| 355 |
+
missed.append(module)
|
| 356 |
+
continue
|
| 357 |
+
# `W` is [out, in]; the adapter must be `lora_B` [out, r] @ `lora_A` [r, in].
|
| 358 |
+
if (b_part.shape[0], a_part.shape[1]) != base:
|
| 359 |
+
raise ValueError(
|
| 360 |
+
f"LoRA delta for `{module}` would be {(b_part.shape[0], a_part.shape[1])}, "
|
| 361 |
+
f"base weight is {base}"
|
| 362 |
+
)
|
| 363 |
+
state_dict[f"{module}.lora_A.weight"] = a_part
|
| 364 |
+
state_dict[f"{module}.lora_B.weight"] = b_part
|
| 365 |
+
if missed:
|
| 366 |
+
raise ValueError(
|
| 367 |
+
f"{len(missed)} LoRA targets matched no transformer weight, e.g. {missed[:5]}. "
|
| 368 |
+
"The LoRA and the diffusers transformer disagree on module naming."
|
| 369 |
+
)
|
| 370 |
+
return state_dict, rank, len(bases)
|
| 371 |
+
|
| 372 |
+
|
| 373 |
+
def load_and_apply_lora(transformer) -> str:
|
| 374 |
+
"""Attach the LoRA as a PEFT adapter so its strength stays a per-request knob."""
|
| 375 |
+
state_dict, rank, targets = build_lora_state_dict(transformer)
|
| 376 |
+
# `prefix=None` because the keys are already the transformer's own module paths, and no `network_alphas` because
|
| 377 |
+
# the file carries no `.alpha` tensors — PEFT then sets alpha == rank, i.e. the adapter's intrinsic scale is 1.0
|
| 378 |
+
# and the `set_adapters` weight *is* the strength the card talks about.
|
| 379 |
+
transformer.load_lora_adapter(state_dict, prefix=None, adapter_name=ADAPTER)
|
| 380 |
+
transformer.set_adapters([ADAPTER], [DEFAULT_WEIGHT])
|
| 381 |
+
return (
|
| 382 |
+
f"LoRA attached · {targets} checkpoint targets -> {len(state_dict) // 2} diffusers modules, rank {rank}, "
|
| 383 |
+
f"alpha == rank · default weight {DEFAULT_WEIGHT:g} · {LORA_FILE}"
|
| 384 |
+
)
|
| 385 |
+
|
| 386 |
+
|
| 387 |
+
# ── Model loading ────────────────────────────────────────────────────────────
|
| 388 |
+
|
| 389 |
+
PIPE = None
|
| 390 |
+
LOAD_ERROR: str | None = None
|
| 391 |
+
LORA_STATUS: str | None = None
|
| 392 |
+
|
| 393 |
+
|
| 394 |
+
def load_models() -> str | None:
|
| 395 |
+
"""Load the denoising half at startup, but *not* onto the card — see `PLACEMENT`.
|
| 396 |
+
|
| 397 |
+
`MiniMaxH3Ref2VAGeneratorBlocks` declares `transformer_ref`, `vae`, `audio_vae`, the two schedulers and
|
| 398 |
+
`video_processor`, so `load_components` fetches exactly those subfolders — `text_encoder/` and the `transformer/`
|
| 399 |
+
partition are never touched. Both autoencoders carry `_keep_in_fp32_modules` over every module and stay float32:
|
| 400 |
+
a bfloat16 audio VAE decodes the soundtrack roughly 20 dB too quiet.
|
| 401 |
+
"""
|
| 402 |
+
global PIPE, LOAD_ERROR, LORA_STATUS
|
| 403 |
+
|
| 404 |
+
if PIPE is not None or LOAD_ERROR is not None:
|
| 405 |
+
return LOAD_ERROR
|
| 406 |
+
|
| 407 |
+
started = time.time()
|
| 408 |
+
try:
|
| 409 |
+
from diffusers import ComponentsManager
|
| 410 |
+
|
| 411 |
+
from h3_split_blocks import MiniMaxH3Ref2VAGeneratorBlocks
|
| 412 |
+
|
| 413 |
+
lower_duration_floor()
|
| 414 |
+
blocks = MiniMaxH3Ref2VAGeneratorBlocks()
|
| 415 |
+
print(f"[{VERSION}] loading {[c.name for c in blocks.expected_components]} from {MODEL_REPO} ...", flush=True)
|
| 416 |
+
pipe = blocks.init_pipeline(MODEL_REPO, components_manager=ComponentsManager(), collection="h3")
|
| 417 |
+
# The pruned DiT is served as remote code (`transformer_ref/modeling_minimax_h3_pruned.py`, reached through
|
| 418 |
+
# the `AutoModel` type hint in `modular_model_index.json`). `load_components` forwards `trust_remote_code`
|
| 419 |
+
# only to components that live in the pipeline's own repo, so the two VAEs and the schedulers — which
|
| 420 |
+
# `modular_model_index.json` still points at `MiniMaxAI/MiniMax-H3` — never see it.
|
| 421 |
+
pipe.load_components(dtype=torch.bfloat16, trust_remote_code=True)
|
| 422 |
+
|
| 423 |
+
# Both VAEs explicitly, and before the transformer: `set_attention_backend` also sets the registry's *global*
|
| 424 |
+
# backend, and the float32 audio VAE has no cuDNN kernel.
|
| 425 |
+
pipe.vae.set_attention_backend("native")
|
| 426 |
+
pipe.audio_vae.set_attention_backend("native")
|
| 427 |
+
pipe.transformer_ref.set_attention_backend(ATTENTION)
|
| 428 |
+
|
| 429 |
+
# Before any GPU placement, so the adapter's parameters travel with the transformer. No AoTI package on this
|
| 430 |
+
# Space on purpose: a compiled block graph is captured around the base linears and would silently bypass the
|
| 431 |
+
# LoRA layers PEFT injects.
|
| 432 |
+
LORA_STATUS = load_and_apply_lora(pipe.transformer_ref)
|
| 433 |
+
print(f"[{VERSION}] {LORA_STATUS}", flush=True)
|
| 434 |
+
|
| 435 |
+
import h3_fbc
|
| 436 |
+
|
| 437 |
+
print(f"[h3-fbc] {h3_fbc.status()}", flush=True)
|
| 438 |
+
|
| 439 |
+
PIPE = pipe
|
| 440 |
+
print(f"[{VERSION}] ready in {time.time() - started:.0f}s", flush=True)
|
| 441 |
+
except Exception as error:
|
| 442 |
+
traceback.print_exc()
|
| 443 |
+
LOAD_ERROR = (
|
| 444 |
+
f"**Loading `{MODEL_REPO}` failed** after {time.time() - started:.0f}s: "
|
| 445 |
+
f"`{type(error).__name__}: {error}`"
|
| 446 |
+
)
|
| 447 |
+
return LOAD_ERROR
|
| 448 |
+
|
| 449 |
+
|
| 450 |
+
# ── Prompt composition ───────────────────────────────────────────────────────
|
| 451 |
+
|
| 452 |
+
|
| 453 |
+
def collect_slots(slots) -> list[tuple[str, str, str]]:
|
| 454 |
+
"""The `(path, role, description)` references of a request, in the order the model reads them.
|
| 455 |
+
|
| 456 |
+
That order numbers `<Picture N>` and advances the shared audio/video rotary clock, so the same references in a
|
| 457 |
+
different order are a different request.
|
| 458 |
+
"""
|
| 459 |
+
return [(path, role, (detail or "").strip()) for path, role, detail in slots if path]
|
| 460 |
+
|
| 461 |
+
|
| 462 |
+
def compose_prompt(references, action, camera, hud, soundscape, music, seconds) -> str:
|
| 463 |
+
"""Assemble MiniMax-H3's structured document in the layout the LoRA's card publishes."""
|
| 464 |
+
action = (action or "").strip().rstrip(".")
|
| 465 |
+
spec = CAMERAS.get(camera, CAMERAS[DEFAULT_CAMERA])
|
| 466 |
+
|
| 467 |
+
entries, subjects, environments = [], 0, 0
|
| 468 |
+
for index, (_, role, detail) in enumerate(references, start=1):
|
| 469 |
+
kind, sheet = ROLES.get(role, ROLES["Character"])
|
| 470 |
+
if kind == "subject":
|
| 471 |
+
subjects += 1
|
| 472 |
+
label = f"<Subject {subjects}>"
|
| 473 |
+
else:
|
| 474 |
+
environments += 1
|
| 475 |
+
label = f"<Environment {environments}>"
|
| 476 |
+
entries.append((label, index, detail, sheet))
|
| 477 |
+
|
| 478 |
+
subject = next((label for label, _, _, sheet in entries if "environment" not in sheet), "the player character")
|
| 479 |
+
environment = next((label for label, _, _, sheet in entries if "environment" in sheet), None)
|
| 480 |
+
place = f" in {environment}" if environment else ""
|
| 481 |
+
|
| 482 |
+
definitions = [
|
| 483 |
+
f"{label} is {detail} from <Picture {index}>."
|
| 484 |
+
if detail
|
| 485 |
+
else f"{label} is the subject shown in <Picture {index}>."
|
| 486 |
+
for label, index, detail, _ in entries
|
| 487 |
+
]
|
| 488 |
+
definitions += [f"<Picture {index}> is the {sheet} for {label}." for label, index, _, sheet in entries]
|
| 489 |
+
|
| 490 |
+
retention = [
|
| 491 |
+
f"{label} (appears in [Shot 1]): fully_preserved - the design, colours and silhouette from "
|
| 492 |
+
f"<Picture {index}> are fully preserved."
|
| 493 |
+
for label, index, _, _ in entries
|
| 494 |
+
]
|
| 495 |
+
|
| 496 |
+
sections = [
|
| 497 |
+
"subject_definitions:\n" + "\n".join(definitions),
|
| 498 |
+
(
|
| 499 |
+
"summary:\n"
|
| 500 |
+
f"[reference generation] A {float(seconds):g}-second photorealistic {spec['genre']} action RPG "
|
| 501 |
+
f"gameplay sequence{place}, where {subject} {action}, {spec['summary']}."
|
| 502 |
+
),
|
| 503 |
+
"retention_analysis:\n" + "\n".join(retention),
|
| 504 |
+
(
|
| 505 |
+
"detailed_description:\n"
|
| 506 |
+
f"The target video features a photorealistic 3D {spec['genre']} action RPG gameplay aesthetic rendered "
|
| 507 |
+
f"in Unreal Engine with real-time game mechanics, {spec['aesthetic']}, deep depth of field"
|
| 508 |
+
f"{HUD_CLAUSE if hud else NO_HUD_CLAUSE}.\n"
|
| 509 |
+
f"[Shot 1] The video opens with {spec['opening']} {subject}{place}. {action.capitalize()}. "
|
| 510 |
+
f"{spec['continues']}"
|
| 511 |
+
),
|
| 512 |
+
]
|
| 513 |
+
if (soundscape or "").strip():
|
| 514 |
+
sections.append("overall_soundscape:\n" + soundscape.strip())
|
| 515 |
+
if (music or "").strip():
|
| 516 |
+
sections.append("non_diegetic_music:\n" + music.strip())
|
| 517 |
+
return "\n\n".join(sections)
|
| 518 |
+
|
| 519 |
+
|
| 520 |
+
# ── Guard, conditioner ───────────────────────────────────────────────────────
|
| 521 |
+
|
| 522 |
+
|
| 523 |
+
def check_prompt(prompt: str) -> None:
|
| 524 |
+
"""The NCII guard. Every request here carries an uploaded image, so every request is checked; it runs before the
|
| 525 |
+
conditioner call and the denoise booking, so a refused prompt costs no GPU time on either half."""
|
| 526 |
+
import ncii_guard
|
| 527 |
+
|
| 528 |
+
try:
|
| 529 |
+
flag = ncii_guard.classify(prompt)
|
| 530 |
+
except Exception as error:
|
| 531 |
+
traceback.print_exc()
|
| 532 |
+
raise gr.Error(f"The content filter is unavailable (`{type(error).__name__}`), so nothing was run.")
|
| 533 |
+
if flag["label"] == "ncii":
|
| 534 |
+
print(f"[guard] prompt refused (ncii {flag['score']:.2f})", flush=True)
|
| 535 |
+
raise gr.Error("This prompt was flagged by a content filter and wasn't run.")
|
| 536 |
+
|
| 537 |
+
|
| 538 |
+
@cache
|
| 539 |
+
def conditioner():
|
| 540 |
+
"""The other half, over the gradio API. `gradio_client` attaches the caller's own ZeroGPU token per call, so the
|
| 541 |
+
conditioner's booking is billed to whoever asked for the video."""
|
| 542 |
+
from gradio_client import Client
|
| 543 |
+
|
| 544 |
+
return Client(CONDITIONER_SPACE)
|
| 545 |
+
|
| 546 |
+
|
| 547 |
+
def encode_remote(prompt, image_paths, canvas, num_frames):
|
| 548 |
+
"""`/encode_ref2va` on the conditioner Space: a safetensors file holding `prompt_embeds` + `text_token_tags`, with
|
| 549 |
+
the resolved `height` / `width` / `num_frames` in its metadata, plus the plan.
|
| 550 |
+
|
| 551 |
+
`canvas` is the label. `media` and `kinds` are parallel and ordered; the references go over because `ref2va`'s
|
| 552 |
+
presentation puts a vision block in front of the prompt for every image.
|
| 553 |
+
"""
|
| 554 |
+
from gradio_client import handle_file
|
| 555 |
+
from safetensors import safe_open
|
| 556 |
+
|
| 557 |
+
path, plan = conditioner().predict(
|
| 558 |
+
prompt=prompt,
|
| 559 |
+
media=[handle_file(image) for image in image_paths],
|
| 560 |
+
kinds=",".join("image" for _ in image_paths),
|
| 561 |
+
canvas=canvas,
|
| 562 |
+
num_frames=num_frames,
|
| 563 |
+
rewrite_prompt=False,
|
| 564 |
+
api_name="/encode_ref2va",
|
| 565 |
+
)
|
| 566 |
+
with safe_open(path, framework="pt") as handle:
|
| 567 |
+
return handle.get_tensor("prompt_embeds"), handle.get_tensor("text_token_tags"), handle.metadata(), plan
|
| 568 |
+
|
| 569 |
+
|
| 570 |
+
# ── Inference ────────────────────────────────────────────────────────────────
|
| 571 |
+
|
| 572 |
+
|
| 573 |
+
@spaces.GPU(duration=get_duration, size=GPU_SIZE)
|
| 574 |
+
def _generate(prompt_embeds, text_token_tags, image_paths, height, width, num_frames, steps, weight, seed):
|
| 575 |
+
"""The only thing on GPU time: the reference encoder, the packed-sequence denoise loop and the decoders.
|
| 576 |
+
|
| 577 |
+
References cross as paths and are decoded here; only the generated outputs come back. A `@spaces.GPU` argument
|
| 578 |
+
crosses a process boundary by pickling, and expanded reference frames are large.
|
| 579 |
+
"""
|
| 580 |
+
from diffusers.modular_pipelines.minimax_h3 import MiniMaxH3ImageReference
|
| 581 |
+
|
| 582 |
+
import h3_fbc
|
| 583 |
+
|
| 584 |
+
if PLACEMENT == "lazy":
|
| 585 |
+
PIPE.to("cuda")
|
| 586 |
+
|
| 587 |
+
# The whole reason the adapter is not folded into the weights: strength is a per-request scale on the LoRA layers.
|
| 588 |
+
PIPE.transformer_ref.set_adapters([ADAPTER], [float(weight)])
|
| 589 |
+
|
| 590 |
+
with h3_fbc.enabled(PIPE.transformer_ref, steps=int(steps)):
|
| 591 |
+
state = PIPE(
|
| 592 |
+
prompt_embeds=prompt_embeds.to("cuda"),
|
| 593 |
+
text_token_tags=text_token_tags,
|
| 594 |
+
references=[MiniMaxH3ImageReference.from_file(path) for path in image_paths],
|
| 595 |
+
height=height,
|
| 596 |
+
width=width,
|
| 597 |
+
num_frames=num_frames,
|
| 598 |
+
num_inference_steps=int(steps),
|
| 599 |
+
generator=torch.Generator("cpu").manual_seed(int(seed)),
|
| 600 |
+
)
|
| 601 |
+
return state.get("videos")[0], state.get("audio")[0].cpu(), state.get("sampling_rate")
|
| 602 |
+
|
| 603 |
+
|
| 604 |
+
def generate(
|
| 605 |
+
# The first nine are the columns `gr.Examples` varies, and they lead the signature for that reason: an example row
|
| 606 |
+
# is applied to `inputs` positionally. Every parameter has a default, which is what lets a nine-column row call
|
| 607 |
+
# this at all.
|
| 608 |
+
action="",
|
| 609 |
+
picture_1=None,
|
| 610 |
+
role_1="Character",
|
| 611 |
+
detail_1="",
|
| 612 |
+
picture_2=None,
|
| 613 |
+
role_2="Environment",
|
| 614 |
+
detail_2="",
|
| 615 |
+
camera=DEFAULT_CAMERA,
|
| 616 |
+
hud=True,
|
| 617 |
+
picture_3=None,
|
| 618 |
+
role_3="Creature / Boss",
|
| 619 |
+
detail_3="",
|
| 620 |
+
canvas=DEFAULT_CANVAS,
|
| 621 |
+
duration=5,
|
| 622 |
+
steps=DEFAULT_STEPS,
|
| 623 |
+
weight=DEFAULT_WEIGHT,
|
| 624 |
+
seed=DEFAULT_SEED,
|
| 625 |
+
soundscape=DEFAULT_SOUNDSCAPE,
|
| 626 |
+
music=DEFAULT_MUSIC,
|
| 627 |
+
override="",
|
| 628 |
+
progress=gr.Progress(track_tqdm=True),
|
| 629 |
+
):
|
| 630 |
+
"""One request: compose the structured prompt, condition it remotely, denoise here."""
|
| 631 |
+
if LOAD_ERROR:
|
| 632 |
+
raise gr.Error(LOAD_ERROR)
|
| 633 |
+
if PIPE is None:
|
| 634 |
+
raise gr.Error("The denoiser is still loading.")
|
| 635 |
+
|
| 636 |
+
from diffusers.utils import encode_video
|
| 637 |
+
|
| 638 |
+
references = collect_slots(
|
| 639 |
+
[(picture_1, role_1, detail_1), (picture_2, role_2, detail_2), (picture_3, role_3, detail_3)]
|
| 640 |
+
)
|
| 641 |
+
if not references:
|
| 642 |
+
raise gr.Error("Add at least one reference sheet — a character, an environment or a creature.")
|
| 643 |
+
if not (override or "").strip() and not (action or "").strip():
|
| 644 |
+
raise gr.Error("Write one line of action, or paste a full structured prompt in the advanced options.")
|
| 645 |
+
|
| 646 |
+
num_frames = snap_frames(duration)
|
| 647 |
+
prompt = (override or "").strip() or compose_prompt(
|
| 648 |
+
references, action, camera, hud, soundscape, music, num_frames / FPS
|
| 649 |
+
)
|
| 650 |
+
check_prompt(prompt)
|
| 651 |
+
|
| 652 |
+
image_paths = [path for path, _, _ in references]
|
| 653 |
+
progress(0.0, desc="Reading the prompt and the reference sheets ...")
|
| 654 |
+
conditioned = time.time()
|
| 655 |
+
try:
|
| 656 |
+
prompt_embeds, text_token_tags, metadata, plan = encode_remote(prompt, image_paths, canvas, num_frames)
|
| 657 |
+
except gr.Error:
|
| 658 |
+
raise
|
| 659 |
+
except Exception as error:
|
| 660 |
+
# gradio only puts the exception *type* on the wire, so the useful half of a conditioner-side failure is in
|
| 661 |
+
# that Space's logs.
|
| 662 |
+
traceback.print_exc()
|
| 663 |
+
raise gr.Error(
|
| 664 |
+
f"The conditioner ({CONDITIONER_SPACE}) failed with `{type(error).__name__}: {error}`. "
|
| 665 |
+
"Its logs carry the full traceback."
|
| 666 |
+
) from error
|
| 667 |
+
condition_seconds = time.time() - conditioned
|
| 668 |
+
height, width, num_frames = (int(metadata[key]) for key in ("height", "width", "num_frames"))
|
| 669 |
+
|
| 670 |
+
progress(0.1, desc=f"Rendering {num_frames / FPS:.1f} s at {width}x{height} ...")
|
| 671 |
+
started = time.time()
|
| 672 |
+
frames, audio, sampling_rate = _generate(
|
| 673 |
+
prompt_embeds, text_token_tags, image_paths, height, width, num_frames, steps, weight, seed
|
| 674 |
+
)
|
| 675 |
+
generate_seconds = time.time() - started
|
| 676 |
+
|
| 677 |
+
directory = os.path.join(tempfile.gettempdir(), "h3-outputs")
|
| 678 |
+
os.makedirs(directory, exist_ok=True)
|
| 679 |
+
path = os.path.join(directory, f"h3-tpv-{int(time.time() * 1000)}.mp4")
|
| 680 |
+
encode_video(frames, fps=FPS, output_path=path, audio=audio, audio_sample_rate=sampling_rate)
|
| 681 |
+
|
| 682 |
+
print(
|
| 683 |
+
f"[{VERSION}] {len(image_paths)} references · {width}x{height}, {num_frames} frames "
|
| 684 |
+
f"({num_frames / FPS:.3f} s), {int(steps)} steps, weight {float(weight):g} · conditioner "
|
| 685 |
+
f"{condition_seconds:.0f}s ({plan['num_text_tokens']} tokens) · denoise + decode {generate_seconds:.0f}s "
|
| 686 |
+
f"({generate_seconds / int(steps):.1f} s/step) · seed {int(seed)}",
|
| 687 |
+
flush=True,
|
| 688 |
+
)
|
| 689 |
+
return path, prompt
|
| 690 |
+
|
| 691 |
+
|
| 692 |
+
# ── UI ───────────────────────────────────────────────────────────────────────
|
| 693 |
+
|
| 694 |
+
try:
|
| 695 |
+
import ncii_guard
|
| 696 |
+
|
| 697 |
+
ncii_guard.start()
|
| 698 |
+
except Exception as guard_error: # the guard revives itself per request; a cold cache should not fail the boot
|
| 699 |
+
print(f"[guard] start failed ({type(guard_error).__name__}: {guard_error}); it will be retried per request")
|
| 700 |
+
load_models()
|
| 701 |
+
|
| 702 |
+
INTRO = """# MiniMax-H3 · Third-Person View LoRA
|
| 703 |
+
|
| 704 |
+
<div align="center">
|
| 705 |
+
<a href="https://huggingface.co/WarmBloodAban/Minimax_H3_LoRAs" target="_blank" rel="noopener"><strong>[ LoRA ]</strong></a>
|
| 706 |
+
<a href="https://huggingface.co/MiniMaxAI/MiniMax-H3" target="_blank" rel="noopener"><strong>[ base model ]</strong></a>
|
| 707 |
+
<a href="https://huggingface.co/spaces/multimodalart/minimax-h3-reference" target="_blank" rel="noopener"><strong>[ base demo ]</strong></a>
|
| 708 |
+
</div>
|
| 709 |
+
|
| 710 |
+
**Game cutscenes out of design reference sheets.** Label each sheet you upload, write one line of action, pick a
|
| 711 |
+
camera, and this composes MiniMax-H3's structured prompt around the LoRA's own vocabulary — third-person
|
| 712 |
+
over-the-shoulder / spring-arm tracking, first-person POV, whip-pan transitions, target-lock reticles and Boss health
|
| 713 |
+
bars — then renders the clip **with its synchronized soundtrack** in a single denoising pass.
|
| 714 |
+
"""
|
| 715 |
+
|
| 716 |
+
CSS = """
|
| 717 |
+
.main.fillable { max-width: 1250px !important; }
|
| 718 |
+
.dark .gradio-container { color: var(--body-text-color); }
|
| 719 |
+
"""
|
| 720 |
+
|
| 721 |
+
with gr.Blocks(title="MiniMax-H3 · Third-Person View LoRA", theme=gr.themes.Citrus(), css=CSS) as demo:
|
| 722 |
+
gr.Markdown(INTRO)
|
| 723 |
+
if LOAD_ERROR:
|
| 724 |
+
gr.Markdown(LOAD_ERROR)
|
| 725 |
+
|
| 726 |
+
with gr.Row():
|
| 727 |
+
with gr.Column():
|
| 728 |
+
gr.Markdown("### 1 · Reference sheets — they become `<Picture 1..3>`, in this order")
|
| 729 |
+
pictures, roles, details, slot_columns = [], [], [], []
|
| 730 |
+
with gr.Row():
|
| 731 |
+
for index in range(MAX_IMAGE_SLOTS):
|
| 732 |
+
with gr.Column(min_width=180, visible=index < OPEN_IMAGE_SLOTS) as column:
|
| 733 |
+
pictures.append(
|
| 734 |
+
gr.Image(label=f"Picture {index + 1}", type="filepath", height=190, show_label=True)
|
| 735 |
+
)
|
| 736 |
+
roles.append(
|
| 737 |
+
gr.Dropdown(
|
| 738 |
+
label="Role",
|
| 739 |
+
choices=list(ROLES),
|
| 740 |
+
value="Character" if index == 0 else ("Environment" if index == 1 else "Creature / Boss"),
|
| 741 |
+
container=True,
|
| 742 |
+
)
|
| 743 |
+
)
|
| 744 |
+
details.append(
|
| 745 |
+
gr.Textbox(
|
| 746 |
+
label="What's in it",
|
| 747 |
+
lines=2,
|
| 748 |
+
placeholder="the mint-green cat mecha with a tattered beige cape",
|
| 749 |
+
)
|
| 750 |
+
)
|
| 751 |
+
slot_columns.append(column)
|
| 752 |
+
add_slot = gr.Button("+ Add a third reference sheet", size="sm", variant="secondary")
|
| 753 |
+
|
| 754 |
+
action = gr.Textbox(
|
| 755 |
+
label="2 · What happens in the shot",
|
| 756 |
+
lines=2,
|
| 757 |
+
placeholder="dodges a leaping wyvern attack, counterattacks with an energy-blade combo, then parries "
|
| 758 |
+
"a wing-claw swipe",
|
| 759 |
+
)
|
| 760 |
+
with gr.Row():
|
| 761 |
+
camera = gr.Dropdown(label="3 · Camera", choices=list(CAMERAS), value=DEFAULT_CAMERA, scale=3)
|
| 762 |
+
hud = gr.Checkbox(label="Combat HUD overlay", value=True, scale=1)
|
| 763 |
+
run = gr.Button("Render the cutscene", variant="primary")
|
| 764 |
+
estimate = gr.Markdown(estimate_label(DEFAULT_CANVAS, 5, DEFAULT_STEPS, None))
|
| 765 |
+
|
| 766 |
+
with gr.Accordion("Advanced options", open=False):
|
| 767 |
+
weight = gr.Slider(
|
| 768 |
+
label="LoRA weight (the card recommends 0.6–0.85)",
|
| 769 |
+
minimum=WEIGHT_MIN,
|
| 770 |
+
maximum=WEIGHT_MAX,
|
| 771 |
+
step=0.05,
|
| 772 |
+
value=DEFAULT_WEIGHT,
|
| 773 |
+
)
|
| 774 |
+
canvas = gr.Dropdown(label="Canvas", choices=list(CANVASES), value=DEFAULT_CANVAS)
|
| 775 |
+
duration = gr.Slider(
|
| 776 |
+
label="Duration (s)", minimum=MIN_DURATION, maximum=MAX_UI_DURATION, step=1, value=5
|
| 777 |
+
)
|
| 778 |
+
steps = gr.Slider(label="Steps", minimum=10, maximum=40, step=1, value=DEFAULT_STEPS)
|
| 779 |
+
seed = gr.Number(label="Seed", value=DEFAULT_SEED, precision=0)
|
| 780 |
+
soundscape = gr.Textbox(label="Diegetic soundscape", lines=2, value=DEFAULT_SOUNDSCAPE)
|
| 781 |
+
music = gr.Textbox(label="Non-diegetic music", lines=2, value=DEFAULT_MUSIC)
|
| 782 |
+
override = gr.Textbox(
|
| 783 |
+
label="Write the structured prompt myself (overrides everything above)",
|
| 784 |
+
lines=6,
|
| 785 |
+
placeholder="subject_definitions:\n<Subject 1> is ... from <Picture 1>.\n\nsummary:\n"
|
| 786 |
+
"[reference generation] A 10-second photorealistic third-person ...",
|
| 787 |
+
)
|
| 788 |
+
|
| 789 |
+
with gr.Column():
|
| 790 |
+
result = gr.Video(label="Cutscene + soundtrack")
|
| 791 |
+
with gr.Accordion("Structured prompt sent to MiniMax-H3", open=False):
|
| 792 |
+
prompt_view = gr.Textbox(show_label=False, lines=18, interactive=False)
|
| 793 |
+
|
| 794 |
+
open_slots = gr.State(OPEN_IMAGE_SLOTS)
|
| 795 |
+
|
| 796 |
+
def reveal_slot(open_count):
|
| 797 |
+
open_count = min(open_count + 1, MAX_IMAGE_SLOTS)
|
| 798 |
+
return [
|
| 799 |
+
open_count,
|
| 800 |
+
*[gr.update(visible=index < open_count) for index in range(MAX_IMAGE_SLOTS)],
|
| 801 |
+
gr.update(visible=open_count < MAX_IMAGE_SLOTS),
|
| 802 |
+
]
|
| 803 |
+
|
| 804 |
+
add_slot.click(reveal_slot, open_slots, [open_slots, *slot_columns, add_slot], api_name=False)
|
| 805 |
+
|
| 806 |
+
for control in (canvas, duration, steps, *pictures):
|
| 807 |
+
control.change(
|
| 808 |
+
estimate_label,
|
| 809 |
+
[canvas, duration, steps, *pictures],
|
| 810 |
+
estimate,
|
| 811 |
+
show_progress="hidden",
|
| 812 |
+
api_name=False,
|
| 813 |
+
)
|
| 814 |
+
|
| 815 |
+
request = [
|
| 816 |
+
action,
|
| 817 |
+
pictures[0],
|
| 818 |
+
roles[0],
|
| 819 |
+
details[0],
|
| 820 |
+
pictures[1],
|
| 821 |
+
roles[1],
|
| 822 |
+
details[1],
|
| 823 |
+
camera,
|
| 824 |
+
hud,
|
| 825 |
+
pictures[2],
|
| 826 |
+
roles[2],
|
| 827 |
+
details[2],
|
| 828 |
+
canvas,
|
| 829 |
+
duration,
|
| 830 |
+
steps,
|
| 831 |
+
weight,
|
| 832 |
+
seed,
|
| 833 |
+
soundscape,
|
| 834 |
+
music,
|
| 835 |
+
override,
|
| 836 |
+
]
|
| 837 |
+
exampled = request[:9]
|
| 838 |
+
|
| 839 |
+
gr.Examples(
|
| 840 |
+
examples=[
|
| 841 |
+
[
|
| 842 |
+
"dodges a leaping wyvern attack, counterattacks with a glowing energy-blade combo, then parries a "
|
| 843 |
+
"wing-claw swipe in a shower of sparks",
|
| 844 |
+
"examples/katana_character.jpg",
|
| 845 |
+
"Character",
|
| 846 |
+
"the young woman with long black hair in two buns, blue eyes, a black mouth mask, an orange utility "
|
| 847 |
+
"jacket and two white-corded katana hilts rising behind her shoulders",
|
| 848 |
+
"examples/ruined_village.jpg",
|
| 849 |
+
"Environment",
|
| 850 |
+
"the foggy ruined village street with wooden houses, a glowing paper lantern, blood-red spider "
|
| 851 |
+
"lilies and wet stone paving",
|
| 852 |
+
"Third-person · over-the-shoulder",
|
| 853 |
+
True,
|
| 854 |
+
],
|
| 855 |
+
[
|
| 856 |
+
"plants her feet into a combat stance and charges an ultimate finisher as the pale wyvern staggers "
|
| 857 |
+
"and opens its ring-like tooth-lined maw",
|
| 858 |
+
"examples/katana_character.jpg",
|
| 859 |
+
"Character",
|
| 860 |
+
"the young woman with long black hair in two buns, an orange utility jacket and two katana hilts "
|
| 861 |
+
"rising behind her shoulders",
|
| 862 |
+
"examples/wyvern_boss.jpg",
|
| 863 |
+
"Creature / Boss",
|
| 864 |
+
"the pale eyeless wyvern with a circular ring-like maw lined with rows of sharp teeth, a long "
|
| 865 |
+
"muscular neck, clawed wings and heavy bipedal limbs",
|
| 866 |
+
"Whip-pan: third-person → first-person",
|
| 867 |
+
True,
|
| 868 |
+
],
|
| 869 |
+
[
|
| 870 |
+
"sprints across a rain-slick rooftop, vaults a neon billboard gap and lands into a slide while "
|
| 871 |
+
"tracer fire streaks past",
|
| 872 |
+
"examples/doorway_character.jpg",
|
| 873 |
+
"Character",
|
| 874 |
+
"the young girl with brown hair, a white shirt and skirt and a guitar case slung across her back",
|
| 875 |
+
"examples/night_city.jpg",
|
| 876 |
+
"Environment",
|
| 877 |
+
"the night city of lit high-rise towers and a faceted glass skyscraper reflected in the dark river "
|
| 878 |
+
"below",
|
| 879 |
+
"First-person POV (FPV)",
|
| 880 |
+
True,
|
| 881 |
+
],
|
| 882 |
+
],
|
| 883 |
+
inputs=exampled,
|
| 884 |
+
outputs=[result, prompt_view],
|
| 885 |
+
fn=generate,
|
| 886 |
+
cache_examples=True,
|
| 887 |
+
cache_mode="lazy",
|
| 888 |
+
)
|
| 889 |
+
|
| 890 |
+
run.click(generate, request, [result, prompt_view], api_name="generate")
|
| 891 |
+
|
| 892 |
+
|
| 893 |
+
if __name__ == "__main__":
|
| 894 |
+
demo.launch(show_error=True, max_threads=1000)
|
examples/doorway_character.jpg
ADDED
|
Git LFS Details
|
examples/katana_character.jpg
ADDED
|
Git LFS Details
|
examples/night_city.jpg
ADDED
|
Git LFS Details
|
examples/ruined_village.jpg
ADDED
|
examples/wyvern_boss.jpg
ADDED
|
Git LFS Details
|
h3_fbc.py
ADDED
|
@@ -0,0 +1,357 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""First-block cache for MiniMax-H3: skip the trunk on steps whose block-0 residual barely moved.
|
| 2 |
+
|
| 3 |
+
Block 0 and the final AdaLN head run at the true timestep on **every** step; only blocks 1..49 are skipped, and only
|
| 4 |
+
while the residual they would have been handed looks like the one from the last step that actually ran. The decision
|
| 5 |
+
signal is the relative L1 between this step's block-0 residual (`block0_out - block0_in`) and the residual of the last
|
| 6 |
+
*computed* step, over the whole packed sequence. On a skip the trunk's contribution is replayed as a cached residual,
|
| 7 |
+
`(final_trunk_out - block0_out)` of that computed step.
|
| 8 |
+
|
| 9 |
+
Ported from `duckyshell/ComfyUI-MiniMaxH3-FirstBlockCache` (`nodes.py` @ 725973c) — same signal, same protected window
|
| 10 |
+
(10%-95% of the schedule, converted through the video shift of 12.0) and the same cap of two consecutive skips, which
|
| 11 |
+
is their "H3 Safe" preset at threshold 0.08.
|
| 12 |
+
|
| 13 |
+
**The audio exemption is not theirs.** duckyshell has no audio term anywhere: the decision is ~98% video by row count
|
| 14 |
+
and every audio row rides the stale trunk residual, which is what costs the soundtrack its energy. Audio runs its own
|
| 15 |
+
schedule (shift 3 against video's 12, diverging by up to 11x in rate across 20 steps), so on a skip the audio rows of
|
| 16 |
+
the trunk output are instead a 2-point **linear** extrapolation of the last two actually-computed audio features, in
|
| 17 |
+
the audio sigma coordinate — `xmarre/ComfyUI-Spectrum-MiniMax-H3`'s `audio_blend_weight=0.0` path, applied where
|
| 18 |
+
Spectrum applies it (the post-trunk hidden feature, ahead of the head that still runs at the true timestep) and fixed
|
| 19 |
+
to the right coordinate. Following ComfyUI PR #15390, no carried audio tensor is ever mutated: every write is a fresh
|
| 20 |
+
tensor out of `index_copy`.
|
| 21 |
+
|
| 22 |
+
It earns its place. Running this Space's own request at threshold 0.08 with `H3_FBC_AUDIO_EXEMPT=0` — the duckyshell
|
| 23 |
+
mechanism verbatim, same 13 skipped forwards, same video to within 0.2 dB — costs the soundtrack **16.7% of its RMS**
|
| 24 |
+
(0.0483 against an uncached 0.0580), while the exemption holds it at 1.04x. That is the same direction the offline
|
| 25 |
+
study measured on a near-silent clip, three times the size on one with real audio energy.
|
| 26 |
+
|
| 27 |
+
**A threshold does not travel across step counts.** The signature shrinks as the schedule is subdivided, so the same
|
| 28 |
+
number gates far more loosely at more steps. Measured on this Space — 960x544, 124 frames, one image reference,
|
| 29 |
+
seed 42, AoTI blocks, at its **default 28 steps** (27 forwards) — against the same request with `H3_FBC=0`:
|
| 30 |
+
|
| 31 |
+
| threshold | skipped | denoise loop | end to end | audio RMS vs uncached |
|
| 32 |
+
|---|---|---|---|---|
|
| 33 |
+
| 0.03 | 0 / 27 | 1.02x | 1.06x | 1.000 (bitwise-identical audio) |
|
| 34 |
+
| 0.05 | 9 / 27 | 1.46x | 1.36x | 0.956 |
|
| 35 |
+
| 0.08 | 13 / 27 | 2.19x | 1.82x | 1.040 |
|
| 36 |
+
|
| 37 |
+
The offline study calibrated 0.08 over a 20-step schedule, where it skipped 7 of 19 forwards. At 28 steps that same
|
| 38 |
+
0.08 skips 13 of 27 — nearly half — and the sampled video visibly re-rolls its background detail. **0.05 is the
|
| 39 |
+
default** because it reproduces the skip fraction that study validated (33% here against 37% there); 0.08 is the
|
| 40 |
+
aggressive setting. Below the signature floor — 0.03 skipped nothing at all here — a threshold buys nothing and still
|
| 41 |
+
pays for the signature, so lower is not safer, it is just slower.
|
| 42 |
+
|
| 43 |
+
A cached request is not the uncached one: the trajectory moves, so the video is a different sample of the same prompt
|
| 44 |
+
(same shot, same subject, same quality — different signage and background detail). `H3_FBC=0` restores today's output
|
| 45 |
+
exactly, and is worth reaching for when a request has to reproduce a specific earlier result.
|
| 46 |
+
|
| 47 |
+
This composes with `h3_aoti`: that module patches each of the 50 blocks' own `forward` and they stay a real
|
| 48 |
+
`ModuleList`, so skipping the trunk simply does not call blocks 1..49 that step. `LazyAOTIModel` rebinds its constants
|
| 49 |
+
whenever the weights dict it is handed changes identity, and the first forward of a request never skips, so every
|
| 50 |
+
block has bound its own weights before any step is cached.
|
| 51 |
+
"""
|
| 52 |
+
|
| 53 |
+
from __future__ import annotations
|
| 54 |
+
|
| 55 |
+
import contextlib
|
| 56 |
+
import os
|
| 57 |
+
import types
|
| 58 |
+
|
| 59 |
+
ENABLED = os.environ.get("H3_FBC", "1") == "1"
|
| 60 |
+
# Relative-L1 gate on the block-0 residual, calibrated at this Space's default 28 steps — see the table above.
|
| 61 |
+
THRESHOLD = float(os.environ.get("H3_FBC_THRESHOLD", "0.05"))
|
| 62 |
+
# duckyshell's cap. Without it the gate compares against an ever-older computed step and drifts away unbounded.
|
| 63 |
+
MAX_CONSECUTIVE_HITS = int(os.environ.get("H3_FBC_MAX_CONSECUTIVE", "2"))
|
| 64 |
+
# The protected head and tail of the schedule, as fractions, converted to sigma through the video shift below.
|
| 65 |
+
START_PERCENT = float(os.environ.get("H3_FBC_START_PERCENT", "0.10"))
|
| 66 |
+
END_PERCENT = float(os.environ.get("H3_FBC_END_PERCENT", "0.95"))
|
| 67 |
+
# `MiniMaxH3SetTimestepsStep` builds the video schedule at shift 12.0 and the audio one at 3.0.
|
| 68 |
+
VIDEO_SHIFT = float(os.environ.get("H3_FBC_VIDEO_SHIFT", "12.0"))
|
| 69 |
+
AUDIO_EXEMPT = os.environ.get("H3_FBC_AUDIO_EXEMPT", "1") == "1"
|
| 70 |
+
|
| 71 |
+
# Every keyword `MiniMaxH3LoopDenoiser` passes. It filters the packed-sequence layout through
|
| 72 |
+
# `inspect.signature(transformer.forward).parameters`, so a replacement forward that drops a name silently stops
|
| 73 |
+
# receiving it; `install` refuses rather than let that happen quietly.
|
| 74 |
+
FORWARD_PARAMETERS = (
|
| 75 |
+
"hidden_states",
|
| 76 |
+
"audio_hidden_states",
|
| 77 |
+
"encoder_hidden_states",
|
| 78 |
+
"timestep",
|
| 79 |
+
"timestep_indices",
|
| 80 |
+
"token_tags",
|
| 81 |
+
"position_ids",
|
| 82 |
+
"video_indices",
|
| 83 |
+
"audio_indices",
|
| 84 |
+
"text_indices",
|
| 85 |
+
"attention_kwargs",
|
| 86 |
+
"return_dict",
|
| 87 |
+
)
|
| 88 |
+
|
| 89 |
+
|
| 90 |
+
def status() -> str:
|
| 91 |
+
return (
|
| 92 |
+
f"first-block cache **on** · threshold `{THRESHOLD}` · audio exemption "
|
| 93 |
+
f"{'on' if AUDIO_EXEMPT else 'off'}"
|
| 94 |
+
if ENABLED
|
| 95 |
+
else "first-block cache **off** (`H3_FBC=1` to skip the trunk on steady steps)"
|
| 96 |
+
)
|
| 97 |
+
|
| 98 |
+
|
| 99 |
+
def _shifted_sigma(u: float, shift: float) -> float:
|
| 100 |
+
return shift * u / (1.0 + (shift - 1.0) * u)
|
| 101 |
+
|
| 102 |
+
|
| 103 |
+
def _rel_l1(current, previous) -> float:
|
| 104 |
+
numerator = (current.float() - previous.float()).abs().mean()
|
| 105 |
+
denominator = previous.float().abs().mean().clamp(min=1e-8)
|
| 106 |
+
return float((numerator / denominator).item())
|
| 107 |
+
|
| 108 |
+
|
| 109 |
+
class _State:
|
| 110 |
+
def __init__(self, threshold: float, steps: int, audio_exempt: bool):
|
| 111 |
+
self.threshold = threshold
|
| 112 |
+
self.steps = steps
|
| 113 |
+
self.audio_exempt = audio_exempt
|
| 114 |
+
# duckyshell reads the window as sigma bounds: a flow model's sigma at `u = 1 - percent`, shifted.
|
| 115 |
+
self.start_sigma = _shifted_sigma(1.0 - START_PERCENT, VIDEO_SHIFT)
|
| 116 |
+
self.end_sigma = _shifted_sigma(1.0 - END_PERCENT, VIDEO_SHIFT)
|
| 117 |
+
self.original = None
|
| 118 |
+
self.failed = False
|
| 119 |
+
self.consecutive_hits = 0
|
| 120 |
+
self.prev_first_residual = None
|
| 121 |
+
self.tail_residual = None
|
| 122 |
+
self.audio_history = [] # [(sigma_audio, audio rows of the trunk output)], newest last, at most two
|
| 123 |
+
self.computed = 0
|
| 124 |
+
self.skipped = 0
|
| 125 |
+
|
| 126 |
+
|
| 127 |
+
def _cached_forward(
|
| 128 |
+
self,
|
| 129 |
+
state,
|
| 130 |
+
hidden_states,
|
| 131 |
+
audio_hidden_states,
|
| 132 |
+
encoder_hidden_states,
|
| 133 |
+
timestep,
|
| 134 |
+
timestep_indices,
|
| 135 |
+
token_tags,
|
| 136 |
+
position_ids,
|
| 137 |
+
video_indices,
|
| 138 |
+
audio_indices,
|
| 139 |
+
text_indices,
|
| 140 |
+
return_dict,
|
| 141 |
+
):
|
| 142 |
+
"""`MiniMaxH3Transformer3DModel.forward` with the block loop split at block 0.
|
| 143 |
+
|
| 144 |
+
Everything outside the loop is that method verbatim, at the `diffusers` commit `requirements.txt` pins; keep the
|
| 145 |
+
two in step when the pin moves.
|
| 146 |
+
"""
|
| 147 |
+
import torch
|
| 148 |
+
|
| 149 |
+
from diffusers.models.transformers.transformer_minimax_h3 import (
|
| 150 |
+
MINIMAX_H3_MODALITY_NUM,
|
| 151 |
+
MiniMaxH3TransformerOutput,
|
| 152 |
+
)
|
| 153 |
+
|
| 154 |
+
sequence_length = position_ids.shape[0]
|
| 155 |
+
rotary_emb = self.rope(position_ids)
|
| 156 |
+
|
| 157 |
+
video_embeds = self.proj_in(hidden_states.to(self.proj_in.weight.dtype))
|
| 158 |
+
audio_embeds = self.audio_proj_in(audio_hidden_states.to(self.audio_proj_in.weight.dtype))
|
| 159 |
+
text_embeds = self.context_embedder(encoder_hidden_states.to(self.context_embedder.weight.dtype))
|
| 160 |
+
text_embeds = self.token_refiner(text_embeds)
|
| 161 |
+
|
| 162 |
+
packed = text_embeds.new_zeros((text_embeds.shape[0], sequence_length, text_embeds.shape[-1]))
|
| 163 |
+
packed = packed.index_copy(1, text_indices, text_embeds)
|
| 164 |
+
packed = packed.index_copy(1, video_indices, video_embeds.to(text_embeds.dtype))
|
| 165 |
+
packed = packed.index_copy(1, audio_indices, audio_embeds.to(text_embeds.dtype))
|
| 166 |
+
|
| 167 |
+
temb = self.time_proj(timestep)
|
| 168 |
+
temb = self.time_embedder(temb.to(self.time_embedder.linear_1.weight.dtype))
|
| 169 |
+
adaln_indices = timestep_indices * MINIMAX_H3_MODALITY_NUM + token_tags.clamp(min=0)
|
| 170 |
+
|
| 171 |
+
attention_mask = None
|
| 172 |
+
is_pad = token_tags < 0
|
| 173 |
+
if bool(is_pad.any()):
|
| 174 |
+
attention_mask = is_pad[None, :] == is_pad[:, None]
|
| 175 |
+
|
| 176 |
+
blocks = self.transformer_blocks
|
| 177 |
+
block0_out = blocks[0](packed, temb, adaln_indices, rotary_emb, attention_mask)
|
| 178 |
+
first_residual = block0_out - packed
|
| 179 |
+
|
| 180 |
+
# The generated rows trail their modality's index list, so the last row of each carries that stream's live noise
|
| 181 |
+
# level. The scheduler exposes `timesteps = 1 - sigmas[:-1]`.
|
| 182 |
+
sigma_video = 1.0 - float(timestep[timestep_indices[video_indices[-1]]].item())
|
| 183 |
+
sigma_audio = 1.0 - float(timestep[timestep_indices[audio_indices[-1]]].item())
|
| 184 |
+
|
| 185 |
+
use_cache = False
|
| 186 |
+
if (
|
| 187 |
+
state.prev_first_residual is not None
|
| 188 |
+
and state.tail_residual is not None
|
| 189 |
+
and state.prev_first_residual.shape == first_residual.shape
|
| 190 |
+
and state.consecutive_hits < MAX_CONSECUTIVE_HITS
|
| 191 |
+
and state.end_sigma <= sigma_video <= state.start_sigma
|
| 192 |
+
):
|
| 193 |
+
use_cache = _rel_l1(first_residual, state.prev_first_residual) <= state.threshold
|
| 194 |
+
|
| 195 |
+
if use_cache:
|
| 196 |
+
state.consecutive_hits += 1
|
| 197 |
+
state.skipped += 1
|
| 198 |
+
trunk_out = block0_out + state.tail_residual
|
| 199 |
+
if state.audio_exempt and state.audio_history:
|
| 200 |
+
sigma_prev, feature_prev = state.audio_history[-1]
|
| 201 |
+
if len(state.audio_history) == 2 and abs(sigma_prev - state.audio_history[-2][0]) > 1e-8:
|
| 202 |
+
sigma_prev2, feature_prev2 = state.audio_history[-2]
|
| 203 |
+
ratio = (sigma_audio - sigma_prev) / (sigma_prev - sigma_prev2)
|
| 204 |
+
audio_feature = feature_prev + (feature_prev - feature_prev2) * ratio
|
| 205 |
+
else:
|
| 206 |
+
audio_feature = feature_prev
|
| 207 |
+
trunk_out = trunk_out.index_copy(1, audio_indices, audio_feature.to(trunk_out.dtype))
|
| 208 |
+
else:
|
| 209 |
+
state.consecutive_hits = 0
|
| 210 |
+
state.computed += 1
|
| 211 |
+
trunk_out = block0_out
|
| 212 |
+
for block in blocks[1:]:
|
| 213 |
+
trunk_out = block(trunk_out, temb, adaln_indices, rotary_emb, attention_mask)
|
| 214 |
+
state.tail_residual = (trunk_out - block0_out).detach()
|
| 215 |
+
state.prev_first_residual = first_residual.detach()
|
| 216 |
+
if state.audio_exempt:
|
| 217 |
+
state.audio_history.append((sigma_audio, trunk_out.index_select(1, audio_indices).detach().float()))
|
| 218 |
+
state.audio_history = state.audio_history[-2:]
|
| 219 |
+
|
| 220 |
+
out = self.norm_out(trunk_out, temb, timestep_indices).to(self.proj_out.weight.dtype)
|
| 221 |
+
video_output = self.proj_out(out).index_select(1, video_indices)
|
| 222 |
+
audio_output = self.audio_proj_out(out).index_select(1, audio_indices)
|
| 223 |
+
|
| 224 |
+
if not return_dict:
|
| 225 |
+
return (video_output, audio_output)
|
| 226 |
+
return MiniMaxH3TransformerOutput(sample=video_output, audio_sample=audio_output)
|
| 227 |
+
|
| 228 |
+
|
| 229 |
+
def _forward(
|
| 230 |
+
self,
|
| 231 |
+
hidden_states,
|
| 232 |
+
audio_hidden_states,
|
| 233 |
+
encoder_hidden_states,
|
| 234 |
+
timestep,
|
| 235 |
+
timestep_indices,
|
| 236 |
+
token_tags,
|
| 237 |
+
position_ids,
|
| 238 |
+
video_indices,
|
| 239 |
+
audio_indices,
|
| 240 |
+
text_indices,
|
| 241 |
+
attention_kwargs=None,
|
| 242 |
+
return_dict: bool = True,
|
| 243 |
+
):
|
| 244 |
+
"""The installed forward. Anything it cannot serve — a LoRA scale, an unexpected layout, a bug — is handed to the
|
| 245 |
+
original forward instead, for this call and every later one, so a cached request can degrade to an uncached one but
|
| 246 |
+
never to a failed one."""
|
| 247 |
+
state = self._h3_fbc
|
| 248 |
+
original = dict(
|
| 249 |
+
hidden_states=hidden_states,
|
| 250 |
+
audio_hidden_states=audio_hidden_states,
|
| 251 |
+
encoder_hidden_states=encoder_hidden_states,
|
| 252 |
+
timestep=timestep,
|
| 253 |
+
timestep_indices=timestep_indices,
|
| 254 |
+
token_tags=token_tags,
|
| 255 |
+
position_ids=position_ids,
|
| 256 |
+
video_indices=video_indices,
|
| 257 |
+
audio_indices=audio_indices,
|
| 258 |
+
text_indices=text_indices,
|
| 259 |
+
attention_kwargs=attention_kwargs,
|
| 260 |
+
return_dict=return_dict,
|
| 261 |
+
)
|
| 262 |
+
# `apply_lora_scale` decorates the real forward and this one is not it, so a request that actually scales a LoRA
|
| 263 |
+
# goes down the original path rather than silently losing its scale.
|
| 264 |
+
if state.failed or (attention_kwargs or {}).get("scale") is not None:
|
| 265 |
+
return state.original(**original)
|
| 266 |
+
try:
|
| 267 |
+
return _cached_forward(
|
| 268 |
+
self,
|
| 269 |
+
state,
|
| 270 |
+
hidden_states,
|
| 271 |
+
audio_hidden_states,
|
| 272 |
+
encoder_hidden_states,
|
| 273 |
+
timestep,
|
| 274 |
+
timestep_indices,
|
| 275 |
+
token_tags,
|
| 276 |
+
position_ids,
|
| 277 |
+
video_indices,
|
| 278 |
+
audio_indices,
|
| 279 |
+
text_indices,
|
| 280 |
+
return_dict,
|
| 281 |
+
)
|
| 282 |
+
except Exception as error:
|
| 283 |
+
state.failed = True
|
| 284 |
+
print(f"[h3-fbc] disabled for this request ({type(error).__name__}: {error}); running uncached", flush=True)
|
| 285 |
+
return state.original(**original)
|
| 286 |
+
|
| 287 |
+
|
| 288 |
+
def install(transformer, steps: int = 0, threshold: float = THRESHOLD, audio_exempt: bool = AUDIO_EXEMPT) -> bool:
|
| 289 |
+
"""Bind the caching forward onto `transformer`. Returns whether it went on.
|
| 290 |
+
|
| 291 |
+
`accelerate`'s `add_hook_to_module` — what `ComponentsManager.enable_auto_cpu_offload` installs — moves the real
|
| 292 |
+
forward to `_old_forward` and puts its own onload wrapper in `forward`. Replacing `forward` there would step over
|
| 293 |
+
the wrapper and run the block stack against weights still on the host, so the replacement goes into `_old_forward`
|
| 294 |
+
whenever the hook is present.
|
| 295 |
+
"""
|
| 296 |
+
import inspect
|
| 297 |
+
|
| 298 |
+
if getattr(transformer, "_h3_fbc", None) is not None:
|
| 299 |
+
return True
|
| 300 |
+
|
| 301 |
+
hooked = hasattr(transformer, "_hf_hook") and hasattr(transformer, "_old_forward")
|
| 302 |
+
current = transformer._old_forward if hooked else transformer.forward
|
| 303 |
+
missing = [name for name in FORWARD_PARAMETERS if name not in inspect.signature(current).parameters]
|
| 304 |
+
if missing:
|
| 305 |
+
print(f"[h3-fbc] this transformer's forward has no {missing}; running uncached", flush=True)
|
| 306 |
+
return False
|
| 307 |
+
if not hasattr(transformer, "transformer_blocks") or len(transformer.transformer_blocks) < 2:
|
| 308 |
+
print("[h3-fbc] no block stack to skip; running uncached", flush=True)
|
| 309 |
+
return False
|
| 310 |
+
|
| 311 |
+
state = _State(threshold, steps, audio_exempt)
|
| 312 |
+
state.original = current
|
| 313 |
+
transformer._h3_fbc = state
|
| 314 |
+
bound = types.MethodType(_forward, transformer)
|
| 315 |
+
if hooked:
|
| 316 |
+
transformer._old_forward = bound
|
| 317 |
+
else:
|
| 318 |
+
transformer.forward = bound
|
| 319 |
+
return True
|
| 320 |
+
|
| 321 |
+
|
| 322 |
+
def uninstall(transformer) -> None:
|
| 323 |
+
state = getattr(transformer, "_h3_fbc", None)
|
| 324 |
+
if state is None:
|
| 325 |
+
return
|
| 326 |
+
if hasattr(transformer, "_hf_hook") and hasattr(transformer, "_old_forward"):
|
| 327 |
+
transformer._old_forward = state.original
|
| 328 |
+
else:
|
| 329 |
+
transformer.__dict__.pop("forward", None)
|
| 330 |
+
del transformer._h3_fbc
|
| 331 |
+
total = state.computed + state.skipped
|
| 332 |
+
if total:
|
| 333 |
+
print(
|
| 334 |
+
f"[h3-fbc] {state.skipped}/{total} forwards served from cache "
|
| 335 |
+
f"(threshold {state.threshold}, audio exemption {'on' if state.audio_exempt else 'off'})",
|
| 336 |
+
flush=True,
|
| 337 |
+
)
|
| 338 |
+
|
| 339 |
+
|
| 340 |
+
@contextlib.contextmanager
|
| 341 |
+
def enabled(transformer, steps: int = 0):
|
| 342 |
+
"""Cache the trunk for the duration of one request. The state is per-request by construction — a residual only ever
|
| 343 |
+
means something within the schedule it was measured on — and nothing in here can raise into the request."""
|
| 344 |
+
installed = False
|
| 345 |
+
if ENABLED:
|
| 346 |
+
try:
|
| 347 |
+
installed = install(transformer, steps=steps)
|
| 348 |
+
except Exception as error:
|
| 349 |
+
print(f"[h3-fbc] install failed ({type(error).__name__}: {error}); running uncached", flush=True)
|
| 350 |
+
try:
|
| 351 |
+
yield installed
|
| 352 |
+
finally:
|
| 353 |
+
if installed:
|
| 354 |
+
try:
|
| 355 |
+
uninstall(transformer)
|
| 356 |
+
except Exception as error:
|
| 357 |
+
print(f"[h3-fbc] uninstall failed ({type(error).__name__}: {error})", flush=True)
|
h3_split_blocks.py
ADDED
|
@@ -0,0 +1,147 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""The halves of a **split** MiniMax-H3 deployment, for both of its checkpoint partitions.
|
| 2 |
+
|
| 3 |
+
MiniMax-H3 is 195.9 GiB in bfloat16 and a ZeroGPU Space is evicted at 150 GB of storage, so `MiniMaxH3Blocks` is cut
|
| 4 |
+
at its `text_encoder` step: the 62.14 GiB Qwen3-VL runs in the conditioner Space, everything else in a generator
|
| 5 |
+
Space, and `prompt_embeds` + `text_token_tags` is the whole wire format between them.
|
| 6 |
+
|
| 7 |
+
`resize` / `setup` run on **both** sides: they own no pretrained component, and each half needs the canvas and the
|
| 8 |
+
prepared keyframes or normalized references. Both conditioner halves also return the resolved `height` / `width` /
|
| 9 |
+
`num_frames`, which the generating half pins rather than re-deriving.
|
| 10 |
+
|
| 11 |
+
Two things the blocks leave to the caller: a keyframe reaches them EXIF-transposed and in RGB, and the `t2va` / `fl2va`
|
| 12 |
+
frame count is aligned to `17 * n + 5` before the call, since that arithmetic lives on the denoising side of the cut.
|
| 13 |
+
"""
|
| 14 |
+
|
| 15 |
+
from diffusers.modular_pipelines.minimax_h3.before_encoder import MiniMaxH3Ref2VASetupStep
|
| 16 |
+
from diffusers.modular_pipelines.minimax_h3.decoders import MiniMaxH3AfterDenoiseStep
|
| 17 |
+
from diffusers.modular_pipelines.minimax_h3.encoders import (
|
| 18 |
+
MiniMaxH3Ref2VAReferenceEncoderStep,
|
| 19 |
+
MiniMaxH3Ref2VATextEncoderStep,
|
| 20 |
+
MiniMaxH3TextEncoderStep,
|
| 21 |
+
)
|
| 22 |
+
from diffusers.modular_pipelines.minimax_h3.modular_blocks_minimax_h3 import (
|
| 23 |
+
MiniMaxH3AutoKeyframeVaeEncoderStep,
|
| 24 |
+
MiniMaxH3AutoResizeStep,
|
| 25 |
+
MiniMaxH3CoreDenoiseStep,
|
| 26 |
+
MiniMaxH3DecodeStep,
|
| 27 |
+
MiniMaxH3Ref2VACoreDenoiseStep,
|
| 28 |
+
_generation_outputs,
|
| 29 |
+
)
|
| 30 |
+
from diffusers.modular_pipelines.modular_pipeline import SequentialPipelineBlocks
|
| 31 |
+
from diffusers.modular_pipelines.modular_pipeline_utils import OutputParam
|
| 32 |
+
|
| 33 |
+
|
| 34 |
+
def _wire_outputs(num_frames: bool = True) -> list[OutputParam]:
|
| 35 |
+
"""The wire format of the split. `num_frames` is declared by the `ref2va` half alone, whose setup resolves one."""
|
| 36 |
+
return [
|
| 37 |
+
OutputParam.template("prompt_embeds"),
|
| 38 |
+
OutputParam("text_token_tags", description="The per-row modality tag of every row of `prompt_embeds`."),
|
| 39 |
+
OutputParam("height", type_hint=int, description="Resolved height of the generated video in pixels."),
|
| 40 |
+
OutputParam("width", type_hint=int, description="Resolved width of the generated video in pixels."),
|
| 41 |
+
*(
|
| 42 |
+
[OutputParam("num_frames", type_hint=int, description="Resolved number of frames, of the form 17 * n + 5.")]
|
| 43 |
+
if num_frames
|
| 44 |
+
else []
|
| 45 |
+
),
|
| 46 |
+
]
|
| 47 |
+
|
| 48 |
+
|
| 49 |
+
class MiniMaxH3ConditionerBlocks(SequentialPipelineBlocks):
|
| 50 |
+
"""The conditioner half of a split MiniMax-H3: the keyframes on the canvas plus the Qwen3-VL read at layer 50."""
|
| 51 |
+
|
| 52 |
+
model_name = "minimax-h3"
|
| 53 |
+
block_classes = [MiniMaxH3AutoResizeStep, MiniMaxH3TextEncoderStep]
|
| 54 |
+
block_names = ["resize", "text_encoder"]
|
| 55 |
+
|
| 56 |
+
@property
|
| 57 |
+
def description(self):
|
| 58 |
+
return (
|
| 59 |
+
"The conditioner half of a split MiniMax-H3 deployment: puts the keyframes onto the target canvas and "
|
| 60 |
+
"encodes MiniMax-H3's presentation of the request into the `prompt_embeds` / `text_token_tags` pair the "
|
| 61 |
+
"denoising half consumes. The frame count is the caller's to align."
|
| 62 |
+
)
|
| 63 |
+
|
| 64 |
+
@property
|
| 65 |
+
def outputs(self):
|
| 66 |
+
return _wire_outputs(num_frames=False)
|
| 67 |
+
|
| 68 |
+
|
| 69 |
+
class MiniMaxH3GeneratorBlocks(SequentialPipelineBlocks):
|
| 70 |
+
"""The denoising half of a split MiniMax-H3: `MiniMaxH3Blocks` with its `text_encoder` step removed."""
|
| 71 |
+
|
| 72 |
+
model_name = "minimax-h3"
|
| 73 |
+
block_classes = [
|
| 74 |
+
MiniMaxH3AutoResizeStep,
|
| 75 |
+
MiniMaxH3AutoKeyframeVaeEncoderStep,
|
| 76 |
+
MiniMaxH3CoreDenoiseStep,
|
| 77 |
+
MiniMaxH3AfterDenoiseStep,
|
| 78 |
+
MiniMaxH3DecodeStep,
|
| 79 |
+
]
|
| 80 |
+
block_names = ["resize", "vae_encoder", "denoise", "after_denoise", "decode"]
|
| 81 |
+
|
| 82 |
+
@property
|
| 83 |
+
def description(self):
|
| 84 |
+
return (
|
| 85 |
+
"The denoising half of a split MiniMax-H3 deployment: the `t2va` / `fl2va` branch of `MiniMaxH3Blocks` "
|
| 86 |
+
"without its text-encoder step, so `prompt_embeds` and `text_token_tags` come in as inputs and the "
|
| 87 |
+
"62.14 GiB Qwen3-VL conditioner is never loaded here."
|
| 88 |
+
)
|
| 89 |
+
|
| 90 |
+
@property
|
| 91 |
+
def outputs(self):
|
| 92 |
+
return _generation_outputs()
|
| 93 |
+
|
| 94 |
+
|
| 95 |
+
class MiniMaxH3Ref2VAConditionerBlocks(SequentialPipelineBlocks):
|
| 96 |
+
"""The conditioner half of a split `ref2va`: the resolved plan plus the Qwen3-VL read at its 50th layer.
|
| 97 |
+
|
| 98 |
+
Component for component this is `MiniMaxH3ConditionerBlocks`, so one conditioner Space serves both partitions.
|
| 99 |
+
What differs is the presentation: `ref2va` prepends a label per reference and a vision block per image and per
|
| 100 |
+
merged video frame pair, so the references themselves have to reach this half.
|
| 101 |
+
"""
|
| 102 |
+
|
| 103 |
+
model_name = "minimax-h3"
|
| 104 |
+
block_classes = [MiniMaxH3Ref2VASetupStep, MiniMaxH3Ref2VATextEncoderStep]
|
| 105 |
+
block_names = ["setup", "text_encoder"]
|
| 106 |
+
|
| 107 |
+
@property
|
| 108 |
+
def description(self):
|
| 109 |
+
return (
|
| 110 |
+
"The conditioner half of a split MiniMax-H3 `ref2va` deployment: resolves the request plan (canvas, frame "
|
| 111 |
+
"count, references normalized onto MiniMax-H3's own rates and resolutions) and encodes MiniMax-H3's "
|
| 112 |
+
"presentation of it into the `prompt_embeds` / `text_token_tags` pair the denoising half consumes."
|
| 113 |
+
)
|
| 114 |
+
|
| 115 |
+
@property
|
| 116 |
+
def outputs(self):
|
| 117 |
+
return _wire_outputs()
|
| 118 |
+
|
| 119 |
+
|
| 120 |
+
class MiniMaxH3Ref2VAGeneratorBlocks(SequentialPipelineBlocks):
|
| 121 |
+
"""The denoising half of a split `ref2va`: the `ref2va` branch with its `text_encoder` step removed.
|
| 122 |
+
|
| 123 |
+
`reference_encoder` stays here, next to the two autoencoders it runs: its output shapes are where every reference
|
| 124 |
+
block's geometry in the packed layout comes from.
|
| 125 |
+
"""
|
| 126 |
+
|
| 127 |
+
model_name = "minimax-h3"
|
| 128 |
+
block_classes = [
|
| 129 |
+
MiniMaxH3Ref2VASetupStep,
|
| 130 |
+
MiniMaxH3Ref2VAReferenceEncoderStep,
|
| 131 |
+
MiniMaxH3Ref2VACoreDenoiseStep,
|
| 132 |
+
MiniMaxH3AfterDenoiseStep,
|
| 133 |
+
MiniMaxH3DecodeStep,
|
| 134 |
+
]
|
| 135 |
+
block_names = ["setup", "reference_encoder", "denoise", "after_denoise", "decode"]
|
| 136 |
+
|
| 137 |
+
@property
|
| 138 |
+
def description(self):
|
| 139 |
+
return (
|
| 140 |
+
"The denoising half of a split MiniMax-H3 `ref2va` deployment: the `ref2va` branch of `MiniMaxH3Blocks` "
|
| 141 |
+
"without its text-encoder step, so `prompt_embeds` and `text_token_tags` come in as inputs and the "
|
| 142 |
+
"62.14 GiB Qwen3-VL conditioner is never loaded here. The transformer is the `transformer_ref` partition."
|
| 143 |
+
)
|
| 144 |
+
|
| 145 |
+
@property
|
| 146 |
+
def outputs(self):
|
| 147 |
+
return _generation_outputs()
|
ncii_guard.py
ADDED
|
@@ -0,0 +1,90 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""The NCII prompt guard, in its own process.
|
| 2 |
+
|
| 3 |
+
[`hfmlsoc/ncii-light-guard-v01`](https://huggingface.co/hfmlsoc/ncii-light-guard-v01) is a 270M CPU text
|
| 4 |
+
classifier scoring the NCII risk of an edit prompt. It cannot live in the main process: with it loaded there,
|
| 5 |
+
every subsequent `@spaces.GPU` worker dies at `worker_init` with `RuntimeError: No CUDA GPUs are available` —
|
| 6 |
+
the fork inherits whatever CUDA driver state the classifier's torch activity left behind, and a factory reboot
|
| 7 |
+
does not clear it. It cannot be a `multiprocessing.spawn` child either: spawn re-imports the parent's main
|
| 8 |
+
module, and on a Space that main module is `app.py` — the child would re-run the whole startup, `start()`
|
| 9 |
+
included. So the classifier runs this file as a plain subprocess — a fresh interpreter that never sees `spaces`
|
| 10 |
+
— and answers over stdin/stdout, one JSON object per line.
|
| 11 |
+
"""
|
| 12 |
+
|
| 13 |
+
from __future__ import annotations
|
| 14 |
+
|
| 15 |
+
import json
|
| 16 |
+
import os
|
| 17 |
+
import select
|
| 18 |
+
import subprocess
|
| 19 |
+
import sys
|
| 20 |
+
import threading
|
| 21 |
+
|
| 22 |
+
GUARD_REPO = "hfmlsoc/ncii-light-guard-v01"
|
| 23 |
+
|
| 24 |
+
_lock = threading.Lock()
|
| 25 |
+
_process: subprocess.Popen | None = None
|
| 26 |
+
|
| 27 |
+
|
| 28 |
+
def _read(timeout: float) -> dict:
|
| 29 |
+
readable, _, _ = select.select([_process.stdout], [], [], timeout)
|
| 30 |
+
if not readable:
|
| 31 |
+
raise TimeoutError(f"the guard did not answer within {timeout}s")
|
| 32 |
+
line = _process.stdout.readline()
|
| 33 |
+
if not line:
|
| 34 |
+
raise EOFError("the guard process died")
|
| 35 |
+
return json.loads(line)
|
| 36 |
+
|
| 37 |
+
|
| 38 |
+
def _spawn() -> None:
|
| 39 |
+
global _process
|
| 40 |
+
_process = subprocess.Popen(
|
| 41 |
+
[sys.executable, os.path.abspath(__file__)],
|
| 42 |
+
stdin=subprocess.PIPE,
|
| 43 |
+
stdout=subprocess.PIPE,
|
| 44 |
+
text=True,
|
| 45 |
+
bufsize=1,
|
| 46 |
+
)
|
| 47 |
+
# Generous: a cold cache downloads the checkpoint first.
|
| 48 |
+
assert _read(300.0) == {"status": "ready"}
|
| 49 |
+
|
| 50 |
+
|
| 51 |
+
def start() -> None:
|
| 52 |
+
"""Launch the worker and block until its model is up. Called once at startup; `classify` revives it if it dies."""
|
| 53 |
+
with _lock:
|
| 54 |
+
_spawn()
|
| 55 |
+
|
| 56 |
+
|
| 57 |
+
def classify(prompt: str, timeout: float = 60.0) -> dict:
|
| 58 |
+
"""`{'label': 'safe' | 'ncii', 'score': ...}` for one prompt, replacing a dead or wedged worker once."""
|
| 59 |
+
with _lock:
|
| 60 |
+
for attempt in (0, 1):
|
| 61 |
+
try:
|
| 62 |
+
if _process is None or _process.poll() is not None:
|
| 63 |
+
_spawn()
|
| 64 |
+
_process.stdin.write(json.dumps({"prompt": prompt}) + "\n")
|
| 65 |
+
_process.stdin.flush()
|
| 66 |
+
return _read(timeout)
|
| 67 |
+
except Exception:
|
| 68 |
+
if attempt:
|
| 69 |
+
raise
|
| 70 |
+
if _process is not None and _process.poll() is None:
|
| 71 |
+
_process.kill()
|
| 72 |
+
|
| 73 |
+
|
| 74 |
+
def _serve() -> None:
|
| 75 |
+
"""The child: plain torch on CPU. The protocol keeps the real stdout to itself — everything else
|
| 76 |
+
(download progress, warnings) is pushed over to stderr so it cannot corrupt a reply."""
|
| 77 |
+
protocol = os.fdopen(os.dup(1), "w", buffering=1)
|
| 78 |
+
os.dup2(2, 1)
|
| 79 |
+
|
| 80 |
+
from transformers import pipeline
|
| 81 |
+
|
| 82 |
+
classifier = pipeline("text-classification", model=GUARD_REPO, device="cpu")
|
| 83 |
+
protocol.write(json.dumps({"status": "ready"}) + "\n")
|
| 84 |
+
for line in sys.stdin:
|
| 85 |
+
result = classifier(json.loads(line)["prompt"], truncation=True)[0]
|
| 86 |
+
protocol.write(json.dumps({"label": result["label"], "score": float(result["score"])}) + "\n")
|
| 87 |
+
|
| 88 |
+
|
| 89 |
+
if __name__ == "__main__":
|
| 90 |
+
_serve()
|
packages.txt
ADDED
|
@@ -0,0 +1 @@
|
|
|
|
|
|
|
| 1 |
+
ffmpeg
|
requirements.txt
ADDED
|
@@ -0,0 +1,18 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# diffusers from the MiniMax-H3 PR (https://github.com/huggingface/diffusers/pull/14371), pinned to a
|
| 2 |
+
# commit on its minimax-h3-refactor branch — the modular MiniMax-H3 blocks are in no release yet, and
|
| 3 |
+
# h3_split_blocks.py subclasses them. main has since renamed those block classes, so the pin is load-bearing.
|
| 4 |
+
# 665f578278365ea4a3318cb8c9b66ce6c01204b9 = refs/pull/14371/head at the time of this deploy
|
| 5 |
+
--extra-index-url https://download.pytorch.org/whl/cu130
|
| 6 |
+
diffusers @ git+https://github.com/huggingface/diffusers.git@665f578278365ea4a3318cb8c9b66ce6c01204b9
|
| 7 |
+
torch==2.11.0
|
| 8 |
+
torchvision==0.26.0
|
| 9 |
+
torchaudio==2.11.0
|
| 10 |
+
transformers==5.8.0
|
| 11 |
+
accelerate==1.14.0
|
| 12 |
+
huggingface-hub==1.24.0
|
| 13 |
+
peft
|
| 14 |
+
av
|
| 15 |
+
pillow
|
| 16 |
+
numpy
|
| 17 |
+
requests
|
| 18 |
+
safetensors>=0.8.0
|