multimodalart HF Staff commited on
Commit
4d565a4
·
verified ·
1 Parent(s): 3963264

MiniMax-H3 Third-Person View LoRA demo

Browse files
.gitattributes CHANGED
@@ -33,3 +33,7 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
 
 
 
 
 
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
+ examples/doorway_character.jpg filter=lfs diff=lfs merge=lfs -text
37
+ examples/katana_character.jpg filter=lfs diff=lfs merge=lfs -text
38
+ examples/night_city.jpg filter=lfs diff=lfs merge=lfs -text
39
+ examples/wyvern_boss.jpg filter=lfs diff=lfs merge=lfs -text
README.md CHANGED
@@ -1,13 +1,110 @@
1
  ---
2
- title: Minimax H3 Third Person View Lora
3
- emoji: 👁
4
- colorFrom: purple
5
- colorTo: purple
6
  sdk: gradio
7
  sdk_version: 6.28.0
8
- python_version: '3.13'
9
  app_file: app.py
 
 
10
  pinned: false
 
 
 
 
 
 
11
  ---
12
 
13
- Check out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  ---
2
+ title: MiniMax-H3 Third-Person View LoRA
3
+ emoji: 🎮
4
+ colorFrom: yellow
5
+ colorTo: pink
6
  sdk: gradio
7
  sdk_version: 6.28.0
 
8
  app_file: app.py
9
+ python_version: "3.12"
10
+ startup_duration_timeout: 1h
11
  pinned: false
12
+ short_description: Design sheets to game cutscenes with HUD and audio
13
+ suggested_hardware: zero-a10g
14
+ models:
15
+ - MiniMaxAI/MiniMax-H3
16
+ - multimodalart/MiniMax-H3-Pruned
17
+ - WarmBloodAban/Minimax_H3_LoRAs
18
  ---
19
 
20
+ # MiniMax-H3 · Third-Person View LoRA
21
+
22
+ A demo of [`WarmBloodAban/Minimax_H3_LoRAs`](https://huggingface.co/WarmBloodAban/Minimax_H3_LoRAs)
23
+ (`Minimax-h3_Third_person_view.safetensors`) on [MiniMax-H3](https://huggingface.co/MiniMaxAI/MiniMax-H3),
24
+ MiniMax's 33B omni-modal model that generates video with natively synchronized audio.
25
+
26
+ The LoRA turns H3 into a **game-cutscene renderer**: over-the-shoulder and FPV camera language,
27
+ Unreal-Engine-looking lighting, and — the part it is actually distinctive for — **legible HUD furniture
28
+ drawn in-frame**: target-lock reticles, floating damage numbers, QTE prompts, translucent status bars.
29
+ The soundtrack (impacts, foley, score) is base H3's, generated in the same denoising pass.
30
+
31
+ ## It is a *reference* LoRA, not a text-to-video one
32
+
33
+ The adapter's own metadata says `ss_base_model_version: minimax_h3_ref2va`, so it was trained on H3's
34
+ **ref2va** partition — the one that conditions on an ordered list of reference media, not the plain
35
+ text-to-video partition. So this Space is built around reference sheets: you upload a character sheet,
36
+ an environment plate, a boss design, and H3 keeps those identities consistent through the shot.
37
+ Running this adapter on the t2v partition would attach cleanly and quietly produce worse output; it
38
+ isn't offered.
39
+
40
+ ## What the demo owes the model card
41
+
42
+ | The card says | Here |
43
+ |---|---|
44
+ | LoRA weight **0.6 – 0.85**, 0.7 recommended | A slider clamped to exactly that window, defaulting to 0.7. The adapter is attached with PEFT and re-scaled per request, so the window is live |
45
+ | A long structured prompt format — `subject_definitions:` / `summary:` / `retention_analysis:` / `detailed_description:` with `[Shot N]` beats / `overall_soundscape:` / `non_diegetic_music:`, cross-referencing `<Subject N>`, `<Environment N>`, `<Picture N>` | The UI **is** that format. Each reference slot you fill becomes a `<Subject N>`/`<Environment N>` bound to its `<Picture N>`; your one action line plus the camera preset become the `detailed_description:` beats; the `summary:` is prefixed `[reference generation]` as the card shows. The assembled document is shown after every run, and an override box takes a hand-written one |
46
+ | Trigger vocabulary in three groups — camera views (*third-person perspective*, *over-the-shoulder camera*, *first-person POV*, *FPV HUD*), game rendering (*rendered in Unreal Engine*, *gameplay sequence*, *combat stance*, *macro close-up*), UI elements (*transparent HUD elements*, *target-lock UI reticle*, *QTE UI prompt*, *floating damage text UI*) | The four camera presets each spend the camera-view + game-rendering tags their shot actually wants; the HUD checkbox spends the UI group. Nothing is sprinkled in decoratively |
47
+ | Base is the **INT8 Reference** architecture | The ref2va DiT, in bf16 on ZeroGPU's 96 GB tier — no quantization needed at this size |
48
+
49
+ ## Architecture
50
+
51
+ MiniMax-H3 is 195.9 GiB in bf16 and a ZeroGPU Space is evicted at 150 GB of storage, so the pipeline is
52
+ cut at its `text_encoder` step — the split every MiniMax-H3 Space uses. The 62 GiB Qwen3-VL conditioner
53
+ runs in [`multimodalart/qwen3vl-conditioner`](https://huggingface.co/spaces/multimodalart/qwen3vl-conditioner)
54
+ and is called over the gradio API with each caller's own ZeroGPU token forwarded, so the booking is
55
+ billed to whoever asked for the video; `prompt_embeds` + `text_token_tags` is the whole wire format.
56
+ This Space holds the DiT and the two autoencoders. `h3_split_blocks.py` holds the denoising-half block
57
+ stack.
58
+
59
+ The DiT is [`multimodalart/MiniMax-H3-Pruned`](https://huggingface.co/multimodalart/MiniMax-H3-Pruned),
60
+ which folds the AdaLN input projections onto a rank-8 subspace: `transformer_ref` drops from 61.7 GiB
61
+ to 37.5 GiB with identical output. This LoRA touches no AdaLN module, so nothing needs projecting.
62
+
63
+ ## How the LoRA is applied
64
+
65
+ Attached as a live **PEFT adapter** (`load_lora_adapter(..., adapter_name="third_person_view")`) and
66
+ re-scaled per request via `set_adapters`, rather than folded into the base weights — folding would be
67
+ faster to serve but would freeze the strength, and 0.6–0.85 is the knob this LoRA is actually steered
68
+ with. For the same reason the pipeline is **not** AoTI-compiled: a captured graph would route around
69
+ the PEFT-injected layers and silently drop the adapter.
70
+
71
+ The export is keyed on the *original* H3 partition (`diffusion_model.blocks.N.…`, 416 tensors, rank 64,
72
+ ai-toolkit), so applying it to the diffusers port means replaying the layout transforms
73
+ [`convert_minimax_h3_to_diffusers.py`](https://github.com/huggingface/diffusers/blob/main/scripts/convert_minimax_h3_to_diffusers.py)
74
+ applied to the base weights — a delta is only valid in the layout of the weight it is added to:
75
+
76
+ | Original LoRA target | diffusers parameter | Transform |
77
+ |---|---|---|
78
+ | `blocks.N.…` / `token_refiner.blocks.N.…` | `transformer_blocks.N.…` / `token_refiner.refiner_blocks.N.…` | rename |
79
+ | `attn.out_proj` | `attn.to_out.0` | rename |
80
+ | `attn.qkv_proj` | `attn.to_q` / `to_k` / `to_v` | split `lora_B`'s rows into contiguous thirds |
81
+ | `mlp.fc2` | `ff.net.2` | rename |
82
+ | `mlp.fc1` | `ff.net.0.proj` | **swap the two fused SwiGLU halves** (`[gate; value]` → `[value; gate]`) |
83
+
84
+ Both of the last two rows are easy to get wrong. The gate/value swap is a real difference between H3's
85
+ native SwiGLU packing and diffusers' `GEGLU`-style ordering — skip it and every MLP delta lands on the
86
+ wrong half of the activation. Conversely the contiguous QKV split is *only* correct for ai-toolkit
87
+ exports like this one; fal-trained H3 LoRAs store per-head-interleaved QKV rows and need de-interleaving
88
+ first. The two are distinguished by key naming, and this repo is unambiguously the former. Any target
89
+ that resolves to no transformer weight is a fatal error rather than a silent skip.
90
+
91
+ ## Examples
92
+
93
+ Three ready-to-run briefs. The wyvern-boss and ruined-village reference sheets are frames lifted from
94
+ the LoRA author's own showcase clip in
95
+ [`WarmBloodAban/Minimax_H3_LoRAs`](https://huggingface.co/WarmBloodAban/Minimax_H3_LoRAs) (Apache-2.0) —
96
+ the closest thing to the inputs the adapter was demonstrated on. The character and city plates come from
97
+ [`linoyts/repo-to-space-example-inputs`](https://huggingface.co/datasets/linoyts/repo-to-space-example-inputs)
98
+ (CC0-1.0).
99
+
100
+ ## Known limits
101
+
102
+ Five seconds of 24 fps video per run, at 960×544 or 1280×720. HUD text is *shaped* like HUD text and
103
+ usually isn't real words — that's a base-H3 limit the LoRA doesn't fix. Faces drift on fast whip-pans.
104
+ The first request after a cold start pays for weight download plus placement.
105
+
106
+ ## License
107
+
108
+ MiniMax-H3 is covered by the **MiniMax H3 Community License Agreement** — note its attribution,
109
+ territorial-scope and revenue clauses. The LoRA is Apache-2.0. Read both before using any output of
110
+ this Space.
app.py ADDED
@@ -0,0 +1,894 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """MiniMax-H3 · Third-Person View LoRA — game cutscenes from design reference sheets.
2
+
3
+ [`WarmBloodAban/Minimax_H3_LoRAs`](https://huggingface.co/WarmBloodAban/Minimax_H3_LoRAs) ships
4
+ `Minimax-h3_Third_person_view.safetensors`, a rank-64 ai-toolkit LoRA whose own metadata names its base as
5
+ `minimax_h3_ref2va` — the **reference** partition of MiniMax-H3, the one that conditions on an ordered list of
6
+ reference images rather than on a first frame. That is why this Space is a `ref2va` deployment: the adapter goes
7
+ onto `transformer_ref`, and the user's uploads arrive as `<Picture 1..3>` in the order they are given.
8
+
9
+ What the LoRA does, from its card: cinematic game cutscenes — third-person over-the-shoulder / spring-arm camera
10
+ tracking, first-person POV, whip-pan view transitions, and native game HUD overlays (target-lock reticles, Boss
11
+ health bars, QTE prompts, floating damage text). Its card is explicit that it wants MiniMax-H3's **structured**
12
+ prompt format (`subject_definitions` / `summary` / `retention_analysis` / `detailed_description` /
13
+ `overall_soundscape` / `non_diegetic_music`) with `<Subject N>` / `<Environment N>` / `<Picture N>` cross-references,
14
+ so this demo is a composer for exactly that document rather than a prompt box: you label each reference sheet, write
15
+ one line of action, pick a camera and whether the HUD is on, and the Space assembles the card's format around the
16
+ LoRA's own trigger vocabulary. The assembled document is shown next to the result, and an override box takes a
17
+ hand-written one.
18
+
19
+ Recipe, all from the card: LoRA weight **0.6–0.85, 0.7 recommended** (a live slider here, which is why the adapter
20
+ is attached with PEFT rather than folded into the weights).
21
+
22
+ Deployment follows the other MiniMax-H3 Spaces. H3 is 195.9 GiB in bfloat16 and a ZeroGPU Space is evicted at 150 GB
23
+ of storage, so `MiniMaxH3Ref2VAGeneratorBlocks` is cut at its `text_encoder` step: the 62 GiB Qwen3-VL conditioner
24
+ runs in [`multimodalart/qwen3vl-conditioner`](https://huggingface.co/spaces/multimodalart/qwen3vl-conditioner) and is
25
+ called over the gradio API, while this Space holds the DiT and the two autoencoders. The DiT is
26
+ [`multimodalart/MiniMax-H3-Pruned`](https://huggingface.co/multimodalart/MiniMax-H3-Pruned)'s `transformer_ref`
27
+ (37.5 GiB — the AdaLN input projections folded onto their reachable rank-8 subspace, 1.5e-5 relative, ~250x below one
28
+ bfloat16 step), which this LoRA does not touch: it adapts `attn.qkv_proj`, `attn.out_proj` and `mlp.fc1`/`fc2` only.
29
+ """
30
+
31
+ from __future__ import annotations
32
+
33
+ import os
34
+ import tempfile
35
+ import time
36
+ import traceback
37
+ from functools import cache
38
+
39
+ # Before anything that could initialize CUDA: `import spaces` patches `torch.cuda` so the weights can be loaded at
40
+ # startup rather than on GPU time.
41
+ import spaces # noqa: F401
42
+ import gradio as gr
43
+ import torch
44
+
45
+ VERSION = "tpv-lora"
46
+ MODEL_REPO = os.environ.get("H3_MODEL_REPO", "multimodalart/MiniMax-H3-Pruned")
47
+ LORA_REPO = os.environ.get("H3_LORA_REPO", "WarmBloodAban/Minimax_H3_LoRAs")
48
+ LORA_FILE = os.environ.get("H3_LORA_FILE", "Minimax-h3_Third_person_view.safetensors")
49
+ ADAPTER = "third_person_view"
50
+ CONDITIONER_SPACE = os.environ.get("H3_CONDITIONER", "multimodalart/qwen3vl-conditioner")
51
+ # `lazy` moves the weights onto the card inside the first GPU call and leaves them there. Startup placement is not an
52
+ # option: `spaces`' startup `torch.pack()` writes a second on-disk copy of every startup-resident CUDA tensor, and
53
+ # 48 GB of weights plus its pack runs at the 150 GB Space storage quota.
54
+ PLACEMENT = os.environ.get("H3_PLACEMENT", "lazy").lower()
55
+ # cuDNN's fused attention is 10-20% faster than the SDPA default on this pool and needs nothing installed.
56
+ # flash-attention 3 is sm90-only and this card is sm120 (the `zero-a10g` flavour name is legacy).
57
+ ATTENTION = os.environ.get("H3_ATTENTION", "_native_cudnn").lower()
58
+ GPU_SIZE = os.environ.get("H3_GPU_SIZE", "xlarge")
59
+ MIN_GPU_DURATION = int(os.environ.get("H3_GPU_DURATION_MIN", "120"))
60
+ MAX_GPU_DURATION = int(os.environ.get("H3_GPU_DURATION_MAX", "1500"))
61
+
62
+ # The card's recommended weight window, verbatim: "LoRA Weight: 0.6 - 0.85 (0.7 is recommended as a starting point)".
63
+ WEIGHT_MIN, WEIGHT_MAX, DEFAULT_WEIGHT = 0.6, 0.85, 0.7
64
+ DEFAULT_STEPS = 20
65
+ DEFAULT_SEED = 42
66
+
67
+ # Must stay identical to the conditioner's table: the *label* goes over the wire, so a canvas that half does not know
68
+ # is rejected there and surfaces as a failure here.
69
+ CANVASES = {
70
+ # 16:9
71
+ "960x544 · 16:9 fast": (544, 960),
72
+ "1024x576 · 16:9 fast": (576, 1024),
73
+ "1152x640 · 16:9": (640, 1152),
74
+ "1280x704 · 16:9": (704, 1280),
75
+ "1344x768 · 16:9 full": (768, 1344),
76
+ # 21:9
77
+ "1152x512 · 21:9 fast": (512, 1152),
78
+ "1536x672 · 21:9 full": (672, 1536),
79
+ # 9:16
80
+ "544x960 · 9:16 fast": (960, 544),
81
+ "640x1152 · 9:16": (1152, 640),
82
+ # 4:3
83
+ "768x576 · 4:3 fast": (576, 768),
84
+ "1024x768 · 4:3 full": (768, 1024),
85
+ }
86
+ # Game cutscenes are widescreen; the cheapest 16:9 bucket keeps a default request inside a few minutes of GPU.
87
+ DEFAULT_CANVAS = "960x544 · 16:9 fast"
88
+ FPS, FRAMES_PER_CHUNK, LATENTS_PER_CHUNK = 24, 17, 5
89
+ MIN_DURATION, MAX_UI_DURATION = 2, 14
90
+ MAX_IMAGE_SLOTS, OPEN_IMAGE_SLOTS = 3, 2
91
+
92
+ # Seconds of GPU one request needs, from the packed sequence it is about to denoise: linear in the rows for the
93
+ # matmuls, quadratic for the attention. Fit shared with the other MiniMax-H3 Spaces on this pool.
94
+ STEP_LINEAR, STEP_QUADRATIC, SAFETY = 1.1745e-4, 3.8396e-9, 1.3
95
+ PLACEMENT_ALLOWANCE = int(os.environ.get("H3_PLACEMENT_ALLOWANCE", "90"))
96
+ AUDIO_LATENTS_PER_SECOND, AUDIO_CHANNELS = 40, 2
97
+ REFERENCE_IMAGE_SHORT_EDGE, CANVAS_MULTIPLE = 2048, 32
98
+ DECODE_BASE, DECODE_PER_DEFAULT_CANVAS, DEFAULT_CANVAS_PIXELS = 15, 25, 960 * 544 * 124
99
+ # A 2048-short-edge reference also becomes Qwen3-VL vision tokens in the text stream. Only the *estimate* label needs
100
+ # a number for them; the booking itself reads the tags the conditioner actually returned.
101
+ ESTIMATED_TEXT_ROWS, ESTIMATED_VISION_ROWS_PER_IMAGE = 900, 1800
102
+
103
+ # ── The LoRA's prompt contract ────────────────────────────────────────────────
104
+ # Roles, camera phrasings and the HUD clause are built out of the card's own "Key Trigger Words & Recommended Tags"
105
+ # (camera views: third-person perspective, over-the-shoulder camera, first-person POV, FPV HUD; game rendering:
106
+ # rendered in Unreal Engine, gameplay sequence, combat stance; UI elements: transparent HUD elements, target-lock UI
107
+ # reticle, QTE UI prompt, floating damage text UI) and out of the structured example prompt it publishes.
108
+
109
+ ROLES = {
110
+ "Character": ("subject", "character design reference sheet"),
111
+ "Creature / Boss": ("subject", "monster design reference sheet"),
112
+ "Prop / Weapon": ("subject", "prop design reference sheet"),
113
+ "Environment": ("environment", "environment design reference sheet"),
114
+ }
115
+
116
+ CAMERAS = {
117
+ "Third-person · over-the-shoulder": {
118
+ "genre": "third-person",
119
+ "summary": "tracked continuously by an over-the-shoulder spring-arm camera",
120
+ "aesthetic": "third-person over-the-shoulder spring-arm camera tracking",
121
+ "opening": "an over-the-shoulder camera locked 3.0 meters directly behind",
122
+ "continues": (
123
+ "The camera holds its over-the-back angle and adjusts its spring-arm distance with every movement, "
124
+ "with hit-impulse micro-shakes on each impact."
125
+ ),
126
+ },
127
+ "Third-person · orbiting combat camera": {
128
+ "genre": "third-person",
129
+ "summary": "tracked by an orbiting third-person combat camera",
130
+ "aesthetic": "orbiting third-person combat camera with a dynamic spring-arm distance",
131
+ "opening": "a third-person combat camera orbiting 4.0 meters around",
132
+ "continues": (
133
+ "The camera orbits around the action and tightens its distance on every impact, keeping the combat "
134
+ "stance centred in frame."
135
+ ),
136
+ },
137
+ "First-person POV (FPV)": {
138
+ "genre": "first-person",
139
+ "summary": "shown entirely in first-person POV with FPV framing",
140
+ "aesthetic": "first-person POV (FPV) camera with weapon-in-hand framing",
141
+ "opening": "a first-person POV camera looking out through the eyes of",
142
+ "continues": (
143
+ "The camera stays in first-person POV throughout, with FPV head-bob, fast view whip-pans and "
144
+ "weapon-in-hand framing at the bottom of frame."
145
+ ),
146
+ },
147
+ "Whip-pan: third-person → first-person": {
148
+ "genre": "third-person",
149
+ "summary": "starting over-the-shoulder and whip-panning into first-person POV",
150
+ "aesthetic": "third-person to first-person perspective transition driven by a fast camera whip-pan",
151
+ "opening": "an over-the-shoulder camera locked 3.0 meters behind",
152
+ "continues": (
153
+ "Mid-sequence the camera whip-pans forward into a first-person POV view and holds it to the end of "
154
+ "the shot."
155
+ ),
156
+ },
157
+ }
158
+ DEFAULT_CAMERA = "Third-person · over-the-shoulder"
159
+
160
+ HUD_CLAUSE = (
161
+ ", and transparent combat HUD elements in the top-right corner: a target-lock UI reticle, a Boss health bar, "
162
+ "floating damage text UI and QTE UI prompts"
163
+ )
164
+ NO_HUD_CLAUSE = ", and a clean cinematic frame with no UI overlays"
165
+ DEFAULT_SOUNDSCAPE = (
166
+ "Footsteps on wet stone, weapon impacts and metallic blade clashes, monstrous roars, thruster and dash bursts, "
167
+ "and crisp action RPG combat UI audio effects."
168
+ )
169
+ DEFAULT_MUSIC = (
170
+ "An intense, high-tempo epic battle track, orchestral-electronic, with heavy industrial percussion and a low "
171
+ "pulsing synth bass that swells through the sequence."
172
+ )
173
+
174
+
175
+ def snap_frames(seconds: float) -> int:
176
+ """The frame count MiniMax-H3's video VAE can decode: the next `17 * n + 5` at 24 fps."""
177
+ frames = max(1, round(float(seconds) * FPS))
178
+ while frames % FRAMES_PER_CHUNK != LATENTS_PER_CHUNK:
179
+ frames += 1
180
+ return frames
181
+
182
+
183
+ def lower_duration_floor(seconds: float = MIN_DURATION) -> None:
184
+ """Let the pipeline generate below its 5 s floor. 56 frames (2.33 s) is fine on this checkpoint."""
185
+ from diffusers.modular_pipelines.minimax_h3.modular_pipeline import MiniMaxH3ModularPipeline
186
+
187
+ MiniMaxH3ModularPipeline.min_duration = property(lambda self: float(seconds))
188
+
189
+
190
+ def video_latent_frames(num_frames: int) -> int:
191
+ """`17 * n + 5` frames become `5 * n + 2` video latents."""
192
+ return 5 * ((num_frames - LATENTS_PER_CHUNK) // FRAMES_PER_CHUNK) + 2
193
+
194
+
195
+ def target_rows(height: int, width: int, num_frames: int) -> int:
196
+ """The generated rows of the packed sequence: video patched `(1, 2, 2)`, plus two audio rows per latent."""
197
+ video = video_latent_frames(num_frames) * (height // CANVAS_MULTIPLE) * (width // CANVAS_MULTIPLE)
198
+ return video + round(num_frames / FPS * AUDIO_LATENTS_PER_SECOND) * AUDIO_CHANNELS
199
+
200
+
201
+ def reference_rows(image_paths: list[str]) -> int:
202
+ """The rows the image reference blocks add, from metadata alone — no decode.
203
+
204
+ A reference image is resized to a 2048-pixel short edge and encoded as a single frame, so a squarer reference is
205
+ a cheaper one.
206
+ """
207
+ from PIL import Image
208
+
209
+ rows = 0
210
+ for path in image_paths:
211
+ if not path:
212
+ continue
213
+ try:
214
+ width, height = Image.open(path).size
215
+ except Exception:
216
+ width, height = 1024, 1024
217
+ scale = REFERENCE_IMAGE_SHORT_EDGE / min(width, height)
218
+ resolved = [
219
+ max(CANVAS_MULTIPLE, round(edge * scale / CANVAS_MULTIPLE) * CANVAS_MULTIPLE) for edge in (height, width)
220
+ ]
221
+ rows += (resolved[0] // CANVAS_MULTIPLE) * (resolved[1] // CANVAS_MULTIPLE)
222
+ return rows
223
+
224
+
225
+ def denoise_seconds(sequence: int, steps: int) -> float:
226
+ return int(steps) * (STEP_LINEAR * sequence + STEP_QUADRATIC * sequence**2) * SAFETY
227
+
228
+
229
+ def get_duration(prompt_embeds, text_token_tags, image_paths, height, width, num_frames, steps, weight, seed, **_):
230
+ """Seconds of GPU to reserve for one request, from the packed sequence it is about to denoise."""
231
+ sequence = int(text_token_tags.shape[0]) + reference_rows(image_paths) + target_rows(height, width, num_frames)
232
+ denoise = denoise_seconds(sequence, steps)
233
+ encode = 5 + reference_rows(image_paths) * 1e-3
234
+ decode = DECODE_BASE + DECODE_PER_DEFAULT_CANVAS * (height * width * num_frames) / DEFAULT_CANVAS_PIXELS
235
+ total = PLACEMENT_ALLOWANCE + encode + denoise + decode + 10
236
+ duration = max(MIN_GPU_DURATION, min(MAX_GPU_DURATION, int(total)))
237
+ print(f"[{VERSION}] S={sequence} -> reserving {duration}s ({denoise:.0f}s of denoise at {steps} steps)", flush=True)
238
+ return duration
239
+
240
+
241
+ def estimate_label(canvas, duration, steps, *image_paths) -> str:
242
+ """The same estimate, phrased for the UI, so the cost of a canvas / duration / steps choice is visible up front."""
243
+ height, width = CANVASES.get(canvas, CANVASES[DEFAULT_CANVAS])
244
+ num_frames = snap_frames(duration)
245
+ paths = [path for path in image_paths if path]
246
+ refs = reference_rows(paths)
247
+ sequence = ESTIMATED_TEXT_ROWS + ESTIMATED_VISION_ROWS_PER_IMAGE * len(paths) + refs
248
+ sequence += target_rows(height, width, num_frames)
249
+ seconds = int(PLACEMENT_ALLOWANCE + 5 + denoise_seconds(sequence, steps) + DECODE_BASE)
250
+ return (
251
+ f"{width}x{height} · {num_frames} frames ({num_frames / FPS:.2f} s) · {int(steps)} steps · "
252
+ f"{len(paths)} reference{'' if len(paths) == 1 else 's'} → roughly "
253
+ f"**{seconds // 60}m {seconds % 60:02d}s** of GPU time"
254
+ )
255
+
256
+
257
+ # ── The LoRA, onto the diffusers port ────────────────────────────────────────
258
+ #
259
+ # `Minimax-h3_Third_person_view.safetensors` is an ai-toolkit export: 416 tensors, rank 64, no `.alpha`, keys
260
+ # `diffusion_model.blocks.N.{attn.qkv_proj,attn.out_proj,mlp.fc1,mlp.fc2}.lora_{A,B}.weight` plus the same four under
261
+ # `token_refiner.blocks.N`, and `ss_base_model_version: minimax_h3_ref2va` in its metadata. diffusers serves the
262
+ # *converted* port, so every name — and, for two of them, the row layout of `lora_B` — has to be pushed through the
263
+ # same transforms `scripts/convert_minimax_h3_to_diffusers.py` applied to the base weights. This reproduces
264
+ # `_convert_non_diffusers_minimax_h3_lora_to_diffusers`:
265
+ #
266
+ # * `blocks.` -> `transformer_blocks.`, `token_refiner.blocks.` -> `token_refiner.refiner_blocks.`,
267
+ # `attn.out_proj` -> `attn.to_out.0`, `mlp.fc2` -> `ff.net.2` (pure renames),
268
+ # * `mlp.fc1` -> `ff.net.0.proj` with its two fused halves swapped: the reference computes `fc2(silu(gate) * value)`
269
+ # from a fused `[gate; value]` while diffusers' `SwiGLU` computes `value * silu(gate)` from a fused
270
+ # `[value; gate]`, so the halves trade places,
271
+ # * `attn.qkv_proj` -> `to_q` / `to_k` / `to_v`: split `lora_B`'s 21504 rows into **contiguous** thirds of 7168
272
+ # (`num_attention_heads * attention_head_dim`).
273
+ #
274
+ # That last split is where MiniMax-H3 LoRAs diverge. DiffSynth-Studio exports run the raw checkpoint's *per-head
275
+ # interleaved* fused QKV and have to be de-interleaved before the split; they are identified by peft's `.default.`
276
+ # infix over these names. ai-toolkit exports under `diffusion_model.` are already `[q_all; k_all; v_all]` and must
277
+ # **not** be reordered — de-interleaving this file anyway would scatter each head's q/k/v across all three
278
+ # projections and turn the adapter into structured noise on all 50 blocks.
279
+ #
280
+ # Row transforms only ever touch `lora_B`, so the three attention projections share one `lora_A`: the rows of `B @ A`
281
+ # are the rows of `B`, which makes the split exact rather than an approximation.
282
+
283
+
284
+ def _lora_target_name(source_name: str) -> str:
285
+ """Map an original-checkpoint module path to the diffusers one (no `.lora_A/B.*` suffix)."""
286
+ if source_name.startswith("token_refiner.blocks."):
287
+ return source_name.replace("token_refiner.blocks.", "token_refiner.refiner_blocks.", 1)
288
+ if source_name.startswith("blocks."):
289
+ return source_name.replace("blocks.", "transformer_blocks.", 1)
290
+ return source_name
291
+
292
+
293
+ def _lora_modules(name: str, a_weight, b_weight, inner_dim: int):
294
+ """Yield `(diffusers_module_path, lora_A, lora_B)` for one original module path."""
295
+ target = _lora_target_name(name)
296
+
297
+ if target.endswith(".attn.qkv_proj"):
298
+ prefix = target.removesuffix("qkv_proj")
299
+ for kind, part in zip(("q", "k", "v"), b_weight.split(inner_dim, dim=0)):
300
+ yield f"{prefix}to_{kind}", a_weight, part.contiguous()
301
+ elif target.endswith(".mlp.fc1"):
302
+ gate, value = b_weight.chunk(2, dim=0)
303
+ yield target.replace(".mlp.fc1", ".ff.net.0.proj"), a_weight, torch.cat([value, gate]).contiguous()
304
+ elif target.endswith(".mlp.fc2"):
305
+ yield target.replace(".mlp.fc2", ".ff.net.2"), a_weight, b_weight
306
+ elif target.endswith(".attn.out_proj"):
307
+ yield target.replace(".attn.out_proj", ".attn.to_out.0"), a_weight, b_weight
308
+ else:
309
+ raise ValueError(
310
+ f"unexpected LoRA target `{name}`: this adapter is documented as attention + feed-forward only, so a "
311
+ "new module type means the checkpoint changed"
312
+ )
313
+
314
+
315
+ def build_lora_state_dict(transformer) -> tuple[dict, int, int]:
316
+ """Download the LoRA and remap it into a PEFT-format state dict for the diffusers transformer.
317
+
318
+ Every target is validated against the transformer's own parameter shapes, and an unresolved one is fatal: a
319
+ silently dropped target means the name mapping is wrong and the Space would serve a half-applied adapter that
320
+ still *looks* like it worked.
321
+ """
322
+ from huggingface_hub import hf_hub_download
323
+ from safetensors.torch import load_file
324
+
325
+ raw = load_file(hf_hub_download(LORA_REPO, LORA_FILE))
326
+
327
+ prefix, suffix_a, suffix_b = "diffusion_model.", ".lora_A.weight", ".lora_B.weight"
328
+ unexpected = [key for key in raw if not (key.startswith(prefix) and key.endswith((suffix_a, suffix_b)))]
329
+ if unexpected:
330
+ raise ValueError(f"{LORA_FILE} holds {len(unexpected)} unexpected tensors, e.g. {unexpected[:5]}")
331
+ bases = sorted({key[len(prefix) : -len(suffix_a)] for key in raw if key.endswith(suffix_a)})
332
+ if not bases:
333
+ raise ValueError(f"No `{prefix}*{suffix_a}` / `{suffix_b}` pairs found in {LORA_FILE}")
334
+
335
+ ranks = set()
336
+ for name in bases:
337
+ if f"{prefix}{name}{suffix_b}" not in raw:
338
+ raise ValueError(f"LoRA is missing the lora_B twin of {prefix}{name}{suffix_a}")
339
+ ranks.add(raw[f"{prefix}{name}{suffix_a}"].shape[0])
340
+ if len(ranks) != 1:
341
+ raise ValueError(f"LoRA mixes ranks {sorted(ranks)}; this loader assumes a single rank")
342
+ rank = ranks.pop()
343
+
344
+ inner_dim = transformer.config.num_attention_heads * transformer.config.attention_head_dim
345
+ base_shapes = {key: tuple(value.shape) for key, value in transformer.state_dict().items()}
346
+
347
+ state_dict: dict[str, torch.Tensor] = {}
348
+ missed: list[str] = []
349
+ for name in bases:
350
+ a_weight = raw[f"{prefix}{name}{suffix_a}"]
351
+ b_weight = raw[f"{prefix}{name}{suffix_b}"]
352
+ for module, a_part, b_part in _lora_modules(name, a_weight, b_weight, inner_dim):
353
+ base = base_shapes.get(f"{module}.weight")
354
+ if base is None:
355
+ missed.append(module)
356
+ continue
357
+ # `W` is [out, in]; the adapter must be `lora_B` [out, r] @ `lora_A` [r, in].
358
+ if (b_part.shape[0], a_part.shape[1]) != base:
359
+ raise ValueError(
360
+ f"LoRA delta for `{module}` would be {(b_part.shape[0], a_part.shape[1])}, "
361
+ f"base weight is {base}"
362
+ )
363
+ state_dict[f"{module}.lora_A.weight"] = a_part
364
+ state_dict[f"{module}.lora_B.weight"] = b_part
365
+ if missed:
366
+ raise ValueError(
367
+ f"{len(missed)} LoRA targets matched no transformer weight, e.g. {missed[:5]}. "
368
+ "The LoRA and the diffusers transformer disagree on module naming."
369
+ )
370
+ return state_dict, rank, len(bases)
371
+
372
+
373
+ def load_and_apply_lora(transformer) -> str:
374
+ """Attach the LoRA as a PEFT adapter so its strength stays a per-request knob."""
375
+ state_dict, rank, targets = build_lora_state_dict(transformer)
376
+ # `prefix=None` because the keys are already the transformer's own module paths, and no `network_alphas` because
377
+ # the file carries no `.alpha` tensors — PEFT then sets alpha == rank, i.e. the adapter's intrinsic scale is 1.0
378
+ # and the `set_adapters` weight *is* the strength the card talks about.
379
+ transformer.load_lora_adapter(state_dict, prefix=None, adapter_name=ADAPTER)
380
+ transformer.set_adapters([ADAPTER], [DEFAULT_WEIGHT])
381
+ return (
382
+ f"LoRA attached · {targets} checkpoint targets -> {len(state_dict) // 2} diffusers modules, rank {rank}, "
383
+ f"alpha == rank · default weight {DEFAULT_WEIGHT:g} · {LORA_FILE}"
384
+ )
385
+
386
+
387
+ # ── Model loading ────────────────────────────────────────────────────────────
388
+
389
+ PIPE = None
390
+ LOAD_ERROR: str | None = None
391
+ LORA_STATUS: str | None = None
392
+
393
+
394
+ def load_models() -> str | None:
395
+ """Load the denoising half at startup, but *not* onto the card — see `PLACEMENT`.
396
+
397
+ `MiniMaxH3Ref2VAGeneratorBlocks` declares `transformer_ref`, `vae`, `audio_vae`, the two schedulers and
398
+ `video_processor`, so `load_components` fetches exactly those subfolders — `text_encoder/` and the `transformer/`
399
+ partition are never touched. Both autoencoders carry `_keep_in_fp32_modules` over every module and stay float32:
400
+ a bfloat16 audio VAE decodes the soundtrack roughly 20 dB too quiet.
401
+ """
402
+ global PIPE, LOAD_ERROR, LORA_STATUS
403
+
404
+ if PIPE is not None or LOAD_ERROR is not None:
405
+ return LOAD_ERROR
406
+
407
+ started = time.time()
408
+ try:
409
+ from diffusers import ComponentsManager
410
+
411
+ from h3_split_blocks import MiniMaxH3Ref2VAGeneratorBlocks
412
+
413
+ lower_duration_floor()
414
+ blocks = MiniMaxH3Ref2VAGeneratorBlocks()
415
+ print(f"[{VERSION}] loading {[c.name for c in blocks.expected_components]} from {MODEL_REPO} ...", flush=True)
416
+ pipe = blocks.init_pipeline(MODEL_REPO, components_manager=ComponentsManager(), collection="h3")
417
+ # The pruned DiT is served as remote code (`transformer_ref/modeling_minimax_h3_pruned.py`, reached through
418
+ # the `AutoModel` type hint in `modular_model_index.json`). `load_components` forwards `trust_remote_code`
419
+ # only to components that live in the pipeline's own repo, so the two VAEs and the schedulers — which
420
+ # `modular_model_index.json` still points at `MiniMaxAI/MiniMax-H3` — never see it.
421
+ pipe.load_components(dtype=torch.bfloat16, trust_remote_code=True)
422
+
423
+ # Both VAEs explicitly, and before the transformer: `set_attention_backend` also sets the registry's *global*
424
+ # backend, and the float32 audio VAE has no cuDNN kernel.
425
+ pipe.vae.set_attention_backend("native")
426
+ pipe.audio_vae.set_attention_backend("native")
427
+ pipe.transformer_ref.set_attention_backend(ATTENTION)
428
+
429
+ # Before any GPU placement, so the adapter's parameters travel with the transformer. No AoTI package on this
430
+ # Space on purpose: a compiled block graph is captured around the base linears and would silently bypass the
431
+ # LoRA layers PEFT injects.
432
+ LORA_STATUS = load_and_apply_lora(pipe.transformer_ref)
433
+ print(f"[{VERSION}] {LORA_STATUS}", flush=True)
434
+
435
+ import h3_fbc
436
+
437
+ print(f"[h3-fbc] {h3_fbc.status()}", flush=True)
438
+
439
+ PIPE = pipe
440
+ print(f"[{VERSION}] ready in {time.time() - started:.0f}s", flush=True)
441
+ except Exception as error:
442
+ traceback.print_exc()
443
+ LOAD_ERROR = (
444
+ f"**Loading `{MODEL_REPO}` failed** after {time.time() - started:.0f}s: "
445
+ f"`{type(error).__name__}: {error}`"
446
+ )
447
+ return LOAD_ERROR
448
+
449
+
450
+ # ── Prompt composition ───────────────────────────────────────────────────────
451
+
452
+
453
+ def collect_slots(slots) -> list[tuple[str, str, str]]:
454
+ """The `(path, role, description)` references of a request, in the order the model reads them.
455
+
456
+ That order numbers `<Picture N>` and advances the shared audio/video rotary clock, so the same references in a
457
+ different order are a different request.
458
+ """
459
+ return [(path, role, (detail or "").strip()) for path, role, detail in slots if path]
460
+
461
+
462
+ def compose_prompt(references, action, camera, hud, soundscape, music, seconds) -> str:
463
+ """Assemble MiniMax-H3's structured document in the layout the LoRA's card publishes."""
464
+ action = (action or "").strip().rstrip(".")
465
+ spec = CAMERAS.get(camera, CAMERAS[DEFAULT_CAMERA])
466
+
467
+ entries, subjects, environments = [], 0, 0
468
+ for index, (_, role, detail) in enumerate(references, start=1):
469
+ kind, sheet = ROLES.get(role, ROLES["Character"])
470
+ if kind == "subject":
471
+ subjects += 1
472
+ label = f"<Subject {subjects}>"
473
+ else:
474
+ environments += 1
475
+ label = f"<Environment {environments}>"
476
+ entries.append((label, index, detail, sheet))
477
+
478
+ subject = next((label for label, _, _, sheet in entries if "environment" not in sheet), "the player character")
479
+ environment = next((label for label, _, _, sheet in entries if "environment" in sheet), None)
480
+ place = f" in {environment}" if environment else ""
481
+
482
+ definitions = [
483
+ f"{label} is {detail} from <Picture {index}>."
484
+ if detail
485
+ else f"{label} is the subject shown in <Picture {index}>."
486
+ for label, index, detail, _ in entries
487
+ ]
488
+ definitions += [f"<Picture {index}> is the {sheet} for {label}." for label, index, _, sheet in entries]
489
+
490
+ retention = [
491
+ f"{label} (appears in [Shot 1]): fully_preserved - the design, colours and silhouette from "
492
+ f"<Picture {index}> are fully preserved."
493
+ for label, index, _, _ in entries
494
+ ]
495
+
496
+ sections = [
497
+ "subject_definitions:\n" + "\n".join(definitions),
498
+ (
499
+ "summary:\n"
500
+ f"[reference generation] A {float(seconds):g}-second photorealistic {spec['genre']} action RPG "
501
+ f"gameplay sequence{place}, where {subject} {action}, {spec['summary']}."
502
+ ),
503
+ "retention_analysis:\n" + "\n".join(retention),
504
+ (
505
+ "detailed_description:\n"
506
+ f"The target video features a photorealistic 3D {spec['genre']} action RPG gameplay aesthetic rendered "
507
+ f"in Unreal Engine with real-time game mechanics, {spec['aesthetic']}, deep depth of field"
508
+ f"{HUD_CLAUSE if hud else NO_HUD_CLAUSE}.\n"
509
+ f"[Shot 1] The video opens with {spec['opening']} {subject}{place}. {action.capitalize()}. "
510
+ f"{spec['continues']}"
511
+ ),
512
+ ]
513
+ if (soundscape or "").strip():
514
+ sections.append("overall_soundscape:\n" + soundscape.strip())
515
+ if (music or "").strip():
516
+ sections.append("non_diegetic_music:\n" + music.strip())
517
+ return "\n\n".join(sections)
518
+
519
+
520
+ # ── Guard, conditioner ───────────────────────────────────────────────────────
521
+
522
+
523
+ def check_prompt(prompt: str) -> None:
524
+ """The NCII guard. Every request here carries an uploaded image, so every request is checked; it runs before the
525
+ conditioner call and the denoise booking, so a refused prompt costs no GPU time on either half."""
526
+ import ncii_guard
527
+
528
+ try:
529
+ flag = ncii_guard.classify(prompt)
530
+ except Exception as error:
531
+ traceback.print_exc()
532
+ raise gr.Error(f"The content filter is unavailable (`{type(error).__name__}`), so nothing was run.")
533
+ if flag["label"] == "ncii":
534
+ print(f"[guard] prompt refused (ncii {flag['score']:.2f})", flush=True)
535
+ raise gr.Error("This prompt was flagged by a content filter and wasn't run.")
536
+
537
+
538
+ @cache
539
+ def conditioner():
540
+ """The other half, over the gradio API. `gradio_client` attaches the caller's own ZeroGPU token per call, so the
541
+ conditioner's booking is billed to whoever asked for the video."""
542
+ from gradio_client import Client
543
+
544
+ return Client(CONDITIONER_SPACE)
545
+
546
+
547
+ def encode_remote(prompt, image_paths, canvas, num_frames):
548
+ """`/encode_ref2va` on the conditioner Space: a safetensors file holding `prompt_embeds` + `text_token_tags`, with
549
+ the resolved `height` / `width` / `num_frames` in its metadata, plus the plan.
550
+
551
+ `canvas` is the label. `media` and `kinds` are parallel and ordered; the references go over because `ref2va`'s
552
+ presentation puts a vision block in front of the prompt for every image.
553
+ """
554
+ from gradio_client import handle_file
555
+ from safetensors import safe_open
556
+
557
+ path, plan = conditioner().predict(
558
+ prompt=prompt,
559
+ media=[handle_file(image) for image in image_paths],
560
+ kinds=",".join("image" for _ in image_paths),
561
+ canvas=canvas,
562
+ num_frames=num_frames,
563
+ rewrite_prompt=False,
564
+ api_name="/encode_ref2va",
565
+ )
566
+ with safe_open(path, framework="pt") as handle:
567
+ return handle.get_tensor("prompt_embeds"), handle.get_tensor("text_token_tags"), handle.metadata(), plan
568
+
569
+
570
+ # ── Inference ────────────────────────────────────────────────────────────────
571
+
572
+
573
+ @spaces.GPU(duration=get_duration, size=GPU_SIZE)
574
+ def _generate(prompt_embeds, text_token_tags, image_paths, height, width, num_frames, steps, weight, seed):
575
+ """The only thing on GPU time: the reference encoder, the packed-sequence denoise loop and the decoders.
576
+
577
+ References cross as paths and are decoded here; only the generated outputs come back. A `@spaces.GPU` argument
578
+ crosses a process boundary by pickling, and expanded reference frames are large.
579
+ """
580
+ from diffusers.modular_pipelines.minimax_h3 import MiniMaxH3ImageReference
581
+
582
+ import h3_fbc
583
+
584
+ if PLACEMENT == "lazy":
585
+ PIPE.to("cuda")
586
+
587
+ # The whole reason the adapter is not folded into the weights: strength is a per-request scale on the LoRA layers.
588
+ PIPE.transformer_ref.set_adapters([ADAPTER], [float(weight)])
589
+
590
+ with h3_fbc.enabled(PIPE.transformer_ref, steps=int(steps)):
591
+ state = PIPE(
592
+ prompt_embeds=prompt_embeds.to("cuda"),
593
+ text_token_tags=text_token_tags,
594
+ references=[MiniMaxH3ImageReference.from_file(path) for path in image_paths],
595
+ height=height,
596
+ width=width,
597
+ num_frames=num_frames,
598
+ num_inference_steps=int(steps),
599
+ generator=torch.Generator("cpu").manual_seed(int(seed)),
600
+ )
601
+ return state.get("videos")[0], state.get("audio")[0].cpu(), state.get("sampling_rate")
602
+
603
+
604
+ def generate(
605
+ # The first nine are the columns `gr.Examples` varies, and they lead the signature for that reason: an example row
606
+ # is applied to `inputs` positionally. Every parameter has a default, which is what lets a nine-column row call
607
+ # this at all.
608
+ action="",
609
+ picture_1=None,
610
+ role_1="Character",
611
+ detail_1="",
612
+ picture_2=None,
613
+ role_2="Environment",
614
+ detail_2="",
615
+ camera=DEFAULT_CAMERA,
616
+ hud=True,
617
+ picture_3=None,
618
+ role_3="Creature / Boss",
619
+ detail_3="",
620
+ canvas=DEFAULT_CANVAS,
621
+ duration=5,
622
+ steps=DEFAULT_STEPS,
623
+ weight=DEFAULT_WEIGHT,
624
+ seed=DEFAULT_SEED,
625
+ soundscape=DEFAULT_SOUNDSCAPE,
626
+ music=DEFAULT_MUSIC,
627
+ override="",
628
+ progress=gr.Progress(track_tqdm=True),
629
+ ):
630
+ """One request: compose the structured prompt, condition it remotely, denoise here."""
631
+ if LOAD_ERROR:
632
+ raise gr.Error(LOAD_ERROR)
633
+ if PIPE is None:
634
+ raise gr.Error("The denoiser is still loading.")
635
+
636
+ from diffusers.utils import encode_video
637
+
638
+ references = collect_slots(
639
+ [(picture_1, role_1, detail_1), (picture_2, role_2, detail_2), (picture_3, role_3, detail_3)]
640
+ )
641
+ if not references:
642
+ raise gr.Error("Add at least one reference sheet — a character, an environment or a creature.")
643
+ if not (override or "").strip() and not (action or "").strip():
644
+ raise gr.Error("Write one line of action, or paste a full structured prompt in the advanced options.")
645
+
646
+ num_frames = snap_frames(duration)
647
+ prompt = (override or "").strip() or compose_prompt(
648
+ references, action, camera, hud, soundscape, music, num_frames / FPS
649
+ )
650
+ check_prompt(prompt)
651
+
652
+ image_paths = [path for path, _, _ in references]
653
+ progress(0.0, desc="Reading the prompt and the reference sheets ...")
654
+ conditioned = time.time()
655
+ try:
656
+ prompt_embeds, text_token_tags, metadata, plan = encode_remote(prompt, image_paths, canvas, num_frames)
657
+ except gr.Error:
658
+ raise
659
+ except Exception as error:
660
+ # gradio only puts the exception *type* on the wire, so the useful half of a conditioner-side failure is in
661
+ # that Space's logs.
662
+ traceback.print_exc()
663
+ raise gr.Error(
664
+ f"The conditioner ({CONDITIONER_SPACE}) failed with `{type(error).__name__}: {error}`. "
665
+ "Its logs carry the full traceback."
666
+ ) from error
667
+ condition_seconds = time.time() - conditioned
668
+ height, width, num_frames = (int(metadata[key]) for key in ("height", "width", "num_frames"))
669
+
670
+ progress(0.1, desc=f"Rendering {num_frames / FPS:.1f} s at {width}x{height} ...")
671
+ started = time.time()
672
+ frames, audio, sampling_rate = _generate(
673
+ prompt_embeds, text_token_tags, image_paths, height, width, num_frames, steps, weight, seed
674
+ )
675
+ generate_seconds = time.time() - started
676
+
677
+ directory = os.path.join(tempfile.gettempdir(), "h3-outputs")
678
+ os.makedirs(directory, exist_ok=True)
679
+ path = os.path.join(directory, f"h3-tpv-{int(time.time() * 1000)}.mp4")
680
+ encode_video(frames, fps=FPS, output_path=path, audio=audio, audio_sample_rate=sampling_rate)
681
+
682
+ print(
683
+ f"[{VERSION}] {len(image_paths)} references · {width}x{height}, {num_frames} frames "
684
+ f"({num_frames / FPS:.3f} s), {int(steps)} steps, weight {float(weight):g} · conditioner "
685
+ f"{condition_seconds:.0f}s ({plan['num_text_tokens']} tokens) · denoise + decode {generate_seconds:.0f}s "
686
+ f"({generate_seconds / int(steps):.1f} s/step) · seed {int(seed)}",
687
+ flush=True,
688
+ )
689
+ return path, prompt
690
+
691
+
692
+ # ── UI ───────────────────────────────────────────────────────────────────────
693
+
694
+ try:
695
+ import ncii_guard
696
+
697
+ ncii_guard.start()
698
+ except Exception as guard_error: # the guard revives itself per request; a cold cache should not fail the boot
699
+ print(f"[guard] start failed ({type(guard_error).__name__}: {guard_error}); it will be retried per request")
700
+ load_models()
701
+
702
+ INTRO = """# MiniMax-H3 · Third-Person View LoRA
703
+
704
+ <div align="center">
705
+ <a href="https://huggingface.co/WarmBloodAban/Minimax_H3_LoRAs" target="_blank" rel="noopener"><strong>[ LoRA ]</strong></a> &nbsp;
706
+ <a href="https://huggingface.co/MiniMaxAI/MiniMax-H3" target="_blank" rel="noopener"><strong>[ base model ]</strong></a> &nbsp;
707
+ <a href="https://huggingface.co/spaces/multimodalart/minimax-h3-reference" target="_blank" rel="noopener"><strong>[ base demo ]</strong></a>
708
+ </div>
709
+
710
+ **Game cutscenes out of design reference sheets.** Label each sheet you upload, write one line of action, pick a
711
+ camera, and this composes MiniMax-H3's structured prompt around the LoRA's own vocabulary — third-person
712
+ over-the-shoulder / spring-arm tracking, first-person POV, whip-pan transitions, target-lock reticles and Boss health
713
+ bars — then renders the clip **with its synchronized soundtrack** in a single denoising pass.
714
+ """
715
+
716
+ CSS = """
717
+ .main.fillable { max-width: 1250px !important; }
718
+ .dark .gradio-container { color: var(--body-text-color); }
719
+ """
720
+
721
+ with gr.Blocks(title="MiniMax-H3 · Third-Person View LoRA", theme=gr.themes.Citrus(), css=CSS) as demo:
722
+ gr.Markdown(INTRO)
723
+ if LOAD_ERROR:
724
+ gr.Markdown(LOAD_ERROR)
725
+
726
+ with gr.Row():
727
+ with gr.Column():
728
+ gr.Markdown("### 1 · Reference sheets — they become `<Picture 1..3>`, in this order")
729
+ pictures, roles, details, slot_columns = [], [], [], []
730
+ with gr.Row():
731
+ for index in range(MAX_IMAGE_SLOTS):
732
+ with gr.Column(min_width=180, visible=index < OPEN_IMAGE_SLOTS) as column:
733
+ pictures.append(
734
+ gr.Image(label=f"Picture {index + 1}", type="filepath", height=190, show_label=True)
735
+ )
736
+ roles.append(
737
+ gr.Dropdown(
738
+ label="Role",
739
+ choices=list(ROLES),
740
+ value="Character" if index == 0 else ("Environment" if index == 1 else "Creature / Boss"),
741
+ container=True,
742
+ )
743
+ )
744
+ details.append(
745
+ gr.Textbox(
746
+ label="What's in it",
747
+ lines=2,
748
+ placeholder="the mint-green cat mecha with a tattered beige cape",
749
+ )
750
+ )
751
+ slot_columns.append(column)
752
+ add_slot = gr.Button("+ Add a third reference sheet", size="sm", variant="secondary")
753
+
754
+ action = gr.Textbox(
755
+ label="2 · What happens in the shot",
756
+ lines=2,
757
+ placeholder="dodges a leaping wyvern attack, counterattacks with an energy-blade combo, then parries "
758
+ "a wing-claw swipe",
759
+ )
760
+ with gr.Row():
761
+ camera = gr.Dropdown(label="3 · Camera", choices=list(CAMERAS), value=DEFAULT_CAMERA, scale=3)
762
+ hud = gr.Checkbox(label="Combat HUD overlay", value=True, scale=1)
763
+ run = gr.Button("Render the cutscene", variant="primary")
764
+ estimate = gr.Markdown(estimate_label(DEFAULT_CANVAS, 5, DEFAULT_STEPS, None))
765
+
766
+ with gr.Accordion("Advanced options", open=False):
767
+ weight = gr.Slider(
768
+ label="LoRA weight (the card recommends 0.6–0.85)",
769
+ minimum=WEIGHT_MIN,
770
+ maximum=WEIGHT_MAX,
771
+ step=0.05,
772
+ value=DEFAULT_WEIGHT,
773
+ )
774
+ canvas = gr.Dropdown(label="Canvas", choices=list(CANVASES), value=DEFAULT_CANVAS)
775
+ duration = gr.Slider(
776
+ label="Duration (s)", minimum=MIN_DURATION, maximum=MAX_UI_DURATION, step=1, value=5
777
+ )
778
+ steps = gr.Slider(label="Steps", minimum=10, maximum=40, step=1, value=DEFAULT_STEPS)
779
+ seed = gr.Number(label="Seed", value=DEFAULT_SEED, precision=0)
780
+ soundscape = gr.Textbox(label="Diegetic soundscape", lines=2, value=DEFAULT_SOUNDSCAPE)
781
+ music = gr.Textbox(label="Non-diegetic music", lines=2, value=DEFAULT_MUSIC)
782
+ override = gr.Textbox(
783
+ label="Write the structured prompt myself (overrides everything above)",
784
+ lines=6,
785
+ placeholder="subject_definitions:\n<Subject 1> is ... from <Picture 1>.\n\nsummary:\n"
786
+ "[reference generation] A 10-second photorealistic third-person ...",
787
+ )
788
+
789
+ with gr.Column():
790
+ result = gr.Video(label="Cutscene + soundtrack")
791
+ with gr.Accordion("Structured prompt sent to MiniMax-H3", open=False):
792
+ prompt_view = gr.Textbox(show_label=False, lines=18, interactive=False)
793
+
794
+ open_slots = gr.State(OPEN_IMAGE_SLOTS)
795
+
796
+ def reveal_slot(open_count):
797
+ open_count = min(open_count + 1, MAX_IMAGE_SLOTS)
798
+ return [
799
+ open_count,
800
+ *[gr.update(visible=index < open_count) for index in range(MAX_IMAGE_SLOTS)],
801
+ gr.update(visible=open_count < MAX_IMAGE_SLOTS),
802
+ ]
803
+
804
+ add_slot.click(reveal_slot, open_slots, [open_slots, *slot_columns, add_slot], api_name=False)
805
+
806
+ for control in (canvas, duration, steps, *pictures):
807
+ control.change(
808
+ estimate_label,
809
+ [canvas, duration, steps, *pictures],
810
+ estimate,
811
+ show_progress="hidden",
812
+ api_name=False,
813
+ )
814
+
815
+ request = [
816
+ action,
817
+ pictures[0],
818
+ roles[0],
819
+ details[0],
820
+ pictures[1],
821
+ roles[1],
822
+ details[1],
823
+ camera,
824
+ hud,
825
+ pictures[2],
826
+ roles[2],
827
+ details[2],
828
+ canvas,
829
+ duration,
830
+ steps,
831
+ weight,
832
+ seed,
833
+ soundscape,
834
+ music,
835
+ override,
836
+ ]
837
+ exampled = request[:9]
838
+
839
+ gr.Examples(
840
+ examples=[
841
+ [
842
+ "dodges a leaping wyvern attack, counterattacks with a glowing energy-blade combo, then parries a "
843
+ "wing-claw swipe in a shower of sparks",
844
+ "examples/katana_character.jpg",
845
+ "Character",
846
+ "the young woman with long black hair in two buns, blue eyes, a black mouth mask, an orange utility "
847
+ "jacket and two white-corded katana hilts rising behind her shoulders",
848
+ "examples/ruined_village.jpg",
849
+ "Environment",
850
+ "the foggy ruined village street with wooden houses, a glowing paper lantern, blood-red spider "
851
+ "lilies and wet stone paving",
852
+ "Third-person · over-the-shoulder",
853
+ True,
854
+ ],
855
+ [
856
+ "plants her feet into a combat stance and charges an ultimate finisher as the pale wyvern staggers "
857
+ "and opens its ring-like tooth-lined maw",
858
+ "examples/katana_character.jpg",
859
+ "Character",
860
+ "the young woman with long black hair in two buns, an orange utility jacket and two katana hilts "
861
+ "rising behind her shoulders",
862
+ "examples/wyvern_boss.jpg",
863
+ "Creature / Boss",
864
+ "the pale eyeless wyvern with a circular ring-like maw lined with rows of sharp teeth, a long "
865
+ "muscular neck, clawed wings and heavy bipedal limbs",
866
+ "Whip-pan: third-person → first-person",
867
+ True,
868
+ ],
869
+ [
870
+ "sprints across a rain-slick rooftop, vaults a neon billboard gap and lands into a slide while "
871
+ "tracer fire streaks past",
872
+ "examples/doorway_character.jpg",
873
+ "Character",
874
+ "the young girl with brown hair, a white shirt and skirt and a guitar case slung across her back",
875
+ "examples/night_city.jpg",
876
+ "Environment",
877
+ "the night city of lit high-rise towers and a faceted glass skyscraper reflected in the dark river "
878
+ "below",
879
+ "First-person POV (FPV)",
880
+ True,
881
+ ],
882
+ ],
883
+ inputs=exampled,
884
+ outputs=[result, prompt_view],
885
+ fn=generate,
886
+ cache_examples=True,
887
+ cache_mode="lazy",
888
+ )
889
+
890
+ run.click(generate, request, [result, prompt_view], api_name="generate")
891
+
892
+
893
+ if __name__ == "__main__":
894
+ demo.launch(show_error=True, max_threads=1000)
examples/doorway_character.jpg ADDED

Git LFS Details

  • SHA256: 2d9a4aee6e9ccf5e2b57653622ab92c3d16acaa7c6d12054974870e162321899
  • Pointer size: 131 Bytes
  • Size of remote file: 122 kB
examples/katana_character.jpg ADDED

Git LFS Details

  • SHA256: b49d114bf121cbab2e8b695a0a60b846fef7d7bfdcaaf3ff0ccbd91cea156cbf
  • Pointer size: 131 Bytes
  • Size of remote file: 106 kB
examples/night_city.jpg ADDED

Git LFS Details

  • SHA256: a5e679cb38b8f85352a4e76ac52e532889bd644616448109f98e698375fa91af
  • Pointer size: 131 Bytes
  • Size of remote file: 189 kB
examples/ruined_village.jpg ADDED
examples/wyvern_boss.jpg ADDED

Git LFS Details

  • SHA256: e52f10ab69cca4e04f016861e84f9099fa220305b805b5fe93e576732c154ecb
  • Pointer size: 131 Bytes
  • Size of remote file: 127 kB
h3_fbc.py ADDED
@@ -0,0 +1,357 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """First-block cache for MiniMax-H3: skip the trunk on steps whose block-0 residual barely moved.
2
+
3
+ Block 0 and the final AdaLN head run at the true timestep on **every** step; only blocks 1..49 are skipped, and only
4
+ while the residual they would have been handed looks like the one from the last step that actually ran. The decision
5
+ signal is the relative L1 between this step's block-0 residual (`block0_out - block0_in`) and the residual of the last
6
+ *computed* step, over the whole packed sequence. On a skip the trunk's contribution is replayed as a cached residual,
7
+ `(final_trunk_out - block0_out)` of that computed step.
8
+
9
+ Ported from `duckyshell/ComfyUI-MiniMaxH3-FirstBlockCache` (`nodes.py` @ 725973c) — same signal, same protected window
10
+ (10%-95% of the schedule, converted through the video shift of 12.0) and the same cap of two consecutive skips, which
11
+ is their "H3 Safe" preset at threshold 0.08.
12
+
13
+ **The audio exemption is not theirs.** duckyshell has no audio term anywhere: the decision is ~98% video by row count
14
+ and every audio row rides the stale trunk residual, which is what costs the soundtrack its energy. Audio runs its own
15
+ schedule (shift 3 against video's 12, diverging by up to 11x in rate across 20 steps), so on a skip the audio rows of
16
+ the trunk output are instead a 2-point **linear** extrapolation of the last two actually-computed audio features, in
17
+ the audio sigma coordinate — `xmarre/ComfyUI-Spectrum-MiniMax-H3`'s `audio_blend_weight=0.0` path, applied where
18
+ Spectrum applies it (the post-trunk hidden feature, ahead of the head that still runs at the true timestep) and fixed
19
+ to the right coordinate. Following ComfyUI PR #15390, no carried audio tensor is ever mutated: every write is a fresh
20
+ tensor out of `index_copy`.
21
+
22
+ It earns its place. Running this Space's own request at threshold 0.08 with `H3_FBC_AUDIO_EXEMPT=0` — the duckyshell
23
+ mechanism verbatim, same 13 skipped forwards, same video to within 0.2 dB — costs the soundtrack **16.7% of its RMS**
24
+ (0.0483 against an uncached 0.0580), while the exemption holds it at 1.04x. That is the same direction the offline
25
+ study measured on a near-silent clip, three times the size on one with real audio energy.
26
+
27
+ **A threshold does not travel across step counts.** The signature shrinks as the schedule is subdivided, so the same
28
+ number gates far more loosely at more steps. Measured on this Space — 960x544, 124 frames, one image reference,
29
+ seed 42, AoTI blocks, at its **default 28 steps** (27 forwards) — against the same request with `H3_FBC=0`:
30
+
31
+ | threshold | skipped | denoise loop | end to end | audio RMS vs uncached |
32
+ |---|---|---|---|---|
33
+ | 0.03 | 0 / 27 | 1.02x | 1.06x | 1.000 (bitwise-identical audio) |
34
+ | 0.05 | 9 / 27 | 1.46x | 1.36x | 0.956 |
35
+ | 0.08 | 13 / 27 | 2.19x | 1.82x | 1.040 |
36
+
37
+ The offline study calibrated 0.08 over a 20-step schedule, where it skipped 7 of 19 forwards. At 28 steps that same
38
+ 0.08 skips 13 of 27 — nearly half — and the sampled video visibly re-rolls its background detail. **0.05 is the
39
+ default** because it reproduces the skip fraction that study validated (33% here against 37% there); 0.08 is the
40
+ aggressive setting. Below the signature floor — 0.03 skipped nothing at all here — a threshold buys nothing and still
41
+ pays for the signature, so lower is not safer, it is just slower.
42
+
43
+ A cached request is not the uncached one: the trajectory moves, so the video is a different sample of the same prompt
44
+ (same shot, same subject, same quality — different signage and background detail). `H3_FBC=0` restores today's output
45
+ exactly, and is worth reaching for when a request has to reproduce a specific earlier result.
46
+
47
+ This composes with `h3_aoti`: that module patches each of the 50 blocks' own `forward` and they stay a real
48
+ `ModuleList`, so skipping the trunk simply does not call blocks 1..49 that step. `LazyAOTIModel` rebinds its constants
49
+ whenever the weights dict it is handed changes identity, and the first forward of a request never skips, so every
50
+ block has bound its own weights before any step is cached.
51
+ """
52
+
53
+ from __future__ import annotations
54
+
55
+ import contextlib
56
+ import os
57
+ import types
58
+
59
+ ENABLED = os.environ.get("H3_FBC", "1") == "1"
60
+ # Relative-L1 gate on the block-0 residual, calibrated at this Space's default 28 steps — see the table above.
61
+ THRESHOLD = float(os.environ.get("H3_FBC_THRESHOLD", "0.05"))
62
+ # duckyshell's cap. Without it the gate compares against an ever-older computed step and drifts away unbounded.
63
+ MAX_CONSECUTIVE_HITS = int(os.environ.get("H3_FBC_MAX_CONSECUTIVE", "2"))
64
+ # The protected head and tail of the schedule, as fractions, converted to sigma through the video shift below.
65
+ START_PERCENT = float(os.environ.get("H3_FBC_START_PERCENT", "0.10"))
66
+ END_PERCENT = float(os.environ.get("H3_FBC_END_PERCENT", "0.95"))
67
+ # `MiniMaxH3SetTimestepsStep` builds the video schedule at shift 12.0 and the audio one at 3.0.
68
+ VIDEO_SHIFT = float(os.environ.get("H3_FBC_VIDEO_SHIFT", "12.0"))
69
+ AUDIO_EXEMPT = os.environ.get("H3_FBC_AUDIO_EXEMPT", "1") == "1"
70
+
71
+ # Every keyword `MiniMaxH3LoopDenoiser` passes. It filters the packed-sequence layout through
72
+ # `inspect.signature(transformer.forward).parameters`, so a replacement forward that drops a name silently stops
73
+ # receiving it; `install` refuses rather than let that happen quietly.
74
+ FORWARD_PARAMETERS = (
75
+ "hidden_states",
76
+ "audio_hidden_states",
77
+ "encoder_hidden_states",
78
+ "timestep",
79
+ "timestep_indices",
80
+ "token_tags",
81
+ "position_ids",
82
+ "video_indices",
83
+ "audio_indices",
84
+ "text_indices",
85
+ "attention_kwargs",
86
+ "return_dict",
87
+ )
88
+
89
+
90
+ def status() -> str:
91
+ return (
92
+ f"first-block cache **on** · threshold `{THRESHOLD}` · audio exemption "
93
+ f"{'on' if AUDIO_EXEMPT else 'off'}"
94
+ if ENABLED
95
+ else "first-block cache **off** (`H3_FBC=1` to skip the trunk on steady steps)"
96
+ )
97
+
98
+
99
+ def _shifted_sigma(u: float, shift: float) -> float:
100
+ return shift * u / (1.0 + (shift - 1.0) * u)
101
+
102
+
103
+ def _rel_l1(current, previous) -> float:
104
+ numerator = (current.float() - previous.float()).abs().mean()
105
+ denominator = previous.float().abs().mean().clamp(min=1e-8)
106
+ return float((numerator / denominator).item())
107
+
108
+
109
+ class _State:
110
+ def __init__(self, threshold: float, steps: int, audio_exempt: bool):
111
+ self.threshold = threshold
112
+ self.steps = steps
113
+ self.audio_exempt = audio_exempt
114
+ # duckyshell reads the window as sigma bounds: a flow model's sigma at `u = 1 - percent`, shifted.
115
+ self.start_sigma = _shifted_sigma(1.0 - START_PERCENT, VIDEO_SHIFT)
116
+ self.end_sigma = _shifted_sigma(1.0 - END_PERCENT, VIDEO_SHIFT)
117
+ self.original = None
118
+ self.failed = False
119
+ self.consecutive_hits = 0
120
+ self.prev_first_residual = None
121
+ self.tail_residual = None
122
+ self.audio_history = [] # [(sigma_audio, audio rows of the trunk output)], newest last, at most two
123
+ self.computed = 0
124
+ self.skipped = 0
125
+
126
+
127
+ def _cached_forward(
128
+ self,
129
+ state,
130
+ hidden_states,
131
+ audio_hidden_states,
132
+ encoder_hidden_states,
133
+ timestep,
134
+ timestep_indices,
135
+ token_tags,
136
+ position_ids,
137
+ video_indices,
138
+ audio_indices,
139
+ text_indices,
140
+ return_dict,
141
+ ):
142
+ """`MiniMaxH3Transformer3DModel.forward` with the block loop split at block 0.
143
+
144
+ Everything outside the loop is that method verbatim, at the `diffusers` commit `requirements.txt` pins; keep the
145
+ two in step when the pin moves.
146
+ """
147
+ import torch
148
+
149
+ from diffusers.models.transformers.transformer_minimax_h3 import (
150
+ MINIMAX_H3_MODALITY_NUM,
151
+ MiniMaxH3TransformerOutput,
152
+ )
153
+
154
+ sequence_length = position_ids.shape[0]
155
+ rotary_emb = self.rope(position_ids)
156
+
157
+ video_embeds = self.proj_in(hidden_states.to(self.proj_in.weight.dtype))
158
+ audio_embeds = self.audio_proj_in(audio_hidden_states.to(self.audio_proj_in.weight.dtype))
159
+ text_embeds = self.context_embedder(encoder_hidden_states.to(self.context_embedder.weight.dtype))
160
+ text_embeds = self.token_refiner(text_embeds)
161
+
162
+ packed = text_embeds.new_zeros((text_embeds.shape[0], sequence_length, text_embeds.shape[-1]))
163
+ packed = packed.index_copy(1, text_indices, text_embeds)
164
+ packed = packed.index_copy(1, video_indices, video_embeds.to(text_embeds.dtype))
165
+ packed = packed.index_copy(1, audio_indices, audio_embeds.to(text_embeds.dtype))
166
+
167
+ temb = self.time_proj(timestep)
168
+ temb = self.time_embedder(temb.to(self.time_embedder.linear_1.weight.dtype))
169
+ adaln_indices = timestep_indices * MINIMAX_H3_MODALITY_NUM + token_tags.clamp(min=0)
170
+
171
+ attention_mask = None
172
+ is_pad = token_tags < 0
173
+ if bool(is_pad.any()):
174
+ attention_mask = is_pad[None, :] == is_pad[:, None]
175
+
176
+ blocks = self.transformer_blocks
177
+ block0_out = blocks[0](packed, temb, adaln_indices, rotary_emb, attention_mask)
178
+ first_residual = block0_out - packed
179
+
180
+ # The generated rows trail their modality's index list, so the last row of each carries that stream's live noise
181
+ # level. The scheduler exposes `timesteps = 1 - sigmas[:-1]`.
182
+ sigma_video = 1.0 - float(timestep[timestep_indices[video_indices[-1]]].item())
183
+ sigma_audio = 1.0 - float(timestep[timestep_indices[audio_indices[-1]]].item())
184
+
185
+ use_cache = False
186
+ if (
187
+ state.prev_first_residual is not None
188
+ and state.tail_residual is not None
189
+ and state.prev_first_residual.shape == first_residual.shape
190
+ and state.consecutive_hits < MAX_CONSECUTIVE_HITS
191
+ and state.end_sigma <= sigma_video <= state.start_sigma
192
+ ):
193
+ use_cache = _rel_l1(first_residual, state.prev_first_residual) <= state.threshold
194
+
195
+ if use_cache:
196
+ state.consecutive_hits += 1
197
+ state.skipped += 1
198
+ trunk_out = block0_out + state.tail_residual
199
+ if state.audio_exempt and state.audio_history:
200
+ sigma_prev, feature_prev = state.audio_history[-1]
201
+ if len(state.audio_history) == 2 and abs(sigma_prev - state.audio_history[-2][0]) > 1e-8:
202
+ sigma_prev2, feature_prev2 = state.audio_history[-2]
203
+ ratio = (sigma_audio - sigma_prev) / (sigma_prev - sigma_prev2)
204
+ audio_feature = feature_prev + (feature_prev - feature_prev2) * ratio
205
+ else:
206
+ audio_feature = feature_prev
207
+ trunk_out = trunk_out.index_copy(1, audio_indices, audio_feature.to(trunk_out.dtype))
208
+ else:
209
+ state.consecutive_hits = 0
210
+ state.computed += 1
211
+ trunk_out = block0_out
212
+ for block in blocks[1:]:
213
+ trunk_out = block(trunk_out, temb, adaln_indices, rotary_emb, attention_mask)
214
+ state.tail_residual = (trunk_out - block0_out).detach()
215
+ state.prev_first_residual = first_residual.detach()
216
+ if state.audio_exempt:
217
+ state.audio_history.append((sigma_audio, trunk_out.index_select(1, audio_indices).detach().float()))
218
+ state.audio_history = state.audio_history[-2:]
219
+
220
+ out = self.norm_out(trunk_out, temb, timestep_indices).to(self.proj_out.weight.dtype)
221
+ video_output = self.proj_out(out).index_select(1, video_indices)
222
+ audio_output = self.audio_proj_out(out).index_select(1, audio_indices)
223
+
224
+ if not return_dict:
225
+ return (video_output, audio_output)
226
+ return MiniMaxH3TransformerOutput(sample=video_output, audio_sample=audio_output)
227
+
228
+
229
+ def _forward(
230
+ self,
231
+ hidden_states,
232
+ audio_hidden_states,
233
+ encoder_hidden_states,
234
+ timestep,
235
+ timestep_indices,
236
+ token_tags,
237
+ position_ids,
238
+ video_indices,
239
+ audio_indices,
240
+ text_indices,
241
+ attention_kwargs=None,
242
+ return_dict: bool = True,
243
+ ):
244
+ """The installed forward. Anything it cannot serve — a LoRA scale, an unexpected layout, a bug — is handed to the
245
+ original forward instead, for this call and every later one, so a cached request can degrade to an uncached one but
246
+ never to a failed one."""
247
+ state = self._h3_fbc
248
+ original = dict(
249
+ hidden_states=hidden_states,
250
+ audio_hidden_states=audio_hidden_states,
251
+ encoder_hidden_states=encoder_hidden_states,
252
+ timestep=timestep,
253
+ timestep_indices=timestep_indices,
254
+ token_tags=token_tags,
255
+ position_ids=position_ids,
256
+ video_indices=video_indices,
257
+ audio_indices=audio_indices,
258
+ text_indices=text_indices,
259
+ attention_kwargs=attention_kwargs,
260
+ return_dict=return_dict,
261
+ )
262
+ # `apply_lora_scale` decorates the real forward and this one is not it, so a request that actually scales a LoRA
263
+ # goes down the original path rather than silently losing its scale.
264
+ if state.failed or (attention_kwargs or {}).get("scale") is not None:
265
+ return state.original(**original)
266
+ try:
267
+ return _cached_forward(
268
+ self,
269
+ state,
270
+ hidden_states,
271
+ audio_hidden_states,
272
+ encoder_hidden_states,
273
+ timestep,
274
+ timestep_indices,
275
+ token_tags,
276
+ position_ids,
277
+ video_indices,
278
+ audio_indices,
279
+ text_indices,
280
+ return_dict,
281
+ )
282
+ except Exception as error:
283
+ state.failed = True
284
+ print(f"[h3-fbc] disabled for this request ({type(error).__name__}: {error}); running uncached", flush=True)
285
+ return state.original(**original)
286
+
287
+
288
+ def install(transformer, steps: int = 0, threshold: float = THRESHOLD, audio_exempt: bool = AUDIO_EXEMPT) -> bool:
289
+ """Bind the caching forward onto `transformer`. Returns whether it went on.
290
+
291
+ `accelerate`'s `add_hook_to_module` — what `ComponentsManager.enable_auto_cpu_offload` installs — moves the real
292
+ forward to `_old_forward` and puts its own onload wrapper in `forward`. Replacing `forward` there would step over
293
+ the wrapper and run the block stack against weights still on the host, so the replacement goes into `_old_forward`
294
+ whenever the hook is present.
295
+ """
296
+ import inspect
297
+
298
+ if getattr(transformer, "_h3_fbc", None) is not None:
299
+ return True
300
+
301
+ hooked = hasattr(transformer, "_hf_hook") and hasattr(transformer, "_old_forward")
302
+ current = transformer._old_forward if hooked else transformer.forward
303
+ missing = [name for name in FORWARD_PARAMETERS if name not in inspect.signature(current).parameters]
304
+ if missing:
305
+ print(f"[h3-fbc] this transformer's forward has no {missing}; running uncached", flush=True)
306
+ return False
307
+ if not hasattr(transformer, "transformer_blocks") or len(transformer.transformer_blocks) < 2:
308
+ print("[h3-fbc] no block stack to skip; running uncached", flush=True)
309
+ return False
310
+
311
+ state = _State(threshold, steps, audio_exempt)
312
+ state.original = current
313
+ transformer._h3_fbc = state
314
+ bound = types.MethodType(_forward, transformer)
315
+ if hooked:
316
+ transformer._old_forward = bound
317
+ else:
318
+ transformer.forward = bound
319
+ return True
320
+
321
+
322
+ def uninstall(transformer) -> None:
323
+ state = getattr(transformer, "_h3_fbc", None)
324
+ if state is None:
325
+ return
326
+ if hasattr(transformer, "_hf_hook") and hasattr(transformer, "_old_forward"):
327
+ transformer._old_forward = state.original
328
+ else:
329
+ transformer.__dict__.pop("forward", None)
330
+ del transformer._h3_fbc
331
+ total = state.computed + state.skipped
332
+ if total:
333
+ print(
334
+ f"[h3-fbc] {state.skipped}/{total} forwards served from cache "
335
+ f"(threshold {state.threshold}, audio exemption {'on' if state.audio_exempt else 'off'})",
336
+ flush=True,
337
+ )
338
+
339
+
340
+ @contextlib.contextmanager
341
+ def enabled(transformer, steps: int = 0):
342
+ """Cache the trunk for the duration of one request. The state is per-request by construction — a residual only ever
343
+ means something within the schedule it was measured on — and nothing in here can raise into the request."""
344
+ installed = False
345
+ if ENABLED:
346
+ try:
347
+ installed = install(transformer, steps=steps)
348
+ except Exception as error:
349
+ print(f"[h3-fbc] install failed ({type(error).__name__}: {error}); running uncached", flush=True)
350
+ try:
351
+ yield installed
352
+ finally:
353
+ if installed:
354
+ try:
355
+ uninstall(transformer)
356
+ except Exception as error:
357
+ print(f"[h3-fbc] uninstall failed ({type(error).__name__}: {error})", flush=True)
h3_split_blocks.py ADDED
@@ -0,0 +1,147 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """The halves of a **split** MiniMax-H3 deployment, for both of its checkpoint partitions.
2
+
3
+ MiniMax-H3 is 195.9 GiB in bfloat16 and a ZeroGPU Space is evicted at 150 GB of storage, so `MiniMaxH3Blocks` is cut
4
+ at its `text_encoder` step: the 62.14 GiB Qwen3-VL runs in the conditioner Space, everything else in a generator
5
+ Space, and `prompt_embeds` + `text_token_tags` is the whole wire format between them.
6
+
7
+ `resize` / `setup` run on **both** sides: they own no pretrained component, and each half needs the canvas and the
8
+ prepared keyframes or normalized references. Both conditioner halves also return the resolved `height` / `width` /
9
+ `num_frames`, which the generating half pins rather than re-deriving.
10
+
11
+ Two things the blocks leave to the caller: a keyframe reaches them EXIF-transposed and in RGB, and the `t2va` / `fl2va`
12
+ frame count is aligned to `17 * n + 5` before the call, since that arithmetic lives on the denoising side of the cut.
13
+ """
14
+
15
+ from diffusers.modular_pipelines.minimax_h3.before_encoder import MiniMaxH3Ref2VASetupStep
16
+ from diffusers.modular_pipelines.minimax_h3.decoders import MiniMaxH3AfterDenoiseStep
17
+ from diffusers.modular_pipelines.minimax_h3.encoders import (
18
+ MiniMaxH3Ref2VAReferenceEncoderStep,
19
+ MiniMaxH3Ref2VATextEncoderStep,
20
+ MiniMaxH3TextEncoderStep,
21
+ )
22
+ from diffusers.modular_pipelines.minimax_h3.modular_blocks_minimax_h3 import (
23
+ MiniMaxH3AutoKeyframeVaeEncoderStep,
24
+ MiniMaxH3AutoResizeStep,
25
+ MiniMaxH3CoreDenoiseStep,
26
+ MiniMaxH3DecodeStep,
27
+ MiniMaxH3Ref2VACoreDenoiseStep,
28
+ _generation_outputs,
29
+ )
30
+ from diffusers.modular_pipelines.modular_pipeline import SequentialPipelineBlocks
31
+ from diffusers.modular_pipelines.modular_pipeline_utils import OutputParam
32
+
33
+
34
+ def _wire_outputs(num_frames: bool = True) -> list[OutputParam]:
35
+ """The wire format of the split. `num_frames` is declared by the `ref2va` half alone, whose setup resolves one."""
36
+ return [
37
+ OutputParam.template("prompt_embeds"),
38
+ OutputParam("text_token_tags", description="The per-row modality tag of every row of `prompt_embeds`."),
39
+ OutputParam("height", type_hint=int, description="Resolved height of the generated video in pixels."),
40
+ OutputParam("width", type_hint=int, description="Resolved width of the generated video in pixels."),
41
+ *(
42
+ [OutputParam("num_frames", type_hint=int, description="Resolved number of frames, of the form 17 * n + 5.")]
43
+ if num_frames
44
+ else []
45
+ ),
46
+ ]
47
+
48
+
49
+ class MiniMaxH3ConditionerBlocks(SequentialPipelineBlocks):
50
+ """The conditioner half of a split MiniMax-H3: the keyframes on the canvas plus the Qwen3-VL read at layer 50."""
51
+
52
+ model_name = "minimax-h3"
53
+ block_classes = [MiniMaxH3AutoResizeStep, MiniMaxH3TextEncoderStep]
54
+ block_names = ["resize", "text_encoder"]
55
+
56
+ @property
57
+ def description(self):
58
+ return (
59
+ "The conditioner half of a split MiniMax-H3 deployment: puts the keyframes onto the target canvas and "
60
+ "encodes MiniMax-H3's presentation of the request into the `prompt_embeds` / `text_token_tags` pair the "
61
+ "denoising half consumes. The frame count is the caller's to align."
62
+ )
63
+
64
+ @property
65
+ def outputs(self):
66
+ return _wire_outputs(num_frames=False)
67
+
68
+
69
+ class MiniMaxH3GeneratorBlocks(SequentialPipelineBlocks):
70
+ """The denoising half of a split MiniMax-H3: `MiniMaxH3Blocks` with its `text_encoder` step removed."""
71
+
72
+ model_name = "minimax-h3"
73
+ block_classes = [
74
+ MiniMaxH3AutoResizeStep,
75
+ MiniMaxH3AutoKeyframeVaeEncoderStep,
76
+ MiniMaxH3CoreDenoiseStep,
77
+ MiniMaxH3AfterDenoiseStep,
78
+ MiniMaxH3DecodeStep,
79
+ ]
80
+ block_names = ["resize", "vae_encoder", "denoise", "after_denoise", "decode"]
81
+
82
+ @property
83
+ def description(self):
84
+ return (
85
+ "The denoising half of a split MiniMax-H3 deployment: the `t2va` / `fl2va` branch of `MiniMaxH3Blocks` "
86
+ "without its text-encoder step, so `prompt_embeds` and `text_token_tags` come in as inputs and the "
87
+ "62.14 GiB Qwen3-VL conditioner is never loaded here."
88
+ )
89
+
90
+ @property
91
+ def outputs(self):
92
+ return _generation_outputs()
93
+
94
+
95
+ class MiniMaxH3Ref2VAConditionerBlocks(SequentialPipelineBlocks):
96
+ """The conditioner half of a split `ref2va`: the resolved plan plus the Qwen3-VL read at its 50th layer.
97
+
98
+ Component for component this is `MiniMaxH3ConditionerBlocks`, so one conditioner Space serves both partitions.
99
+ What differs is the presentation: `ref2va` prepends a label per reference and a vision block per image and per
100
+ merged video frame pair, so the references themselves have to reach this half.
101
+ """
102
+
103
+ model_name = "minimax-h3"
104
+ block_classes = [MiniMaxH3Ref2VASetupStep, MiniMaxH3Ref2VATextEncoderStep]
105
+ block_names = ["setup", "text_encoder"]
106
+
107
+ @property
108
+ def description(self):
109
+ return (
110
+ "The conditioner half of a split MiniMax-H3 `ref2va` deployment: resolves the request plan (canvas, frame "
111
+ "count, references normalized onto MiniMax-H3's own rates and resolutions) and encodes MiniMax-H3's "
112
+ "presentation of it into the `prompt_embeds` / `text_token_tags` pair the denoising half consumes."
113
+ )
114
+
115
+ @property
116
+ def outputs(self):
117
+ return _wire_outputs()
118
+
119
+
120
+ class MiniMaxH3Ref2VAGeneratorBlocks(SequentialPipelineBlocks):
121
+ """The denoising half of a split `ref2va`: the `ref2va` branch with its `text_encoder` step removed.
122
+
123
+ `reference_encoder` stays here, next to the two autoencoders it runs: its output shapes are where every reference
124
+ block's geometry in the packed layout comes from.
125
+ """
126
+
127
+ model_name = "minimax-h3"
128
+ block_classes = [
129
+ MiniMaxH3Ref2VASetupStep,
130
+ MiniMaxH3Ref2VAReferenceEncoderStep,
131
+ MiniMaxH3Ref2VACoreDenoiseStep,
132
+ MiniMaxH3AfterDenoiseStep,
133
+ MiniMaxH3DecodeStep,
134
+ ]
135
+ block_names = ["setup", "reference_encoder", "denoise", "after_denoise", "decode"]
136
+
137
+ @property
138
+ def description(self):
139
+ return (
140
+ "The denoising half of a split MiniMax-H3 `ref2va` deployment: the `ref2va` branch of `MiniMaxH3Blocks` "
141
+ "without its text-encoder step, so `prompt_embeds` and `text_token_tags` come in as inputs and the "
142
+ "62.14 GiB Qwen3-VL conditioner is never loaded here. The transformer is the `transformer_ref` partition."
143
+ )
144
+
145
+ @property
146
+ def outputs(self):
147
+ return _generation_outputs()
ncii_guard.py ADDED
@@ -0,0 +1,90 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """The NCII prompt guard, in its own process.
2
+
3
+ [`hfmlsoc/ncii-light-guard-v01`](https://huggingface.co/hfmlsoc/ncii-light-guard-v01) is a 270M CPU text
4
+ classifier scoring the NCII risk of an edit prompt. It cannot live in the main process: with it loaded there,
5
+ every subsequent `@spaces.GPU` worker dies at `worker_init` with `RuntimeError: No CUDA GPUs are available` —
6
+ the fork inherits whatever CUDA driver state the classifier's torch activity left behind, and a factory reboot
7
+ does not clear it. It cannot be a `multiprocessing.spawn` child either: spawn re-imports the parent's main
8
+ module, and on a Space that main module is `app.py` — the child would re-run the whole startup, `start()`
9
+ included. So the classifier runs this file as a plain subprocess — a fresh interpreter that never sees `spaces`
10
+ — and answers over stdin/stdout, one JSON object per line.
11
+ """
12
+
13
+ from __future__ import annotations
14
+
15
+ import json
16
+ import os
17
+ import select
18
+ import subprocess
19
+ import sys
20
+ import threading
21
+
22
+ GUARD_REPO = "hfmlsoc/ncii-light-guard-v01"
23
+
24
+ _lock = threading.Lock()
25
+ _process: subprocess.Popen | None = None
26
+
27
+
28
+ def _read(timeout: float) -> dict:
29
+ readable, _, _ = select.select([_process.stdout], [], [], timeout)
30
+ if not readable:
31
+ raise TimeoutError(f"the guard did not answer within {timeout}s")
32
+ line = _process.stdout.readline()
33
+ if not line:
34
+ raise EOFError("the guard process died")
35
+ return json.loads(line)
36
+
37
+
38
+ def _spawn() -> None:
39
+ global _process
40
+ _process = subprocess.Popen(
41
+ [sys.executable, os.path.abspath(__file__)],
42
+ stdin=subprocess.PIPE,
43
+ stdout=subprocess.PIPE,
44
+ text=True,
45
+ bufsize=1,
46
+ )
47
+ # Generous: a cold cache downloads the checkpoint first.
48
+ assert _read(300.0) == {"status": "ready"}
49
+
50
+
51
+ def start() -> None:
52
+ """Launch the worker and block until its model is up. Called once at startup; `classify` revives it if it dies."""
53
+ with _lock:
54
+ _spawn()
55
+
56
+
57
+ def classify(prompt: str, timeout: float = 60.0) -> dict:
58
+ """`{'label': 'safe' | 'ncii', 'score': ...}` for one prompt, replacing a dead or wedged worker once."""
59
+ with _lock:
60
+ for attempt in (0, 1):
61
+ try:
62
+ if _process is None or _process.poll() is not None:
63
+ _spawn()
64
+ _process.stdin.write(json.dumps({"prompt": prompt}) + "\n")
65
+ _process.stdin.flush()
66
+ return _read(timeout)
67
+ except Exception:
68
+ if attempt:
69
+ raise
70
+ if _process is not None and _process.poll() is None:
71
+ _process.kill()
72
+
73
+
74
+ def _serve() -> None:
75
+ """The child: plain torch on CPU. The protocol keeps the real stdout to itself — everything else
76
+ (download progress, warnings) is pushed over to stderr so it cannot corrupt a reply."""
77
+ protocol = os.fdopen(os.dup(1), "w", buffering=1)
78
+ os.dup2(2, 1)
79
+
80
+ from transformers import pipeline
81
+
82
+ classifier = pipeline("text-classification", model=GUARD_REPO, device="cpu")
83
+ protocol.write(json.dumps({"status": "ready"}) + "\n")
84
+ for line in sys.stdin:
85
+ result = classifier(json.loads(line)["prompt"], truncation=True)[0]
86
+ protocol.write(json.dumps({"label": result["label"], "score": float(result["score"])}) + "\n")
87
+
88
+
89
+ if __name__ == "__main__":
90
+ _serve()
packages.txt ADDED
@@ -0,0 +1 @@
 
 
1
+ ffmpeg
requirements.txt ADDED
@@ -0,0 +1,18 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # diffusers from the MiniMax-H3 PR (https://github.com/huggingface/diffusers/pull/14371), pinned to a
2
+ # commit on its minimax-h3-refactor branch — the modular MiniMax-H3 blocks are in no release yet, and
3
+ # h3_split_blocks.py subclasses them. main has since renamed those block classes, so the pin is load-bearing.
4
+ # 665f578278365ea4a3318cb8c9b66ce6c01204b9 = refs/pull/14371/head at the time of this deploy
5
+ --extra-index-url https://download.pytorch.org/whl/cu130
6
+ diffusers @ git+https://github.com/huggingface/diffusers.git@665f578278365ea4a3318cb8c9b66ce6c01204b9
7
+ torch==2.11.0
8
+ torchvision==0.26.0
9
+ torchaudio==2.11.0
10
+ transformers==5.8.0
11
+ accelerate==1.14.0
12
+ huggingface-hub==1.24.0
13
+ peft
14
+ av
15
+ pillow
16
+ numpy
17
+ requests
18
+ safetensors>=0.8.0