Spaces:
Running on Zero
Running on Zero
Add a 28-step vs 8-step turbo LoRA sampling choice (lightx2v/Minimax-h3-Turbo)
Browse files- README.md +39 -0
- app.py +80 -12
- h3_turbo_lora.py +144 -0
README.md
CHANGED
|
@@ -12,6 +12,7 @@ short_description: Drive a world model with WASD and camera keys
|
|
| 12 |
models:
|
| 13 |
- MiniMaxAI/MiniMax-H3
|
| 14 |
- DANNY621/H3-World
|
|
|
|
| 15 |
---
|
| 16 |
|
| 17 |
# H3-World — action-conditioned world model
|
|
@@ -122,6 +123,39 @@ gate/value halves swapped, and the fused `attn.qkv_proj` de-interleaved per head
|
|
| 122 |
`to_q` / `to_k` / `to_v`. 104 LoRA pairs become 208 merged weight deltas; **any target that fails to resolve is
|
| 123 |
fatal**, never skipped.
|
| 124 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 125 |
## Generation constraints
|
| 126 |
|
| 127 |
Fixed by the checkpoint: 24 fps, `num_frames` snapped to `17n + 5`, no CFG and no negative prompt (it is
|
|
@@ -160,6 +194,11 @@ manifest prompt.
|
|
| 160 |
| `H3_MODEL_REPO` | `MiniMaxAI/MiniMax-H3` | The diffusers-layout base checkpoint. |
|
| 161 |
| `H3_LORA_REPO` | `DANNY621/H3-World` | The LoRA. |
|
| 162 |
| `H3_LORA_FILE` | `step-10000.safetensors` | The checkpoint the author's own test runs used. |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 163 |
| `H3_CONDITIONER` | `multimodalart/qwen3vl-conditioner` | The Space this one asks for embeddings. |
|
| 164 |
| `H3_ATTENTION` | `_native_cudnn` | cuDNN's fused kernel. flash-attention 3 is sm90-only; this pool is sm120. |
|
| 165 |
| `H3_GPU_SIZE` | `xlarge` | ZeroGPU allocation size. `large` does not fit. |
|
|
|
|
| 12 |
models:
|
| 13 |
- MiniMaxAI/MiniMax-H3
|
| 14 |
- DANNY621/H3-World
|
| 15 |
+
- lightx2v/Minimax-h3-Turbo
|
| 16 |
---
|
| 17 |
|
| 18 |
# H3-World — action-conditioned world model
|
|
|
|
| 123 |
`to_q` / `to_k` / `to_v`. 104 LoRA pairs become 208 merged weight deltas; **any target that fails to resolve is
|
| 124 |
fatal**, never skipped.
|
| 125 |
|
| 126 |
+
## 28 steps vs the 8-step turbo LoRA
|
| 127 |
+
|
| 128 |
+
**Sampling** in the UI picks between two configurations of the same request — same seed, same action
|
| 129 |
+
script, same directed mask — so the quality cost of the distillation is directly visible:
|
| 130 |
+
|
| 131 |
+
| mode | steps | transformer |
|
| 132 |
+
|---|---|---|
|
| 133 |
+
| `28 steps · no turbo LoRA` | 28 (MiniMax-H3's default) | H3-World only |
|
| 134 |
+
| `8 steps · turbo LoRA` | 8 | H3-World **+** [`lightx2v/Minimax-h3-Turbo`](https://huggingface.co/lightx2v/Minimax-h3-Turbo) `minimax_h3_fl2v_turbo_8step_v1.0_bf16.safetensors` |
|
| 135 |
+
|
| 136 |
+
`h3_turbo_lora.py` mirrors the `h3_lora.py` of the Spaces that already run this adapter —
|
| 137 |
+
[`MiniMaxAI/MiniMax-H3-Turbo-Lora`](https://huggingface.co/spaces/MiniMaxAI/MiniMax-H3-Turbo-Lora)
|
| 138 |
+
(its `lightx` set) and
|
| 139 |
+
[`hugging-apps/minimax-h3-turbo-sla-demo`](https://huggingface.co/spaces/hugging-apps/minimax-h3-turbo-sla-demo):
|
| 140 |
+
|
| 141 |
+
* the file is a **PEFT checkpoint against the diffusers module tree itself**
|
| 142 |
+
(`transformer_blocks.N.attn.to_q.lora_A.default.weight`), so unlike H3-World it needs no key
|
| 143 |
+
conversion — 312 targets (50 transformer blocks + 2 token-refiner blocks x
|
| 144 |
+
`to_q`/`to_k`/`to_v`/`to_out.0`/`ff.net.0.proj`/`ff.net.2`) map name-for-name;
|
| 145 |
+
* rank 128 with `alpha: 8` in the file's own safetensors metadata, so the fold scale is
|
| 146 |
+
`alpha / rank = 0.0625` — what `set_adapters(weights=1.0)` applies in lightx2v's reference script;
|
| 147 |
+
* the step count is overridden to the distillation's own **8 NFE**. No scheduler swap and no CFG
|
| 148 |
+
change: MiniMax-H3 is already guidance-distilled and every Space above keeps its native
|
| 149 |
+
`MiniMaxH3Scheduler` with the turbo LoRA folded.
|
| 150 |
+
|
| 151 |
+
It is *folded* into the bf16 weights, like H3-World, because this Space patches
|
| 152 |
+
`MiniMaxH3AttnProcessor` and drives the transformer's live weights. Fold and unfold are the same
|
| 153 |
+
operation with a sign, so the low-rank factors stay resident and `set_active` flips the mode in
|
| 154 |
+
place inside the `@spaces.GPU` call, through one bf16 rounding.
|
| 155 |
+
|
| 156 |
+
The Steps slider still overrides the mode's count (4–50), so `50 steps + turbo LoRA` or
|
| 157 |
+
`8 steps without it` are both reachable for the sake of the comparison.
|
| 158 |
+
|
| 159 |
## Generation constraints
|
| 160 |
|
| 161 |
Fixed by the checkpoint: 24 fps, `num_frames` snapped to `17n + 5`, no CFG and no negative prompt (it is
|
|
|
|
| 194 |
| `H3_MODEL_REPO` | `MiniMaxAI/MiniMax-H3` | The diffusers-layout base checkpoint. |
|
| 195 |
| `H3_LORA_REPO` | `DANNY621/H3-World` | The LoRA. |
|
| 196 |
| `H3_LORA_FILE` | `step-10000.safetensors` | The checkpoint the author's own test runs used. |
|
| 197 |
+
| `H3_TURBO_REPO` | `lightx2v/Minimax-h3-Turbo` | The turbo-LoRA repo. |
|
| 198 |
+
| `H3_TURBO_FILE` | `minimax_h3_fl2v_turbo_8step_v1.0_bf16.safetensors` | The 8-step distillation. |
|
| 199 |
+
| `H3_TURBO_STEPS` | `8` | Steps the turbo mode asks for (the card also offers 4). |
|
| 200 |
+
| `H3_TURBO_ALPHA` | `0` | `0` reads `alpha` out of the file's metadata (`8`). |
|
| 201 |
+
| `H3_TURBO_STRENGTH` | `1.0` | Extra multiplier on the turbo delta. |
|
| 202 |
| `H3_CONDITIONER` | `multimodalart/qwen3vl-conditioner` | The Space this one asks for embeddings. |
|
| 203 |
| `H3_ATTENTION` | `_native_cudnn` | cuDNN's fused kernel. flash-attention 3 is sm90-only; this pool is sm120. |
|
| 204 |
| `H3_GPU_SIZE` | `xlarge` | ZeroGPU allocation size. `large` does not fit. |
|
app.py
CHANGED
|
@@ -48,6 +48,8 @@ import spaces
|
|
| 48 |
import gradio as gr
|
| 49 |
import torch
|
| 50 |
|
|
|
|
|
|
|
| 51 |
MODEL_REPO = os.environ.get("H3_MODEL_REPO", "MiniMaxAI/MiniMax-H3")
|
| 52 |
LORA_REPO = os.environ.get("H3_LORA_REPO", "DANNY621/H3-World")
|
| 53 |
LORA_FILE = os.environ.get("H3_LORA_FILE", "step-10000.safetensors")
|
|
@@ -59,11 +61,21 @@ GPU_SIZE = os.environ.get("H3_GPU_SIZE", "xlarge")
|
|
| 59 |
FPS, FRAMES_PER_CHUNK, LATENTS_PER_CHUNK = 24, 17, 5
|
| 60 |
MIN_UI_DURATION, MAX_UI_DURATION = 2, 8
|
| 61 |
# The H3-World inference runs recorded in the checkpoint's own manifests: 124 frames, 480x832, 24 fps,
|
| 62 |
-
# 50 steps, seed 0.
|
| 63 |
DEFAULT_DURATION = 5
|
| 64 |
-
DEFAULT_STEPS =
|
| 65 |
DEFAULT_SUBJECT = "the man"
|
| 66 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 67 |
# The conditioner offers canvases up to 1344x768, but H3-World was trained at 832x480 and the directed
|
| 68 |
# mask costs O(sequence x captions) — at 1344x768 / 8s a request would need ~35 GPU-minutes, far past
|
| 69 |
# what any visitor's ZeroGPU quota can book. So the list is the cheap tier of each aspect ratio, which
|
|
@@ -89,6 +101,7 @@ os.makedirs(OUTPUT_DIR, exist_ok=True)
|
|
| 89 |
PIPE = None
|
| 90 |
LOAD_ERROR: str | None = None
|
| 91 |
LOADED_IN: float | None = None
|
|
|
|
| 92 |
|
| 93 |
|
| 94 |
# ── The action vocabulary ────────────────────────────────────────────────────
|
|
@@ -575,7 +588,7 @@ def lower_duration_floor(seconds: float = MIN_UI_DURATION) -> None:
|
|
| 575 |
|
| 576 |
def load_models() -> str | None:
|
| 577 |
"""Load the denoising half at startup: transformer + VAEs, fold H3-World in, patch attention."""
|
| 578 |
-
global PIPE, LOAD_ERROR, LOADED_IN
|
| 579 |
|
| 580 |
if PIPE is not None or LOAD_ERROR is not None:
|
| 581 |
return LOAD_ERROR
|
|
@@ -601,6 +614,15 @@ def load_models() -> str | None:
|
|
| 601 |
|
| 602 |
print(f"[gen] {load_and_apply_lora(pipe.transformer)}", flush=True)
|
| 603 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 604 |
if PLACEMENT == "pack":
|
| 605 |
pipe.transformer.to("cuda")
|
| 606 |
|
|
@@ -740,9 +762,14 @@ _TOKENS_PER_CAPTION = 8
|
|
| 740 |
_DECODE_BASE, _DECODE_PER_DEFAULT_CANVAS, _DEFAULT_CANVAS_PIXELS = 15, 15, 960 * 544 * 124
|
| 741 |
# ~12% over the fit, plus one cold worker's weight placement.
|
| 742 |
_MARGIN, _PLACEMENT_ALLOWANCE = 1.12, 12
|
|
|
|
|
|
|
|
|
|
| 743 |
|
| 744 |
|
| 745 |
-
def get_duration(
|
|
|
|
|
|
|
| 746 |
"""Estimate GPU seconds for one ``_generate`` call, from measurements rather than guesswork."""
|
| 747 |
height, width, num_frames, steps = int(height), int(width), int(num_frames), int(steps)
|
| 748 |
patches = (height // 32) * (width // 32)
|
|
@@ -753,14 +780,17 @@ def get_duration(prompt_embeds, tags, plan, image, height, width, num_frames, st
|
|
| 753 |
if plan is not None:
|
| 754 |
per_step += _MASK * rows * len(plan["spans"]) * _TOKENS_PER_CAPTION
|
| 755 |
decode = _DECODE_BASE + _DECODE_PER_DEFAULT_CANVAS * (height * width * num_frames) / _DEFAULT_CANVAS_PIXELS
|
| 756 |
-
|
|
|
|
|
|
|
|
|
|
| 757 |
|
| 758 |
|
| 759 |
# ── Inference ────────────────────────────────────────────────────────────────
|
| 760 |
|
| 761 |
|
| 762 |
@spaces.GPU(duration=get_duration, size=GPU_SIZE)
|
| 763 |
-
def _generate(prompt_embeds, tags, plan, image, height, width, num_frames, steps, seed):
|
| 764 |
"""Denoise + decode on GPU time, with the directed mask active for this request's geometry."""
|
| 765 |
if PLACEMENT == "pack":
|
| 766 |
PIPE.vae.to("cuda")
|
|
@@ -768,6 +798,10 @@ def _generate(prompt_embeds, tags, plan, image, height, width, num_frames, steps
|
|
| 768 |
elif PLACEMENT == "lazy":
|
| 769 |
PIPE.to("cuda")
|
| 770 |
|
|
|
|
|
|
|
|
|
|
|
|
|
| 771 |
DIRECTED["plan"] = plan
|
| 772 |
try:
|
| 773 |
state = PIPE(
|
|
@@ -801,10 +835,11 @@ def generate(
|
|
| 801 |
script: str = "forward",
|
| 802 |
canvas: str = DEFAULT_CANVAS,
|
| 803 |
duration: float = DEFAULT_DURATION,
|
| 804 |
-
steps: int =
|
| 805 |
seed: int = 0,
|
| 806 |
subject: str = DEFAULT_SUBJECT,
|
| 807 |
directed: bool = True,
|
|
|
|
| 808 |
progress=gr.Progress(track_tqdm=True),
|
| 809 |
):
|
| 810 |
"""Roll a world forward from one frame under a keyboard action script.
|
|
@@ -818,11 +853,15 @@ def generate(
|
|
| 818 |
tilt-down, or a raw key combination out of W A S D I J K L F.
|
| 819 |
canvas: Output resolution; snapped to the first frame's aspect ratio when one is given.
|
| 820 |
duration: Seconds of video, rounded up to the next frame count the VAE can decode.
|
| 821 |
-
steps: Denoising steps
|
|
|
|
| 822 |
seed: Random seed.
|
| 823 |
subject: How the per-frame sentences refer to the character.
|
| 824 |
directed: Bind each sentence to its own latent frame with H3-World's directed attention mask.
|
| 825 |
Turning it off is the ablation: the sentences become one global prompt.
|
|
|
|
|
|
|
|
|
|
| 826 |
|
| 827 |
Returns:
|
| 828 |
The generated video (with MiniMax-H3's native soundtrack) and a markdown report holding the
|
|
@@ -835,6 +874,11 @@ def generate(
|
|
| 835 |
if not prompt or not prompt.strip():
|
| 836 |
raise gr.Error("H3-World still needs a scene description alongside the actions.")
|
| 837 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 838 |
from PIL import Image, ImageOps
|
| 839 |
from diffusers.utils import encode_video
|
| 840 |
|
|
@@ -872,7 +916,7 @@ def generate(
|
|
| 872 |
progress(0.1, desc=f"Generating {num_frames / FPS:.1f}s at {width}x{height} in {int(steps)} steps ...")
|
| 873 |
started = time.time()
|
| 874 |
frames, audio, sampling_rate = _generate(
|
| 875 |
-
prompt_embeds, tags, plan, keyframe, height, width, num_frames, int(steps), int(seed)
|
| 876 |
)
|
| 877 |
generate_seconds = time.time() - started
|
| 878 |
|
|
@@ -884,8 +928,10 @@ def generate(
|
|
| 884 |
if plan is not None
|
| 885 |
else "directed mask **off** (sentences act as one global prompt)"
|
| 886 |
)
|
|
|
|
| 887 |
report = (
|
| 888 |
f"{width}x{height} · {num_frames} frames ({num_frames / FPS:.2f}s) · {int(steps)} steps · "
|
|
|
|
| 889 |
f"{mask_note} · conditioner {condition_seconds:.0f}s · denoise + decode {generate_seconds:.0f}s · "
|
| 890 |
f"seed {int(seed)}\n\n"
|
| 891 |
f"<details><summary>Action timeline</summary>\n\n{summarize(sequence, subject)}\n\n</details>"
|
|
@@ -924,7 +970,12 @@ Each frame's key state becomes one short sentence — *"the man walks forward, c
|
|
| 924 |
sharply"* — and a **directed attention mask** binds that sentence to that frame inside MiniMax-H3's
|
| 925 |
packed sequence, which is what makes the control per-frame rather than a global prompt.
|
| 926 |
|
| 927 |
-
`W`/`A`/`S`/`D` move · `I`/`J`/`K`/`L` aim the camera · `F` makes the camera move sharp.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 928 |
|
| 929 |
CSS = """
|
| 930 |
.main.fillable { max-width: 1150px !important; }
|
|
@@ -984,6 +1035,15 @@ with gr.Blocks(title="H3-World") as demo:
|
|
| 984 |
+ " — or raw keys like WA, LF."
|
| 985 |
),
|
| 986 |
)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 987 |
run = gr.Button("Roll the world forward", variant="primary", size="lg")
|
| 988 |
|
| 989 |
with gr.Accordion("Advanced", open=False):
|
|
@@ -995,7 +1055,12 @@ with gr.Blocks(title="H3-World") as demo:
|
|
| 995 |
value=DEFAULT_DURATION,
|
| 996 |
)
|
| 997 |
steps_slider = gr.Slider(
|
| 998 |
-
label="Steps",
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 999 |
)
|
| 1000 |
canvas = gr.Dropdown(label="Canvas", choices=list(CANVASES), value=DEFAULT_CANVAS)
|
| 1001 |
seed = gr.Number(label="Seed", value=0, precision=0)
|
|
@@ -1013,12 +1078,15 @@ with gr.Blocks(title="H3-World") as demo:
|
|
| 1013 |
preview_md = gr.Markdown(preview("forward*20, forward-right*10, pan-right-fast*7",
|
| 1014 |
DEFAULT_DURATION, DEFAULT_SUBJECT))
|
| 1015 |
|
| 1016 |
-
INPUTS = [prompt, image, script, canvas, duration_slider, steps_slider, seed, subject, directed]
|
| 1017 |
OUTPUTS = [result, report]
|
| 1018 |
|
| 1019 |
image.upload(_fit_keyframe_ui, [image, canvas], [image, canvas])
|
| 1020 |
for control in (script, duration_slider, subject):
|
| 1021 |
control.change(preview, [script, duration_slider, subject], [preview_md])
|
|
|
|
|
|
|
|
|
|
| 1022 |
|
| 1023 |
gr.Examples(
|
| 1024 |
examples=EXAMPLES,
|
|
|
|
| 48 |
import gradio as gr
|
| 49 |
import torch
|
| 50 |
|
| 51 |
+
import h3_turbo_lora
|
| 52 |
+
|
| 53 |
MODEL_REPO = os.environ.get("H3_MODEL_REPO", "MiniMaxAI/MiniMax-H3")
|
| 54 |
LORA_REPO = os.environ.get("H3_LORA_REPO", "DANNY621/H3-World")
|
| 55 |
LORA_FILE = os.environ.get("H3_LORA_FILE", "step-10000.safetensors")
|
|
|
|
| 61 |
FPS, FRAMES_PER_CHUNK, LATENTS_PER_CHUNK = 24, 17, 5
|
| 62 |
MIN_UI_DURATION, MAX_UI_DURATION = 2, 8
|
| 63 |
# The H3-World inference runs recorded in the checkpoint's own manifests: 124 frames, 480x832, 24 fps,
|
| 64 |
+
# 50 steps, seed 0. 50 is still reachable on the Steps slider (it is its maximum).
|
| 65 |
DEFAULT_DURATION = 5
|
| 66 |
+
DEFAULT_STEPS = 28
|
| 67 |
DEFAULT_SUBJECT = "the man"
|
| 68 |
|
| 69 |
+
# The two sampling modes the quality comparison is between: MiniMax-H3's own 28-step default with
|
| 70 |
+
# H3-World alone, and `lightx2v/Minimax-h3-Turbo`'s 8-NFE distillation folded in on top of it.
|
| 71 |
+
MODE_BASE = "28 steps · no turbo LoRA"
|
| 72 |
+
MODE_TURBO = "8 steps · turbo LoRA"
|
| 73 |
+
MODES: dict[str, tuple[int, bool]] = {
|
| 74 |
+
MODE_BASE: (DEFAULT_STEPS, False),
|
| 75 |
+
MODE_TURBO: (h3_turbo_lora.TURBO_STEPS, True),
|
| 76 |
+
}
|
| 77 |
+
DEFAULT_MODE = MODE_BASE
|
| 78 |
+
|
| 79 |
# The conditioner offers canvases up to 1344x768, but H3-World was trained at 832x480 and the directed
|
| 80 |
# mask costs O(sequence x captions) — at 1344x768 / 8s a request would need ~35 GPU-minutes, far past
|
| 81 |
# what any visitor's ZeroGPU quota can book. So the list is the cheap tier of each aspect ratio, which
|
|
|
|
| 101 |
PIPE = None
|
| 102 |
LOAD_ERROR: str | None = None
|
| 103 |
LOADED_IN: float | None = None
|
| 104 |
+
TURBO_ERROR: str | None = None
|
| 105 |
|
| 106 |
|
| 107 |
# ── The action vocabulary ────────────────────────────────────────────────────
|
|
|
|
| 588 |
|
| 589 |
def load_models() -> str | None:
|
| 590 |
"""Load the denoising half at startup: transformer + VAEs, fold H3-World in, patch attention."""
|
| 591 |
+
global PIPE, LOAD_ERROR, LOADED_IN, TURBO_ERROR
|
| 592 |
|
| 593 |
if PIPE is not None or LOAD_ERROR is not None:
|
| 594 |
return LOAD_ERROR
|
|
|
|
| 614 |
|
| 615 |
print(f"[gen] {load_and_apply_lora(pipe.transformer)}", flush=True)
|
| 616 |
|
| 617 |
+
# The turbo LoRA is only *prepared* here — its factors stay resident and are folded in (and
|
| 618 |
+
# back out) per request, so both sampling modes are one click apart.
|
| 619 |
+
try:
|
| 620 |
+
print(f"[gen] {h3_turbo_lora.prepare(pipe.transformer)}", flush=True)
|
| 621 |
+
except Exception as error: # noqa: BLE001
|
| 622 |
+
traceback.print_exc()
|
| 623 |
+
TURBO_ERROR = f"{type(error).__name__}: {error}"
|
| 624 |
+
print(f"[gen] WARNING: turbo LoRA unavailable ({TURBO_ERROR})", flush=True)
|
| 625 |
+
|
| 626 |
if PLACEMENT == "pack":
|
| 627 |
pipe.transformer.to("cuda")
|
| 628 |
|
|
|
|
| 762 |
_DECODE_BASE, _DECODE_PER_DEFAULT_CANVAS, _DEFAULT_CANVAS_PIXELS = 15, 15, 960 * 544 * 124
|
| 763 |
# ~12% over the fit, plus one cold worker's weight placement.
|
| 764 |
_MARGIN, _PLACEMENT_ALLOWANCE = 1.12, 12
|
| 765 |
+
# Folding (or unfolding) the turbo LoRA rewrites 312 weights through a rank-128 matmul each — a few
|
| 766 |
+
# seconds of card, but budget generously since a worker may have to unfold the other state first.
|
| 767 |
+
_TURBO_FOLD_ALLOWANCE = 60
|
| 768 |
|
| 769 |
|
| 770 |
+
def get_duration(
|
| 771 |
+
prompt_embeds, tags, plan, image, height, width, num_frames, steps, seed, turbo=False, *args, **kwargs
|
| 772 |
+
):
|
| 773 |
"""Estimate GPU seconds for one ``_generate`` call, from measurements rather than guesswork."""
|
| 774 |
height, width, num_frames, steps = int(height), int(width), int(num_frames), int(steps)
|
| 775 |
patches = (height // 32) * (width // 32)
|
|
|
|
| 780 |
if plan is not None:
|
| 781 |
per_step += _MASK * rows * len(plan["spans"]) * _TOKENS_PER_CAPTION
|
| 782 |
decode = _DECODE_BASE + _DECODE_PER_DEFAULT_CANVAS * (height * width * num_frames) / _DEFAULT_CANVAS_PIXELS
|
| 783 |
+
# Either direction can cost a fold: a turbo request folds it in, and the next base request on a
|
| 784 |
+
# warm worker folds it back out. So the allowance rides along whenever the LoRA is loaded at all.
|
| 785 |
+
fold = _TURBO_FOLD_ALLOWANCE if TURBO_ERROR is None else 0
|
| 786 |
+
return max(60, int((steps * per_step + decode) * _MARGIN) + _PLACEMENT_ALLOWANCE) + fold
|
| 787 |
|
| 788 |
|
| 789 |
# ── Inference ────────────────────────────────────────────────────────────────
|
| 790 |
|
| 791 |
|
| 792 |
@spaces.GPU(duration=get_duration, size=GPU_SIZE)
|
| 793 |
+
def _generate(prompt_embeds, tags, plan, image, height, width, num_frames, steps, seed, turbo=False):
|
| 794 |
"""Denoise + decode on GPU time, with the directed mask active for this request's geometry."""
|
| 795 |
if PLACEMENT == "pack":
|
| 796 |
PIPE.vae.to("cuda")
|
|
|
|
| 798 |
elif PLACEMENT == "lazy":
|
| 799 |
PIPE.to("cuda")
|
| 800 |
|
| 801 |
+
# In or out, in place, on the card the weights already sit on.
|
| 802 |
+
folded = h3_turbo_lora.set_active(PIPE.transformer, turbo)
|
| 803 |
+
print(f"[gen] turbo LoRA {'folded in' if folded else 'off'}", flush=True)
|
| 804 |
+
|
| 805 |
DIRECTED["plan"] = plan
|
| 806 |
try:
|
| 807 |
state = PIPE(
|
|
|
|
| 835 |
script: str = "forward",
|
| 836 |
canvas: str = DEFAULT_CANVAS,
|
| 837 |
duration: float = DEFAULT_DURATION,
|
| 838 |
+
steps: int = 0,
|
| 839 |
seed: int = 0,
|
| 840 |
subject: str = DEFAULT_SUBJECT,
|
| 841 |
directed: bool = True,
|
| 842 |
+
mode: str = DEFAULT_MODE,
|
| 843 |
progress=gr.Progress(track_tqdm=True),
|
| 844 |
):
|
| 845 |
"""Roll a world forward from one frame under a keyboard action script.
|
|
|
|
| 853 |
tilt-down, or a raw key combination out of W A S D I J K L F.
|
| 854 |
canvas: Output resolution; snapped to the first frame's aspect ratio when one is given.
|
| 855 |
duration: Seconds of video, rounded up to the next frame count the VAE can decode.
|
| 856 |
+
steps: Denoising steps, or ``0`` for whatever ``mode`` asks for (28 without the turbo LoRA,
|
| 857 |
+
8 with it). 28 is MiniMax-H3's default; the released H3-World results use 50.
|
| 858 |
seed: Random seed.
|
| 859 |
subject: How the per-frame sentences refer to the character.
|
| 860 |
directed: Bind each sentence to its own latent frame with H3-World's directed attention mask.
|
| 861 |
Turning it off is the ablation: the sentences become one global prompt.
|
| 862 |
+
mode: ``"28 steps · no turbo LoRA"`` or ``"8 steps · turbo LoRA"`` — the second folds
|
| 863 |
+
``lightx2v/Minimax-h3-Turbo``'s 8-step distillation in on top of H3-World and drops the
|
| 864 |
+
step count to 8, for a quality-versus-speed comparison at the same seed.
|
| 865 |
|
| 866 |
Returns:
|
| 867 |
The generated video (with MiniMax-H3's native soundtrack) and a markdown report holding the
|
|
|
|
| 874 |
if not prompt or not prompt.strip():
|
| 875 |
raise gr.Error("H3-World still needs a scene description alongside the actions.")
|
| 876 |
|
| 877 |
+
mode_steps, turbo = MODES.get(str(mode), MODES[DEFAULT_MODE])
|
| 878 |
+
if turbo and TURBO_ERROR:
|
| 879 |
+
raise gr.Error(f"The turbo LoRA failed to load, so only {MODE_BASE} is available: {TURBO_ERROR}")
|
| 880 |
+
steps = int(steps) or mode_steps
|
| 881 |
+
|
| 882 |
from PIL import Image, ImageOps
|
| 883 |
from diffusers.utils import encode_video
|
| 884 |
|
|
|
|
| 916 |
progress(0.1, desc=f"Generating {num_frames / FPS:.1f}s at {width}x{height} in {int(steps)} steps ...")
|
| 917 |
started = time.time()
|
| 918 |
frames, audio, sampling_rate = _generate(
|
| 919 |
+
prompt_embeds, tags, plan, keyframe, height, width, num_frames, int(steps), int(seed), turbo
|
| 920 |
)
|
| 921 |
generate_seconds = time.time() - started
|
| 922 |
|
|
|
|
| 928 |
if plan is not None
|
| 929 |
else "directed mask **off** (sentences act as one global prompt)"
|
| 930 |
)
|
| 931 |
+
turbo_note = "turbo LoRA **on** (8-step distillation)" if turbo else "no turbo LoRA"
|
| 932 |
report = (
|
| 933 |
f"{width}x{height} · {num_frames} frames ({num_frames / FPS:.2f}s) · {int(steps)} steps · "
|
| 934 |
+
f"{turbo_note} · "
|
| 935 |
f"{mask_note} · conditioner {condition_seconds:.0f}s · denoise + decode {generate_seconds:.0f}s · "
|
| 936 |
f"seed {int(seed)}\n\n"
|
| 937 |
f"<details><summary>Action timeline</summary>\n\n{summarize(sequence, subject)}\n\n</details>"
|
|
|
|
| 970 |
sharply"* — and a **directed attention mask** binds that sentence to that frame inside MiniMax-H3's
|
| 971 |
packed sequence, which is what makes the control per-frame rather than a global prompt.
|
| 972 |
|
| 973 |
+
`W`/`A`/`S`/`D` move · `I`/`J`/`K`/`L` aim the camera · `F` makes the camera move sharp.
|
| 974 |
+
|
| 975 |
+
**Sampling** picks between MiniMax-H3's 28-step default and
|
| 976 |
+
[`lightx2v/Minimax-h3-Turbo`](https://huggingface.co/lightx2v/Minimax-h3-Turbo)'s **8-step**
|
| 977 |
+
distillation (`minimax_h3_fl2v_turbo_8step_v1.0_bf16`) folded in on top of H3-World — same seed,
|
| 978 |
+
same action script, ~3.5x fewer steps, so you can judge the quality difference yourself."""
|
| 979 |
|
| 980 |
CSS = """
|
| 981 |
.main.fillable { max-width: 1150px !important; }
|
|
|
|
| 1035 |
+ " — or raw keys like WA, LF."
|
| 1036 |
),
|
| 1037 |
)
|
| 1038 |
+
mode = gr.Radio(
|
| 1039 |
+
label="Sampling",
|
| 1040 |
+
choices=list(MODES),
|
| 1041 |
+
value=DEFAULT_MODE,
|
| 1042 |
+
info=(
|
| 1043 |
+
"Same seed, same script — compare MiniMax-H3's 28-step default against "
|
| 1044 |
+
"lightx2v/Minimax-h3-Turbo's 8-step distillation folded in on top of H3-World."
|
| 1045 |
+
),
|
| 1046 |
+
)
|
| 1047 |
run = gr.Button("Roll the world forward", variant="primary", size="lg")
|
| 1048 |
|
| 1049 |
with gr.Accordion("Advanced", open=False):
|
|
|
|
| 1055 |
value=DEFAULT_DURATION,
|
| 1056 |
)
|
| 1057 |
steps_slider = gr.Slider(
|
| 1058 |
+
label="Steps",
|
| 1059 |
+
minimum=4,
|
| 1060 |
+
maximum=50,
|
| 1061 |
+
step=1,
|
| 1062 |
+
value=DEFAULT_STEPS,
|
| 1063 |
+
info="Follows the sampling mode; 50 is the released H3-World configuration.",
|
| 1064 |
)
|
| 1065 |
canvas = gr.Dropdown(label="Canvas", choices=list(CANVASES), value=DEFAULT_CANVAS)
|
| 1066 |
seed = gr.Number(label="Seed", value=0, precision=0)
|
|
|
|
| 1078 |
preview_md = gr.Markdown(preview("forward*20, forward-right*10, pan-right-fast*7",
|
| 1079 |
DEFAULT_DURATION, DEFAULT_SUBJECT))
|
| 1080 |
|
| 1081 |
+
INPUTS = [prompt, image, script, canvas, duration_slider, steps_slider, seed, subject, directed, mode]
|
| 1082 |
OUTPUTS = [result, report]
|
| 1083 |
|
| 1084 |
image.upload(_fit_keyframe_ui, [image, canvas], [image, canvas])
|
| 1085 |
for control in (script, duration_slider, subject):
|
| 1086 |
control.change(preview, [script, duration_slider, subject], [preview_md])
|
| 1087 |
+
# Picking a sampling mode moves the Steps slider to that mode's own count; the slider stays an
|
| 1088 |
+
# override, so 50 (the released configuration) is still reachable with the LoRA either way.
|
| 1089 |
+
mode.change(lambda choice: gr.update(value=MODES.get(choice, MODES[DEFAULT_MODE])[0]), [mode], [steps_slider])
|
| 1090 |
|
| 1091 |
gr.Examples(
|
| 1092 |
examples=EXAMPLES,
|
h3_turbo_lora.py
ADDED
|
@@ -0,0 +1,144 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""The lightx2v 8-step Turbo LoRA, folded into the same bf16 weights H3-World is merged into.
|
| 2 |
+
|
| 3 |
+
This mirrors the `h3_lora.py` of the Spaces that already run this adapter on MiniMax-H3 —
|
| 4 |
+
[`MiniMaxAI/MiniMax-H3-Turbo-Lora`](https://huggingface.co/spaces/MiniMaxAI/MiniMax-H3-Turbo-Lora)
|
| 5 |
+
(its `lightx` set) and
|
| 6 |
+
[`hugging-apps/minimax-h3-turbo-sla-demo`](https://huggingface.co/spaces/hugging-apps/minimax-h3-turbo-sla-demo):
|
| 7 |
+
|
| 8 |
+
* `lightx2v/Minimax-h3-Turbo` is a **PEFT checkpoint against the diffusers module tree itself** —
|
| 9 |
+
`transformer_blocks.N.attn.to_q.lora_A.default.weight` and friends — so unlike H3-World it needs
|
| 10 |
+
no key conversion at all: 312 targets (50 transformer blocks + 2 token-refiner blocks x
|
| 11 |
+
`to_q/to_k/to_v/to_out.0/ff.net.0.proj/ff.net.2`) map name-for-name.
|
| 12 |
+
* Rank is 128 and the file's own safetensors metadata records `alpha: 8`, so the fold scale is
|
| 13 |
+
`alpha / rank = 0.0625` — exactly what `set_adapters(weights=1.0)` applies in lightx2v's
|
| 14 |
+
reference inference script, and what the Spaces above use for this repo.
|
| 15 |
+
* `minimax_h3_fl2v_turbo_8step_v1.0_bf16.safetensors` is an 8-NFE distillation, so the step count
|
| 16 |
+
is overridden to **8** when it is active. No scheduler swap and no CFG change: MiniMax-H3 is
|
| 17 |
+
already guidance-distilled and its own `MiniMaxH3Scheduler` is what every one of those Spaces
|
| 18 |
+
keeps using with the turbo LoRA folded.
|
| 19 |
+
|
| 20 |
+
The adapter is *folded* into the weights rather than wrapped as a runtime module, for the same
|
| 21 |
+
reason H3-World is: this Space patches `MiniMaxH3AttnProcessor` and drives the transformer's live
|
| 22 |
+
weights. Fold and unfold are the same operation with a sign, so the active set can be flipped per
|
| 23 |
+
request (`set_active`) through one bf16 rounding, with the low-rank factors staying resident.
|
| 24 |
+
"""
|
| 25 |
+
|
| 26 |
+
from __future__ import annotations
|
| 27 |
+
|
| 28 |
+
import os
|
| 29 |
+
|
| 30 |
+
import torch
|
| 31 |
+
|
| 32 |
+
TURBO_REPO = os.environ.get("H3_TURBO_REPO", "lightx2v/Minimax-h3-Turbo")
|
| 33 |
+
TURBO_FILE = os.environ.get("H3_TURBO_FILE", "minimax_h3_fl2v_turbo_8step_v1.0_bf16.safetensors")
|
| 34 |
+
# The distillation's own NFE count; the card offers 4 as the faster, softer alternative.
|
| 35 |
+
TURBO_STEPS = int(os.environ.get("H3_TURBO_STEPS", "8"))
|
| 36 |
+
# 0 = read `alpha` out of the file's safetensors metadata (it is `8` for every lightx2v file).
|
| 37 |
+
TURBO_ALPHA = float(os.environ.get("H3_TURBO_ALPHA", "0"))
|
| 38 |
+
TURBO_STRENGTH = float(os.environ.get("H3_TURBO_STRENGTH", "1.0"))
|
| 39 |
+
|
| 40 |
+
SUFFIX_A, SUFFIX_B = ".lora_A.default.weight", ".lora_B.default.weight"
|
| 41 |
+
|
| 42 |
+
|
| 43 |
+
def load() -> dict:
|
| 44 |
+
"""Download the adapter and return ``{"label", "scale", "entries"}``; entries are ``(param_key, A, B)``."""
|
| 45 |
+
from huggingface_hub import hf_hub_download
|
| 46 |
+
from safetensors import safe_open
|
| 47 |
+
from safetensors.torch import load_file
|
| 48 |
+
|
| 49 |
+
path = hf_hub_download(TURBO_REPO, TURBO_FILE)
|
| 50 |
+
with safe_open(path, framework="pt") as handle:
|
| 51 |
+
metadata = handle.metadata() or {}
|
| 52 |
+
alpha = TURBO_ALPHA or float(metadata.get("alpha", 8))
|
| 53 |
+
|
| 54 |
+
lora = load_file(path)
|
| 55 |
+
bases = sorted({key[: -len(SUFFIX_A)] for key in lora if key.endswith(SUFFIX_A)})
|
| 56 |
+
if not bases:
|
| 57 |
+
raise ValueError(f"No lora_A/lora_B pairs found in {TURBO_FILE}")
|
| 58 |
+
unexpected = [key for key in lora if not key.endswith((SUFFIX_A, SUFFIX_B))]
|
| 59 |
+
if unexpected:
|
| 60 |
+
raise ValueError(f"{TURBO_FILE} holds {len(unexpected)} non-LoRA tensors, e.g. {unexpected[:5]}")
|
| 61 |
+
for name in bases:
|
| 62 |
+
if f"{name}{SUFFIX_B}" not in lora:
|
| 63 |
+
raise ValueError(f"LoRA is missing the lora_B twin of {name}{SUFFIX_A}")
|
| 64 |
+
|
| 65 |
+
ranks = {lora[f"{name}{SUFFIX_A}"].shape[0] for name in bases}
|
| 66 |
+
if len(ranks) != 1:
|
| 67 |
+
raise ValueError(f"LoRA mixes ranks {sorted(ranks)}")
|
| 68 |
+
rank = ranks.pop()
|
| 69 |
+
|
| 70 |
+
entries = [(f"{name}.weight", lora[f"{name}{SUFFIX_A}"], lora[f"{name}{SUFFIX_B}"]) for name in bases]
|
| 71 |
+
return {
|
| 72 |
+
"label": f"{TURBO_REPO}/{TURBO_FILE}",
|
| 73 |
+
"scale": alpha / rank * TURBO_STRENGTH,
|
| 74 |
+
"entries": entries,
|
| 75 |
+
"rank": rank,
|
| 76 |
+
"alpha": alpha,
|
| 77 |
+
}
|
| 78 |
+
|
| 79 |
+
|
| 80 |
+
def _fold(transformer, spec: dict, sign: float) -> None:
|
| 81 |
+
"""Add ``sign * scale * (B @ A)`` to every target weight, in place.
|
| 82 |
+
|
| 83 |
+
The factors ride to the weight's own device before the matmul — the deltas are 5 TFLOP in
|
| 84 |
+
total, which is seconds on the card and minutes on a Space's two vCPUs.
|
| 85 |
+
"""
|
| 86 |
+
params = dict(transformer.named_parameters())
|
| 87 |
+
scale = sign * spec["scale"]
|
| 88 |
+
with torch.no_grad():
|
| 89 |
+
for key, a, b in spec["entries"]:
|
| 90 |
+
param = params[key]
|
| 91 |
+
delta = b.to(param.device, torch.float32) @ a.to(param.device, torch.float32)
|
| 92 |
+
if delta.shape != param.shape:
|
| 93 |
+
raise ValueError(
|
| 94 |
+
f"Turbo LoRA delta for `{key}` has shape {tuple(delta.shape)}, "
|
| 95 |
+
f"base weight is {tuple(param.shape)}"
|
| 96 |
+
)
|
| 97 |
+
param.data = (param.data.float() + scale * delta).to(param.dtype)
|
| 98 |
+
del delta
|
| 99 |
+
|
| 100 |
+
|
| 101 |
+
def prepare(transformer) -> str:
|
| 102 |
+
"""Load the adapter and validate it against ``transformer``, without folding it yet."""
|
| 103 |
+
spec = load()
|
| 104 |
+
params = dict(transformer.named_parameters())
|
| 105 |
+
missed = [key for key, _, _ in spec["entries"] if key not in params]
|
| 106 |
+
if missed:
|
| 107 |
+
raise ValueError(
|
| 108 |
+
f"{len(missed)} turbo-LoRA targets matched no transformer weight, e.g. {missed[:5]}. "
|
| 109 |
+
"The LoRA and the diffusers transformer disagree on module naming."
|
| 110 |
+
)
|
| 111 |
+
for key, a, b in spec["entries"]: # cheaper here than discovering it mid-request
|
| 112 |
+
shape = (b.shape[0], a.shape[1])
|
| 113 |
+
if shape != tuple(params[key].shape):
|
| 114 |
+
raise ValueError(
|
| 115 |
+
f"Turbo LoRA delta for `{key}` would be {shape}, base weight is {tuple(params[key].shape)}"
|
| 116 |
+
)
|
| 117 |
+
transformer._turbo_state = {"active": False, "spec": spec}
|
| 118 |
+
return (
|
| 119 |
+
f"Turbo LoRA ready · {spec['label']} · {len(spec['entries'])} targets, "
|
| 120 |
+
f"rank {spec['rank']}, alpha {spec['alpha']:g} -> scale {spec['scale']:.4f} · "
|
| 121 |
+
f"{TURBO_STEPS} steps when active"
|
| 122 |
+
)
|
| 123 |
+
|
| 124 |
+
|
| 125 |
+
def is_available(transformer) -> bool:
|
| 126 |
+
return getattr(transformer, "_turbo_state", None) is not None
|
| 127 |
+
|
| 128 |
+
|
| 129 |
+
def set_active(transformer, on: bool) -> bool:
|
| 130 |
+
"""Fold the turbo LoRA in (or back out). No-op when the weights already match. Returns the state.
|
| 131 |
+
|
| 132 |
+
Called from inside the ``@spaces.GPU`` function: the transformer's weights live on the card
|
| 133 |
+
there, and a ZeroGPU worker that is reused keeps both them and this state, while a freshly
|
| 134 |
+
forked one starts from the unfolded weights *and* the unfolded state — consistent either way.
|
| 135 |
+
"""
|
| 136 |
+
state = getattr(transformer, "_turbo_state", None)
|
| 137 |
+
if state is None:
|
| 138 |
+
return False
|
| 139 |
+
on = bool(on)
|
| 140 |
+
if state["active"] == on:
|
| 141 |
+
return on
|
| 142 |
+
_fold(transformer, state["spec"], 1.0 if on else -1.0)
|
| 143 |
+
state["active"] = on
|
| 144 |
+
return on
|