multimodalart HF Staff commited on
Commit
5a00791
·
verified ·
1 Parent(s): 1889dd6

Add a 28-step vs 8-step turbo LoRA sampling choice (lightx2v/Minimax-h3-Turbo)

Browse files
Files changed (3) hide show
  1. README.md +39 -0
  2. app.py +80 -12
  3. h3_turbo_lora.py +144 -0
README.md CHANGED
@@ -12,6 +12,7 @@ short_description: Drive a world model with WASD and camera keys
12
  models:
13
  - MiniMaxAI/MiniMax-H3
14
  - DANNY621/H3-World
 
15
  ---
16
 
17
  # H3-World — action-conditioned world model
@@ -122,6 +123,39 @@ gate/value halves swapped, and the fused `attn.qkv_proj` de-interleaved per head
122
  `to_q` / `to_k` / `to_v`. 104 LoRA pairs become 208 merged weight deltas; **any target that fails to resolve is
123
  fatal**, never skipped.
124
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
125
  ## Generation constraints
126
 
127
  Fixed by the checkpoint: 24 fps, `num_frames` snapped to `17n + 5`, no CFG and no negative prompt (it is
@@ -160,6 +194,11 @@ manifest prompt.
160
  | `H3_MODEL_REPO` | `MiniMaxAI/MiniMax-H3` | The diffusers-layout base checkpoint. |
161
  | `H3_LORA_REPO` | `DANNY621/H3-World` | The LoRA. |
162
  | `H3_LORA_FILE` | `step-10000.safetensors` | The checkpoint the author's own test runs used. |
 
 
 
 
 
163
  | `H3_CONDITIONER` | `multimodalart/qwen3vl-conditioner` | The Space this one asks for embeddings. |
164
  | `H3_ATTENTION` | `_native_cudnn` | cuDNN's fused kernel. flash-attention 3 is sm90-only; this pool is sm120. |
165
  | `H3_GPU_SIZE` | `xlarge` | ZeroGPU allocation size. `large` does not fit. |
 
12
  models:
13
  - MiniMaxAI/MiniMax-H3
14
  - DANNY621/H3-World
15
+ - lightx2v/Minimax-h3-Turbo
16
  ---
17
 
18
  # H3-World — action-conditioned world model
 
123
  `to_q` / `to_k` / `to_v`. 104 LoRA pairs become 208 merged weight deltas; **any target that fails to resolve is
124
  fatal**, never skipped.
125
 
126
+ ## 28 steps vs the 8-step turbo LoRA
127
+
128
+ **Sampling** in the UI picks between two configurations of the same request — same seed, same action
129
+ script, same directed mask — so the quality cost of the distillation is directly visible:
130
+
131
+ | mode | steps | transformer |
132
+ |---|---|---|
133
+ | `28 steps · no turbo LoRA` | 28 (MiniMax-H3's default) | H3-World only |
134
+ | `8 steps · turbo LoRA` | 8 | H3-World **+** [`lightx2v/Minimax-h3-Turbo`](https://huggingface.co/lightx2v/Minimax-h3-Turbo) `minimax_h3_fl2v_turbo_8step_v1.0_bf16.safetensors` |
135
+
136
+ `h3_turbo_lora.py` mirrors the `h3_lora.py` of the Spaces that already run this adapter —
137
+ [`MiniMaxAI/MiniMax-H3-Turbo-Lora`](https://huggingface.co/spaces/MiniMaxAI/MiniMax-H3-Turbo-Lora)
138
+ (its `lightx` set) and
139
+ [`hugging-apps/minimax-h3-turbo-sla-demo`](https://huggingface.co/spaces/hugging-apps/minimax-h3-turbo-sla-demo):
140
+
141
+ * the file is a **PEFT checkpoint against the diffusers module tree itself**
142
+ (`transformer_blocks.N.attn.to_q.lora_A.default.weight`), so unlike H3-World it needs no key
143
+ conversion — 312 targets (50 transformer blocks + 2 token-refiner blocks x
144
+ `to_q`/`to_k`/`to_v`/`to_out.0`/`ff.net.0.proj`/`ff.net.2`) map name-for-name;
145
+ * rank 128 with `alpha: 8` in the file's own safetensors metadata, so the fold scale is
146
+ `alpha / rank = 0.0625` — what `set_adapters(weights=1.0)` applies in lightx2v's reference script;
147
+ * the step count is overridden to the distillation's own **8 NFE**. No scheduler swap and no CFG
148
+ change: MiniMax-H3 is already guidance-distilled and every Space above keeps its native
149
+ `MiniMaxH3Scheduler` with the turbo LoRA folded.
150
+
151
+ It is *folded* into the bf16 weights, like H3-World, because this Space patches
152
+ `MiniMaxH3AttnProcessor` and drives the transformer's live weights. Fold and unfold are the same
153
+ operation with a sign, so the low-rank factors stay resident and `set_active` flips the mode in
154
+ place inside the `@spaces.GPU` call, through one bf16 rounding.
155
+
156
+ The Steps slider still overrides the mode's count (4–50), so `50 steps + turbo LoRA` or
157
+ `8 steps without it` are both reachable for the sake of the comparison.
158
+
159
  ## Generation constraints
160
 
161
  Fixed by the checkpoint: 24 fps, `num_frames` snapped to `17n + 5`, no CFG and no negative prompt (it is
 
194
  | `H3_MODEL_REPO` | `MiniMaxAI/MiniMax-H3` | The diffusers-layout base checkpoint. |
195
  | `H3_LORA_REPO` | `DANNY621/H3-World` | The LoRA. |
196
  | `H3_LORA_FILE` | `step-10000.safetensors` | The checkpoint the author's own test runs used. |
197
+ | `H3_TURBO_REPO` | `lightx2v/Minimax-h3-Turbo` | The turbo-LoRA repo. |
198
+ | `H3_TURBO_FILE` | `minimax_h3_fl2v_turbo_8step_v1.0_bf16.safetensors` | The 8-step distillation. |
199
+ | `H3_TURBO_STEPS` | `8` | Steps the turbo mode asks for (the card also offers 4). |
200
+ | `H3_TURBO_ALPHA` | `0` | `0` reads `alpha` out of the file's metadata (`8`). |
201
+ | `H3_TURBO_STRENGTH` | `1.0` | Extra multiplier on the turbo delta. |
202
  | `H3_CONDITIONER` | `multimodalart/qwen3vl-conditioner` | The Space this one asks for embeddings. |
203
  | `H3_ATTENTION` | `_native_cudnn` | cuDNN's fused kernel. flash-attention 3 is sm90-only; this pool is sm120. |
204
  | `H3_GPU_SIZE` | `xlarge` | ZeroGPU allocation size. `large` does not fit. |
app.py CHANGED
@@ -48,6 +48,8 @@ import spaces
48
  import gradio as gr
49
  import torch
50
 
 
 
51
  MODEL_REPO = os.environ.get("H3_MODEL_REPO", "MiniMaxAI/MiniMax-H3")
52
  LORA_REPO = os.environ.get("H3_LORA_REPO", "DANNY621/H3-World")
53
  LORA_FILE = os.environ.get("H3_LORA_FILE", "step-10000.safetensors")
@@ -59,11 +61,21 @@ GPU_SIZE = os.environ.get("H3_GPU_SIZE", "xlarge")
59
  FPS, FRAMES_PER_CHUNK, LATENTS_PER_CHUNK = 24, 17, 5
60
  MIN_UI_DURATION, MAX_UI_DURATION = 2, 8
61
  # The H3-World inference runs recorded in the checkpoint's own manifests: 124 frames, 480x832, 24 fps,
62
- # 50 steps, seed 0. Keep those as the defaults so the Space reproduces the released results.
63
  DEFAULT_DURATION = 5
64
- DEFAULT_STEPS = 50
65
  DEFAULT_SUBJECT = "the man"
66
 
 
 
 
 
 
 
 
 
 
 
67
  # The conditioner offers canvases up to 1344x768, but H3-World was trained at 832x480 and the directed
68
  # mask costs O(sequence x captions) — at 1344x768 / 8s a request would need ~35 GPU-minutes, far past
69
  # what any visitor's ZeroGPU quota can book. So the list is the cheap tier of each aspect ratio, which
@@ -89,6 +101,7 @@ os.makedirs(OUTPUT_DIR, exist_ok=True)
89
  PIPE = None
90
  LOAD_ERROR: str | None = None
91
  LOADED_IN: float | None = None
 
92
 
93
 
94
  # ── The action vocabulary ────────────────────────────────────────────────────
@@ -575,7 +588,7 @@ def lower_duration_floor(seconds: float = MIN_UI_DURATION) -> None:
575
 
576
  def load_models() -> str | None:
577
  """Load the denoising half at startup: transformer + VAEs, fold H3-World in, patch attention."""
578
- global PIPE, LOAD_ERROR, LOADED_IN
579
 
580
  if PIPE is not None or LOAD_ERROR is not None:
581
  return LOAD_ERROR
@@ -601,6 +614,15 @@ def load_models() -> str | None:
601
 
602
  print(f"[gen] {load_and_apply_lora(pipe.transformer)}", flush=True)
603
 
 
 
 
 
 
 
 
 
 
604
  if PLACEMENT == "pack":
605
  pipe.transformer.to("cuda")
606
 
@@ -740,9 +762,14 @@ _TOKENS_PER_CAPTION = 8
740
  _DECODE_BASE, _DECODE_PER_DEFAULT_CANVAS, _DEFAULT_CANVAS_PIXELS = 15, 15, 960 * 544 * 124
741
  # ~12% over the fit, plus one cold worker's weight placement.
742
  _MARGIN, _PLACEMENT_ALLOWANCE = 1.12, 12
 
 
 
743
 
744
 
745
- def get_duration(prompt_embeds, tags, plan, image, height, width, num_frames, steps, seed, *args, **kwargs):
 
 
746
  """Estimate GPU seconds for one ``_generate`` call, from measurements rather than guesswork."""
747
  height, width, num_frames, steps = int(height), int(width), int(num_frames), int(steps)
748
  patches = (height // 32) * (width // 32)
@@ -753,14 +780,17 @@ def get_duration(prompt_embeds, tags, plan, image, height, width, num_frames, st
753
  if plan is not None:
754
  per_step += _MASK * rows * len(plan["spans"]) * _TOKENS_PER_CAPTION
755
  decode = _DECODE_BASE + _DECODE_PER_DEFAULT_CANVAS * (height * width * num_frames) / _DEFAULT_CANVAS_PIXELS
756
- return max(60, int((steps * per_step + decode) * _MARGIN) + _PLACEMENT_ALLOWANCE)
 
 
 
757
 
758
 
759
  # ── Inference ────────────────────────────────────────────────────────────────
760
 
761
 
762
  @spaces.GPU(duration=get_duration, size=GPU_SIZE)
763
- def _generate(prompt_embeds, tags, plan, image, height, width, num_frames, steps, seed):
764
  """Denoise + decode on GPU time, with the directed mask active for this request's geometry."""
765
  if PLACEMENT == "pack":
766
  PIPE.vae.to("cuda")
@@ -768,6 +798,10 @@ def _generate(prompt_embeds, tags, plan, image, height, width, num_frames, steps
768
  elif PLACEMENT == "lazy":
769
  PIPE.to("cuda")
770
 
 
 
 
 
771
  DIRECTED["plan"] = plan
772
  try:
773
  state = PIPE(
@@ -801,10 +835,11 @@ def generate(
801
  script: str = "forward",
802
  canvas: str = DEFAULT_CANVAS,
803
  duration: float = DEFAULT_DURATION,
804
- steps: int = DEFAULT_STEPS,
805
  seed: int = 0,
806
  subject: str = DEFAULT_SUBJECT,
807
  directed: bool = True,
 
808
  progress=gr.Progress(track_tqdm=True),
809
  ):
810
  """Roll a world forward from one frame under a keyboard action script.
@@ -818,11 +853,15 @@ def generate(
818
  tilt-down, or a raw key combination out of W A S D I J K L F.
819
  canvas: Output resolution; snapped to the first frame's aspect ratio when one is given.
820
  duration: Seconds of video, rounded up to the next frame count the VAE can decode.
821
- steps: Denoising steps. The released H3-World results use 50.
 
822
  seed: Random seed.
823
  subject: How the per-frame sentences refer to the character.
824
  directed: Bind each sentence to its own latent frame with H3-World's directed attention mask.
825
  Turning it off is the ablation: the sentences become one global prompt.
 
 
 
826
 
827
  Returns:
828
  The generated video (with MiniMax-H3's native soundtrack) and a markdown report holding the
@@ -835,6 +874,11 @@ def generate(
835
  if not prompt or not prompt.strip():
836
  raise gr.Error("H3-World still needs a scene description alongside the actions.")
837
 
 
 
 
 
 
838
  from PIL import Image, ImageOps
839
  from diffusers.utils import encode_video
840
 
@@ -872,7 +916,7 @@ def generate(
872
  progress(0.1, desc=f"Generating {num_frames / FPS:.1f}s at {width}x{height} in {int(steps)} steps ...")
873
  started = time.time()
874
  frames, audio, sampling_rate = _generate(
875
- prompt_embeds, tags, plan, keyframe, height, width, num_frames, int(steps), int(seed)
876
  )
877
  generate_seconds = time.time() - started
878
 
@@ -884,8 +928,10 @@ def generate(
884
  if plan is not None
885
  else "directed mask **off** (sentences act as one global prompt)"
886
  )
 
887
  report = (
888
  f"{width}x{height} · {num_frames} frames ({num_frames / FPS:.2f}s) · {int(steps)} steps · "
 
889
  f"{mask_note} · conditioner {condition_seconds:.0f}s · denoise + decode {generate_seconds:.0f}s · "
890
  f"seed {int(seed)}\n\n"
891
  f"<details><summary>Action timeline</summary>\n\n{summarize(sequence, subject)}\n\n</details>"
@@ -924,7 +970,12 @@ Each frame's key state becomes one short sentence — *"the man walks forward, c
924
  sharply"* — and a **directed attention mask** binds that sentence to that frame inside MiniMax-H3's
925
  packed sequence, which is what makes the control per-frame rather than a global prompt.
926
 
927
- `W`/`A`/`S`/`D` move · `I`/`J`/`K`/`L` aim the camera · `F` makes the camera move sharp."""
 
 
 
 
 
928
 
929
  CSS = """
930
  .main.fillable { max-width: 1150px !important; }
@@ -984,6 +1035,15 @@ with gr.Blocks(title="H3-World") as demo:
984
  + " — or raw keys like WA, LF."
985
  ),
986
  )
 
 
 
 
 
 
 
 
 
987
  run = gr.Button("Roll the world forward", variant="primary", size="lg")
988
 
989
  with gr.Accordion("Advanced", open=False):
@@ -995,7 +1055,12 @@ with gr.Blocks(title="H3-World") as demo:
995
  value=DEFAULT_DURATION,
996
  )
997
  steps_slider = gr.Slider(
998
- label="Steps", minimum=16, maximum=50, step=1, value=DEFAULT_STEPS
 
 
 
 
 
999
  )
1000
  canvas = gr.Dropdown(label="Canvas", choices=list(CANVASES), value=DEFAULT_CANVAS)
1001
  seed = gr.Number(label="Seed", value=0, precision=0)
@@ -1013,12 +1078,15 @@ with gr.Blocks(title="H3-World") as demo:
1013
  preview_md = gr.Markdown(preview("forward*20, forward-right*10, pan-right-fast*7",
1014
  DEFAULT_DURATION, DEFAULT_SUBJECT))
1015
 
1016
- INPUTS = [prompt, image, script, canvas, duration_slider, steps_slider, seed, subject, directed]
1017
  OUTPUTS = [result, report]
1018
 
1019
  image.upload(_fit_keyframe_ui, [image, canvas], [image, canvas])
1020
  for control in (script, duration_slider, subject):
1021
  control.change(preview, [script, duration_slider, subject], [preview_md])
 
 
 
1022
 
1023
  gr.Examples(
1024
  examples=EXAMPLES,
 
48
  import gradio as gr
49
  import torch
50
 
51
+ import h3_turbo_lora
52
+
53
  MODEL_REPO = os.environ.get("H3_MODEL_REPO", "MiniMaxAI/MiniMax-H3")
54
  LORA_REPO = os.environ.get("H3_LORA_REPO", "DANNY621/H3-World")
55
  LORA_FILE = os.environ.get("H3_LORA_FILE", "step-10000.safetensors")
 
61
  FPS, FRAMES_PER_CHUNK, LATENTS_PER_CHUNK = 24, 17, 5
62
  MIN_UI_DURATION, MAX_UI_DURATION = 2, 8
63
  # The H3-World inference runs recorded in the checkpoint's own manifests: 124 frames, 480x832, 24 fps,
64
+ # 50 steps, seed 0. 50 is still reachable on the Steps slider (it is its maximum).
65
  DEFAULT_DURATION = 5
66
+ DEFAULT_STEPS = 28
67
  DEFAULT_SUBJECT = "the man"
68
 
69
+ # The two sampling modes the quality comparison is between: MiniMax-H3's own 28-step default with
70
+ # H3-World alone, and `lightx2v/Minimax-h3-Turbo`'s 8-NFE distillation folded in on top of it.
71
+ MODE_BASE = "28 steps · no turbo LoRA"
72
+ MODE_TURBO = "8 steps · turbo LoRA"
73
+ MODES: dict[str, tuple[int, bool]] = {
74
+ MODE_BASE: (DEFAULT_STEPS, False),
75
+ MODE_TURBO: (h3_turbo_lora.TURBO_STEPS, True),
76
+ }
77
+ DEFAULT_MODE = MODE_BASE
78
+
79
  # The conditioner offers canvases up to 1344x768, but H3-World was trained at 832x480 and the directed
80
  # mask costs O(sequence x captions) — at 1344x768 / 8s a request would need ~35 GPU-minutes, far past
81
  # what any visitor's ZeroGPU quota can book. So the list is the cheap tier of each aspect ratio, which
 
101
  PIPE = None
102
  LOAD_ERROR: str | None = None
103
  LOADED_IN: float | None = None
104
+ TURBO_ERROR: str | None = None
105
 
106
 
107
  # ── The action vocabulary ────────────────────────────────────────────────────
 
588
 
589
  def load_models() -> str | None:
590
  """Load the denoising half at startup: transformer + VAEs, fold H3-World in, patch attention."""
591
+ global PIPE, LOAD_ERROR, LOADED_IN, TURBO_ERROR
592
 
593
  if PIPE is not None or LOAD_ERROR is not None:
594
  return LOAD_ERROR
 
614
 
615
  print(f"[gen] {load_and_apply_lora(pipe.transformer)}", flush=True)
616
 
617
+ # The turbo LoRA is only *prepared* here — its factors stay resident and are folded in (and
618
+ # back out) per request, so both sampling modes are one click apart.
619
+ try:
620
+ print(f"[gen] {h3_turbo_lora.prepare(pipe.transformer)}", flush=True)
621
+ except Exception as error: # noqa: BLE001
622
+ traceback.print_exc()
623
+ TURBO_ERROR = f"{type(error).__name__}: {error}"
624
+ print(f"[gen] WARNING: turbo LoRA unavailable ({TURBO_ERROR})", flush=True)
625
+
626
  if PLACEMENT == "pack":
627
  pipe.transformer.to("cuda")
628
 
 
762
  _DECODE_BASE, _DECODE_PER_DEFAULT_CANVAS, _DEFAULT_CANVAS_PIXELS = 15, 15, 960 * 544 * 124
763
  # ~12% over the fit, plus one cold worker's weight placement.
764
  _MARGIN, _PLACEMENT_ALLOWANCE = 1.12, 12
765
+ # Folding (or unfolding) the turbo LoRA rewrites 312 weights through a rank-128 matmul each — a few
766
+ # seconds of card, but budget generously since a worker may have to unfold the other state first.
767
+ _TURBO_FOLD_ALLOWANCE = 60
768
 
769
 
770
+ def get_duration(
771
+ prompt_embeds, tags, plan, image, height, width, num_frames, steps, seed, turbo=False, *args, **kwargs
772
+ ):
773
  """Estimate GPU seconds for one ``_generate`` call, from measurements rather than guesswork."""
774
  height, width, num_frames, steps = int(height), int(width), int(num_frames), int(steps)
775
  patches = (height // 32) * (width // 32)
 
780
  if plan is not None:
781
  per_step += _MASK * rows * len(plan["spans"]) * _TOKENS_PER_CAPTION
782
  decode = _DECODE_BASE + _DECODE_PER_DEFAULT_CANVAS * (height * width * num_frames) / _DEFAULT_CANVAS_PIXELS
783
+ # Either direction can cost a fold: a turbo request folds it in, and the next base request on a
784
+ # warm worker folds it back out. So the allowance rides along whenever the LoRA is loaded at all.
785
+ fold = _TURBO_FOLD_ALLOWANCE if TURBO_ERROR is None else 0
786
+ return max(60, int((steps * per_step + decode) * _MARGIN) + _PLACEMENT_ALLOWANCE) + fold
787
 
788
 
789
  # ── Inference ────────────────────────────────────────────────────────────────
790
 
791
 
792
  @spaces.GPU(duration=get_duration, size=GPU_SIZE)
793
+ def _generate(prompt_embeds, tags, plan, image, height, width, num_frames, steps, seed, turbo=False):
794
  """Denoise + decode on GPU time, with the directed mask active for this request's geometry."""
795
  if PLACEMENT == "pack":
796
  PIPE.vae.to("cuda")
 
798
  elif PLACEMENT == "lazy":
799
  PIPE.to("cuda")
800
 
801
+ # In or out, in place, on the card the weights already sit on.
802
+ folded = h3_turbo_lora.set_active(PIPE.transformer, turbo)
803
+ print(f"[gen] turbo LoRA {'folded in' if folded else 'off'}", flush=True)
804
+
805
  DIRECTED["plan"] = plan
806
  try:
807
  state = PIPE(
 
835
  script: str = "forward",
836
  canvas: str = DEFAULT_CANVAS,
837
  duration: float = DEFAULT_DURATION,
838
+ steps: int = 0,
839
  seed: int = 0,
840
  subject: str = DEFAULT_SUBJECT,
841
  directed: bool = True,
842
+ mode: str = DEFAULT_MODE,
843
  progress=gr.Progress(track_tqdm=True),
844
  ):
845
  """Roll a world forward from one frame under a keyboard action script.
 
853
  tilt-down, or a raw key combination out of W A S D I J K L F.
854
  canvas: Output resolution; snapped to the first frame's aspect ratio when one is given.
855
  duration: Seconds of video, rounded up to the next frame count the VAE can decode.
856
+ steps: Denoising steps, or ``0`` for whatever ``mode`` asks for (28 without the turbo LoRA,
857
+ 8 with it). 28 is MiniMax-H3's default; the released H3-World results use 50.
858
  seed: Random seed.
859
  subject: How the per-frame sentences refer to the character.
860
  directed: Bind each sentence to its own latent frame with H3-World's directed attention mask.
861
  Turning it off is the ablation: the sentences become one global prompt.
862
+ mode: ``"28 steps · no turbo LoRA"`` or ``"8 steps · turbo LoRA"`` — the second folds
863
+ ``lightx2v/Minimax-h3-Turbo``'s 8-step distillation in on top of H3-World and drops the
864
+ step count to 8, for a quality-versus-speed comparison at the same seed.
865
 
866
  Returns:
867
  The generated video (with MiniMax-H3's native soundtrack) and a markdown report holding the
 
874
  if not prompt or not prompt.strip():
875
  raise gr.Error("H3-World still needs a scene description alongside the actions.")
876
 
877
+ mode_steps, turbo = MODES.get(str(mode), MODES[DEFAULT_MODE])
878
+ if turbo and TURBO_ERROR:
879
+ raise gr.Error(f"The turbo LoRA failed to load, so only {MODE_BASE} is available: {TURBO_ERROR}")
880
+ steps = int(steps) or mode_steps
881
+
882
  from PIL import Image, ImageOps
883
  from diffusers.utils import encode_video
884
 
 
916
  progress(0.1, desc=f"Generating {num_frames / FPS:.1f}s at {width}x{height} in {int(steps)} steps ...")
917
  started = time.time()
918
  frames, audio, sampling_rate = _generate(
919
+ prompt_embeds, tags, plan, keyframe, height, width, num_frames, int(steps), int(seed), turbo
920
  )
921
  generate_seconds = time.time() - started
922
 
 
928
  if plan is not None
929
  else "directed mask **off** (sentences act as one global prompt)"
930
  )
931
+ turbo_note = "turbo LoRA **on** (8-step distillation)" if turbo else "no turbo LoRA"
932
  report = (
933
  f"{width}x{height} · {num_frames} frames ({num_frames / FPS:.2f}s) · {int(steps)} steps · "
934
+ f"{turbo_note} · "
935
  f"{mask_note} · conditioner {condition_seconds:.0f}s · denoise + decode {generate_seconds:.0f}s · "
936
  f"seed {int(seed)}\n\n"
937
  f"<details><summary>Action timeline</summary>\n\n{summarize(sequence, subject)}\n\n</details>"
 
970
  sharply"* — and a **directed attention mask** binds that sentence to that frame inside MiniMax-H3's
971
  packed sequence, which is what makes the control per-frame rather than a global prompt.
972
 
973
+ `W`/`A`/`S`/`D` move · `I`/`J`/`K`/`L` aim the camera · `F` makes the camera move sharp.
974
+
975
+ **Sampling** picks between MiniMax-H3's 28-step default and
976
+ [`lightx2v/Minimax-h3-Turbo`](https://huggingface.co/lightx2v/Minimax-h3-Turbo)'s **8-step**
977
+ distillation (`minimax_h3_fl2v_turbo_8step_v1.0_bf16`) folded in on top of H3-World — same seed,
978
+ same action script, ~3.5x fewer steps, so you can judge the quality difference yourself."""
979
 
980
  CSS = """
981
  .main.fillable { max-width: 1150px !important; }
 
1035
  + " — or raw keys like WA, LF."
1036
  ),
1037
  )
1038
+ mode = gr.Radio(
1039
+ label="Sampling",
1040
+ choices=list(MODES),
1041
+ value=DEFAULT_MODE,
1042
+ info=(
1043
+ "Same seed, same script — compare MiniMax-H3's 28-step default against "
1044
+ "lightx2v/Minimax-h3-Turbo's 8-step distillation folded in on top of H3-World."
1045
+ ),
1046
+ )
1047
  run = gr.Button("Roll the world forward", variant="primary", size="lg")
1048
 
1049
  with gr.Accordion("Advanced", open=False):
 
1055
  value=DEFAULT_DURATION,
1056
  )
1057
  steps_slider = gr.Slider(
1058
+ label="Steps",
1059
+ minimum=4,
1060
+ maximum=50,
1061
+ step=1,
1062
+ value=DEFAULT_STEPS,
1063
+ info="Follows the sampling mode; 50 is the released H3-World configuration.",
1064
  )
1065
  canvas = gr.Dropdown(label="Canvas", choices=list(CANVASES), value=DEFAULT_CANVAS)
1066
  seed = gr.Number(label="Seed", value=0, precision=0)
 
1078
  preview_md = gr.Markdown(preview("forward*20, forward-right*10, pan-right-fast*7",
1079
  DEFAULT_DURATION, DEFAULT_SUBJECT))
1080
 
1081
+ INPUTS = [prompt, image, script, canvas, duration_slider, steps_slider, seed, subject, directed, mode]
1082
  OUTPUTS = [result, report]
1083
 
1084
  image.upload(_fit_keyframe_ui, [image, canvas], [image, canvas])
1085
  for control in (script, duration_slider, subject):
1086
  control.change(preview, [script, duration_slider, subject], [preview_md])
1087
+ # Picking a sampling mode moves the Steps slider to that mode's own count; the slider stays an
1088
+ # override, so 50 (the released configuration) is still reachable with the LoRA either way.
1089
+ mode.change(lambda choice: gr.update(value=MODES.get(choice, MODES[DEFAULT_MODE])[0]), [mode], [steps_slider])
1090
 
1091
  gr.Examples(
1092
  examples=EXAMPLES,
h3_turbo_lora.py ADDED
@@ -0,0 +1,144 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """The lightx2v 8-step Turbo LoRA, folded into the same bf16 weights H3-World is merged into.
2
+
3
+ This mirrors the `h3_lora.py` of the Spaces that already run this adapter on MiniMax-H3 —
4
+ [`MiniMaxAI/MiniMax-H3-Turbo-Lora`](https://huggingface.co/spaces/MiniMaxAI/MiniMax-H3-Turbo-Lora)
5
+ (its `lightx` set) and
6
+ [`hugging-apps/minimax-h3-turbo-sla-demo`](https://huggingface.co/spaces/hugging-apps/minimax-h3-turbo-sla-demo):
7
+
8
+ * `lightx2v/Minimax-h3-Turbo` is a **PEFT checkpoint against the diffusers module tree itself** —
9
+ `transformer_blocks.N.attn.to_q.lora_A.default.weight` and friends — so unlike H3-World it needs
10
+ no key conversion at all: 312 targets (50 transformer blocks + 2 token-refiner blocks x
11
+ `to_q/to_k/to_v/to_out.0/ff.net.0.proj/ff.net.2`) map name-for-name.
12
+ * Rank is 128 and the file's own safetensors metadata records `alpha: 8`, so the fold scale is
13
+ `alpha / rank = 0.0625` — exactly what `set_adapters(weights=1.0)` applies in lightx2v's
14
+ reference inference script, and what the Spaces above use for this repo.
15
+ * `minimax_h3_fl2v_turbo_8step_v1.0_bf16.safetensors` is an 8-NFE distillation, so the step count
16
+ is overridden to **8** when it is active. No scheduler swap and no CFG change: MiniMax-H3 is
17
+ already guidance-distilled and its own `MiniMaxH3Scheduler` is what every one of those Spaces
18
+ keeps using with the turbo LoRA folded.
19
+
20
+ The adapter is *folded* into the weights rather than wrapped as a runtime module, for the same
21
+ reason H3-World is: this Space patches `MiniMaxH3AttnProcessor` and drives the transformer's live
22
+ weights. Fold and unfold are the same operation with a sign, so the active set can be flipped per
23
+ request (`set_active`) through one bf16 rounding, with the low-rank factors staying resident.
24
+ """
25
+
26
+ from __future__ import annotations
27
+
28
+ import os
29
+
30
+ import torch
31
+
32
+ TURBO_REPO = os.environ.get("H3_TURBO_REPO", "lightx2v/Minimax-h3-Turbo")
33
+ TURBO_FILE = os.environ.get("H3_TURBO_FILE", "minimax_h3_fl2v_turbo_8step_v1.0_bf16.safetensors")
34
+ # The distillation's own NFE count; the card offers 4 as the faster, softer alternative.
35
+ TURBO_STEPS = int(os.environ.get("H3_TURBO_STEPS", "8"))
36
+ # 0 = read `alpha` out of the file's safetensors metadata (it is `8` for every lightx2v file).
37
+ TURBO_ALPHA = float(os.environ.get("H3_TURBO_ALPHA", "0"))
38
+ TURBO_STRENGTH = float(os.environ.get("H3_TURBO_STRENGTH", "1.0"))
39
+
40
+ SUFFIX_A, SUFFIX_B = ".lora_A.default.weight", ".lora_B.default.weight"
41
+
42
+
43
+ def load() -> dict:
44
+ """Download the adapter and return ``{"label", "scale", "entries"}``; entries are ``(param_key, A, B)``."""
45
+ from huggingface_hub import hf_hub_download
46
+ from safetensors import safe_open
47
+ from safetensors.torch import load_file
48
+
49
+ path = hf_hub_download(TURBO_REPO, TURBO_FILE)
50
+ with safe_open(path, framework="pt") as handle:
51
+ metadata = handle.metadata() or {}
52
+ alpha = TURBO_ALPHA or float(metadata.get("alpha", 8))
53
+
54
+ lora = load_file(path)
55
+ bases = sorted({key[: -len(SUFFIX_A)] for key in lora if key.endswith(SUFFIX_A)})
56
+ if not bases:
57
+ raise ValueError(f"No lora_A/lora_B pairs found in {TURBO_FILE}")
58
+ unexpected = [key for key in lora if not key.endswith((SUFFIX_A, SUFFIX_B))]
59
+ if unexpected:
60
+ raise ValueError(f"{TURBO_FILE} holds {len(unexpected)} non-LoRA tensors, e.g. {unexpected[:5]}")
61
+ for name in bases:
62
+ if f"{name}{SUFFIX_B}" not in lora:
63
+ raise ValueError(f"LoRA is missing the lora_B twin of {name}{SUFFIX_A}")
64
+
65
+ ranks = {lora[f"{name}{SUFFIX_A}"].shape[0] for name in bases}
66
+ if len(ranks) != 1:
67
+ raise ValueError(f"LoRA mixes ranks {sorted(ranks)}")
68
+ rank = ranks.pop()
69
+
70
+ entries = [(f"{name}.weight", lora[f"{name}{SUFFIX_A}"], lora[f"{name}{SUFFIX_B}"]) for name in bases]
71
+ return {
72
+ "label": f"{TURBO_REPO}/{TURBO_FILE}",
73
+ "scale": alpha / rank * TURBO_STRENGTH,
74
+ "entries": entries,
75
+ "rank": rank,
76
+ "alpha": alpha,
77
+ }
78
+
79
+
80
+ def _fold(transformer, spec: dict, sign: float) -> None:
81
+ """Add ``sign * scale * (B @ A)`` to every target weight, in place.
82
+
83
+ The factors ride to the weight's own device before the matmul — the deltas are 5 TFLOP in
84
+ total, which is seconds on the card and minutes on a Space's two vCPUs.
85
+ """
86
+ params = dict(transformer.named_parameters())
87
+ scale = sign * spec["scale"]
88
+ with torch.no_grad():
89
+ for key, a, b in spec["entries"]:
90
+ param = params[key]
91
+ delta = b.to(param.device, torch.float32) @ a.to(param.device, torch.float32)
92
+ if delta.shape != param.shape:
93
+ raise ValueError(
94
+ f"Turbo LoRA delta for `{key}` has shape {tuple(delta.shape)}, "
95
+ f"base weight is {tuple(param.shape)}"
96
+ )
97
+ param.data = (param.data.float() + scale * delta).to(param.dtype)
98
+ del delta
99
+
100
+
101
+ def prepare(transformer) -> str:
102
+ """Load the adapter and validate it against ``transformer``, without folding it yet."""
103
+ spec = load()
104
+ params = dict(transformer.named_parameters())
105
+ missed = [key for key, _, _ in spec["entries"] if key not in params]
106
+ if missed:
107
+ raise ValueError(
108
+ f"{len(missed)} turbo-LoRA targets matched no transformer weight, e.g. {missed[:5]}. "
109
+ "The LoRA and the diffusers transformer disagree on module naming."
110
+ )
111
+ for key, a, b in spec["entries"]: # cheaper here than discovering it mid-request
112
+ shape = (b.shape[0], a.shape[1])
113
+ if shape != tuple(params[key].shape):
114
+ raise ValueError(
115
+ f"Turbo LoRA delta for `{key}` would be {shape}, base weight is {tuple(params[key].shape)}"
116
+ )
117
+ transformer._turbo_state = {"active": False, "spec": spec}
118
+ return (
119
+ f"Turbo LoRA ready · {spec['label']} · {len(spec['entries'])} targets, "
120
+ f"rank {spec['rank']}, alpha {spec['alpha']:g} -> scale {spec['scale']:.4f} · "
121
+ f"{TURBO_STEPS} steps when active"
122
+ )
123
+
124
+
125
+ def is_available(transformer) -> bool:
126
+ return getattr(transformer, "_turbo_state", None) is not None
127
+
128
+
129
+ def set_active(transformer, on: bool) -> bool:
130
+ """Fold the turbo LoRA in (or back out). No-op when the weights already match. Returns the state.
131
+
132
+ Called from inside the ``@spaces.GPU`` function: the transformer's weights live on the card
133
+ there, and a ZeroGPU worker that is reused keeps both them and this state, while a freshly
134
+ forked one starts from the unfolded weights *and* the unfolded state — consistent either way.
135
+ """
136
+ state = getattr(transformer, "_turbo_state", None)
137
+ if state is None:
138
+ return False
139
+ on = bool(on)
140
+ if state["active"] == on:
141
+ return on
142
+ _fold(transformer, state["spec"], 1.0 if on else -1.0)
143
+ state["active"] = on
144
+ return on