multimodalart's picture
multimodalart HF Staff
Add a 28-step vs 8-step turbo LoRA sampling choice (lightx2v/Minimax-h3-Turbo)
5a00791 verified
|
Raw History Blame
11 kB
metadata
title: H3-World
emoji: ๐ŸŽฎ
colorFrom: indigo
colorTo: green
sdk: gradio
sdk_version: 6.26.0
app_file: app.py
python_version: '3.12'
startup_duration_timeout: 1h
short_description: Drive a world model with WASD and camera keys
models:
  - MiniMaxAI/MiniMax-H3
  - DANNY621/H3-World
  - lightx2v/Minimax-h3-Turbo

H3-World โ€” action-conditioned world model

DANNY621/H3-World is a rank-32 LoRA for MiniMaxAI/MiniMax-H3 that turns it into a playable world model: you give it a first frame and a key sequence, and it renders what happens as those keys are held.

forward*20, forward-right*10, pan-right-fast*7
keys meaning
W A S D walk forward / strafe left / walk backward / strafe right
J L camera pans left / right
K I camera tilts up / down
F modifier โ€” the camera move is sharp rather than slow

How the actions actually reach the model

The checkpoint's own run manifests are unambiguous about this, and it is the thing that makes the LoRA demo-able at all:

"checkpoint": {"action_dim": 0, "action_mode": "text", "action_tensors": 0,
               "lora_tensors": 208, "lora_pairs": 104}
"config": {"num_frames": 124, "height": 480, "width": 832, "fps": 24, "steps": 50,
           "conditioning": "first_frame+8d_actions",
           "action_columns": ["W", "A", "S", "D", "I", "J", "K", "L"]}

action_dim: 0 and action_tensors: 0 โ€” there is no action encoder and no action embedding. The 8-dimensional key state is carried through the text channel: one short English sentence per latent video frame, appended to the scene prompt, in the register the author's own runs use.

So a 124-frame request (37 latent frames) is conditioned on a prompt that looks like:

A third-person view of a man walking through a city intersection...
the man walks forward
the man walks forward
...
the man walks forward and strafes right
...
the man stands still, camera pans right sharply

This Space builds that text from the key script for you โ€” caption_for() in app.py โ€” and shows you the resolved per-frame captions under Conditioning after every run.

The directed attention mask

The model card is explicit that the LoRA weights alone do not reproduce the reported behavior: the training run used a directed attention mask that binds each per-frame caption to the latents of its frame. Without it every video row attends to all 37 sentences at once and the sequence collapses into an average action.

MiniMax-H3 is a single packed 1-D sequence under full self-attention โ€” [text | keyframe anchors | audio | video], no cross-attention โ€” so the mask is a constraint inside one attention call, not a separate cross-attention mask.

app.py reimplements it as an exact log-sum-exp merge rather than a dense [S, S] mask (which would force SDPA off its flash kernel for the whole 21k-row sequence). The keys are split into three regions:

region keys kernel
A text rows before the caption block flash, unmasked
C the ~700 caption rows fp32 masked matmul, chunked over queries
B everything after the text block (~99% of keys) flash, unmasked

Each returns its output and its log-sum-exp; the three are recombined with the online-softmax identity, which is numerically identical to one masked softmax over the full row. Only video queries are restricted (frame i's rows see only sentence i); the captions themselves still see everything โ€” that is the "directed" part.

The mask is toggleable in Advanced so you can see the difference. With it off, the same script produces a video that drifts through a blur of every action at once.

Two guards, both in app.py:

  • The token spans are located by re-tokenizing prefixes of the prompt. If a BPE merge straddles a sentence boundary the spans would be wrong, so build_conditioning_text() verifies cuts[-1] == total and refuses to mask rather than mask the wrong rows.
  • MiniMaxH3TokenRefinerBlock runs the same attention module over the text stream alone. The processor detects that (the sequence is too short to contain the video block) and falls straight through to the stock path.

Architecture โ€” why two Spaces

MiniMax-H3 is 195.9 GiB in bfloat16 and a ZeroGPU Space is evicted at 150 GB of storage, so the pipeline is split at its text_encoder step, exactly as in multimodalart/minimax-h3:

  • the 62.14 GiB Qwen3-VL conditioner runs in multimodalart/qwen3vl-conditioner, which this Space calls over the gradio API for every request;
  • this Space holds the 61.73 GiB transformer (with H3-World merged into it) plus the video and audio autoencoders.

The wire format between them is prompt_embeds (1, N, 5120) bf16 + text_token_tags (N,) int64 in a single safetensors file. The caption spans this Space needs are recovered from it by offset arithmetic, because MiniMaxH3TextEncoderStep appends the prompt verbatim โ€” no chat template, no special tokens โ€” so offset = num_text_tokens - num_prompt_tokens.

h3_split_blocks.py is the blockset with the text_encoder step removed, copied from that Space.

LoRA merge

The LoRA is published against the original MiniMax-H3 layout, not the diffusers port, so load_lora_weights() does not apply. load_and_apply_lora() replays convert_minimax_h3_to_diffusers.py's renames on the way in โ€” blocks. โ†’ transformer_blocks., attn.out_proj โ†’ attn.to_out.0, mlp.fc1 โ†’ ff.net.0.proj with the SwiGLU gate/value halves swapped, and the fused attn.qkv_proj de-interleaved per head before being split into to_q / to_k / to_v. 104 LoRA pairs become 208 merged weight deltas; any target that fails to resolve is fatal, never skipped.

28 steps vs the 8-step turbo LoRA

Sampling in the UI picks between two configurations of the same request โ€” same seed, same action script, same directed mask โ€” so the quality cost of the distillation is directly visible:

mode steps transformer
28 steps ยท no turbo LoRA 28 (MiniMax-H3's default) H3-World only
8 steps ยท turbo LoRA 8 H3-World + lightx2v/Minimax-h3-Turbo minimax_h3_fl2v_turbo_8step_v1.0_bf16.safetensors

h3_turbo_lora.py mirrors the h3_lora.py of the Spaces that already run this adapter โ€” MiniMaxAI/MiniMax-H3-Turbo-Lora (its lightx set) and hugging-apps/minimax-h3-turbo-sla-demo:

  • the file is a PEFT checkpoint against the diffusers module tree itself (transformer_blocks.N.attn.to_q.lora_A.default.weight), so unlike H3-World it needs no key conversion โ€” 312 targets (50 transformer blocks + 2 token-refiner blocks x to_q/to_k/to_v/to_out.0/ff.net.0.proj/ff.net.2) map name-for-name;
  • rank 128 with alpha: 8 in the file's own safetensors metadata, so the fold scale is alpha / rank = 0.0625 โ€” what set_adapters(weights=1.0) applies in lightx2v's reference script;
  • the step count is overridden to the distillation's own 8 NFE. No scheduler swap and no CFG change: MiniMax-H3 is already guidance-distilled and every Space above keeps its native MiniMaxH3Scheduler with the turbo LoRA folded.

It is folded into the bf16 weights, like H3-World, because this Space patches MiniMaxH3AttnProcessor and drives the transformer's live weights. Fold and unfold are the same operation with a sign, so the low-rank factors stay resident and set_active flips the mode in place inside the @spaces.GPU call, through one bf16 rounding.

The Steps slider still overrides the mode's count (4โ€“50), so 50 steps + turbo LoRA or 8 steps without it are both reachable for the sake of the comparison.

Generation constraints

Fixed by the checkpoint: 24 fps, num_frames snapped to 17n + 5, no CFG and no negative prompt (it is guidance-distilled). H3-World was trained at 832x480; the conditioner's canvas list does not offer that exact size, so the default here is its nearest neighbour, 960x544.

The offered canvases are the cheap tier of each aspect ratio rather than the conditioner's full list. The mask term scales as sequence x caption rows, so a 1344x768 / 8 s request would want ~35 GPU-minutes โ€” past what any visitor could book โ€” and it is off-distribution for a LoRA trained at 832x480 anyway.

Measured

On this Space, driven over gradio_client, at 960x544 with a keyframe:

Request Conditioner Denoise + decode Round trip
16 steps, 56 frames, directed 2 s 45 s 50 s
16 steps, 56 frames, no mask 2 s 32 s 36 s
50 steps, 124 frames, directed (the default) 10 s 303 s 316 s

Startup is 95 s: the 66.3 GB download, the load, the LoRA merge, and the ZeroGPU pack. get_duration is fitted to exactly these three points โ€” the unmasked block cost linear + quadratic in the packed sequence, the mask's own term linear in sequence x captions โ€” and books ~15% over the fit. The default request books 348 s.

Examples

The three bundled first frames are extracted from acvlab/ABot-World-Explorer-500h (Apache-2.0), which is the same kind of third-person game footage H3-World was trained on. Each is paired with that clip's own manifest prompt.

Space variables

Variable Default Meaning
H3_MODEL_REPO MiniMaxAI/MiniMax-H3 The diffusers-layout base checkpoint.
H3_LORA_REPO DANNY621/H3-World The LoRA.
H3_LORA_FILE step-10000.safetensors The checkpoint the author's own test runs used.
H3_TURBO_REPO lightx2v/Minimax-h3-Turbo The turbo-LoRA repo.
H3_TURBO_FILE minimax_h3_fl2v_turbo_8step_v1.0_bf16.safetensors The 8-step distillation.
H3_TURBO_STEPS 8 Steps the turbo mode asks for (the card also offers 4).
H3_TURBO_ALPHA 0 0 reads alpha out of the file's metadata (8).
H3_TURBO_STRENGTH 1.0 Extra multiplier on the turbo delta.
H3_CONDITIONER multimodalart/qwen3vl-conditioner The Space this one asks for embeddings.
H3_ATTENTION _native_cudnn cuDNN's fused kernel. flash-attention 3 is sm90-only; this pool is sm120.
H3_GPU_SIZE xlarge ZeroGPU allocation size. large does not fit.

License

The LoRA is Apache-2.0, but usage is governed by the base model's license (MiniMaxAI/MiniMax-H3).