--- title: H3-World emoji: ๐ŸŽฎ colorFrom: indigo colorTo: green sdk: gradio sdk_version: 6.26.0 app_file: app.py python_version: "3.12" startup_duration_timeout: 1h short_description: Drive a world model with WASD and camera keys models: - MiniMaxAI/MiniMax-H3 - DANNY621/H3-World - lightx2v/Minimax-h3-Turbo --- # H3-World โ€” action-conditioned world model [`DANNY621/H3-World`](https://huggingface.co/DANNY621/H3-World) is a rank-32 LoRA for [`MiniMaxAI/MiniMax-H3`](https://huggingface.co/MiniMaxAI/MiniMax-H3) that turns it into a **playable** world model: you give it a first frame and a **key sequence**, and it renders what happens as those keys are held. ``` forward*20, forward-right*10, pan-right-fast*7 ``` | keys | meaning | |---|---| | `W` `A` `S` `D` | walk forward / strafe left / walk backward / strafe right | | `J` `L` | camera pans left / right | | `K` `I` | camera tilts up / down | | `F` | modifier โ€” the camera move is *sharp* rather than *slow* | ## How the actions actually reach the model The checkpoint's own run manifests are unambiguous about this, and it is the thing that makes the LoRA demo-able at all: ```json "checkpoint": {"action_dim": 0, "action_mode": "text", "action_tensors": 0, "lora_tensors": 208, "lora_pairs": 104} "config": {"num_frames": 124, "height": 480, "width": 832, "fps": 24, "steps": 50, "conditioning": "first_frame+8d_actions", "action_columns": ["W", "A", "S", "D", "I", "J", "K", "L"]} ``` `action_dim: 0` and `action_tensors: 0` โ€” there is **no action encoder and no action embedding**. The 8-dimensional key state is carried through the **text channel**: one short English sentence per *latent* video frame, appended to the scene prompt, in the register the author's own runs use. So a 124-frame request (37 latent frames) is conditioned on a prompt that looks like: ``` A third-person view of a man walking through a city intersection... the man walks forward the man walks forward ... the man walks forward and strafes right ... the man stands still, camera pans right sharply ``` This Space builds that text from the key script for you โ€” `caption_for()` in `app.py` โ€” and shows you the resolved per-frame captions under **Conditioning** after every run. ## The directed attention mask The model card is explicit that the LoRA weights alone do not reproduce the reported behavior: the training run used a **directed attention mask** that binds each per-frame caption to the latents of *its* frame. Without it every video row attends to all 37 sentences at once and the sequence collapses into an average action. MiniMax-H3 is a single packed 1-D sequence under full self-attention โ€” `[text | keyframe anchors | audio | video]`, no cross-attention โ€” so the mask is a constraint inside one attention call, not a separate cross-attention mask. `app.py` reimplements it as an **exact log-sum-exp merge** rather than a dense `[S, S]` mask (which would force SDPA off its flash kernel for the whole 21k-row sequence). The keys are split into three regions: | region | keys | kernel | |---|---|---| | **A** | text rows *before* the caption block | flash, unmasked | | **C** | the ~700 caption rows | fp32 masked matmul, chunked over queries | | **B** | everything after the text block (~99% of keys) | flash, unmasked | Each returns its output and its log-sum-exp; the three are recombined with the online-softmax identity, which is numerically identical to one masked softmax over the full row. Only **video** queries are restricted (frame *i*'s rows see only sentence *i*); the captions themselves still see everything โ€” that is the "directed" part. The mask is toggleable in **Advanced** so you can see the difference. With it off, the same script produces a video that drifts through a blur of every action at once. Two guards, both in `app.py`: * The token spans are located by re-tokenizing prefixes of the prompt. If a BPE merge straddles a sentence boundary the spans would be wrong, so `build_conditioning_text()` verifies `cuts[-1] == total` and **refuses to mask** rather than mask the wrong rows. * `MiniMaxH3TokenRefinerBlock` runs the same attention module over the *text stream alone*. The processor detects that (the sequence is too short to contain the video block) and falls straight through to the stock path. ## Architecture โ€” why two Spaces MiniMax-H3 is 195.9 GiB in bfloat16 and a ZeroGPU Space is evicted at 150 GB of storage, so the pipeline is split at its `text_encoder` step, exactly as in [`multimodalart/minimax-h3`](https://huggingface.co/spaces/multimodalart/minimax-h3): * the 62.14 GiB Qwen3-VL conditioner runs in [`multimodalart/qwen3vl-conditioner`](https://huggingface.co/spaces/multimodalart/qwen3vl-conditioner), which this Space calls over the gradio API for every request; * this Space holds the 61.73 GiB transformer (with H3-World merged into it) plus the video and audio autoencoders. The wire format between them is `prompt_embeds` `(1, N, 5120)` bf16 + `text_token_tags` `(N,)` int64 in a single safetensors file. The caption spans this Space needs are recovered from it by offset arithmetic, because `MiniMaxH3TextEncoderStep` appends the prompt **verbatim** โ€” no chat template, no special tokens โ€” so `offset = num_text_tokens - num_prompt_tokens`. `h3_split_blocks.py` is the blockset with the `text_encoder` step removed, copied from that Space. ## LoRA merge The LoRA is published against the **original** MiniMax-H3 layout, not the diffusers port, so `load_lora_weights()` does not apply. `load_and_apply_lora()` replays `convert_minimax_h3_to_diffusers.py`'s renames on the way in โ€” `blocks.` โ†’ `transformer_blocks.`, `attn.out_proj` โ†’ `attn.to_out.0`, `mlp.fc1` โ†’ `ff.net.0.proj` with the SwiGLU gate/value halves swapped, and the fused `attn.qkv_proj` de-interleaved per head before being split into `to_q` / `to_k` / `to_v`. 104 LoRA pairs become 208 merged weight deltas; **any target that fails to resolve is fatal**, never skipped. ## 28 steps vs the 8-step turbo LoRA **Sampling** in the UI picks between two configurations of the same request โ€” same seed, same action script, same directed mask โ€” so the quality cost of the distillation is directly visible: | mode | steps | transformer | |---|---|---| | `28 steps ยท no turbo LoRA` | 28 (MiniMax-H3's default) | H3-World only | | `8 steps ยท turbo LoRA` | 8 | H3-World **+** [`lightx2v/Minimax-h3-Turbo`](https://huggingface.co/lightx2v/Minimax-h3-Turbo) `minimax_h3_fl2v_turbo_8step_v1.0_bf16.safetensors` | `h3_turbo_lora.py` mirrors the `h3_lora.py` of the Spaces that already run this adapter โ€” [`MiniMaxAI/MiniMax-H3-Turbo-Lora`](https://huggingface.co/spaces/MiniMaxAI/MiniMax-H3-Turbo-Lora) (its `lightx` set) and [`hugging-apps/minimax-h3-turbo-sla-demo`](https://huggingface.co/spaces/hugging-apps/minimax-h3-turbo-sla-demo): * the file is a **PEFT checkpoint against the diffusers module tree itself** (`transformer_blocks.N.attn.to_q.lora_A.default.weight`), so unlike H3-World it needs no key conversion โ€” 312 targets (50 transformer blocks + 2 token-refiner blocks x `to_q`/`to_k`/`to_v`/`to_out.0`/`ff.net.0.proj`/`ff.net.2`) map name-for-name; * rank 128 with `alpha: 8` in the file's own safetensors metadata, so the fold scale is `alpha / rank = 0.0625` โ€” what `set_adapters(weights=1.0)` applies in lightx2v's reference script; * the step count is overridden to the distillation's own **8 NFE**. No scheduler swap and no CFG change: MiniMax-H3 is already guidance-distilled and every Space above keeps its native `MiniMaxH3Scheduler` with the turbo LoRA folded. It is *folded* into the bf16 weights, like H3-World, because this Space patches `MiniMaxH3AttnProcessor` and drives the transformer's live weights. Fold and unfold are the same operation with a sign, so the low-rank factors stay resident and `set_active` flips the mode in place inside the `@spaces.GPU` call, through one bf16 rounding. The Steps slider still overrides the mode's count (4โ€“50), so `50 steps + turbo LoRA` or `8 steps without it` are both reachable for the sake of the comparison. ## Generation constraints Fixed by the checkpoint: 24 fps, `num_frames` snapped to `17n + 5`, no CFG and no negative prompt (it is guidance-distilled). H3-World was trained at **832x480**; the conditioner's canvas list does not offer that exact size, so the default here is its nearest neighbour, **960x544**. The offered canvases are the cheap tier of each aspect ratio rather than the conditioner's full list. The mask term scales as *sequence x caption rows*, so a 1344x768 / 8 s request would want ~35 GPU-minutes โ€” past what any visitor could book โ€” and it is off-distribution for a LoRA trained at 832x480 anyway. ## Measured On this Space, driven over `gradio_client`, at 960x544 with a keyframe: | Request | Conditioner | Denoise + decode | Round trip | |---|---|---|---| | 16 steps, 56 frames, directed | 2 s | 45 s | 50 s | | 16 steps, 56 frames, **no mask** | 2 s | 32 s | 36 s | | 50 steps, 124 frames, directed (the default) | 10 s | 303 s | 316 s | Startup is 95 s: the 66.3 GB download, the load, the LoRA merge, and the ZeroGPU pack. `get_duration` is fitted to exactly these three points โ€” the unmasked block cost linear + quadratic in the packed sequence, the mask's own term linear in *sequence x captions* โ€” and books ~15% over the fit. The default request books 348 s. ## Examples The three bundled first frames are extracted from [`acvlab/ABot-World-Explorer-500h`](https://huggingface.co/datasets/acvlab/ABot-World-Explorer-500h) (Apache-2.0), which is the same kind of third-person game footage H3-World was trained on. Each is paired with that clip's own manifest prompt. ## Space variables | Variable | Default | Meaning | |---|---|---| | `H3_MODEL_REPO` | `MiniMaxAI/MiniMax-H3` | The diffusers-layout base checkpoint. | | `H3_LORA_REPO` | `DANNY621/H3-World` | The LoRA. | | `H3_LORA_FILE` | `step-10000.safetensors` | The checkpoint the author's own test runs used. | | `H3_TURBO_REPO` | `lightx2v/Minimax-h3-Turbo` | The turbo-LoRA repo. | | `H3_TURBO_FILE` | `minimax_h3_fl2v_turbo_8step_v1.0_bf16.safetensors` | The 8-step distillation. | | `H3_TURBO_STEPS` | `8` | Steps the turbo mode asks for (the card also offers 4). | | `H3_TURBO_ALPHA` | `0` | `0` reads `alpha` out of the file's metadata (`8`). | | `H3_TURBO_STRENGTH` | `1.0` | Extra multiplier on the turbo delta. | | `H3_CONDITIONER` | `multimodalart/qwen3vl-conditioner` | The Space this one asks for embeddings. | | `H3_ATTENTION` | `_native_cudnn` | cuDNN's fused kernel. flash-attention 3 is sm90-only; this pool is sm120. | | `H3_GPU_SIZE` | `xlarge` | ZeroGPU allocation size. `large` does not fit. | ## License The LoRA is Apache-2.0, but usage is governed by the **base model's** license ([`MiniMaxAI/MiniMax-H3`](https://huggingface.co/MiniMaxAI/MiniMax-H3)).