Spaces:
Running on Zero
Running on Zero
File size: 10,969 Bytes
6749e6b 2c82f76 6749e6b 2c82f76 5a00791 6749e6b 2c82f76 68b5806 2c82f76 5a00791 2c82f76 68b5806 2c82f76 5a00791 2c82f76 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 | ---
title: H3-World
emoji: ๐ฎ
colorFrom: indigo
colorTo: green
sdk: gradio
sdk_version: 6.26.0
app_file: app.py
python_version: "3.12"
startup_duration_timeout: 1h
short_description: Drive a world model with WASD and camera keys
models:
- MiniMaxAI/MiniMax-H3
- DANNY621/H3-World
- lightx2v/Minimax-h3-Turbo
---
# H3-World โ action-conditioned world model
[`DANNY621/H3-World`](https://huggingface.co/DANNY621/H3-World) is a rank-32 LoRA for
[`MiniMaxAI/MiniMax-H3`](https://huggingface.co/MiniMaxAI/MiniMax-H3) that turns it into a **playable** world model:
you give it a first frame and a **key sequence**, and it renders what happens as those keys are held.
```
forward*20, forward-right*10, pan-right-fast*7
```
| keys | meaning |
|---|---|
| `W` `A` `S` `D` | walk forward / strafe left / walk backward / strafe right |
| `J` `L` | camera pans left / right |
| `K` `I` | camera tilts up / down |
| `F` | modifier โ the camera move is *sharp* rather than *slow* |
## How the actions actually reach the model
The checkpoint's own run manifests are unambiguous about this, and it is the thing that makes the LoRA
demo-able at all:
```json
"checkpoint": {"action_dim": 0, "action_mode": "text", "action_tensors": 0,
"lora_tensors": 208, "lora_pairs": 104}
"config": {"num_frames": 124, "height": 480, "width": 832, "fps": 24, "steps": 50,
"conditioning": "first_frame+8d_actions",
"action_columns": ["W", "A", "S", "D", "I", "J", "K", "L"]}
```
`action_dim: 0` and `action_tensors: 0` โ there is **no action encoder and no action embedding**. The 8-dimensional
key state is carried through the **text channel**: one short English sentence per *latent* video frame, appended to the
scene prompt, in the register the author's own runs use.
So a 124-frame request (37 latent frames) is conditioned on a prompt that looks like:
```
A third-person view of a man walking through a city intersection...
the man walks forward
the man walks forward
...
the man walks forward and strafes right
...
the man stands still, camera pans right sharply
```
This Space builds that text from the key script for you โ `caption_for()` in `app.py` โ and shows you the resolved
per-frame captions under **Conditioning** after every run.
## The directed attention mask
The model card is explicit that the LoRA weights alone do not reproduce the reported behavior: the training run used
a **directed attention mask** that binds each per-frame caption to the latents of *its* frame. Without it every video
row attends to all 37 sentences at once and the sequence collapses into an average action.
MiniMax-H3 is a single packed 1-D sequence under full self-attention โ `[text | keyframe anchors | audio | video]`,
no cross-attention โ so the mask is a constraint inside one attention call, not a separate cross-attention mask.
`app.py` reimplements it as an **exact log-sum-exp merge** rather than a dense `[S, S]` mask (which would force SDPA
off its flash kernel for the whole 21k-row sequence). The keys are split into three regions:
| region | keys | kernel |
|---|---|---|
| **A** | text rows *before* the caption block | flash, unmasked |
| **C** | the ~700 caption rows | fp32 masked matmul, chunked over queries |
| **B** | everything after the text block (~99% of keys) | flash, unmasked |
Each returns its output and its log-sum-exp; the three are recombined with the online-softmax identity, which is
numerically identical to one masked softmax over the full row. Only **video** queries are restricted (frame *i*'s rows
see only sentence *i*); the captions themselves still see everything โ that is the "directed" part.
The mask is toggleable in **Advanced** so you can see the difference. With it off, the same script produces a video
that drifts through a blur of every action at once.
Two guards, both in `app.py`:
* The token spans are located by re-tokenizing prefixes of the prompt. If a BPE merge straddles a sentence boundary
the spans would be wrong, so `build_conditioning_text()` verifies `cuts[-1] == total` and **refuses to mask** rather
than mask the wrong rows.
* `MiniMaxH3TokenRefinerBlock` runs the same attention module over the *text stream alone*. The processor detects that
(the sequence is too short to contain the video block) and falls straight through to the stock path.
## Architecture โ why two Spaces
MiniMax-H3 is 195.9 GiB in bfloat16 and a ZeroGPU Space is evicted at 150 GB of storage, so the pipeline is split at
its `text_encoder` step, exactly as in [`multimodalart/minimax-h3`](https://huggingface.co/spaces/multimodalart/minimax-h3):
* the 62.14 GiB Qwen3-VL conditioner runs in
[`multimodalart/qwen3vl-conditioner`](https://huggingface.co/spaces/multimodalart/qwen3vl-conditioner), which this
Space calls over the gradio API for every request;
* this Space holds the 61.73 GiB transformer (with H3-World merged into it) plus the video and audio autoencoders.
The wire format between them is `prompt_embeds` `(1, N, 5120)` bf16 + `text_token_tags` `(N,)` int64 in a single
safetensors file. The caption spans this Space needs are recovered from it by offset arithmetic, because
`MiniMaxH3TextEncoderStep` appends the prompt **verbatim** โ no chat template, no special tokens โ so
`offset = num_text_tokens - num_prompt_tokens`.
`h3_split_blocks.py` is the blockset with the `text_encoder` step removed, copied from that Space.
## LoRA merge
The LoRA is published against the **original** MiniMax-H3 layout, not the diffusers port, so `load_lora_weights()`
does not apply. `load_and_apply_lora()` replays `convert_minimax_h3_to_diffusers.py`'s renames on the way in โ
`blocks.` โ `transformer_blocks.`, `attn.out_proj` โ `attn.to_out.0`, `mlp.fc1` โ `ff.net.0.proj` with the SwiGLU
gate/value halves swapped, and the fused `attn.qkv_proj` de-interleaved per head before being split into
`to_q` / `to_k` / `to_v`. 104 LoRA pairs become 208 merged weight deltas; **any target that fails to resolve is
fatal**, never skipped.
## 28 steps vs the 8-step turbo LoRA
**Sampling** in the UI picks between two configurations of the same request โ same seed, same action
script, same directed mask โ so the quality cost of the distillation is directly visible:
| mode | steps | transformer |
|---|---|---|
| `28 steps ยท no turbo LoRA` | 28 (MiniMax-H3's default) | H3-World only |
| `8 steps ยท turbo LoRA` | 8 | H3-World **+** [`lightx2v/Minimax-h3-Turbo`](https://huggingface.co/lightx2v/Minimax-h3-Turbo) `minimax_h3_fl2v_turbo_8step_v1.0_bf16.safetensors` |
`h3_turbo_lora.py` mirrors the `h3_lora.py` of the Spaces that already run this adapter โ
[`MiniMaxAI/MiniMax-H3-Turbo-Lora`](https://huggingface.co/spaces/MiniMaxAI/MiniMax-H3-Turbo-Lora)
(its `lightx` set) and
[`hugging-apps/minimax-h3-turbo-sla-demo`](https://huggingface.co/spaces/hugging-apps/minimax-h3-turbo-sla-demo):
* the file is a **PEFT checkpoint against the diffusers module tree itself**
(`transformer_blocks.N.attn.to_q.lora_A.default.weight`), so unlike H3-World it needs no key
conversion โ 312 targets (50 transformer blocks + 2 token-refiner blocks x
`to_q`/`to_k`/`to_v`/`to_out.0`/`ff.net.0.proj`/`ff.net.2`) map name-for-name;
* rank 128 with `alpha: 8` in the file's own safetensors metadata, so the fold scale is
`alpha / rank = 0.0625` โ what `set_adapters(weights=1.0)` applies in lightx2v's reference script;
* the step count is overridden to the distillation's own **8 NFE**. No scheduler swap and no CFG
change: MiniMax-H3 is already guidance-distilled and every Space above keeps its native
`MiniMaxH3Scheduler` with the turbo LoRA folded.
It is *folded* into the bf16 weights, like H3-World, because this Space patches
`MiniMaxH3AttnProcessor` and drives the transformer's live weights. Fold and unfold are the same
operation with a sign, so the low-rank factors stay resident and `set_active` flips the mode in
place inside the `@spaces.GPU` call, through one bf16 rounding.
The Steps slider still overrides the mode's count (4โ50), so `50 steps + turbo LoRA` or
`8 steps without it` are both reachable for the sake of the comparison.
## Generation constraints
Fixed by the checkpoint: 24 fps, `num_frames` snapped to `17n + 5`, no CFG and no negative prompt (it is
guidance-distilled). H3-World was trained at **832x480**; the conditioner's canvas list does not offer that exact
size, so the default here is its nearest neighbour, **960x544**.
The offered canvases are the cheap tier of each aspect ratio rather than the conditioner's full list. The mask term
scales as *sequence x caption rows*, so a 1344x768 / 8 s request would want ~35 GPU-minutes โ past what any visitor
could book โ and it is off-distribution for a LoRA trained at 832x480 anyway.
## Measured
On this Space, driven over `gradio_client`, at 960x544 with a keyframe:
| Request | Conditioner | Denoise + decode | Round trip |
|---|---|---|---|
| 16 steps, 56 frames, directed | 2 s | 45 s | 50 s |
| 16 steps, 56 frames, **no mask** | 2 s | 32 s | 36 s |
| 50 steps, 124 frames, directed (the default) | 10 s | 303 s | 316 s |
Startup is 95 s: the 66.3 GB download, the load, the LoRA merge, and the ZeroGPU pack. `get_duration` is fitted to
exactly these three points โ the unmasked block cost linear + quadratic in the packed sequence, the mask's own term
linear in *sequence x captions* โ and books ~15% over the fit. The default request books 348 s.
## Examples
The three bundled first frames are extracted from
[`acvlab/ABot-World-Explorer-500h`](https://huggingface.co/datasets/acvlab/ABot-World-Explorer-500h) (Apache-2.0),
which is the same kind of third-person game footage H3-World was trained on. Each is paired with that clip's own
manifest prompt.
## Space variables
| Variable | Default | Meaning |
|---|---|---|
| `H3_MODEL_REPO` | `MiniMaxAI/MiniMax-H3` | The diffusers-layout base checkpoint. |
| `H3_LORA_REPO` | `DANNY621/H3-World` | The LoRA. |
| `H3_LORA_FILE` | `step-10000.safetensors` | The checkpoint the author's own test runs used. |
| `H3_TURBO_REPO` | `lightx2v/Minimax-h3-Turbo` | The turbo-LoRA repo. |
| `H3_TURBO_FILE` | `minimax_h3_fl2v_turbo_8step_v1.0_bf16.safetensors` | The 8-step distillation. |
| `H3_TURBO_STEPS` | `8` | Steps the turbo mode asks for (the card also offers 4). |
| `H3_TURBO_ALPHA` | `0` | `0` reads `alpha` out of the file's metadata (`8`). |
| `H3_TURBO_STRENGTH` | `1.0` | Extra multiplier on the turbo delta. |
| `H3_CONDITIONER` | `multimodalart/qwen3vl-conditioner` | The Space this one asks for embeddings. |
| `H3_ATTENTION` | `_native_cudnn` | cuDNN's fused kernel. flash-attention 3 is sm90-only; this pool is sm120. |
| `H3_GPU_SIZE` | `xlarge` | ZeroGPU allocation size. `large` does not fit. |
## License
The LoRA is Apache-2.0, but usage is governed by the **base model's** license
([`MiniMaxAI/MiniMax-H3`](https://huggingface.co/MiniMaxAI/MiniMax-H3)).
|