File size: 10,969 Bytes
6749e6b
2c82f76
 
 
 
6749e6b
 
 
2c82f76
 
 
 
 
 
5a00791
6749e6b
 
2c82f76
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
68b5806
2c82f76
 
5a00791
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
2c82f76
 
 
 
 
 
68b5806
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
2c82f76
 
 
 
 
 
 
 
 
 
 
 
 
 
5a00791
 
 
 
 
2c82f76
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
---
title: H3-World
emoji: ๐ŸŽฎ
colorFrom: indigo
colorTo: green
sdk: gradio
sdk_version: 6.26.0
app_file: app.py
python_version: "3.12"
startup_duration_timeout: 1h
short_description: Drive a world model with WASD and camera keys
models:
  - MiniMaxAI/MiniMax-H3
  - DANNY621/H3-World
  - lightx2v/Minimax-h3-Turbo
---

# H3-World โ€” action-conditioned world model

[`DANNY621/H3-World`](https://huggingface.co/DANNY621/H3-World) is a rank-32 LoRA for
[`MiniMaxAI/MiniMax-H3`](https://huggingface.co/MiniMaxAI/MiniMax-H3) that turns it into a **playable** world model:
you give it a first frame and a **key sequence**, and it renders what happens as those keys are held.

```
forward*20, forward-right*10, pan-right-fast*7
```

| keys | meaning |
|---|---|
| `W` `A` `S` `D` | walk forward / strafe left / walk backward / strafe right |
| `J` `L` | camera pans left / right |
| `K` `I` | camera tilts up / down |
| `F` | modifier โ€” the camera move is *sharp* rather than *slow* |

## How the actions actually reach the model

The checkpoint's own run manifests are unambiguous about this, and it is the thing that makes the LoRA
demo-able at all:

```json
"checkpoint": {"action_dim": 0, "action_mode": "text", "action_tensors": 0,
               "lora_tensors": 208, "lora_pairs": 104}
"config": {"num_frames": 124, "height": 480, "width": 832, "fps": 24, "steps": 50,
           "conditioning": "first_frame+8d_actions",
           "action_columns": ["W", "A", "S", "D", "I", "J", "K", "L"]}
```

`action_dim: 0` and `action_tensors: 0` โ€” there is **no action encoder and no action embedding**. The 8-dimensional
key state is carried through the **text channel**: one short English sentence per *latent* video frame, appended to the
scene prompt, in the register the author's own runs use.

So a 124-frame request (37 latent frames) is conditioned on a prompt that looks like:

```
A third-person view of a man walking through a city intersection...
the man walks forward
the man walks forward
...
the man walks forward and strafes right
...
the man stands still, camera pans right sharply
```

This Space builds that text from the key script for you โ€” `caption_for()` in `app.py` โ€” and shows you the resolved
per-frame captions under **Conditioning** after every run.

## The directed attention mask

The model card is explicit that the LoRA weights alone do not reproduce the reported behavior: the training run used
a **directed attention mask** that binds each per-frame caption to the latents of *its* frame. Without it every video
row attends to all 37 sentences at once and the sequence collapses into an average action.

MiniMax-H3 is a single packed 1-D sequence under full self-attention โ€” `[text | keyframe anchors | audio | video]`,
no cross-attention โ€” so the mask is a constraint inside one attention call, not a separate cross-attention mask.

`app.py` reimplements it as an **exact log-sum-exp merge** rather than a dense `[S, S]` mask (which would force SDPA
off its flash kernel for the whole 21k-row sequence). The keys are split into three regions:

| region | keys | kernel |
|---|---|---|
| **A** | text rows *before* the caption block | flash, unmasked |
| **C** | the ~700 caption rows | fp32 masked matmul, chunked over queries |
| **B** | everything after the text block (~99% of keys) | flash, unmasked |

Each returns its output and its log-sum-exp; the three are recombined with the online-softmax identity, which is
numerically identical to one masked softmax over the full row. Only **video** queries are restricted (frame *i*'s rows
see only sentence *i*); the captions themselves still see everything โ€” that is the "directed" part.

The mask is toggleable in **Advanced** so you can see the difference. With it off, the same script produces a video
that drifts through a blur of every action at once.

Two guards, both in `app.py`:

* The token spans are located by re-tokenizing prefixes of the prompt. If a BPE merge straddles a sentence boundary
  the spans would be wrong, so `build_conditioning_text()` verifies `cuts[-1] == total` and **refuses to mask** rather
  than mask the wrong rows.
* `MiniMaxH3TokenRefinerBlock` runs the same attention module over the *text stream alone*. The processor detects that
  (the sequence is too short to contain the video block) and falls straight through to the stock path.

## Architecture โ€” why two Spaces

MiniMax-H3 is 195.9 GiB in bfloat16 and a ZeroGPU Space is evicted at 150 GB of storage, so the pipeline is split at
its `text_encoder` step, exactly as in [`multimodalart/minimax-h3`](https://huggingface.co/spaces/multimodalart/minimax-h3):

* the 62.14 GiB Qwen3-VL conditioner runs in
  [`multimodalart/qwen3vl-conditioner`](https://huggingface.co/spaces/multimodalart/qwen3vl-conditioner), which this
  Space calls over the gradio API for every request;
* this Space holds the 61.73 GiB transformer (with H3-World merged into it) plus the video and audio autoencoders.

The wire format between them is `prompt_embeds` `(1, N, 5120)` bf16 + `text_token_tags` `(N,)` int64 in a single
safetensors file. The caption spans this Space needs are recovered from it by offset arithmetic, because
`MiniMaxH3TextEncoderStep` appends the prompt **verbatim** โ€” no chat template, no special tokens โ€” so
`offset = num_text_tokens - num_prompt_tokens`.

`h3_split_blocks.py` is the blockset with the `text_encoder` step removed, copied from that Space.

## LoRA merge

The LoRA is published against the **original** MiniMax-H3 layout, not the diffusers port, so `load_lora_weights()`
does not apply. `load_and_apply_lora()` replays `convert_minimax_h3_to_diffusers.py`'s renames on the way in โ€”
`blocks.` โ†’ `transformer_blocks.`, `attn.out_proj` โ†’ `attn.to_out.0`, `mlp.fc1` โ†’ `ff.net.0.proj` with the SwiGLU
gate/value halves swapped, and the fused `attn.qkv_proj` de-interleaved per head before being split into
`to_q` / `to_k` / `to_v`. 104 LoRA pairs become 208 merged weight deltas; **any target that fails to resolve is
fatal**, never skipped.

## 28 steps vs the 8-step turbo LoRA

**Sampling** in the UI picks between two configurations of the same request โ€” same seed, same action
script, same directed mask โ€” so the quality cost of the distillation is directly visible:

| mode | steps | transformer |
|---|---|---|
| `28 steps ยท no turbo LoRA` | 28 (MiniMax-H3's default) | H3-World only |
| `8 steps ยท turbo LoRA` | 8 | H3-World **+** [`lightx2v/Minimax-h3-Turbo`](https://huggingface.co/lightx2v/Minimax-h3-Turbo) `minimax_h3_fl2v_turbo_8step_v1.0_bf16.safetensors` |

`h3_turbo_lora.py` mirrors the `h3_lora.py` of the Spaces that already run this adapter โ€”
[`MiniMaxAI/MiniMax-H3-Turbo-Lora`](https://huggingface.co/spaces/MiniMaxAI/MiniMax-H3-Turbo-Lora)
(its `lightx` set) and
[`hugging-apps/minimax-h3-turbo-sla-demo`](https://huggingface.co/spaces/hugging-apps/minimax-h3-turbo-sla-demo):

* the file is a **PEFT checkpoint against the diffusers module tree itself**
  (`transformer_blocks.N.attn.to_q.lora_A.default.weight`), so unlike H3-World it needs no key
  conversion โ€” 312 targets (50 transformer blocks + 2 token-refiner blocks x
  `to_q`/`to_k`/`to_v`/`to_out.0`/`ff.net.0.proj`/`ff.net.2`) map name-for-name;
* rank 128 with `alpha: 8` in the file's own safetensors metadata, so the fold scale is
  `alpha / rank = 0.0625` โ€” what `set_adapters(weights=1.0)` applies in lightx2v's reference script;
* the step count is overridden to the distillation's own **8 NFE**. No scheduler swap and no CFG
  change: MiniMax-H3 is already guidance-distilled and every Space above keeps its native
  `MiniMaxH3Scheduler` with the turbo LoRA folded.

It is *folded* into the bf16 weights, like H3-World, because this Space patches
`MiniMaxH3AttnProcessor` and drives the transformer's live weights. Fold and unfold are the same
operation with a sign, so the low-rank factors stay resident and `set_active` flips the mode in
place inside the `@spaces.GPU` call, through one bf16 rounding.

The Steps slider still overrides the mode's count (4โ€“50), so `50 steps + turbo LoRA` or
`8 steps without it` are both reachable for the sake of the comparison.

## Generation constraints

Fixed by the checkpoint: 24 fps, `num_frames` snapped to `17n + 5`, no CFG and no negative prompt (it is
guidance-distilled). H3-World was trained at **832x480**; the conditioner's canvas list does not offer that exact
size, so the default here is its nearest neighbour, **960x544**.

The offered canvases are the cheap tier of each aspect ratio rather than the conditioner's full list. The mask term
scales as *sequence x caption rows*, so a 1344x768 / 8 s request would want ~35 GPU-minutes โ€” past what any visitor
could book โ€” and it is off-distribution for a LoRA trained at 832x480 anyway.

## Measured

On this Space, driven over `gradio_client`, at 960x544 with a keyframe:

| Request | Conditioner | Denoise + decode | Round trip |
|---|---|---|---|
| 16 steps, 56 frames, directed | 2 s | 45 s | 50 s |
| 16 steps, 56 frames, **no mask** | 2 s | 32 s | 36 s |
| 50 steps, 124 frames, directed (the default) | 10 s | 303 s | 316 s |

Startup is 95 s: the 66.3 GB download, the load, the LoRA merge, and the ZeroGPU pack. `get_duration` is fitted to
exactly these three points โ€” the unmasked block cost linear + quadratic in the packed sequence, the mask's own term
linear in *sequence x captions* โ€” and books ~15% over the fit. The default request books 348 s.

## Examples

The three bundled first frames are extracted from
[`acvlab/ABot-World-Explorer-500h`](https://huggingface.co/datasets/acvlab/ABot-World-Explorer-500h) (Apache-2.0),
which is the same kind of third-person game footage H3-World was trained on. Each is paired with that clip's own
manifest prompt.

## Space variables

| Variable | Default | Meaning |
|---|---|---|
| `H3_MODEL_REPO` | `MiniMaxAI/MiniMax-H3` | The diffusers-layout base checkpoint. |
| `H3_LORA_REPO` | `DANNY621/H3-World` | The LoRA. |
| `H3_LORA_FILE` | `step-10000.safetensors` | The checkpoint the author's own test runs used. |
| `H3_TURBO_REPO` | `lightx2v/Minimax-h3-Turbo` | The turbo-LoRA repo. |
| `H3_TURBO_FILE` | `minimax_h3_fl2v_turbo_8step_v1.0_bf16.safetensors` | The 8-step distillation. |
| `H3_TURBO_STEPS` | `8` | Steps the turbo mode asks for (the card also offers 4). |
| `H3_TURBO_ALPHA` | `0` | `0` reads `alpha` out of the file's metadata (`8`). |
| `H3_TURBO_STRENGTH` | `1.0` | Extra multiplier on the turbo delta. |
| `H3_CONDITIONER` | `multimodalart/qwen3vl-conditioner` | The Space this one asks for embeddings. |
| `H3_ATTENTION` | `_native_cudnn` | cuDNN's fused kernel. flash-attention 3 is sm90-only; this pool is sm120. |
| `H3_GPU_SIZE` | `xlarge` | ZeroGPU allocation size. `large` does not fit. |

## License

The LoRA is Apache-2.0, but usage is governed by the **base model's** license
([`MiniMaxAI/MiniMax-H3`](https://huggingface.co/MiniMaxAI/MiniMax-H3)).