Spaces:
Running on Zero
Download README.md from hugging-apps/h3-world-action-demo: direct link, hf CLI and curl.
- Browser
- Download file 11 kB
-
https://huggingface.co/spaces/hugging-apps/h3-world-action-demo/resolve/d53b4b0673c31df2a995c05d5b7bf445dffbefa2/README.md
- Command line
-
hf download hf://spaces/hugging-apps/h3-world-action-demo@d53b4b0673c31df2a995c05d5b7bf445dffbefa2/README.md
-
curl -L -o README.md https://huggingface.co/spaces/hugging-apps/h3-world-action-demo/resolve/d53b4b0673c31df2a995c05d5b7bf445dffbefa2/README.md
title: H3-World
emoji: ๐ฎ
colorFrom: indigo
colorTo: green
sdk: gradio
sdk_version: 6.26.0
app_file: app.py
python_version: '3.12'
startup_duration_timeout: 1h
short_description: Drive a world model with WASD and camera keys
models:
- MiniMaxAI/MiniMax-H3
- DANNY621/H3-World
- lightx2v/Minimax-h3-Turbo
H3-World โ action-conditioned world model
DANNY621/H3-World is a rank-32 LoRA for
MiniMaxAI/MiniMax-H3 that turns it into a playable world model:
you give it a first frame and a key sequence, and it renders what happens as those keys are held.
forward*20, forward-right*10, pan-right-fast*7
| keys | meaning |
|---|---|
W A S D |
walk forward / strafe left / walk backward / strafe right |
J L |
camera pans left / right |
K I |
camera tilts up / down |
F |
modifier โ the camera move is sharp rather than slow |
How the actions actually reach the model
The checkpoint's own run manifests are unambiguous about this, and it is the thing that makes the LoRA demo-able at all:
"checkpoint": {"action_dim": 0, "action_mode": "text", "action_tensors": 0,
"lora_tensors": 208, "lora_pairs": 104}
"config": {"num_frames": 124, "height": 480, "width": 832, "fps": 24, "steps": 50,
"conditioning": "first_frame+8d_actions",
"action_columns": ["W", "A", "S", "D", "I", "J", "K", "L"]}
action_dim: 0 and action_tensors: 0 โ there is no action encoder and no action embedding. The 8-dimensional
key state is carried through the text channel: one short English sentence per latent video frame, appended to the
scene prompt, in the register the author's own runs use.
So a 124-frame request (37 latent frames) is conditioned on a prompt that looks like:
A third-person view of a man walking through a city intersection...
the man walks forward
the man walks forward
...
the man walks forward and strafes right
...
the man stands still, camera pans right sharply
This Space builds that text from the key script for you โ caption_for() in app.py โ and shows you the resolved
per-frame captions under Conditioning after every run.
The directed attention mask
The model card is explicit that the LoRA weights alone do not reproduce the reported behavior: the training run used a directed attention mask that binds each per-frame caption to the latents of its frame. Without it every video row attends to all 37 sentences at once and the sequence collapses into an average action.
MiniMax-H3 is a single packed 1-D sequence under full self-attention โ [text | keyframe anchors | audio | video],
no cross-attention โ so the mask is a constraint inside one attention call, not a separate cross-attention mask.
app.py reimplements it as an exact log-sum-exp merge rather than a dense [S, S] mask (which would force SDPA
off its flash kernel for the whole 21k-row sequence). The keys are split into three regions:
| region | keys | kernel |
|---|---|---|
| A | text rows before the caption block | flash, unmasked |
| C | the ~700 caption rows | fp32 masked matmul, chunked over queries |
| B | everything after the text block (~99% of keys) | flash, unmasked |
Each returns its output and its log-sum-exp; the three are recombined with the online-softmax identity, which is numerically identical to one masked softmax over the full row. Only video queries are restricted (frame i's rows see only sentence i); the captions themselves still see everything โ that is the "directed" part.
The mask is toggleable in Advanced so you can see the difference. With it off, the same script produces a video that drifts through a blur of every action at once.
Two guards, both in app.py:
- The token spans are located by re-tokenizing prefixes of the prompt. If a BPE merge straddles a sentence boundary
the spans would be wrong, so
build_conditioning_text()verifiescuts[-1] == totaland refuses to mask rather than mask the wrong rows. MiniMaxH3TokenRefinerBlockruns the same attention module over the text stream alone. The processor detects that (the sequence is too short to contain the video block) and falls straight through to the stock path.
Architecture โ why two Spaces
MiniMax-H3 is 195.9 GiB in bfloat16 and a ZeroGPU Space is evicted at 150 GB of storage, so the pipeline is split at
its text_encoder step, exactly as in multimodalart/minimax-h3:
- the 62.14 GiB Qwen3-VL conditioner runs in
multimodalart/qwen3vl-conditioner, which this Space calls over the gradio API for every request; - this Space holds the 61.73 GiB transformer (with H3-World merged into it) plus the video and audio autoencoders.
The wire format between them is prompt_embeds (1, N, 5120) bf16 + text_token_tags (N,) int64 in a single
safetensors file. The caption spans this Space needs are recovered from it by offset arithmetic, because
MiniMaxH3TextEncoderStep appends the prompt verbatim โ no chat template, no special tokens โ so
offset = num_text_tokens - num_prompt_tokens.
h3_split_blocks.py is the blockset with the text_encoder step removed, copied from that Space.
LoRA merge
The LoRA is published against the original MiniMax-H3 layout, not the diffusers port, so load_lora_weights()
does not apply. load_and_apply_lora() replays convert_minimax_h3_to_diffusers.py's renames on the way in โ
blocks. โ transformer_blocks., attn.out_proj โ attn.to_out.0, mlp.fc1 โ ff.net.0.proj with the SwiGLU
gate/value halves swapped, and the fused attn.qkv_proj de-interleaved per head before being split into
to_q / to_k / to_v. 104 LoRA pairs become 208 merged weight deltas; any target that fails to resolve is
fatal, never skipped.
28 steps vs the 8-step turbo LoRA
Sampling in the UI picks between two configurations of the same request โ same seed, same action script, same directed mask โ so the quality cost of the distillation is directly visible:
| mode | steps | transformer |
|---|---|---|
28 steps ยท no turbo LoRA |
28 (MiniMax-H3's default) | H3-World only |
8 steps ยท turbo LoRA |
8 | H3-World + lightx2v/Minimax-h3-Turbo minimax_h3_fl2v_turbo_8step_v1.0_bf16.safetensors |
h3_turbo_lora.py mirrors the h3_lora.py of the Spaces that already run this adapter โ
MiniMaxAI/MiniMax-H3-Turbo-Lora
(its lightx set) and
hugging-apps/minimax-h3-turbo-sla-demo:
- the file is a PEFT checkpoint against the diffusers module tree itself
(
transformer_blocks.N.attn.to_q.lora_A.default.weight), so unlike H3-World it needs no key conversion โ 312 targets (50 transformer blocks + 2 token-refiner blocks xto_q/to_k/to_v/to_out.0/ff.net.0.proj/ff.net.2) map name-for-name; - rank 128 with
alpha: 8in the file's own safetensors metadata, so the fold scale isalpha / rank = 0.0625โ whatset_adapters(weights=1.0)applies in lightx2v's reference script; - the step count is overridden to the distillation's own 8 NFE. No scheduler swap and no CFG
change: MiniMax-H3 is already guidance-distilled and every Space above keeps its native
MiniMaxH3Schedulerwith the turbo LoRA folded.
It is folded into the bf16 weights, like H3-World, because this Space patches
MiniMaxH3AttnProcessor and drives the transformer's live weights. Fold and unfold are the same
operation with a sign, so the low-rank factors stay resident and set_active flips the mode in
place inside the @spaces.GPU call, through one bf16 rounding.
The Steps slider still overrides the mode's count (4โ50), so 50 steps + turbo LoRA or
8 steps without it are both reachable for the sake of the comparison.
Generation constraints
Fixed by the checkpoint: 24 fps, num_frames snapped to 17n + 5, no CFG and no negative prompt (it is
guidance-distilled). H3-World was trained at 832x480; the conditioner's canvas list does not offer that exact
size, so the default here is its nearest neighbour, 960x544.
The offered canvases are the cheap tier of each aspect ratio rather than the conditioner's full list. The mask term scales as sequence x caption rows, so a 1344x768 / 8 s request would want ~35 GPU-minutes โ past what any visitor could book โ and it is off-distribution for a LoRA trained at 832x480 anyway.
Measured
On this Space, driven over gradio_client, at 960x544 with a keyframe:
| Request | Conditioner | Denoise + decode | Round trip |
|---|---|---|---|
| 16 steps, 56 frames, directed | 2 s | 45 s | 50 s |
| 16 steps, 56 frames, no mask | 2 s | 32 s | 36 s |
| 50 steps, 124 frames, directed (the default) | 10 s | 303 s | 316 s |
Startup is 95 s: the 66.3 GB download, the load, the LoRA merge, and the ZeroGPU pack. get_duration is fitted to
exactly these three points โ the unmasked block cost linear + quadratic in the packed sequence, the mask's own term
linear in sequence x captions โ and books ~15% over the fit. The default request books 348 s.
Examples
The three bundled first frames are extracted from
acvlab/ABot-World-Explorer-500h (Apache-2.0),
which is the same kind of third-person game footage H3-World was trained on. Each is paired with that clip's own
manifest prompt.
Space variables
| Variable | Default | Meaning |
|---|---|---|
H3_MODEL_REPO |
MiniMaxAI/MiniMax-H3 |
The diffusers-layout base checkpoint. |
H3_LORA_REPO |
DANNY621/H3-World |
The LoRA. |
H3_LORA_FILE |
step-10000.safetensors |
The checkpoint the author's own test runs used. |
H3_TURBO_REPO |
lightx2v/Minimax-h3-Turbo |
The turbo-LoRA repo. |
H3_TURBO_FILE |
minimax_h3_fl2v_turbo_8step_v1.0_bf16.safetensors |
The 8-step distillation. |
H3_TURBO_STEPS |
8 |
Steps the turbo mode asks for (the card also offers 4). |
H3_TURBO_ALPHA |
0 |
0 reads alpha out of the file's metadata (8). |
H3_TURBO_STRENGTH |
1.0 |
Extra multiplier on the turbo delta. |
H3_CONDITIONER |
multimodalart/qwen3vl-conditioner |
The Space this one asks for embeddings. |
H3_ATTENTION |
_native_cudnn |
cuDNN's fused kernel. flash-attention 3 is sm90-only; this pool is sm120. |
H3_GPU_SIZE |
xlarge |
ZeroGPU allocation size. large does not fit. |
License
The LoRA is Apache-2.0, but usage is governed by the base model's license
(MiniMaxAI/MiniMax-H3).