Spaces:
Running on Zero
title: MiniMax-H3 Third-Person View LoRA
emoji: 🎮
colorFrom: yellow
colorTo: pink
sdk: gradio
sdk_version: 6.28.0
app_file: app.py
python_version: '3.12'
startup_duration_timeout: 1h
pinned: false
short_description: Design sheets to game cutscenes with HUD and audio
suggested_hardware: zero-a10g
models:
- MiniMaxAI/MiniMax-H3
- multimodalart/MiniMax-H3-Pruned
- WarmBloodAban/Minimax_H3_LoRAs
MiniMax-H3 · Third-Person View LoRA
A demo of WarmBloodAban/Minimax_H3_LoRAs
(Minimax-h3_Third_person_view.safetensors) on MiniMax-H3,
MiniMax's 33B omni-modal model that generates video with natively synchronized audio.
The LoRA turns H3 into a game-cutscene renderer: over-the-shoulder and FPV camera language, Unreal-Engine-looking lighting, and — the part it is actually distinctive for — legible HUD furniture drawn in-frame: target-lock reticles, floating damage numbers, QTE prompts, translucent status bars. The soundtrack (impacts, foley, score) is base H3's, generated in the same denoising pass.
It is a reference LoRA, not a text-to-video one
The adapter's own metadata says ss_base_model_version: minimax_h3_ref2va, so it was trained on H3's
ref2va partition — the one that conditions on an ordered list of reference media, not the plain
text-to-video partition. So this Space is built around reference sheets: you upload a character sheet,
an environment plate, a boss design, and H3 keeps those identities consistent through the shot.
Running this adapter on the t2v partition would attach cleanly and quietly produce worse output; it
isn't offered.
What the demo owes the model card
| The card says | Here |
|---|---|
| LoRA weight 0.6 – 0.85, 0.7 recommended | A slider clamped to exactly that window, defaulting to 0.7. The adapter is attached with PEFT and re-scaled per request, so the window is live |
A long structured prompt format — subject_definitions: / summary: / retention_analysis: / detailed_description: with [Shot N] beats / overall_soundscape: / non_diegetic_music:, cross-referencing <Subject N>, <Environment N>, <Picture N> |
The UI is that format. Each reference slot you fill becomes a <Subject N>/<Environment N> bound to its <Picture N>; your one action line plus the camera preset become the detailed_description: beats; the summary: is prefixed [reference generation] as the card shows. The assembled document is shown after every run, and an override box takes a hand-written one |
| Trigger vocabulary in three groups — camera views (third-person perspective, over-the-shoulder camera, first-person POV, FPV HUD), game rendering (rendered in Unreal Engine, gameplay sequence, combat stance, macro close-up), UI elements (transparent HUD elements, target-lock UI reticle, QTE UI prompt, floating damage text UI) | The four camera presets each spend the camera-view + game-rendering tags their shot actually wants; the HUD checkbox spends the UI group. Nothing is sprinkled in decoratively |
| Base is the INT8 Reference architecture | The ref2va DiT, in bf16 on ZeroGPU's 96 GB tier — no quantization needed at this size |
Architecture
MiniMax-H3 is 195.9 GiB in bf16 and a ZeroGPU Space is evicted at 150 GB of storage, so the pipeline is
cut at its text_encoder step — the split every MiniMax-H3 Space uses. The 62 GiB Qwen3-VL conditioner
runs in multimodalart/qwen3vl-conditioner
and is called over the gradio API with each caller's own ZeroGPU token forwarded, so the booking is
billed to whoever asked for the video; prompt_embeds + text_token_tags is the whole wire format.
This Space holds the DiT and the two autoencoders. h3_split_blocks.py holds the denoising-half block
stack.
The DiT is multimodalart/MiniMax-H3-Pruned,
which folds the AdaLN input projections onto a rank-8 subspace: transformer_ref drops from 61.7 GiB
to 37.5 GiB with identical output. This LoRA touches no AdaLN module, so nothing needs projecting.
How the LoRA is applied
Attached as a live PEFT adapter (load_lora_adapter(..., adapter_name="third_person_view")) and
re-scaled per request via set_adapters, rather than folded into the base weights — folding would be
faster to serve but would freeze the strength, and 0.6–0.85 is the knob this LoRA is actually steered
with. For the same reason the pipeline is not AoTI-compiled: a captured graph would route around
the PEFT-injected layers and silently drop the adapter.
The export is keyed on the original H3 partition (diffusion_model.blocks.N.…, 416 tensors, rank 64,
ai-toolkit), so applying it to the diffusers port means replaying the layout transforms
convert_minimax_h3_to_diffusers.py
applied to the base weights — a delta is only valid in the layout of the weight it is added to:
| Original LoRA target | diffusers parameter | Transform |
|---|---|---|
blocks.N.… / token_refiner.blocks.N.… |
transformer_blocks.N.… / token_refiner.refiner_blocks.N.… |
rename |
attn.out_proj |
attn.to_out.0 |
rename |
attn.qkv_proj |
attn.to_q / to_k / to_v |
split lora_B's rows into contiguous thirds |
mlp.fc2 |
ff.net.2 |
rename |
mlp.fc1 |
ff.net.0.proj |
swap the two fused SwiGLU halves ([gate; value] → [value; gate]) |
Both of the last two rows are easy to get wrong. The gate/value swap is a real difference between H3's
native SwiGLU packing and diffusers' GEGLU-style ordering — skip it and every MLP delta lands on the
wrong half of the activation. Conversely the contiguous QKV split is only correct for ai-toolkit
exports like this one; fal-trained H3 LoRAs store per-head-interleaved QKV rows and need de-interleaving
first. The two are distinguished by key naming, and this repo is unambiguously the former. Any target
that resolves to no transformer weight is a fatal error rather than a silent skip.
Examples
Three ready-to-run briefs. The wyvern-boss and ruined-village reference sheets are frames lifted from
the LoRA author's own showcase clip in
WarmBloodAban/Minimax_H3_LoRAs (Apache-2.0) —
the closest thing to the inputs the adapter was demonstrated on. The character and city plates come from
linoyts/repo-to-space-example-inputs
(CC0-1.0).
Known limits
A default run is 5.17 s of 24 fps video at 960×544 and takes about four minutes end to end (16 s of conditioning, ~190 s of denoise + decode); the larger canvases and longer durations in the advanced accordion cost proportionally more, and the estimate above the button says how much before you spend it. HUD text is shaped like HUD text and usually isn't real words — that's a base-H3 limit the LoRA doesn't fix. Faces drift on fast whip-pans. The first request after a cold start also pays for weight placement.
License
MiniMax-H3 is covered by the MiniMax H3 Community License Agreement — note its attribution, territorial-scope and revenue clauses. The LoRA is Apache-2.0. Read both before using any output of this Space.