--- title: MiniMax-H3 T2V Demo emoji: ๐ŸŽฌ colorFrom: purple colorTo: indigo sdk: gradio sdk_version: 6.20.0 app_file: app.py pinned: false short_description: Text-to-video with synchronized audio suggested_hardware: zero-a10g models: - MiniMaxAI/MiniMax-H3 --- # MiniMax-H3 ยท Text-to-Video demo Text-to-video (with **synchronized stereo audio**) for [MiniMaxAI/MiniMax-H3](https://huggingface.co/MiniMaxAI/MiniMax-H3), scoped to the `t2va` text-to-video path. Ported from [@mrfakename/minimax-h3-ultra-fast](https://huggingface.co/spaces/mrfakename/minimax-h3-ultra-fast) โ€” all credit for the engine and the optimization stack goes to that Space's author. This port replaces its React studio with an ordinary Gradio Blocks UI and removes everything outside text-to-video: keyframe/FL2VA inputs, Ref2VA omni-references, the storyboard editor and stitcher, custom-LoRA URL downloads, and the VDN federation. The inference engine is byte-identical. ## Engine (inherited from the source Space) | layer | what runs here | |---|---| | Weights | 12.5 GB pruned NVFP4 transformer ([`lilcheaty/MiniMax-H3-NVFP4`](https://huggingface.co/lilcheaty/MiniMax-H3-NVFP4)): 20.1B effective parameters instead of 33.1B/61.7 GiB BF16 | | Compute | Native CUDA 13 NVFP4 tensor-core GEMMs through `comfy-kitchen`; higher-precision norms, embeddings and output heads | | Conditioner | Local 15.7 GB Qwen3-VL NVFP4-AWQ checkpoint ([Comfy-Org/MiniMax-H3](https://huggingface.co/Comfy-Org/MiniMax-H3)) holding only the 50 language layers H3 uses | | Autoencoders | Both VAEs stay full precision (a bf16 audio VAE decodes the soundtrack ~20 dB too quiet) | | Speed | Cache-DiT F1B0 + TaylorSeer block reuse, NVIDIA Sol-Attn sparse routing above 24,576 packed tokens, Turbo LoRA presets (4-step / 8-step distillations) | | Fallback | `H3_ENGINE=bf16` runs the unquantized 33B transformer (with optional AoTI-compiled blocks) instead of the NVFP4 engine | ## Usage 1. Type a prompt (describe motion, camera and what should be heard โ€” H3 generates the soundtrack too). 2. Pick a preset โ€” **Balanced** (28 steps with conservative caching) is the recommended default; Turbo presets trade fidelity for a 3โ€“7ร— shorter ZeroGPU booking. 3. Choose a canvas (aspect ratio) and a duration between 2 and 14 s (snapped to the VAE-decodable `17ยทn+5` frames). 4. Generate. The run report shows canvas, denoiser evaluations, conditioning time and seed. Advanced controls (steps, cache engine, Turbo LoRA) appear when you pick the *Custom* preset. ## Notes - Runs on ZeroGPU. Generation bills your own ZeroGPU quota per request; the transformer and conditioner are loaded at startup, outside GPU time. - Every request env var (`H3_*`) from the source Space is honored โ€” see the header of `app.py`. - Use of MiniMax-H3 is subject to the [MiniMax H3 Community License](https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/LICENSE).