minimax-h3-t2v-demo / README.md
mobiusfr's picture
Re-add suggested_hardware hint (ZeroGPU)
2a6f723 verified
|
Raw History Blame Contribute Delete
2.92 kB

A newer version of the Gradio SDK is available: 6.29.0

Upgrade
metadata
title: MiniMax-H3 T2V Demo
emoji: 🎬
colorFrom: purple
colorTo: indigo
sdk: gradio
sdk_version: 6.20.0
app_file: app.py
pinned: false
short_description: Text-to-video with synchronized audio
suggested_hardware: zero-a10g
models:
  - MiniMaxAI/MiniMax-H3

MiniMax-H3 · Text-to-Video demo

Text-to-video (with synchronized stereo audio) for MiniMaxAI/MiniMax-H3, scoped to the t2va text-to-video path.

Ported from @mrfakename/minimax-h3-ultra-fast — all credit for the engine and the optimization stack goes to that Space's author. This port replaces its React studio with an ordinary Gradio Blocks UI and removes everything outside text-to-video: keyframe/FL2VA inputs, Ref2VA omni-references, the storyboard editor and stitcher, custom-LoRA URL downloads, and the VDN federation. The inference engine is byte-identical.

Engine (inherited from the source Space)

layer what runs here
Weights 12.5 GB pruned NVFP4 transformer (lilcheaty/MiniMax-H3-NVFP4): 20.1B effective parameters instead of 33.1B/61.7 GiB BF16
Compute Native CUDA 13 NVFP4 tensor-core GEMMs through comfy-kitchen; higher-precision norms, embeddings and output heads
Conditioner Local 15.7 GB Qwen3-VL NVFP4-AWQ checkpoint (Comfy-Org/MiniMax-H3) holding only the 50 language layers H3 uses
Autoencoders Both VAEs stay full precision (a bf16 audio VAE decodes the soundtrack ~20 dB too quiet)
Speed Cache-DiT F1B0 + TaylorSeer block reuse, NVIDIA Sol-Attn sparse routing above 24,576 packed tokens, Turbo LoRA presets (4-step / 8-step distillations)
Fallback H3_ENGINE=bf16 runs the unquantized 33B transformer (with optional AoTI-compiled blocks) instead of the NVFP4 engine

Usage

  1. Type a prompt (describe motion, camera and what should be heard — H3 generates the soundtrack too).
  2. Pick a preset — Balanced (28 steps with conservative caching) is the recommended default; Turbo presets trade fidelity for a 3–7× shorter ZeroGPU booking.
  3. Choose a canvas (aspect ratio) and a duration between 2 and 14 s (snapped to the VAE-decodable 17·n+5 frames).
  4. Generate. The run report shows canvas, denoiser evaluations, conditioning time and seed.

Advanced controls (steps, cache engine, Turbo LoRA) appear when you pick the Custom preset.

Notes

  • Runs on ZeroGPU. Generation bills your own ZeroGPU quota per request; the transformer and conditioner are loaded at startup, outside GPU time.
  • Every request env var (H3_*) from the source Space is honored — see the header of app.py.
  • Use of MiniMax-H3 is subject to the MiniMax H3 Community License.