--- title: MiniMax-H3 ยท reference ยท Custom lora + CivitAI links, GPU cost, profiles, clip stitching emoji: ๐ŸŽญ colorFrom: pink colorTo: purple sdk: gradio sdk_version: 6.20.0 app_file: app.py pinned: true short_description: Video + soundtrack, your lora + CivitAI, cost, profiles suggested_hardware: zero-a10g tags: - video - image-to-video - minimax - lora - civitai - audio - zerogpu --- # MiniMax-H3 โ€” omni-references, unquantized, split across two Spaces Joint video **and** soundtrack out of a single denoising pass, conditioned on an ordered list of image, video and audio references, at **bfloat16 with no quantization anywhere**. This Space is the denoising half of the `ref2va` task: the 61.73 GiB `transformer_ref` partition and the two autoencoders. The 62.14 GiB Qwen3-VL conditioner runs in [`qwen3vl-conditioner`](https://huggingface.co/spaces/multimodalart/qwen3vl-conditioner), which this Space calls over the gradio API for every request โ€” the same conditioner Space, and the same resident weights, that the keyframe half [`minimax-h3`](https://huggingface.co/spaces/multimodalart/minimax-h3) uses. ## What this fork adds **Custom lora, three slots.** A Hugging Face repo, a file inside one, a **CivitAI download link**, or an uploaded file. Adapters have to be trained against the `transformer_ref/` partition โ€” a `transformer/` adapter is a different partition and will not match. **kohya / CivitAI files are converted on the fly.** Most of what CivitAI carries is kohya-named, and it differs from diffusers in ways that quietly ruin a result rather than raise: flat underscored names, a fused QKV interleaved *per attention head* (a plain three-way split hands q's rows to k), a gated MLP whose two halves sit in the opposite order, and an `alpha` that sets the scale. All four are handled at load time, and `.pt` files are read as well as `.safetensors`. **Turbo presets.** Few-step distillations that render joint video + soundtrack in 4โ€“8 steps instead of the usual ~28. Picking one fills a free slot at strength `1.0` and moves the steps slider to the count it was distilled for. **Clip stitching, soundtrack included.** Queue several clips and they are joined into a single file; tick the box and every new generation joins on its own. It runs on the CPU, so it costs nothing from the GPU allowance. Everything is scaled to the first clip's frame, and a clip without audio gets silence rather than breaking the join. **Named profiles.** Save every setting under a name and bring it back from a dropdown, or export it as `.json`. **A live GPU cost readout** above the button, using the same `budget()` the pre-flight check and `get_duration` use, so a request that will not fit is visible as a number rather than as a failed run. **Randomize seed**, drawn fresh per press and written back into the box, so the number shown is always the one the clip was made with. Set `CIVITAI_TOKEN` under *Settings โ†’ Variables and secrets* for gated CivitAI models: the Space downloads on its own machine, so being signed in to CivitAI in a browser does not authenticate it โ€” a gated model answers a server with an HTML login page, which is caught and reported rather than saved as weights. ## Why split MiniMax-H3 is 195.9 GiB in bfloat16 and a ZeroGPU Space is evicted at **150 GB of storage**. An unquantized single Space is therefore impossible. Cut the `ref2va` branch of `MiniMaxH3Blocks` at its `text_encoder` step and both halves fit unquantized: | Space | Subfolders | Download | Resident | |---|---|---|---| | [`qwen3vl-conditioner`](https://huggingface.co/spaces/multimodalart/qwen3vl-conditioner) | `text_encoder/` + `tokenizer/` + `processor/` | 66.7 GB | 62.15 GiB bf16 | | this one | `transformer_ref/` + `vae/` + `audio_vae/` | 77.3 GB | 61.73 GiB bf16 + 10.43 GiB float32 | ## References A request carries up to **12** references โ€” at most 9 images, 3 videos and 3 audio clips โ€” **in the order the model reads them**. The order is semantic: it numbers the labels of MiniMax-H3's prompt presentation (``, `