--- title: MiniMax-H3 Studio Β· Simple & Efficient emoji: 🎬 colorFrom: blue colorTo: indigo sdk: gradio sdk_version: 6.24.0 app_file: app.py pinned: true short_description: 'Simple video, smart scenes, identity & efficient Turbo' suggested_hardware: zero-a10g tags: - video - image-to-video - minimax - lora - civitai - audio - zerogpu --- # MiniMax-H3 Studio **One image. One idea. Video with sound β€” with fewer controls in your way.** In **Simple**, upload a picture, describe the shot, choose its length and press **Generate**. **Balanced** is selected for you. Optional tools stay in closed sections; **Everything** reveals the manual settings. For several actions, open the scene planner and review the proposed clips and timings. **πŸ†• Simple quality presets Β· πŸ“ Auto picture shape Β· 🎬 AI Scene Planner Β· πŸ”’ Identity Lock Β· πŸ”„ Looping LoRA Refresh Β· πŸ”Š Soundtrack mixer** ## Start here 1. Upload your picture and describe the action, sounds and any dialogue. 2. Keep **Balanced** for six real denoising steps. Choose **Draft** for four or **Quality** for eight. The default length is three seconds; changing quality keeps your chosen length and custom effects. 3. Press **Generate**. The Auto canvas follows the picture's proportions using a supported size. 4. For several actions, open **AI Scene Planner**, describe the scene, split it, review it and generate. One Turbo adapter is already active. Prompt writing, scene planning and prompt rewriting are optional; ordinary Generate does not automatically call the writing model. More references, Identity Lock, LoRA tools, profiles and finishing controls remain available in closed sections. The supplied `h3_quick_test.json` is a small, one-clip setup. Import it under Profiles and upload your own image. Importing a profile never starts inference. ## AI Scene Planner and visible progress - **1–64 clips**, with automatic count and a separate duration for each action. H3 uses 24 FPS; durations are aligned to its `17 Γ— n + 5` frame grid. For example, a requested 3 seconds becomes 73 frames, about 3.04 seconds. The UI supports requests from 2 to 14 seconds per clip. - **Two clips per writer request**, to keep scene JSON short. Long plans need several requests. Malformed, truncated or failed responses leave the existing plan intact; failed jobs are not automatically replayed. Planning does not start video generation. - The planner preserves action order and carries the ending pose into the next stage. Descriptions are written in English; explicitly requested dialogue should remain in its original language. - The progress panel appears **above the video** and names the current clip before processing starts. It counts completed clips; it does not invent a percentage for unfinished denoising. - Each clip uses a separate browser event. The next is scheduled only when work remains, replacing the old fixed chain of eight callbacks. Planned prompts and timings are held constant during a run. - **Stop** blocks the next clip. A GPU or remote job already running can take time to return and can still consume quota. Completed clips stay queued. **Start at clip** resumes from the preceding completed video's last frame while session files remain available. - A fresh scene starts a fresh run queue. **Just one more clip** continues the current video. The default writer is `amisima/Qwen3.8-27B-Uncensored-Demo`. It receives the scene idea and reference image. Its own admission limits can still reject a request; this app cannot lower another Space's GPU reservation. ## Identity Lock **CPU face protection** uses the conservative alignment and blending helpers from the reviewed WAN/LTX builds. It corrects only a suitable face region on the starting image, with restrained blending, lighting matching and protection for eye and mouth detail. It skips unsuitable head angles or missing faces. The original portrait stays fixed across a scene and on resume. This mode **does not add a second H3 reference image**, avoiding that reference's extra conditioning and encoding work. A missing portrait on a fresh scene uses the original starting image. Strength 0 or **Off** disables correction. It is a starting-frame aid, not an identity guarantee for every frame. **Extra H3 reference** retains the previous method: the portrait is an additional model reference. It consumes one of the nine image slots and adds GPU work. The strength slider controls CPU blending; for the extra-reference method, zero disables it and nonzero enables it, without a model-side weight. OpenCV is included in the updated requirements. Optional YuNet/SFace files download on the CPU with bounded requests; no image is sent to a face-recognition API. Set `YUNET_ONNX` and `SFACE_ONNX` to local files for offline setup. A detector failure keeps the frame and reports a skipped correction. ## LoRA Studio β€” relevant choices, looping Refresh and trigger sync Prompt preparation and scene planning share their choices. **Refresh** rotates through up to five relevant suggestions, wraps to the beginning and preserves up to two selected effects. General NSFW all-in-one entries remain pinned when compatible metadata and size limits allow them. They are not activated automatically merely because they are pinned. Refresh and adapter matching use local name/trigger matching and metadata checks, **not an AI or GPU request**. Initial prompt preparation can choose entries with strong local matches; ambiguous ones remain suggestions. The library's saved strengths are used. Swapping two selected effects replaces the oldest when a new one is ticked. Turbo and unrelated manual slots are preserved. Known activation words are synchronized to the main prompt and every nonempty clip prompt, including changes through custom slots, presets and the shared library. Only tracked insertions are removed; your original wording is preserved. Missing trigger metadata is not proof that an adapter is broken. **Preparation finishes before the video GPU request:** - Default limit: **1536 MiB per source/converted file**, **2048 MiB across selected adapters**. Turbo counts too; a zero-strength adapter does not download or load. These are practical size guards, not guarantees of memory fit or model compatibility. - Oversized files are hidden from automatic suggestions. Manual selections pass the same checks. - CivitAI/direct links and Hugging Face files are validated, including cached files. Incomplete new downloads are not published as reusable cache entries. Explicit HF file URL revisions are kept. - Compatible ComfyUI, kohya and LoKr adapter data is converted **on CPU**, cast to bf16 and cached before GPU allocation. The existing H3 conversion logic is retained; LoKr rank remains configurable. - Confirmed incompatible metadata or attributable invalid-adapter failures are excluded temporarily. Size limits and temporary access errors do not permanently blacklist files. Use `.safetensors` adapters for this `transformer_ref` pipeline. Arbitrary `.pt`/`.bin` checkpoints are not accepted by the new bounded safetensors preparation path. WAN/LTX weights are not H3 adapters. ## Lower compute without complicating Simple | Quality | Real denoising evaluations | Intended use | | --- | ---: | --- | | Draft | 4 | Short tests and simple motion | | Balanced β€” default | 6 | Everyday starting point | | Quality | 8 | More demanding motion; costs more | The default adapter remains Larry v4 step600 EMA. Its [model card](https://huggingface.co/larryvrh/MiniMax-H3-Turbo-Lora) recommends six to eight steps when fast motion smears at four. Quality is a setting, not a guarantee for every scene. **Correct step accounting:** the pinned scheduler counts its terminal zero as a schedule point. The app now passes `N + 1` points for `N` real evaluations, matching the [Turbo inference implementation](https://github.com/ModelTC/Minimax-H3-Turbo/blob/main/inference_minimax_h3.py). Thus an older β€œ8 steps” actually performed seven evaluations; the new Balanced performs six. Keeping an old numeric step value now performs one more evaluation. Exported profiles use version 3. **Smaller reference processing β€” requires the companion conditioner update.** With the new endpoint, both the Qwen encoder and H3 reference encoder use the same source-preserving policy. A small image is no longer enlarged to a 2048-pixel short edge; large images are limited to the output canvas area and aligned to the model's 32-pixel grid. For a 512Γ—768 picture and a 512Γ—768 canvas, reference image rows drop from 6,144 to 384. That is **16Γ— fewer image-reference rows**, not 16Γ— faster video generation. Target video, text, audio, model placement and decoding still cost compute. Smaller references can lose fine detail; Everything retains larger output canvases for comparison. The policy is inspired by the Turbo project's [reference-resizing guidance](https://github.com/ModelTC/Minimax-H3-Turbo#note-on-reference-image-resizing). This package keeps the Ref2VA model and native separate video/audio schedulers. It does not load FL2VA weights or apply a second audio shift. The generator discovers the companion's versioned endpoint from public API metadata, validates the returned plan and tensor metadata, and uses matching dimensions in both halves. With the old public conditioner it retains the old resize policy and canvas table. **Uploading the generator alone does not enable the full reference reduction.** A failed efficient request is not silently retried. Other savings: CPU LoRA download/conversion/cache before GPU allocation, one speed adapter when switching presets, and a session-scoped conditioning cache. Identical conditioning inputs can reuse it when only seed or steps change. Prompt, references, resolved dimensions, frame count, rewriting, session, endpoint or resize policy changes invalidate it. The cache holds eight entries for 30 minutes, at most 256 MiB each. A changed quality preset may select a different Auto canvas and require encoding. An optional **LightX Ref2VA 4-step** adapter appears in Everything for comparison. It is about 1.29 GiB; its CPU download and the total LoRA limit still apply. It is not activated by default and is not stacked with another known speed preset. Native sparse attention, a second LTX pass, upscaling and interpolation are not added to default generation. **INT8 is an owner opt-in experiment**, described below. The default stays BF16 on xlarge. No GPU quality, memory or speed benchmark has been run for this update, and no fixed percentage saving is promised. Lower step counts do not translate proportionally to total request time. The estimate includes the video GPU allocation multiplier. Conditioner and optional writer requests are additional. A reservation estimate is not a bill; actual usage is governed by [Hugging Face ZeroGPU](https://huggingface.co/docs/hub/spaces-zerogpu). ## Audio, motion and finished-video tools H3 generates its own video and soundtrack jointly. Keep the existing image, audio and video reference inputs: up to nine images, one audio file and one motion/camera video in this interface. A reference video must be 2–15 seconds and may carry sound. Audio alone is insufficient; include an image or video. The structured builder exposes shot type, camera motion, soundscape, music and verbatim dialogue. Select Bulgarian or English, or enter another language tag. A language tag is a request, not a guarantee of pronunciation quality. **Quick tags** remain in Everything mode. - **Soundtrack mixer:** mix or replace generated audio with an uploaded track; adjust added volume. - **Loop:** repeat a finished clip or joined scene up to eight times, including its audio. - **Final Cut:** queue clips and join once at the end of a scene, or press **Join queued clips / retry**. Normalization runs one clip at a time, and the final join copies the prepared streams. Silence for a clip without audio has a finite duration. All subprocesses have timeouts. A failed join preserves original clips, and can be retried without generating them again. These finished-video operations run on CPU. Native H3 audio is used instead of loading WAN's MMAudio or HunyuanVideo-Foley models. LTX-specific attention anchors, keyframe/ControlNet nodes and WAN High/Low adapters are not transplanted into a different model pipeline. H3's own reference controls remain available. ## Profiles and diagnostics JSON export/import and named profiles include the scene idea, per-clip prompts and times, count, automatic planning options, identity method/strength, dialogue language, LoRAs and generation settings. Older profiles load with an empty scene plan and reset the resume position. Uploaded files are not embedded; re-upload references after a session or container is replaced. Named profiles and the adapter library are shared with visitors; downloaded JSON can be kept privately. **Runtime diagnostics** reports a boot ID, processing stages, RAM and available cgroup OOM counters. It includes no prompt text or photos. GPU worker PIDs may change without a web-app reboot. Diagnostic history and caches are temporary in this build; download a report before a restart if possible. ## Owner setup and token policy Upload the **contents of `generator/`** to the root of the existing MiniMax-H3 Space. In addition to `app.py`, `requirements.txt` and `README.md`, the new `h3_efficiency.py` is required; include `h3_quantization.py` and the optional dependency file too. Keep `h3_split_blocks.py`, `h3_aoti.py`, `lora_library.py` and the existing runtime files. The generator folder is an update, not a full repository. To enable efficient reference processing: 1. Duplicate the public `multimodalart/qwen3vl-conditioner` Space under your account and select ZeroGPU. 2. Upload the **contents of `conditioner/`** to that Space's root. Its README selects Gradio 6.24.0. 3. Wait for it to be ready, then set the generator's `H3_CONDITIONER` variable to the new public `your-account/your-conditioner-space` ID. It must be publicly callable without an owner token. 4. Restart the generator after the conditioner is ready. Capability discovery is cached until restart. Both copies of `h3_efficiency.py` must be identical. The conditioner's old endpoints remain available. The companion loads the original Qwen conditioner; it is not a paid MiniMax API. It still uses ZeroGPU quota, and hosting eligibility/limits are set by Hugging Face. If you cannot host the companion, the updated generator can keep using the public legacy conditioner with reduced reference savings. The dependency file is based on the matching live Space and preserves its torch/diffusers pins; it adds OpenCV and filelock. Gradio is selected by `sdk_version` above, not by a second requirements pin. Remote compute clients explicitly use `token=False`. **HF_TOKEN, PLANNER_HF_TOKEN, EXPANDER_HF_TOKEN and H3_CONDITIONER_HF_TOKEN are not forwarded to writer/conditioner requests.** The visitor's Gradio/ZeroGPU request context remains intact. There is no new mandatory login gate. Use a public writer and conditioner; private endpoints needing an owner token will not work with this policy. This prevents these code paths from silently using owner secrets; it is not a promise of unlimited free compute or proof of the origin of any previous account charges. | Variable | Default | Purpose | | --- | --- | --- | | `H3_CONDITIONER` | `multimodalart/qwen3vl-conditioner` | Public remote input encoder | | `EXPANDER_SPACE` | `amisima/Qwen3.8-27B-Uncensored-Demo` | Optional scene/prompt writer | | `H3_MODEL_REPO` | `MiniMaxAI/MiniMax-H3` | Existing H3 checkpoint | | `H3_GPU_SIZE` | `xlarge` | Existing large model allocation | | `H3_GPU_DURATION_MIN` / `H3_GPU_DURATION_MAX` | `120` / `1500` | Existing reservation bounds | | `H3_PLACEMENT_ALLOWANCE` | `90` | Existing cold-worker allowance; not silently lowered | | `H3_PLACEMENT` | `lazy` for BF16; `offload` for INT8 | Placement; offload reserves 10 GB for working memory | | `H3_QUANTIZATION` | `bf16` | `int8` is an optional, unbenchmarked owner experiment | | `H3_ATTENTION` | `_native_cudnn` | Base attention; adapter runs use `native` | | `H3_AOTI` | `0` | Existing compilation option; cannot be combined with LoRAs | | `H3_LOKR_RANK` | `32` | CPU LoKr conversion target rank; an approximation when truncated | | `H3_LORA_MAX_FILE_MB` / `H3_LORA_MAX_RUN_MB` | `1536` / `2048` | LoRA limits in MiB | | `H3_LORA_DOWNLOAD_SECONDS` | `600` | Bounded streamed direct download; HF native retries remain separate | | `HF_TOKEN` | unset | Model downloads and existing library/storage access only | | `CIVITAI_TOKEN` | unset | Model downloads and CivitAI metadata/search | | `LORA_LIBRARY_BUCKET` / `LORA_LIBRARY_DIR` | existing configuration | Preserve the shared library location | | `CIVITAI_API_HOST` | unset | Existing search host override | ## Optional INT8 / large trial β€” owner only The helper follows the official [Diffusers MiniMax-H3 INT8 recipe](https://huggingface.co/docs/diffusers/main/en/api/pipelines/minimax_h3): INT8 weight-only transformer, protected input/output/timing modules, and the existing float32 VAEs. Quantization runs at startup, outside the request's GPU job. It still needs host RAM and disk for the original checkpoint and conversion. No pre-quantized checkpoint is downloaded automatically. To opt in, append `-r requirements-int8.txt` to the generator's `requirements.txt` and set: ```text H3_QUANTIZATION=int8 H3_PLACEMENT=offload H3_GPU_SIZE=large H3_AOTI=0 ``` `requirements-int8.txt` pins torchao 0.17.0 for the existing PyTorch 2.11 build. Test this configuration on your own small clip before offering it publicly: TorchAO/PEFT adapters, ZeroGPU workers, total memory and performance have not been validated on a real GPU here. The 48 GB large allocation has a lower quota multiplier, but slower kernels/offloading could offset that benefit. The conditioner still uses xlarge; this switch affects only the generator. BF16 plus large is rejected at startup. Rollback: set `H3_QUANTIZATION=bf16`, `H3_GPU_SIZE=xlarge`, `H3_PLACEMENT=lazy` and restart. There is no automatic GPU fallback, extra retry or silent switch to a larger allocation. ## Validation and attribution 39 offline tests passed. Checks cover both Gradio interfaces and their API schemas, exact scheduler point counts, shared reference resizing, metadata agreement, quality presets, Simple visibility, and the previous Gradio construction and request injection, planner JSON and timing, LoRA rotation and trigger preservation, token isolation, conditioner cache separation, early rejection before compute, no GPU retry, adapter cleanup, Stop/Resume and real ffmpeg joins, mixing and looping. No production GPU generation or paid inference was run. Model quality, actual GPU runtime and the live browser/ZeroGPU path still need validation in the Space. Based on the supplied MiniMax-H3 application and the matching [MiniMax Space](https://huggingface.co/spaces/amisima/minimax-h3-reference-4-step-lora), derived from [multimodalart's implementation](https://huggingface.co/spaces/multimodalart/minimax-h3). The existing pinned diffusers integration and MiniMax-H3 model license remain unchanged.