Spaces:
Running on Zero
Running on Zero
| title: MiniMax-H3 Studio · Simple & Efficient | |
| emoji: 🎬 | |
| colorFrom: blue | |
| colorTo: indigo | |
| sdk: gradio | |
| sdk_version: 6.24.0 | |
| app_file: app.py | |
| pinned: true | |
| short_description: 'Simple video, smart scenes, identity & efficient Turbo' | |
| suggested_hardware: zero-a10g | |
| tags: | |
| - not-for-all-audiences | |
| - video | |
| - image-to-video | |
| - minimax | |
| - lora | |
| - civitai | |
| - audio | |
| - zerogpu | |
| # MiniMax-H3 Studio | |
| **One image. One idea. Video with sound — with fewer controls in your way.** | |
| In **Simple**, upload a picture, describe the shot, choose its length and press **Generate**. | |
| **Balanced** is selected for you. Optional tools stay in closed sections; **Everything** reveals the | |
| manual settings. For several actions, open the scene planner and review the proposed clips and timings. | |
| **🆕 Simple quality presets · 📐 Auto picture shape · 🎬 AI Scene Planner · 🔒 Identity Lock · 🔄 Looping LoRA Refresh · 🔊 Soundtrack mixer** | |
| ## Start here | |
| 1. Upload your picture and describe the action, sounds and any dialogue. | |
| 2. Keep **Balanced** for six real denoising steps. Choose **Draft** for four or **Quality** for eight. | |
| The default length is three seconds; changing quality keeps your chosen length and custom effects. | |
| 3. Press **Generate**. The Auto canvas follows the picture's proportions using a supported size. | |
| 4. For several actions, open **AI Scene Planner**, describe the scene, split it, review it and generate. | |
| One Turbo adapter is already active. Prompt writing, scene planning and prompt rewriting are optional; | |
| ordinary Generate does not automatically call the writing model. More references, Identity Lock, | |
| LoRA tools, profiles and finishing controls remain available in closed sections. | |
| The supplied `h3_quick_test.json` is a small, one-clip setup. Import it under Profiles and upload your | |
| own image. Importing a profile never starts inference. | |
| ## AI Scene Planner and visible progress | |
| - **1–64 clips**, with automatic count and a separate duration for each action. H3 uses 24 FPS; | |
| durations are aligned to its `17 × n + 5` frame grid. For example, a requested 3 seconds becomes | |
| 73 frames, about 3.04 seconds. The UI supports requests from 2 to 14 seconds per clip. | |
| - **Two clips per writer request**, to keep scene JSON short. Long plans need several requests. | |
| Malformed, truncated or failed responses leave the existing plan intact; failed jobs are not | |
| automatically replayed. Planning does not start video generation. | |
| - The planner preserves action order and carries the ending pose into the next stage. Descriptions | |
| are written in English; explicitly requested dialogue should remain in its original language. | |
| - The progress panel appears **above the video** and names the current clip before processing starts. | |
| It counts completed clips; it does not invent a percentage for unfinished denoising. | |
| - Each clip uses a separate browser event. The next is scheduled only when work remains, replacing | |
| the old fixed chain of eight callbacks. Planned prompts and timings are held constant during a run. | |
| - **Stop** blocks the next clip. A GPU or remote job already running can take time to return and can | |
| still consume quota. Completed clips stay queued. **Start at clip** resumes from the preceding | |
| completed video's last frame while session files remain available. | |
| - A fresh scene starts a fresh run queue. **Just one more clip** continues the current video. | |
| The default writer is `amisima/Qwen3.8-27B-Uncensored-Demo`. It receives the scene idea and reference | |
| image. Its own admission limits can still reject a request; this app cannot lower another Space's | |
| GPU reservation. | |
| ## Identity Lock | |
| **CPU face protection** uses the conservative alignment and blending helpers from the reviewed WAN/LTX | |
| builds. It corrects only a suitable face region on the starting image, with restrained blending, | |
| lighting matching and protection for eye and mouth detail. It skips unsuitable head angles or missing | |
| faces. The original portrait stays fixed across a scene and on resume. | |
| This mode **does not add a second H3 reference image**, avoiding that reference's extra conditioning | |
| and encoding work. A missing portrait on a fresh scene uses the original starting image. Strength 0 | |
| or **Off** disables correction. It is a starting-frame aid, not an identity guarantee for every frame. | |
| **Extra H3 reference** retains the previous method: the portrait is an additional model reference. | |
| It consumes one of the nine image slots and adds GPU work. The strength slider controls CPU blending; | |
| for the extra-reference method, zero disables it and nonzero enables it, without a model-side weight. | |
| OpenCV is included in the updated requirements. Optional YuNet/SFace files download on the CPU with | |
| bounded requests; no image is sent to a face-recognition API. Set `YUNET_ONNX` and `SFACE_ONNX` to local | |
| files for offline setup. A detector failure keeps the frame and reports a skipped correction. | |
| ## LoRA Studio — relevant choices, looping Refresh and trigger sync | |
| Prompt preparation and scene planning share their choices. **Refresh** rotates through up to five | |
| relevant suggestions, wraps to the beginning and preserves up to two selected effects. General NSFW | |
| all-in-one entries remain pinned when compatible metadata and size limits allow them. They are not | |
| activated automatically merely because they are pinned. | |
| Refresh and adapter matching use local name/trigger matching and metadata checks, **not an AI or GPU | |
| request**. Initial prompt preparation can choose entries with strong local matches; ambiguous ones | |
| remain suggestions. The library's saved strengths are used. Swapping two selected effects replaces | |
| the oldest when a new one is ticked. Turbo and unrelated manual slots are preserved. | |
| Known activation words are synchronized to the main prompt and every nonempty clip prompt, including | |
| changes through custom slots, presets and the shared library. Only tracked insertions are removed; | |
| your original wording is preserved. Missing trigger metadata is not proof that an adapter is broken. | |
| **Preparation finishes before the video GPU request:** | |
| - Default limit: **1536 MiB per source/converted file**, **2048 MiB across selected adapters**. Turbo | |
| counts too; a zero-strength adapter does not download or load. These are practical size guards, | |
| not guarantees of memory fit or model compatibility. | |
| - Oversized files are hidden from automatic suggestions. Manual selections pass the same checks. | |
| - CivitAI/direct links and Hugging Face files are validated, including cached files. Incomplete new | |
| downloads are not published as reusable cache entries. Explicit HF file URL revisions are kept. | |
| - Compatible ComfyUI, kohya and LoKr adapter data is converted **on CPU**, cast to bf16 and cached | |
| before GPU allocation. The existing H3 conversion logic is retained; LoKr rank remains configurable. | |
| - Confirmed incompatible metadata or attributable invalid-adapter failures are excluded temporarily. | |
| Size limits and temporary access errors do not permanently blacklist files. | |
| Use `.safetensors` adapters for this `transformer_ref` pipeline. Arbitrary `.pt`/`.bin` checkpoints are | |
| not accepted by the new bounded safetensors preparation path. WAN/LTX weights are not H3 adapters. | |
| ## Lower compute without complicating Simple | |
| | Quality | Real denoising evaluations | Intended use | | |
| | --- | ---: | --- | | |
| | Draft | 4 | Short tests and simple motion | | |
| | Balanced — default | 6 | Everyday starting point | | |
| | Quality | 8 | More demanding motion; costs more | | |
| The default adapter remains Larry v4 step600 EMA. Its | |
| [model card](https://huggingface.co/larryvrh/MiniMax-H3-Turbo-Lora) recommends six to eight steps when | |
| fast motion smears at four. Quality is a setting, not a guarantee for every scene. | |
| **Correct step accounting:** the pinned scheduler counts its terminal zero as a schedule point. The | |
| app now passes `N + 1` points for `N` real evaluations, matching the | |
| [Turbo inference implementation](https://github.com/ModelTC/Minimax-H3-Turbo/blob/main/inference_minimax_h3.py). | |
| Thus an older “8 steps” actually performed seven evaluations; the new Balanced performs six. Keeping | |
| an old numeric step value now performs one more evaluation. Exported profiles use version 3. | |
| **Smaller reference processing — requires the companion conditioner update.** With the new endpoint, | |
| both the Qwen encoder and H3 reference encoder use the same source-preserving policy. A small image | |
| is no longer enlarged to a 2048-pixel short edge; large images are limited to the output canvas area | |
| and aligned to the model's 32-pixel grid. For a 512×768 picture and a 512×768 canvas, reference image | |
| rows drop from 6,144 to 384. That is **16× fewer image-reference rows**, not 16× faster video generation. | |
| Target video, text, audio, model placement and decoding still cost compute. Smaller references can | |
| lose fine detail; Everything retains larger output canvases for comparison. | |
| The policy is inspired by the Turbo project's | |
| [reference-resizing guidance](https://github.com/ModelTC/Minimax-H3-Turbo#note-on-reference-image-resizing). | |
| This package keeps the Ref2VA model and native separate video/audio schedulers. It does not load | |
| FL2VA weights or apply a second audio shift. | |
| The generator discovers the companion's versioned endpoint from public API metadata, validates the | |
| returned plan and tensor metadata, and uses matching dimensions in both halves. With the old public | |
| conditioner it retains the old resize policy and canvas table. **Uploading the generator alone does | |
| not enable the full reference reduction.** A failed efficient request is not silently retried. | |
| Other savings: CPU LoRA download/conversion/cache before GPU allocation, one speed adapter when | |
| switching presets, and a session-scoped conditioning cache. Identical conditioning inputs can reuse | |
| it when only seed or steps change. Prompt, references, resolved dimensions, frame count, rewriting, | |
| session, endpoint or resize policy changes invalidate it. The cache holds eight entries for 30 minutes, | |
| at most 256 MiB each. A changed quality preset may select a different Auto canvas and require encoding. | |
| An optional **LightX Ref2VA 4-step** adapter appears in Everything for comparison. It is about 1.29 GiB; | |
| its CPU download and the total LoRA limit still apply. It is not activated by default and is not stacked | |
| with another known speed preset. Native sparse attention, a second LTX pass, upscaling and interpolation | |
| are not added to default generation. | |
| **INT8 is an owner opt-in experiment**, described below. The default stays BF16 on xlarge. No GPU | |
| quality, memory or speed benchmark has been run for this update, and no fixed percentage saving is | |
| promised. Lower step counts do not translate proportionally to total request time. | |
| The estimate includes the video GPU allocation multiplier. Conditioner and optional writer requests | |
| are additional. A reservation estimate is not a bill; actual usage is governed by | |
| [Hugging Face ZeroGPU](https://huggingface.co/docs/hub/spaces-zerogpu). | |
| ## Audio, motion and finished-video tools | |
| H3 generates its own video and soundtrack jointly. Keep the existing image, audio and video reference | |
| inputs: up to nine images, one audio file and one motion/camera video in this interface. A reference | |
| video must be 2–15 seconds and may carry sound. Audio alone is insufficient; include an image or video. | |
| The structured builder exposes shot type, camera motion, soundscape, music and verbatim dialogue. | |
| Select Bulgarian or English, or enter another language tag. A language tag is a request, not a guarantee | |
| of pronunciation quality. **Quick tags** remain in Everything mode. | |
| - **Soundtrack mixer:** mix or replace generated audio with an uploaded track; adjust added volume. | |
| - **Loop:** repeat a finished clip or joined scene up to eight times, including its audio. | |
| - **Final Cut:** queue clips and join once at the end of a scene, or press **Join queued clips / retry**. | |
| Normalization runs one clip at a time, and the final join copies the prepared streams. Silence | |
| for a clip without audio has a finite duration. All subprocesses have timeouts. A failed join | |
| preserves original clips, and can be retried without generating them again. | |
| These finished-video operations run on CPU. Native H3 audio is used instead of loading WAN's MMAudio | |
| or HunyuanVideo-Foley models. LTX-specific attention anchors, keyframe/ControlNet nodes and WAN High/Low | |
| adapters are not transplanted into a different model pipeline. H3's own reference controls remain available. | |
| ## Profiles and diagnostics | |
| JSON export/import and named profiles include the scene idea, per-clip prompts and times, count, | |
| automatic planning options, identity method/strength, dialogue language, LoRAs and generation settings. | |
| Older profiles load with an empty scene plan and reset the resume position. Uploaded files are not | |
| embedded; re-upload references after a session or container is replaced. Named profiles and the adapter | |
| library are shared with visitors; downloaded JSON can be kept privately. | |
| **Runtime diagnostics** reports a boot ID, processing stages, RAM and available cgroup OOM counters. | |
| It includes no prompt text or photos. GPU worker PIDs may change without a web-app reboot. Diagnostic | |
| history and caches are temporary in this build; download a report before a restart if possible. | |
| ## Owner setup and token policy | |
| Upload the **contents of `generator/`** to the root of the existing MiniMax-H3 Space. In addition to | |
| `app.py`, `requirements.txt` and `README.md`, the new `h3_efficiency.py` is required; include | |
| `h3_quantization.py` and the optional dependency file too. Keep `h3_split_blocks.py`, `h3_aoti.py`, | |
| `lora_library.py` and the existing runtime files. The generator folder is an update, not a full repository. | |
| To enable efficient reference processing: | |
| 1. Duplicate the public `multimodalart/qwen3vl-conditioner` Space under your account and select ZeroGPU. | |
| 2. Upload the **contents of `conditioner/`** to that Space's root. Its README selects Gradio 6.24.0. | |
| 3. Wait for it to be ready, then set the generator's `H3_CONDITIONER` variable to the new public | |
| `your-account/your-conditioner-space` ID. It must be publicly callable without an owner token. | |
| 4. Restart the generator after the conditioner is ready. Capability discovery is cached until restart. | |
| Both copies of `h3_efficiency.py` must be identical. The conditioner's old endpoints remain available. | |
| The companion loads the original Qwen conditioner; it is not a paid MiniMax API. It still uses ZeroGPU | |
| quota, and hosting eligibility/limits are set by Hugging Face. If you cannot host the companion, the | |
| updated generator can keep using the public legacy conditioner with reduced reference savings. | |
| The dependency file is based on the matching live Space and preserves its torch/diffusers pins; it | |
| adds OpenCV and filelock. Gradio is selected by `sdk_version` above, not by a second requirements pin. | |
| Remote compute clients explicitly use `token=False`. **HF_TOKEN, PLANNER_HF_TOKEN, | |
| EXPANDER_HF_TOKEN and H3_CONDITIONER_HF_TOKEN are not forwarded to writer/conditioner requests.** | |
| The visitor's Gradio/ZeroGPU request context remains intact. There is no new mandatory login gate. | |
| Use a public writer and conditioner; private endpoints needing an owner token will not work with | |
| this policy. This prevents these code paths from silently using owner secrets; it is not a promise | |
| of unlimited free compute or proof of the origin of any previous account charges. | |
| | Variable | Default | Purpose | | |
| | --- | --- | --- | | |
| | `H3_CONDITIONER` | `multimodalart/qwen3vl-conditioner` | Public remote input encoder | | |
| | `EXPANDER_SPACE` | `amisima/Qwen3.8-27B-Uncensored-Demo` | Optional scene/prompt writer | | |
| | `H3_MODEL_REPO` | `MiniMaxAI/MiniMax-H3` | Existing H3 checkpoint | | |
| | `H3_GPU_SIZE` | `xlarge` | Existing large model allocation | | |
| | `H3_GPU_DURATION_MIN` / `H3_GPU_DURATION_MAX` | `120` / `1500` | Existing reservation bounds | | |
| | `H3_PLACEMENT_ALLOWANCE` | `90` | Existing cold-worker allowance; not silently lowered | | |
| | `H3_PLACEMENT` | `lazy` for BF16; `offload` for INT8 | Placement; offload reserves 10 GB for working memory | | |
| | `H3_QUANTIZATION` | `bf16` | `int8` is an optional, unbenchmarked owner experiment | | |
| | `H3_ATTENTION` | `_native_cudnn` | Base attention; adapter runs use `native` | | |
| | `H3_AOTI` | `0` | Existing compilation option; cannot be combined with LoRAs | | |
| | `H3_LOKR_RANK` | `32` | CPU LoKr conversion target rank; an approximation when truncated | | |
| | `H3_LORA_MAX_FILE_MB` / `H3_LORA_MAX_RUN_MB` | `1536` / `2048` | LoRA limits in MiB | | |
| | `H3_LORA_DOWNLOAD_SECONDS` | `600` | Bounded streamed direct download; HF native retries remain separate | | |
| | `HF_TOKEN` | unset | Model downloads and existing library/storage access only | | |
| | `CIVITAI_TOKEN` | unset | Model downloads and CivitAI metadata/search | | |
| | `LORA_LIBRARY_BUCKET` / `LORA_LIBRARY_DIR` | existing configuration | Preserve the shared library location | | |
| | `CIVITAI_API_HOST` | unset | Existing search host override | | |
| ## Optional INT8 / large trial — owner only | |
| The helper follows the official | |
| [Diffusers MiniMax-H3 INT8 recipe](https://huggingface.co/docs/diffusers/main/en/api/pipelines/minimax_h3): | |
| INT8 weight-only transformer, protected input/output/timing modules, and the existing float32 VAEs. | |
| Quantization runs at startup, outside the request's GPU job. It still needs host RAM and disk for the | |
| original checkpoint and conversion. No pre-quantized checkpoint is downloaded automatically. | |
| To opt in, append `-r requirements-int8.txt` to the generator's `requirements.txt` and set: | |
| ```text | |
| H3_QUANTIZATION=int8 | |
| H3_PLACEMENT=offload | |
| H3_GPU_SIZE=large | |
| H3_AOTI=0 | |
| ``` | |
| `requirements-int8.txt` pins torchao 0.17.0 for the existing PyTorch 2.11 build. Test this configuration | |
| on your own small clip before offering it publicly: TorchAO/PEFT adapters, ZeroGPU workers, total | |
| memory and performance have not been validated on a real GPU here. The 48 GB large allocation has a | |
| lower quota multiplier, but slower kernels/offloading could offset that benefit. The conditioner | |
| still uses xlarge; this switch affects only the generator. BF16 plus large is rejected at startup. | |
| Rollback: set `H3_QUANTIZATION=bf16`, `H3_GPU_SIZE=xlarge`, `H3_PLACEMENT=lazy` and restart. There is no | |
| automatic GPU fallback, extra retry or silent switch to a larger allocation. | |
| ## Validation and attribution | |
| 39 offline tests passed. Checks cover both Gradio interfaces and their API schemas, exact scheduler | |
| point counts, shared reference resizing, metadata agreement, quality presets, Simple visibility, | |
| and the previous Gradio construction and request injection, planner JSON and timing, LoRA rotation | |
| and trigger preservation, token isolation, conditioner cache separation, early rejection before | |
| compute, no GPU retry, adapter cleanup, Stop/Resume and real ffmpeg joins, mixing and looping. | |
| No production GPU generation or paid inference was run. Model quality, actual GPU runtime and the | |
| live browser/ZeroGPU path still need validation in the Space. | |
| Based on the supplied MiniMax-H3 application and the matching | |
| [MiniMax Space](https://huggingface.co/spaces/amisima/minimax-h3-reference-4-step-lora), derived from | |
| [multimodalart's implementation](https://huggingface.co/spaces/multimodalart/minimax-h3). | |
| The existing pinned diffusers integration and MiniMax-H3 model license remain unchanged. | |