amisima's picture
Upload 2 files
179ef37 verified
|
Raw
History Blame
13.7 kB
metadata
title: MiniMax-H3 · reference · One button, photo to video
emoji: 🎭
colorFrom: pink
colorTo: purple
sdk: gradio
sdk_version: 6.20.0
app_file: app.py
pinned: true
short_description: One button — it writes the prompt and picks the loras
suggested_hardware: zero-a10g
tags:
  - not-for-all-audiences
  - video
  - image-to-video
  - minimax
  - lora
  - civitai
  - audio
  - zerogpu

MiniMax-H3 — one button, or every dial

Picture in, a few words, one press. The description is written from your first reference picture, wrapped in the labelled sections H3 was actually trained on, and the shared library is searched for loras that match it. Everything this Space can do is still here — one tick at the top swaps the whole page over to it.

Underneath: joint video and soundtrack out of a single denoising pass, conditioned on an ordered list of image, video and audio references, at bfloat16 with no quantization anywhere.

🟢 Simple — the button does the lot

One press, and:

  • The description is written from your first reference picture. Not from your words alone — the picture is shown to a chat Space of your choosing, so what comes back describes the subject and the setting actually in front of it.
  • The structure goes on around it. The prose is wrapped in the labelled sections below, the part MiniMax call critical to the quality of the final output — so the one thing that most changes the result is no longer something you have to remember to press.
  • The loras are chosen. The shared library is searched, the handful that match your words are shortlisted, and the two that fit best go into the slots at their own strengths.
  • The trigger words go into the prompt. An adapter whose token is missing does nothing at all, and that is the commonest reason one seems to be ignored.

Whatever it picked is listed underneath with the rest of the shortlist. Tick a different one and it swaps — the slots refill, the old trigger words come out and the new ones go in. Two at a time is the limit, because stacking more only means each one shows less.

Leave the Space box empty and the button still builds the structured prompt out of your own words.

🔧 Everything — what this fork adds

This Space is the denoising half of the ref2va task: the 61.73 GiB transformer_ref partition and the two autoencoders. The 62.14 GiB Qwen3-VL conditioner runs in qwen3vl-conditioner, which this Space calls over the gradio API for every request — the same conditioner Space, and the same resident weights, that the keyframe half minimax-h3 uses.

A structured prompt builder. H3 was trained on the output of H3-Context-IR, a preprocessor that rewrites a plain request into labelled sections, and MiniMax's own model card calls that structure critical to the quality of the final output. Nothing in this pipeline adds it — the string reaches the transformer as typed. The builder writes those sections for you: the reference line naming <Picture 1>, <Picture 2> in connection order, then integrated_multimodal_description with the shot type, the camera move written as type + amplitude + speed, and the dialogue verbatim inside <d> tags, then overall_soundscape and non_diegetic_music. Speech is generated together with the picture, so naming that someone speaks without giving the words produces correct mouth shapes with nothing in them; the builder makes that hard to get wrong.

A shared lora library that survives a restart. Every adapter anyone adds is kept in a storage bucket — name, link, trigger words and strength — in two lists, ordinary and nsfw. Tick what you want and press one button: the slots are filled, their strengths are set, and the trigger words are appended to the prompt, so an adapter that needs its token to do anything never silently does nothing. Paste a CivitAI or Hugging Face link and the title and the trigger words are read off it for you. The list is common to everyone who opens the Space and is still there after it sleeps and wakes, so nobody has to hunt down last week's link a second time.

Custom lora, five slots. A Hugging Face repo, a file inside one, a CivitAI download link, or an uploaded file. Adapters have to be trained against the transformer_ref/ partition — a transformer/ adapter is a different partition and will not match.

CivitAI, searched from inside the Space. Type a word, get the model's title, version, file, size, downloads and trigger words, and drop the one you want straight into a slot. A link that is already in a slot can be named with one press, so a row reads Hairy dad butt · v1.0 · hairy.safetensors rather than a number. Gated entries need CIVITAI_TOKEN under Settings → Variables and secrets: the Space downloads on its own machine, so being signed in to CivitAI in a browser does not authenticate it — a gated model answers a server with an HTML login page, which is caught and reported rather than saved as weights.

kohya, CivitAI and LoKr files are converted on the fly. Most of what CivitAI carries is kohya-named, and it differs from diffusers in ways that quietly ruin a result rather than raise: flat underscored names, a fused QKV interleaved per attention head (a plain three-way split hands q's rows to k), a gated MLP whose two halves sit in the opposite order, and an alpha that sets the scale. LyCORIS LoKr goes further and stores each layer as the Kronecker product of two small factors, which PEFT cannot load at all; it is rebuilt into an ordinary low-rank pair at load time — exactly, since the SVD of a Kronecker product is the outer product of the factors' SVDs, so nothing the size of the full 7168×7168 layer is ever built. .pt files are read as well as .safetensors.

Turbo presets. Few-step distillations that render joint video + soundtrack in 4–8 steps instead of the usual ~28. Picking one fills a free slot at that build's own strength and moves the steps slider to the count it was distilled for.

Continue the scene. One press takes the last frame of the clip you have, makes it the first reference, generates again with the settings untouched, and joins the two into a single file — soundtrack included. Press again for a third clip. Each press costs one normal generation; the joining is free.

Clip stitching, soundtrack included. Queue several clips and they are joined into one file; tick the box and every new generation joins on its own. It runs on the CPU, so it costs nothing from the GPU allowance. Everything is scaled to the first clip's frame, and a clip without audio gets silence rather than breaking the join.

Named profiles. Save every setting under a name and bring it back from a dropdown, or export it as .json.

A live GPU cost readout above the button, using the same budget() the pre-flight check and get_duration use, so a request that will not fit is visible as a number rather than as a failed run.

Randomize seed, drawn fresh per press and written back into the box, so the number shown is always the one the clip was made with.

Why split

MiniMax-H3 is 195.9 GiB in bfloat16 and a ZeroGPU Space is evicted at 150 GB of storage. An unquantized single Space is therefore impossible. Cut the ref2va branch of MiniMaxH3Blocks at its text_encoder step and both halves fit unquantized:

Space Subfolders Download Resident
qwen3vl-conditioner text_encoder/ + tokenizer/ + processor/ 66.7 GB 62.15 GiB bf16
this one transformer_ref/ + vae/ + audio_vae/ 77.3 GB 61.73 GiB bf16 + 10.43 GiB float32

References

A request carries up to 12 references — at most 9 images, 3 videos and 3 audio clips — in the order the model reads them. The order is semantic: it numbers the labels of MiniMax-H3's prompt presentation (<Picture 1>, <Video 1>, <Audio 1>) and it advances the shared audio/video rotary clock, so the same references in a different order are a different request. This demo lays the slots out as one tab per modality in reading order — images, then audio, then video — and assembles the request that way. The Images tab opens with two slots and + Add another image reveals the rest, up to the model's own nine.

Rules the model imposes, enforced here before anything is uploaded:

  • an audio reference cannot be the only one; it needs an image or a video alongside it,
  • a reference video runs 2 to 15 seconds, and brings its own soundtrack with it,
  • the generated duration may be left to the references, but only when exactly one of them carries a soundtrack — which is why the duration slider disappears when a single reference can set it, and comes back when two can or when the one that could is out of range.

Generation constraints

Fixed by the checkpoint: 24 fps, a 768 pixel short edge, 5 to 15 s, num_frames snapped up to the next 17 * n + 5, no CFG and no negative prompt (it is guidance-distilled, so every step is one forward pass). The duration slider stops at 14 s because it is the snapped count that has to hold for the ceiling: 15 s is 360 frames, which rounds up to 362, i.e. 15.083 s, and is refused.

GPU time is reserved per request, not per Space

MiniMax-H3 attends over one packed sequence, so what a step costs is a function of that sequence's length alone — and on this half the references dominate it. A single 1344x768 image reference is ~7168 conditioning rows plus the vision block it puts in front of the prompt; a 2.5 s video reference is another ~17000. The same 960x544, 124-frame request runs 2.4 s/step with no references and 16 s/step with an image and a video.

get_duration prices that before the call instead of reserving a flat ceiling for everything. It takes the arguments of the @spaces.GPU function, so it has the conditioner's own text_token_tags (exact) and the reference files (measured from metadata, no decode), and evaluates

S = text rows + reference rows + target rows
seconds = placement + reference encode + steps * (LINEAR * S + QUADRATIC * S**2) * SAFETY + decode + pad

fitted on the t2va half and checked against live ref2va requests to about 10%. It matters beyond tidiness: the pool reserves whatever number it is given, and a flat 900 s is what makes a busy account fail admission with "You have too many ZeroGPU credits allocated to running tasks." A typical single-image request now reserves ~460 s.

Nothing is paid for with GPU time that does not have to be

The 77.3 GB download and the load happen at startup: import spaces at module top patches torch.cuda before any GPU is attached, so nothing about the load needs a card. The conditioner round trip is a network call on this Space's CPU. So is every lora download, every CivitAI lookup, and all clip stitching. A @spaces.GPU call is only the placement (once), the two reference encoders, the denoise loop and the two decoders.

Space variables

Variable Default Meaning
H3_CONDITIONER multimodalart/qwen3vl-conditioner The public Space this one asks for embeddings; the client passes no token, so the call runs on the caller's own quota.
H3_MODEL_REPO MiniMaxAI/MiniMax-H3 The diffusers-layout checkpoint. Public.
H3_AOTI 0 1 loads the compiled block package.
H3_PLACEMENT lazy lazy moves all 72.16 GiB onto the card on the first GPU call and leaves it there; offload hands placement to ComponentsManager.enable_auto_cpu_offload instead.
H3_ATTENTION _native_cudnn cuDNN's fused kernel, 10–20% faster than the SDPA default. flash-attention 3 is sm90-only and this pool is sm120.
H3_GPU_DURATION_MIN / _MAX 120 / 1500 Bounds on what get_duration may reserve.
H3_PLACEMENT_ALLOWANCE 90 Seconds of the reservation set aside for a cold worker's placement.
H3_GPU_SIZE xlarge ZeroGPU allocation size. large does not fit.
H3_LOKR_RANK 32 Rank the LoKr converter targets. Raise it if the log reports a weak layer.
CIVITAI_TOKEN unset Used for CivitAI downloads and search, and required for gated or adult entries.
LORA_LIBRARY_BUCKET amisima/minimax-h3-reference-4-step-lora-storage The storage bucket the shared lora library is kept in. Attach it to the Space under Settings → Storage Buckets and no token is needed; otherwise an HF_TOKEN with write access is.
LORA_LIBRARY_DIR unset Overrides where the library file is read and written, when the mount path is not found on its own.
CIVITAI_API_HOST unset Pins the search host; otherwise civitai.red is asked first and civitai.com is the fallback.

Where diffusers comes from

MiniMax-H3 is modular-only and not in a released diffusers, so requirements.txt installs it from the canonical pull request, huggingface/diffusers#14371, pinned to the commit 665f5782 rather than to the moving minimax-h3-refactor branch. That PR is a WIP, so it needs re-pinning whenever it updates, and h3_split_blocks.py has to be re-checked against the new head at the same time.

Attribution

An optimized derivative of multimodalart/minimax-h3. MiniMax-H3 weights remain governed by the MiniMax-H3 Community License Agreement.