multimodalart's picture
multimodalart HF Staff
Re-encode example clips to browser-playable yuv420p H.264
f4aa38e verified
|
Raw History Blame Contribute Delete
9.59 kB

A newer version of the Gradio SDK is available: 6.29.1

Upgrade
metadata
title: MiniMax-H3 Character Swap LoRA
emoji: 🎭
colorFrom: red
colorTo: yellow
sdk: gradio
sdk_version: 6.28.0
app_file: app.py
pinned: false
short_description: Swap one character in a clip for a reference character
python_version: '3.12'
startup_duration_timeout: 1h
models:
  - akatz-ai/MiniMax-H3-Character-Swap-LoRA
  - multimodalart/MiniMax-H3-Pruned
  - MiniMaxAI/MiniMax-H3
datasets:
  - akatz-ai/H3-Character-Swap-v1

MiniMax-H3 Character Swap LoRA

A demo of akatz-ai/MiniMax-H3-Character-Swap-LoRA, Akatz Labs' experimental character-replacement adapter for MiniMax-H3's ref2va partition. Hand it a scene clip and a character reference, name who to replace, and it puts the reference character into the shot β€” identity, outfit and art style carried over β€” while the background, camera, lighting and everyone else stay where they were. Video and its synchronized soundtrack come out of a single denoising pass.

The request

Two references, in the order the model reads them, and a short targeting instruction:

Slot Label in the prompt What it is
reference 1 <Picture 1> the replacement character β€” a portrait or a full character sheet
reference 2 <Video 1> the clip to edit, 5 frames to 15 s

Swap the man in the purple shirt in <Video 1> with the character in <Picture 1>.

That order is the one thing about this request worth being careful with, and it is not the order the prompt names them in. MiniMax-H3 numbers a reference's label per modality, so the first image is <Picture 1> and the first video is <Video 1> whichever way round they are packed β€” but the packed order still fixes the shared audio/video rotary clock, so the same two references swapped round are a different request. The LoRA was trained with the character sheet first: control_path: [character_references, scene_videos] in configs/trained-run-1000.json of akatz-ai/H3-Character-Swap-v1. collect() builds the list that way and nothing else in the app reorders it.

No trigger word was trained. Strength 1.0 is the card's recommendation; 0 is the base ref2va model, which is the comparison the adapter was judged against, so the slider doubles as an A/B.

What it is good at, and what it is not

The card is candid, and this demo does not oversell it. It is a 1,000-update experimental adapter:

  • background and scene preservation improved over the base model in the author's local comparisons β€” qualitative, not a benchmark,
  • motion timing, facial expressions and hard cuts remain unreliable; a hard cut can become a zoom or a gradual reposition,
  • long windows drift in framing and placement. Short continuous shots of roughly 3–5 s are the promising range,
  • two-character inference was tested but multi-character replacement was never supervised,
  • the soundtrack is generated, not carried over. Audio preservation in the author's later review was a remux, which does not repair lip-sync drift.

Its 94 training edits targeted single still frames, with five-frame static clips standing in for <Video 1> β€” which is exactly what the examples below are. Real moving footage goes in the same slot and is what the 40 preservation clips regularized, but it is the harder case.

Defaults, and where they come from

1344x768 at 24 fps, 73 frames (3.04 s), 28 steps, seed 904231 β€” the sample block of the LoRA's own configs/trained-run-1000.json. The canvas dropdown follows the scene clip's aspect ratio on upload, because a character swap is asked to keep the source framing and a portrait clip generated on a landscape canvas is a recomposed shot before the model has done anything. "Match the scene clip's length" generates for as long as the clip runs whenever that is a length MiniMax-H3 generates (2–14 s); the five-frame example clips are not, so they fall through to the slider.

How it is deployed

MiniMax-H3 is 195.9 GiB in bfloat16 and a ZeroGPU Space is evicted at 150 GB of storage, so MiniMaxH3Blocks is cut at its text_encoder step. This Space is the denoising half of ref2va β€” the transformer_ref partition and the two autoencoders β€” and the 62 GiB Qwen3-VL conditioner runs in multimodalart/qwen3vl-conditioner, which this Space calls over the gradio API for every request. prompt_embeds + text_token_tags is the whole wire format; reference_encoder stays here, next to the autoencoders it runs. h3_split_blocks.py is the subclass that removes the step.

The DiT is multimodalart/MiniMax-H3-Pruned's transformer_ref: the released partition with its AdaLN input projections folded onto their reachable rank, 37.5 GiB instead of 61.7. That is the same checkpoint family the LoRA was trained against β€” ai-toolkit trained it on Comfy-Org's minimax_h3_ref2va_pruned_int8_convrot, an int8 ConvRot quantization of these weights β€” and everything the adapter touches is identical between the pruned and the released partition. The run's network_kwargs.ignore_if_contains = ["adaln_proj"] kept it off the timestep path, which is the only place the two differ.

Measured. The default request β€” 1344x768, 73 frames, 28 steps, one character sheet and one scene clip, a 41,408-row packed sequence β€” runs 228 s of denoise + decode on a warm worker, 10 of its 27 forwards served from the first-block cache, and 242 s end to end including the conditioner round trip. The reservation is larger than that on purpose: it has to cover a cold worker's placement and a request the cache skips nothing on.

GPU time is priced per request, not per Space. MiniMax-H3 attends over one packed sequence, and on this half the references dominate its length: a 2048-short-edge character sheet is thousands of conditioning rows on top of the generated ones. get_duration evaluates a fitted cost model over the sequence it is about to denoise β€” the conditioner's exact token count plus the references measured from metadata β€” instead of reserving a flat ceiling for everything, because the pool reserves whatever number it is given.

A first-block cache (h3_fbc.py, ported from duckyshell/ComfyUI-MiniMaxH3-FirstBlockCache with an audio exemption) skips blocks 1–49 on steps whose block-0 residual has barely moved. H3_FBC=0 restores the uncached trajectory exactly.

There is no AoTI on this Space, unlike its siblings: a compiled block package binds the base module's weights by fully qualified name and would run straight past the PEFT branch the adapter lives in.

Space variables

Variable Default Meaning
H3_LORA_SCALE 1.0 Default adapter strength.
H3_CONDITIONER multimodalart/qwen3vl-conditioner The public Space this one asks for embeddings; the client passes no token, so the call runs on the caller's own quota.
H3_MODEL_REPO multimodalart/MiniMax-H3-Pruned The diffusers-layout DiT.
H3_ATTENTION _native_cudnn cuDNN's fused kernel, 10–20% faster than the SDPA default. The two float32 VAEs are pinned to torch SDPA, which cuDNN has no kernel for.
H3_FBC / H3_FBC_THRESHOLD 1 / 0.05 First-block cache and its relative-L1 gate.
H3_GPU_SIZE xlarge ZeroGPU allocation size. large does not fit.
H3_PLACEMENT lazy Moves the partition onto the card on the first GPU call and leaves it there.

Safety

Every request here carries an image and a video reference β€” the edit-on-a-real-photo case β€” so hfmlsoc/ncii-light-guard-v01 screens the prompt on all of them, before the conditioner call and before any GPU is booked. It runs in its own subprocess (ncii_guard.py): loaded in the main process, its torch activity poisons every later ZeroGPU fork.

Example assets

All five files in examples/ are from akatz-ai/H3-Character-Swap-v1, the LoRA's own training set β€” Apache-2.0 for Akatz Labs' synthetic contributions. They are the dataset's CS001, CS051 and CS090 edits, with the instructions their own captions carry:

File Dataset path
cafe_scene.mp4 checks/smoke-data/edits/train/scene_videos/CS001.mp4 (shared with CS051)
character_hiker.png .../character_references/CS001.png
character_anime.png .../character_references/CS051.png β€” a cross-style swap
workshop_scene.mp4 .../scene_videos/CS090.mp4
character_sheet_orin.png .../character_references/CS090.png β€” a multi-view character sheet

The two clips are re-encoded from the dataset's H.264 High 4:4:4 (yuv444p) to H.264 High yuv420p with the same five frames. Browsers cannot decode 4:4:4 H.264, so the originals would not play in the examples.

License

The adapter is distributed under the MiniMax H3 Community License Agreement, not Apache-2.0, and that agreement excludes the US, EU, UK and Republic of Korea from its standard territorial grant. Read the upstream terms; nothing here extends them. The dataset's own Apache-2.0 covers the example assets only and does not replace the model's terms.