multimodalart's picture
multimodalart HF Staff
Re-encode example clips to browser-playable yuv420p H.264
f4aa38e verified
|
Raw History Blame Contribute Delete
9.59 kB
---
title: MiniMax-H3 Character Swap LoRA
emoji: 🎭
colorFrom: red
colorTo: yellow
sdk: gradio
sdk_version: 6.28.0
app_file: app.py
pinned: false
short_description: Swap one character in a clip for a reference character
python_version: "3.12"
startup_duration_timeout: 1h
models:
- akatz-ai/MiniMax-H3-Character-Swap-LoRA
- multimodalart/MiniMax-H3-Pruned
- MiniMaxAI/MiniMax-H3
datasets:
- akatz-ai/H3-Character-Swap-v1
---
# MiniMax-H3 Character Swap LoRA
A demo of [`akatz-ai/MiniMax-H3-Character-Swap-LoRA`](https://huggingface.co/akatz-ai/MiniMax-H3-Character-Swap-LoRA),
Akatz Labs' experimental character-replacement adapter for
[MiniMax-H3](https://huggingface.co/MiniMaxAI/MiniMax-H3)'s `ref2va` partition. Hand it a scene clip and a character
reference, name who to replace, and it puts the reference character into the shot β€” identity, outfit and art style
carried over β€” while the background, camera, lighting and everyone else stay where they were. Video and its
synchronized soundtrack come out of a single denoising pass.
## The request
Two references, **in the order the model reads them**, and a short targeting instruction:
| Slot | Label in the prompt | What it is |
|---|---|---|
| reference 1 | `<Picture 1>` | the replacement character β€” a portrait or a full character sheet |
| reference 2 | `<Video 1>` | the clip to edit, 5 frames to 15 s |
> Swap the man in the purple shirt in \<Video 1\> with the character in \<Picture 1\>.
That order is the one thing about this request worth being careful with, and it is **not** the order the prompt
names them in. MiniMax-H3 numbers a reference's label *per modality*, so the first image is `<Picture 1>` and the
first video is `<Video 1>` whichever way round they are packed β€” but the packed order still fixes the shared
audio/video rotary clock, so the same two references swapped round are a different request. The LoRA was trained
with the character sheet first: `control_path: [character_references, scene_videos]` in
`configs/trained-run-1000.json` of [`akatz-ai/H3-Character-Swap-v1`](https://huggingface.co/datasets/akatz-ai/H3-Character-Swap-v1).
`collect()` builds the list that way and nothing else in the app reorders it.
No trigger word was trained. Strength **1.0** is the card's recommendation; **0** is the base `ref2va` model, which
is the comparison the adapter was judged against, so the slider doubles as an A/B.
## What it is good at, and what it is not
The card is candid, and this demo does not oversell it. It is a **1,000-update experimental** adapter:
* background and scene preservation improved over the base model in the author's local comparisons β€” qualitative,
not a benchmark,
* motion timing, facial expressions and hard cuts remain unreliable; a hard cut can become a zoom or a gradual
reposition,
* long windows drift in framing and placement. **Short continuous shots of roughly 3–5 s** are the promising range,
* two-character inference was tested but multi-character replacement was never supervised,
* the soundtrack is generated, not carried over. Audio preservation in the author's later review was a remux, which
does not repair lip-sync drift.
Its 94 training edits targeted **single still frames**, with five-frame static clips standing in for `<Video 1>` β€”
which is exactly what the examples below are. Real moving footage goes in the same slot and is what the 40
preservation clips regularized, but it is the harder case.
## Defaults, and where they come from
`1344x768` at 24 fps, 73 frames (3.04 s), 28 steps, seed 904231 β€” the `sample` block of the LoRA's own
`configs/trained-run-1000.json`. The canvas dropdown follows the scene clip's aspect ratio on upload, because a
character swap is asked to keep the source framing and a portrait clip generated on a landscape canvas is a
recomposed shot before the model has done anything. "Match the scene clip's length" generates for as long as the
clip runs whenever that is a length MiniMax-H3 generates (2–14 s); the five-frame example clips are not, so they
fall through to the slider.
## How it is deployed
MiniMax-H3 is 195.9 GiB in bfloat16 and a ZeroGPU Space is evicted at 150 GB of storage, so `MiniMaxH3Blocks` is
cut at its `text_encoder` step. This Space is the **denoising half** of `ref2va` β€” the `transformer_ref` partition
and the two autoencoders β€” and the 62 GiB Qwen3-VL conditioner runs in
[`multimodalart/qwen3vl-conditioner`](https://huggingface.co/spaces/multimodalart/qwen3vl-conditioner), which this
Space calls over the gradio API for every request. `prompt_embeds` + `text_token_tags` is the whole wire format;
`reference_encoder` stays here, next to the autoencoders it runs. `h3_split_blocks.py` is the subclass that removes
the step.
The DiT is [`multimodalart/MiniMax-H3-Pruned`](https://huggingface.co/multimodalart/MiniMax-H3-Pruned)'s
`transformer_ref`: the released partition with its AdaLN input projections folded onto their reachable rank, 37.5
GiB instead of 61.7. That is the **same checkpoint family the LoRA was trained against** β€” ai-toolkit trained it on
Comfy-Org's `minimax_h3_ref2va_pruned_int8_convrot`, an int8 ConvRot quantization of these weights β€” and everything
the adapter touches is identical between the pruned and the released partition. The run's
`network_kwargs.ignore_if_contains = ["adaln_proj"]` kept it off the timestep path, which is the only place the two
differ.
**Measured.** The default request β€” 1344x768, 73 frames, 28 steps, one character sheet and one scene clip, a
41,408-row packed sequence β€” runs **228 s** of denoise + decode on a warm worker, 10 of its 27 forwards served from
the first-block cache, and 242 s end to end including the conditioner round trip. The reservation is larger than
that on purpose: it has to cover a cold worker's placement and a request the cache skips nothing on.
**GPU time is priced per request, not per Space.** MiniMax-H3 attends over one packed sequence, and on this half the
references dominate its length: a 2048-short-edge character sheet is thousands of conditioning rows on top of the
generated ones. `get_duration` evaluates a fitted cost model over the sequence it is about to denoise β€” the
conditioner's exact token count plus the references measured from metadata β€” instead of reserving a flat ceiling for
everything, because the pool reserves whatever number it is given.
A first-block cache (`h3_fbc.py`, ported from `duckyshell/ComfyUI-MiniMaxH3-FirstBlockCache` with an audio
exemption) skips blocks 1–49 on steps whose block-0 residual has barely moved. `H3_FBC=0` restores the uncached
trajectory exactly.
There is no AoTI on this Space, unlike its siblings: a compiled block package binds the base module's weights by
fully qualified name and would run straight past the PEFT branch the adapter lives in.
## Space variables
| Variable | Default | Meaning |
|---|---|---|
| `H3_LORA_SCALE` | `1.0` | Default adapter strength. |
| `H3_CONDITIONER` | `multimodalart/qwen3vl-conditioner` | The public Space this one asks for embeddings; the client passes no token, so the call runs on the caller's own quota. |
| `H3_MODEL_REPO` | `multimodalart/MiniMax-H3-Pruned` | The diffusers-layout DiT. |
| `H3_ATTENTION` | `_native_cudnn` | cuDNN's fused kernel, 10–20% faster than the SDPA default. The two float32 VAEs are pinned to torch SDPA, which cuDNN has no kernel for. |
| `H3_FBC` / `H3_FBC_THRESHOLD` | `1` / `0.05` | First-block cache and its relative-L1 gate. |
| `H3_GPU_SIZE` | `xlarge` | ZeroGPU allocation size. `large` does not fit. |
| `H3_PLACEMENT` | `lazy` | Moves the partition onto the card on the first GPU call and leaves it there. |
## Safety
Every request here carries an image and a video reference β€” the edit-on-a-real-photo case β€” so
[`hfmlsoc/ncii-light-guard-v01`](https://huggingface.co/hfmlsoc/ncii-light-guard-v01) screens the prompt on all of
them, before the conditioner call and before any GPU is booked. It runs in its own subprocess
(`ncii_guard.py`): loaded in the main process, its torch activity poisons every later ZeroGPU fork.
## Example assets
All five files in `examples/` are from
[`akatz-ai/H3-Character-Swap-v1`](https://huggingface.co/datasets/akatz-ai/H3-Character-Swap-v1), the LoRA's own
training set β€” Apache-2.0 for Akatz Labs' synthetic contributions. They are the dataset's `CS001`, `CS051` and
`CS090` edits, with the instructions their own captions carry:
| File | Dataset path |
|---|---|
| `cafe_scene.mp4` | `checks/smoke-data/edits/train/scene_videos/CS001.mp4` (shared with `CS051`) |
| `character_hiker.png` | `.../character_references/CS001.png` |
| `character_anime.png` | `.../character_references/CS051.png` β€” a cross-style swap |
| `workshop_scene.mp4` | `.../scene_videos/CS090.mp4` |
| `character_sheet_orin.png` | `.../character_references/CS090.png` β€” a multi-view character sheet |
The two clips are re-encoded from the dataset's H.264 High 4:4:4 (`yuv444p`) to H.264 High `yuv420p` with the same
five frames. Browsers cannot decode 4:4:4 H.264, so the originals would not play in the examples.
## License
The adapter is distributed under the
[MiniMax H3 Community License Agreement](https://huggingface.co/akatz-ai/MiniMax-H3-Character-Swap-LoRA/blob/main/LICENSE),
**not** Apache-2.0, and that agreement excludes the US, EU, UK and Republic of Korea from its standard territorial
grant. Read the upstream terms; nothing here extends them. The dataset's own Apache-2.0 covers the example assets
only and does not replace the model's terms.