Spaces:
Running on Zero
Running on Zero
|
Download README.md from hugging-apps/minimax-h3-character-swap-lora: direct link, hf CLI and curl.
- Browser
- Download file 9.59 kB
-
https://huggingface.co/spaces/hugging-apps/minimax-h3-character-swap-lora/resolve/main/README.md
- Command line
-
hf download hf://spaces/hugging-apps/minimax-h3-character-swap-lora/README.md
-
curl -L -o README.md https://huggingface.co/spaces/hugging-apps/minimax-h3-character-swap-lora/resolve/main/README.md
9.59 kB
| title: MiniMax-H3 Character Swap LoRA | |
| emoji: π | |
| colorFrom: red | |
| colorTo: yellow | |
| sdk: gradio | |
| sdk_version: 6.28.0 | |
| app_file: app.py | |
| pinned: false | |
| short_description: Swap one character in a clip for a reference character | |
| python_version: "3.12" | |
| startup_duration_timeout: 1h | |
| models: | |
| - akatz-ai/MiniMax-H3-Character-Swap-LoRA | |
| - multimodalart/MiniMax-H3-Pruned | |
| - MiniMaxAI/MiniMax-H3 | |
| datasets: | |
| - akatz-ai/H3-Character-Swap-v1 | |
| # MiniMax-H3 Character Swap LoRA | |
| A demo of [`akatz-ai/MiniMax-H3-Character-Swap-LoRA`](https://huggingface.co/akatz-ai/MiniMax-H3-Character-Swap-LoRA), | |
| Akatz Labs' experimental character-replacement adapter for | |
| [MiniMax-H3](https://huggingface.co/MiniMaxAI/MiniMax-H3)'s `ref2va` partition. Hand it a scene clip and a character | |
| reference, name who to replace, and it puts the reference character into the shot β identity, outfit and art style | |
| carried over β while the background, camera, lighting and everyone else stay where they were. Video and its | |
| synchronized soundtrack come out of a single denoising pass. | |
| ## The request | |
| Two references, **in the order the model reads them**, and a short targeting instruction: | |
| | Slot | Label in the prompt | What it is | | |
| |---|---|---| | |
| | reference 1 | `<Picture 1>` | the replacement character β a portrait or a full character sheet | | |
| | reference 2 | `<Video 1>` | the clip to edit, 5 frames to 15 s | | |
| > Swap the man in the purple shirt in \<Video 1\> with the character in \<Picture 1\>. | |
| That order is the one thing about this request worth being careful with, and it is **not** the order the prompt | |
| names them in. MiniMax-H3 numbers a reference's label *per modality*, so the first image is `<Picture 1>` and the | |
| first video is `<Video 1>` whichever way round they are packed β but the packed order still fixes the shared | |
| audio/video rotary clock, so the same two references swapped round are a different request. The LoRA was trained | |
| with the character sheet first: `control_path: [character_references, scene_videos]` in | |
| `configs/trained-run-1000.json` of [`akatz-ai/H3-Character-Swap-v1`](https://huggingface.co/datasets/akatz-ai/H3-Character-Swap-v1). | |
| `collect()` builds the list that way and nothing else in the app reorders it. | |
| No trigger word was trained. Strength **1.0** is the card's recommendation; **0** is the base `ref2va` model, which | |
| is the comparison the adapter was judged against, so the slider doubles as an A/B. | |
| ## What it is good at, and what it is not | |
| The card is candid, and this demo does not oversell it. It is a **1,000-update experimental** adapter: | |
| * background and scene preservation improved over the base model in the author's local comparisons β qualitative, | |
| not a benchmark, | |
| * motion timing, facial expressions and hard cuts remain unreliable; a hard cut can become a zoom or a gradual | |
| reposition, | |
| * long windows drift in framing and placement. **Short continuous shots of roughly 3β5 s** are the promising range, | |
| * two-character inference was tested but multi-character replacement was never supervised, | |
| * the soundtrack is generated, not carried over. Audio preservation in the author's later review was a remux, which | |
| does not repair lip-sync drift. | |
| Its 94 training edits targeted **single still frames**, with five-frame static clips standing in for `<Video 1>` β | |
| which is exactly what the examples below are. Real moving footage goes in the same slot and is what the 40 | |
| preservation clips regularized, but it is the harder case. | |
| ## Defaults, and where they come from | |
| `1344x768` at 24 fps, 73 frames (3.04 s), 28 steps, seed 904231 β the `sample` block of the LoRA's own | |
| `configs/trained-run-1000.json`. The canvas dropdown follows the scene clip's aspect ratio on upload, because a | |
| character swap is asked to keep the source framing and a portrait clip generated on a landscape canvas is a | |
| recomposed shot before the model has done anything. "Match the scene clip's length" generates for as long as the | |
| clip runs whenever that is a length MiniMax-H3 generates (2β14 s); the five-frame example clips are not, so they | |
| fall through to the slider. | |
| ## How it is deployed | |
| MiniMax-H3 is 195.9 GiB in bfloat16 and a ZeroGPU Space is evicted at 150 GB of storage, so `MiniMaxH3Blocks` is | |
| cut at its `text_encoder` step. This Space is the **denoising half** of `ref2va` β the `transformer_ref` partition | |
| and the two autoencoders β and the 62 GiB Qwen3-VL conditioner runs in | |
| [`multimodalart/qwen3vl-conditioner`](https://huggingface.co/spaces/multimodalart/qwen3vl-conditioner), which this | |
| Space calls over the gradio API for every request. `prompt_embeds` + `text_token_tags` is the whole wire format; | |
| `reference_encoder` stays here, next to the autoencoders it runs. `h3_split_blocks.py` is the subclass that removes | |
| the step. | |
| The DiT is [`multimodalart/MiniMax-H3-Pruned`](https://huggingface.co/multimodalart/MiniMax-H3-Pruned)'s | |
| `transformer_ref`: the released partition with its AdaLN input projections folded onto their reachable rank, 37.5 | |
| GiB instead of 61.7. That is the **same checkpoint family the LoRA was trained against** β ai-toolkit trained it on | |
| Comfy-Org's `minimax_h3_ref2va_pruned_int8_convrot`, an int8 ConvRot quantization of these weights β and everything | |
| the adapter touches is identical between the pruned and the released partition. The run's | |
| `network_kwargs.ignore_if_contains = ["adaln_proj"]` kept it off the timestep path, which is the only place the two | |
| differ. | |
| **Measured.** The default request β 1344x768, 73 frames, 28 steps, one character sheet and one scene clip, a | |
| 41,408-row packed sequence β runs **228 s** of denoise + decode on a warm worker, 10 of its 27 forwards served from | |
| the first-block cache, and 242 s end to end including the conditioner round trip. The reservation is larger than | |
| that on purpose: it has to cover a cold worker's placement and a request the cache skips nothing on. | |
| **GPU time is priced per request, not per Space.** MiniMax-H3 attends over one packed sequence, and on this half the | |
| references dominate its length: a 2048-short-edge character sheet is thousands of conditioning rows on top of the | |
| generated ones. `get_duration` evaluates a fitted cost model over the sequence it is about to denoise β the | |
| conditioner's exact token count plus the references measured from metadata β instead of reserving a flat ceiling for | |
| everything, because the pool reserves whatever number it is given. | |
| A first-block cache (`h3_fbc.py`, ported from `duckyshell/ComfyUI-MiniMaxH3-FirstBlockCache` with an audio | |
| exemption) skips blocks 1β49 on steps whose block-0 residual has barely moved. `H3_FBC=0` restores the uncached | |
| trajectory exactly. | |
| There is no AoTI on this Space, unlike its siblings: a compiled block package binds the base module's weights by | |
| fully qualified name and would run straight past the PEFT branch the adapter lives in. | |
| ## Space variables | |
| | Variable | Default | Meaning | | |
| |---|---|---| | |
| | `H3_LORA_SCALE` | `1.0` | Default adapter strength. | | |
| | `H3_CONDITIONER` | `multimodalart/qwen3vl-conditioner` | The public Space this one asks for embeddings; the client passes no token, so the call runs on the caller's own quota. | | |
| | `H3_MODEL_REPO` | `multimodalart/MiniMax-H3-Pruned` | The diffusers-layout DiT. | | |
| | `H3_ATTENTION` | `_native_cudnn` | cuDNN's fused kernel, 10β20% faster than the SDPA default. The two float32 VAEs are pinned to torch SDPA, which cuDNN has no kernel for. | | |
| | `H3_FBC` / `H3_FBC_THRESHOLD` | `1` / `0.05` | First-block cache and its relative-L1 gate. | | |
| | `H3_GPU_SIZE` | `xlarge` | ZeroGPU allocation size. `large` does not fit. | | |
| | `H3_PLACEMENT` | `lazy` | Moves the partition onto the card on the first GPU call and leaves it there. | | |
| ## Safety | |
| Every request here carries an image and a video reference β the edit-on-a-real-photo case β so | |
| [`hfmlsoc/ncii-light-guard-v01`](https://huggingface.co/hfmlsoc/ncii-light-guard-v01) screens the prompt on all of | |
| them, before the conditioner call and before any GPU is booked. It runs in its own subprocess | |
| (`ncii_guard.py`): loaded in the main process, its torch activity poisons every later ZeroGPU fork. | |
| ## Example assets | |
| All five files in `examples/` are from | |
| [`akatz-ai/H3-Character-Swap-v1`](https://huggingface.co/datasets/akatz-ai/H3-Character-Swap-v1), the LoRA's own | |
| training set β Apache-2.0 for Akatz Labs' synthetic contributions. They are the dataset's `CS001`, `CS051` and | |
| `CS090` edits, with the instructions their own captions carry: | |
| | File | Dataset path | | |
| |---|---| | |
| | `cafe_scene.mp4` | `checks/smoke-data/edits/train/scene_videos/CS001.mp4` (shared with `CS051`) | | |
| | `character_hiker.png` | `.../character_references/CS001.png` | | |
| | `character_anime.png` | `.../character_references/CS051.png` β a cross-style swap | | |
| | `workshop_scene.mp4` | `.../scene_videos/CS090.mp4` | | |
| | `character_sheet_orin.png` | `.../character_references/CS090.png` β a multi-view character sheet | | |
| The two clips are re-encoded from the dataset's H.264 High 4:4:4 (`yuv444p`) to H.264 High `yuv420p` with the same | |
| five frames. Browsers cannot decode 4:4:4 H.264, so the originals would not play in the examples. | |
| ## License | |
| The adapter is distributed under the | |
| [MiniMax H3 Community License Agreement](https://huggingface.co/akatz-ai/MiniMax-H3-Character-Swap-LoRA/blob/main/LICENSE), | |
| **not** Apache-2.0, and that agreement excludes the US, EU, UK and Republic of Korea from its standard territorial | |
| grant. Read the upstream terms; nothing here extends them. The dataset's own Apache-2.0 covers the example assets | |
| only and does not replace the model's terms. | |