chaswatkins67 commited on
Commit
681b14a
Β·
verified Β·
1 Parent(s): 4a2d6a7

Upload 3 files

Browse files
Files changed (3) hide show
  1. README.md +57 -7
  2. app.py +186 -0
  3. requirements.txt +2 -0
README.md CHANGED
@@ -1,10 +1,60 @@
1
  ---
2
- title: MiniMax H3 Character Swap LoRA
3
- emoji: πŸ–ΌοΈ
4
- colorFrom: yellow
5
- colorTo: red
6
- sdk: static
7
- pinned: false
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
8
  ---
9
 
10
- Check out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  ---
2
+ title: MiniMax-H3 Character Swap LoRA
3
+ emoji: 🎭
4
+ colorFrom: purple
5
+ colorTo: pink
6
+ sdk: gradio
7
+ sdk_version: 5.49.1
8
+ app_file: app.py
9
+ short_description: Prompt + workflow demo for H3 character swap
10
+ python_version: "3.12"
11
+ license: other
12
+ tags:
13
+ - video-to-video
14
+ - character-swap
15
+ - minimax-h3
16
+ - lora
17
+ - comfyui
18
+ models:
19
+ - akatz-ai/MiniMax-H3-Character-Swap-LoRA
20
+ - Comfy-Org/MiniMax-H3
21
+ datasets:
22
+ - akatz-ai/H3-Character-Swap-v1
23
  ---
24
 
25
+ # MiniMax H3 Character Swap LoRA β€” demo Space
26
+
27
+ Interactive prompt builder and ComfyUI checklist for
28
+ [akatz-ai/MiniMax-H3-Character-Swap-LoRA](https://huggingface.co/akatz-ai/MiniMax-H3-Character-Swap-LoRA).
29
+
30
+ This Space does **not** run the H3 Ref2VA base model. That adapter is not a
31
+ standalone generator: you still need the pruned INT8 Ref2VA checkpoint, the
32
+ video/audio VAEs, and an H3 Ref2VA runtime (typically ComfyUI). Inference
33
+ Providers do not currently host this model.
34
+
35
+ What this demo does:
36
+
37
+ 1. Take a source-clip description and a replacement-character description
38
+ 2. Emit the card-recommended Ref2VA prompt (`<Video 1>` / `<Picture 1>`)
39
+ 3. Print a one-page ComfyUI setup sheet at LoRA strength **1.0**
40
+
41
+ ## Weights
42
+
43
+ - LoRA: `h3_character_swap_pro4500_1000.safetensors` (final 1,000-step only)
44
+ - Strength: **1.0**
45
+ - No trained trigger word
46
+ - Base: `minimax_h3_ref2va_pruned_int8_convrot.safetensors` from
47
+ [Comfy-Org/MiniMax-H3](https://huggingface.co/Comfy-Org/MiniMax-H3)
48
+ - Training assistant (not merged): [ostris/minimax_h3_training_adapter](https://huggingface.co/ostris/minimax_h3_training_adapter)
49
+
50
+ ## License
51
+
52
+ Distributed under the MiniMax H3 Community License Agreement. It is **not**
53
+ Apache-2.0. Read the full terms on the model repo before commercial or
54
+ territorial use.
55
+
56
+ ## Author notes from the card
57
+
58
+ Short continuous 4–5 s shots at 24 fps work better than long 14 s tests.
59
+ Hard cuts, facial-expression lock, and multi-character swaps are unreliable.
60
+ This LoRA does not require a Turbo LoRA, Spectrum, or Sol attention.
app.py ADDED
@@ -0,0 +1,186 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """MiniMax H3 Character Swap LoRA β€” prompt + workflow demo.
2
+
3
+ This Space does not load the H3 base weights. It builds the Ref2VA prompt
4
+ and ComfyUI checklist from the published model card.
5
+ """
6
+
7
+ from __future__ import annotations
8
+
9
+ import gradio as gr
10
+
11
+ LORA_REPO = "https://huggingface.co/akatz-ai/MiniMax-H3-Character-Swap-LoRA"
12
+ LORA_FILE = "h3_character_swap_pro4500_1000.safetensors"
13
+ DATASET = "https://huggingface.co/datasets/akatz-ai/H3-Character-Swap-v1"
14
+ BASE = "https://huggingface.co/Comfy-Org/MiniMax-H3"
15
+
16
+ DEFAULT_TARGET = "the man in the purple shirt"
17
+ DEFAULT_KEEP = (
18
+ "Keep the replacement character's identity, outfit, and art style from <Picture 1>. "
19
+ "Preserve the source video's camera, background, lighting, objects, and all other people. "
20
+ "Match the target person's position, scale, pose, and movement. "
21
+ "Do not show the reference sheet or its background."
22
+ )
23
+
24
+ SHORT_CAPTION = (
25
+ "Swap {target} in <Video 1> with the character in <Picture 1>."
26
+ )
27
+
28
+ LONG_PROMPT = (
29
+ "Replace only {target} in <Video 1> with the character in <Picture 1>. "
30
+ "{keep}"
31
+ )
32
+
33
+
34
+ def build_pack(
35
+ target: str,
36
+ extra_keep: str,
37
+ duration_s: float,
38
+ fps: int,
39
+ strength: float,
40
+ use_long: bool,
41
+ ) -> tuple[str, str, str]:
42
+ """Build the Ref2VA prompt and a ComfyUI run sheet."""
43
+ target = (target or DEFAULT_TARGET).strip()
44
+ keep = (extra_keep or DEFAULT_KEEP).strip()
45
+ prompt = LONG_PROMPT.format(target=target, keep=keep) if use_long else SHORT_CAPTION.format(target=target)
46
+
47
+ sheet = f"""# H3 Character Swap β€” run sheet
48
+
49
+ ## Weights
50
+ - LoRA: `{LORA_FILE}`
51
+ - Repo: {LORA_REPO}
52
+ - Strength: {strength:.2f} (card default: 1.00)
53
+ - Trigger word: none
54
+ - Turbo / Spectrum / Sol: not required
55
+
56
+ ## Base runtime (not bundled in this Space)
57
+ - Diffusion: `minimax_h3_ref2va_pruned_int8_convrot.safetensors`
58
+ - Source: {BASE}
59
+ - Video + audio VAEs from the same org pack
60
+ - Loader: model-only LoRA loader on the H3 transformer
61
+
62
+ ## Inputs
63
+ - `<Video 1>` = source clip (prefer one continuous shot)
64
+ - `<Picture 1>` = replacement character or sheet (no sheet background in the output)
65
+ - Prompt below, verbatim
66
+
67
+ ## Timing
68
+ - Duration request: {duration_s:.1f} s
69
+ - Frame rate: {fps} fps
70
+ - Card finding: 4–5 s continuous shots beat ~14 s tests
71
+ - Hard cuts often become zooms or slides β€” trim cuts first
72
+
73
+ ## Limits (from the trainer)
74
+ - Background preservation is the advertised win vs base Ref2VA
75
+ - Expressions and cut timing are still unreliable
76
+ - Multi-character replacement was not supervised
77
+ - Audio from the model is not trustworthy; remux source audio after if needed
78
+ """
79
+
80
+ checklist = f"""1. Drop `{LORA_FILE}` into `ComfyUI/models/loras/`
81
+ 2. Load H3 Ref2VA INT8 + VAEs (not this LoRA alone)
82
+ 3. Apply LoRA at strength {strength:.2f} with a model-only loader
83
+ 4. Wire source video β†’ `<Video 1>`, character image β†’ `<Picture 1>`
84
+ 5. Paste the prompt
85
+ 6. Generate a short 24 fps window, then continue if the join holds
86
+ 7. If lips drift, copy original audio onto the new video in post
87
+ """
88
+ return prompt, sheet, checklist
89
+
90
+
91
+ def preview_note(video, image) -> str:
92
+ v = "source video attached" if video else "no source video yet"
93
+ i = "reference image attached" if image else "no reference image yet"
94
+ return (
95
+ f"Local preview only β€” this Space does **not** run MiniMax H3.\n"
96
+ f"- {v}\n- {i}\n\n"
97
+ f"Take both files into ComfyUI with the generated prompt."
98
+ )
99
+
100
+
101
+ with gr.Blocks(title="MiniMax H3 Character Swap LoRA") as demo:
102
+ gr.Markdown(
103
+ f"""
104
+ # MiniMax H3 Character Swap LoRA
105
+ Experimental **one-character** replacement adapter by Akatz Labs
106
+ ([model card]({LORA_REPO}), [dataset]({DATASET})).
107
+
108
+ Upload a clip and a character still, describe who to replace, and this demo
109
+ writes the Ref2VA prompt plus a ComfyUI checklist. **Live H3 inference is not
110
+ hosted here** β€” the adapter is 155 MB, the base Ref2VA stack is not, and no
111
+ Inference Provider serves it yet.
112
+ """
113
+ )
114
+
115
+ with gr.Row():
116
+ with gr.Column():
117
+ video = gr.Video(label="Source clip β†’ <Video 1>")
118
+ image = gr.Image(type="filepath", label="Replacement character β†’ <Picture 1>")
119
+ target = gr.Textbox(
120
+ label="Who to replace in the clip",
121
+ value=DEFAULT_TARGET,
122
+ placeholder="the woman in the red coat / the boy on the left / …",
123
+ )
124
+ extra_keep = gr.Textbox(
125
+ label="Preservation instructions",
126
+ value=DEFAULT_KEEP,
127
+ lines=5,
128
+ )
129
+ with gr.Row():
130
+ duration = gr.Slider(2, 14, value=5, step=0.5, label="Clip window (seconds)")
131
+ fps = gr.Slider(8, 30, value=24, step=1, label="FPS")
132
+ strength = gr.Slider(0.2, 1.2, value=1.0, step=0.05, label="LoRA strength")
133
+ use_long = gr.Checkbox(
134
+ value=True,
135
+ label="Use the long preservation prompt (recommended in local evals)",
136
+ )
137
+ go = gr.Button("Build prompt + run sheet", variant="primary")
138
+ with gr.Column():
139
+ status = gr.Markdown("Attach a clip and a still, then build the pack.")
140
+ prompt_out = gr.Textbox(label="Ref2VA prompt", lines=8)
141
+ sheet_out = gr.Markdown()
142
+ check_out = gr.Textbox(label="ComfyUI checklist", lines=10)
143
+
144
+ go.click(
145
+ preview_note,
146
+ inputs=[video, image],
147
+ outputs=status,
148
+ ).then(
149
+ build_pack,
150
+ inputs=[target, extra_keep, duration, fps, strength, use_long],
151
+ outputs=[prompt_out, sheet_out, check_out],
152
+ )
153
+
154
+ gr.Examples(
155
+ examples=[
156
+ [
157
+ "the man in the purple shirt",
158
+ DEFAULT_KEEP,
159
+ 5.0,
160
+ 24,
161
+ 1.0,
162
+ True,
163
+ ],
164
+ [
165
+ "the woman speaking on the left",
166
+ DEFAULT_KEEP,
167
+ 4.0,
168
+ 24,
169
+ 1.0,
170
+ True,
171
+ ],
172
+ ],
173
+ inputs=[target, extra_keep, duration, fps, strength, use_long],
174
+ label="Starter targets",
175
+ )
176
+
177
+ gr.Markdown(
178
+ f"""
179
+ ---
180
+ **Base model** is [Comfy-Org/MiniMax-H3]({BASE}), not merged into the LoRA.
181
+ License: MiniMax H3 Community License β€” read it before you ship anything.
182
+ """
183
+ )
184
+
185
+ if __name__ == "__main__":
186
+ demo.launch(mcp_server=True)
requirements.txt ADDED
@@ -0,0 +1,2 @@
 
 
 
1
+ # Gradio / spaces / huggingface_hub are preinstalled on Spaces.
2
+ # Keep this file tiny so a CPU or ZeroGPU Space can boot.