alibaybay commited on
Commit
21d3176
·
verified ·
1 Parent(s): 7dd75bb

Deploy DualSpace H3 Generator Hub & Async REST API (Comfy2API)

Browse files
README.md CHANGED
@@ -1,13 +1,32 @@
1
  ---
2
- title: Dualspace H3 Generator
3
- emoji: 📚
4
- colorFrom: gray
5
- colorTo: indigo
6
  sdk: gradio
7
- sdk_version: 6.29.1
8
- python_version: '3.12'
9
  app_file: app.py
10
  pinned: false
 
 
 
 
 
 
 
 
11
  ---
12
 
13
- Check out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  ---
2
+ title: DualSpace MiniMax-H3 Generator
3
+ emoji: ⚡
4
+ colorFrom: purple
5
+ colorTo: pink
6
  sdk: gradio
7
+ sdk_version: 5.44.1
 
8
  app_file: app.py
9
  pinned: false
10
+ license: apache-2.0
11
+ tags:
12
+ - comfyui
13
+ - zerogpu
14
+ - minimax-h3
15
+ - video-generator
16
+ - image-to-video
17
+ - rest-api
18
  ---
19
 
20
+ # ⚡ MiniMax-H3 FL2VA — Dual-Space Video Studio & Asynchronous REST API
21
+
22
+ Aplikasi Video Generative berbasis **MiniMax-H3 (FL2VA)** dengan arsitektur **Dual-Space Split (ComfyUI Core Native Backend)** pada Hugging Face ZeroGPU.
23
+ Dilengkapi dengan antarmuka web interaktif Gradio 5 dan **100% Asynchronous Twin REST API** standar Fal.ai/Replicate (`/api/generate`, `/api/status/{job_id}`, `/api/result/{job_id}`).
24
+
25
+ ## 🚀 Fitur Utama:
26
+ - **100% Asynchronous REST API**: Integrasi mudah dengan cURL, Python, Node.js/Cloudflare Workers, dan Bruno.
27
+ - **Dual-Space Split**: Memisahkan Qwen3-VL 32B + Keyframe Encode (Space 1) dan Denoising UNet INT8 + Video & Audio VAE (Space 2).
28
+ - **TaoMate 3-Step LoRA**: Inferensi ultra-cepat hanya dalam 3 sampling steps dengan kualitas visual yang tajam.
29
+ - **Video VAE INT8 ConvRot**: Efisiensi komputasi decode dengan konsumsi VRAM minimal.
30
+ - **ComfyUI Core Native**: Seluruh grafis komputasi mengandalkan node bawaan resmi (`EmptyMiniMaxH3LatentAV`, `MiniMaxH3SigmaShift`, dll).
31
+ - **Thin Wire Protocol**: Mengirimkan file `.safetensors` ramping dengan transfer instan antar-Space (~0.05s).
32
+ - **Audio-Video Synchronized**: Menghasilkan video MP4 lengkap dengan track audio tersinkronisasi.
__pycache__/app.cpython-311.pyc ADDED
Binary file (57.2 kB). View file
 
app.py ADDED
@@ -0,0 +1,969 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ from __future__ import annotations
2
+
3
+ import os
4
+ import sys
5
+ import subprocess
6
+ import pathlib
7
+ import shutil
8
+ import re
9
+ import uuid
10
+ import json
11
+ import glob
12
+ import random
13
+ import time
14
+ import base64
15
+ import asyncio
16
+ from functools import lru_cache
17
+ from typing import Any
18
+ import requests as http_requests
19
+
20
+ # ============================================================
21
+ # 1. HUGGINGFACE_HUB SELF-HEALING REPAIR
22
+ # ============================================================
23
+ def _hf_hub_version() -> str:
24
+ try:
25
+ from importlib.metadata import version as _pkg_version
26
+ return _pkg_version("huggingface_hub")
27
+ except Exception:
28
+ return ""
29
+
30
+ def _hf_hub_is_broken() -> bool:
31
+ import importlib
32
+ import importlib.util
33
+ for module_name in ("huggingface_hub._snapshot_download", "huggingface_hub._tree_cache"):
34
+ try:
35
+ if importlib.util.find_spec(module_name) is None:
36
+ continue
37
+ except Exception:
38
+ return True
39
+ try:
40
+ importlib.import_module(module_name)
41
+ except ImportError:
42
+ return True
43
+ except Exception:
44
+ continue
45
+ return False
46
+
47
+ def _hf_hub_reinstall(upgrade: bool) -> None:
48
+ cmd = [
49
+ sys.executable,
50
+ "-m",
51
+ "pip",
52
+ "install",
53
+ "--no-cache-dir",
54
+ "--force-reinstall",
55
+ "--no-deps",
56
+ ]
57
+ if upgrade:
58
+ cmd += ["--upgrade", "huggingface_hub"]
59
+ else:
60
+ pinned = _hf_hub_version()
61
+ cmd.append(f"huggingface_hub=={pinned}" if pinned else "huggingface_hub")
62
+ print(f"[hf-repair] {' '.join(cmd)}", flush=True)
63
+ subprocess.run(cmd, check=False)
64
+
65
+ def _repair_huggingface_hub_and_restart() -> None:
66
+ stage = int(os.environ.get("_HF_HUB_REPAIR_STAGE", "0") or "0")
67
+ if stage >= 2 or not _hf_hub_is_broken():
68
+ return
69
+ _hf_hub_reinstall(upgrade=stage == 1)
70
+ os.environ["_HF_HUB_REPAIR_STAGE"] = str(stage + 1)
71
+ os.execv(sys.executable, [sys.executable, *sys.argv])
72
+
73
+ _repair_huggingface_hub_and_restart()
74
+
75
+ # ============================================================
76
+ # 2. IMPORTS UTAMA (SPACES WAJIB PERTAMA SEBELUM TORCH)
77
+ # ============================================================
78
+ import spaces # WAJIB PERTAMA sebelum torch!
79
+ import torch
80
+
81
+ # ============================================================
82
+ # ZEROGPU COMPATIBILITY PATCH FOR PYTORCH CUDA MOCK PROPERTIES
83
+ # ============================================================
84
+ if hasattr(torch, "cuda") and hasattr(torch.cuda, "get_device_properties"):
85
+ _orig_cuda_get_device_properties = torch.cuda.get_device_properties
86
+ def _safe_cuda_get_device_properties(device=None):
87
+ props = _orig_cuda_get_device_properties(device)
88
+ if not hasattr(props, "is_integrated"):
89
+ try:
90
+ setattr(props, "is_integrated", False)
91
+ except Exception:
92
+ class _PropsProxy:
93
+ def __init__(self, p):
94
+ self._p = p
95
+ self.is_integrated = False
96
+ def __getattr__(self, name):
97
+ return getattr(self._p, name)
98
+ return _PropsProxy(props)
99
+ return props
100
+ torch.cuda.get_device_properties = _safe_cuda_get_device_properties
101
+
102
+ from fastapi import Request, Response, HTTPException
103
+ import gradio as gr
104
+ import gradio_client.utils
105
+ from gradio_client import Client, handle_file
106
+ from huggingface_hub import hf_hub_download
107
+
108
+ # ============================================================
109
+ # 2.1 MONKEY-PATCH GRADIO_CLIENT OPENAPI SCHEMA BUG
110
+ # ============================================================
111
+ _orig_get_type = gradio_client.utils.get_type
112
+ def _safe_get_type(schema):
113
+ if isinstance(schema, bool):
114
+ return "boolean"
115
+ if not isinstance(schema, dict):
116
+ return "str"
117
+ return _orig_get_type(schema)
118
+ gradio_client.utils.get_type = _safe_get_type
119
+
120
+ _orig_json_schema = gradio_client.utils._json_schema_to_python_type
121
+ def _safe_json_schema(schema, defs=None):
122
+ if isinstance(schema, bool):
123
+ return "bool"
124
+ if not isinstance(schema, dict):
125
+ return "str"
126
+ return _orig_json_schema(schema, defs)
127
+ gradio_client.utils._json_schema_to_python_type = _safe_json_schema
128
+
129
+ # ============================================================
130
+ # 3. KONFIGURASI PATH & DIREKTORI
131
+ # ============================================================
132
+ ROOT = pathlib.Path(__file__).resolve().parent
133
+ COMFY = ROOT / "ComfyUI"
134
+ MODELS = COMFY / "models"
135
+ INPUT = COMFY / "input"
136
+ OUTPUT = COMFY / "output"
137
+ LOCAL_CUSTOM_NODES = ROOT / "custom_nodes"
138
+
139
+ WORKFLOW_FILE = ROOT / "workflow_generator.json"
140
+ NODE_OUTPUT_ID = "92"
141
+
142
+ CONDITIONER_SPACE = os.environ.get("H3_CONDITIONER_SPACE", "alibaybay/dualspace-h3-clip")
143
+
144
+ # ============================================================
145
+ # 4. DAFTAR MODEL GENERATOR MINIMAX-H3 (INT8 + TAOMATE 3-STEP)
146
+ # ============================================================
147
+ DOWNLOADS = [
148
+ {
149
+ "repo": "Comfy-Org/MiniMax-H3",
150
+ "file": "diffusion_models/minimax_h3_fl2va_pruned_int8_convrot.safetensors",
151
+ "dest": MODELS / "diffusion_models" / "minimax_h3_fl2va_pruned_int8_convrot.safetensors",
152
+ "alt_dest": MODELS / "unet" / "minimax_h3_fl2va_pruned_int8_convrot.safetensors",
153
+ "label": "Diffusion Model (FL2VA Pruned INT8 ~21.0GB)",
154
+ },
155
+ {
156
+ "repo": "Kijai/MiniMax-H3_comfy",
157
+ "file": "loras/minimax_h3_taomate_3step_lora_avg_rank_19_bf16.safetensors",
158
+ "dest": MODELS / "loras" / "minimax_h3_taomate_3step_lora_avg_rank_19_bf16.safetensors",
159
+ "alt_dest": None,
160
+ "label": "TaoMate 3-Step LoRA BF16 (~1.96GB)",
161
+ },
162
+ {
163
+ "repo": "Comfy-Org/MiniMax-H3",
164
+ "file": "vae/minimax_h3_video_vae_int8_convrot.safetensors",
165
+ "dest": MODELS / "vae" / "minimax_h3_video_vae_int8_convrot.safetensors",
166
+ "alt_dest": None,
167
+ "label": "Video VAE INT8 ConvRot (~2.6GB)",
168
+ },
169
+ {
170
+ "repo": "Comfy-Org/MiniMax-H3",
171
+ "file": "vae/minimax_h3_audio_vae_fp32.safetensors",
172
+ "dest": MODELS / "vae" / "minimax_h3_audio_vae_fp32.safetensors",
173
+ "alt_dest": None,
174
+ "label": "Audio VAE FP32 (~605MB)",
175
+ },
176
+ ]
177
+
178
+ CUSTOM_NODES: list[tuple[str, str]] = []
179
+
180
+ _comfy_ready = False
181
+ _nodes_ready = False
182
+ server_instance = None
183
+ gpu_lock = asyncio.Lock()
184
+
185
+ # Job Tracking Asynchronous REST API
186
+ JOBS: dict[str, dict[str, Any]] = {}
187
+
188
+ # ============================================================
189
+ # 5. HELPER & MODEL DOWNLOADER
190
+ # ============================================================
191
+ def _run_cmd(cmd: list[str], cwd: pathlib.Path = ROOT, check: bool = True) -> None:
192
+ print(f"[*] Menjalankan: {' '.join(cmd)} di {cwd}", flush=True)
193
+ subprocess.run(cmd, cwd=cwd, check=check)
194
+
195
+ def _link_or_copy(src: pathlib.Path, dest: pathlib.Path) -> None:
196
+ dest.parent.mkdir(parents=True, exist_ok=True)
197
+ if dest.is_symlink():
198
+ dest.unlink()
199
+ if dest.exists() and dest.stat().st_size > 1000:
200
+ return
201
+ try:
202
+ os.link(src, dest)
203
+ return
204
+ except OSError:
205
+ pass
206
+ shutil.copy2(src, dest)
207
+
208
+ def _download_to_dest(repo: str, file_path: str, dest: pathlib.Path, token: str | None) -> None:
209
+ dest.parent.mkdir(parents=True, exist_ok=True)
210
+ if dest.is_symlink():
211
+ dest.unlink()
212
+ if dest.exists() and dest.stat().st_size > 1000:
213
+ return
214
+
215
+ p = pathlib.Path(file_path)
216
+ filename = p.name
217
+ subfolder = str(p.parent) if str(p.parent) != "." else None
218
+
219
+ print(f"[*] Mengunduh {filename} dari {repo} ke {dest.parent}...", flush=True)
220
+ downloaded_str = hf_hub_download(
221
+ repo_id=repo,
222
+ filename=filename,
223
+ subfolder=subfolder,
224
+ local_dir=str(dest.parent),
225
+ token=token,
226
+ )
227
+ downloaded = pathlib.Path(downloaded_str)
228
+
229
+ if downloaded.resolve() == dest.resolve():
230
+ return
231
+
232
+ if dest.exists() or dest.is_symlink():
233
+ dest.unlink()
234
+ dest.parent.mkdir(parents=True, exist_ok=True)
235
+ try:
236
+ os.replace(downloaded, dest)
237
+ except OSError:
238
+ shutil.copy2(downloaded, dest)
239
+ if downloaded.exists():
240
+ downloaded.unlink()
241
+
242
+ sub_dir = dest.parent / "split_files"
243
+ if sub_dir.exists():
244
+ shutil.rmtree(sub_dir, ignore_errors=True)
245
+
246
+ def _install_filtered_requirements(req_path: pathlib.Path, cwd: pathlib.Path) -> None:
247
+ if not req_path.exists():
248
+ return
249
+ blocked = {"torch", "torchvision", "torchaudio", "transformers", "huggingface-hub", "accelerate", "xformers"}
250
+ safe: list[str] = []
251
+ for line in req_path.read_text(encoding="utf-8", errors="ignore").splitlines():
252
+ item = line.strip()
253
+ if not item or item.startswith("#"):
254
+ continue
255
+ low = item.lower().replace("_", "-")
256
+ package = re.split(r"[<>=!~;\[\s]", low, maxsplit=1)[0]
257
+ if package in blocked:
258
+ continue
259
+ safe.append(item)
260
+ if safe:
261
+ filtered_file = cwd / "requirements_filtered.txt"
262
+ filtered_file.write_text("\n".join(safe) + "\n", encoding="utf-8")
263
+ _run_cmd([sys.executable, "-m", "pip", "install", "-r", "requirements_filtered.txt", "--no-cache-dir"], cwd=cwd, check=False)
264
+
265
+ def _apply_comfy_utils_namespace_fix() -> None:
266
+ utils_path = COMFY / "utils"
267
+ utilities_path = COMFY / "utilities"
268
+ if utils_path.exists() and not utilities_path.exists():
269
+ try:
270
+ utils_path.rename(utilities_path)
271
+ except OSError:
272
+ pass
273
+
274
+ replacements = [
275
+ (re.compile(r"(^|\n)(\s*)from utils(\s|\.)"), r"\1\2from utilities\3"),
276
+ (re.compile(r"(^|\n)(\s*)import utils(\s|\.|$)"), r"\1\2import utilities\3"),
277
+ ]
278
+ for path in COMFY.rglob("*.py"):
279
+ if "__pycache__" in path.parts:
280
+ continue
281
+ try:
282
+ text = path.read_text(encoding="utf-8")
283
+ except UnicodeDecodeError:
284
+ continue
285
+ updated = text
286
+ for pattern, repl in replacements:
287
+ updated = pattern.sub(repl, updated)
288
+ updated = updated.replace("from utils import", "from utilities import")
289
+ if updated != text:
290
+ path.write_text(updated, encoding="utf-8")
291
+
292
+ def _ensure_comfy() -> None:
293
+ global _comfy_ready
294
+ if _comfy_ready:
295
+ return
296
+
297
+ print("[1/3] Menyiapkan ComfyUI Runtime untuk H3 Generator...", flush=True)
298
+ if not COMFY.exists():
299
+ _run_cmd(["git", "clone", "--depth", "1", "--branch", "v0.38.2", "https://github.com/comfyanonymous/ComfyUI.git", str(COMFY)])
300
+ _install_filtered_requirements(COMFY / "requirements.txt", COMFY)
301
+
302
+ custom_root = COMFY / "custom_nodes"
303
+ custom_root.mkdir(parents=True, exist_ok=True)
304
+
305
+ # Pasang local thin wire node: external_h3_conditioning
306
+ if LOCAL_CUSTOM_NODES.exists():
307
+ for src_node in LOCAL_CUSTOM_NODES.iterdir():
308
+ if src_node.is_dir() and not src_node.name.startswith("."):
309
+ target_node = custom_root / src_node.name
310
+ if target_node.exists():
311
+ shutil.rmtree(target_node, ignore_errors=True)
312
+ shutil.copytree(src_node, target_node)
313
+ print(f"[*] Terpasang local custom node: {src_node.name}", flush=True)
314
+
315
+ _apply_comfy_utils_namespace_fix()
316
+
317
+ for folder in ("diffusion_models", "unet", "loras", "vae"):
318
+ (MODELS / folder).mkdir(parents=True, exist_ok=True)
319
+ INPUT.mkdir(parents=True, exist_ok=True)
320
+ OUTPUT.mkdir(parents=True, exist_ok=True)
321
+
322
+ _comfy_ready = True
323
+ print("[1/3] ComfyUI Runtime H3 Generator Siap.", flush=True)
324
+
325
+ def _ensure_models(progress=None) -> None:
326
+ print("[2/3] Memeriksa & Mengunduh Model Generator MiniMax-H3...", flush=True)
327
+ token = os.environ.get("HF_TOKEN") or os.environ.get("HUGGINGFACE_HUB_TOKEN")
328
+ for row in DOWNLOADS:
329
+ dest = pathlib.Path(row["dest"])
330
+ dest.parent.mkdir(parents=True, exist_ok=True)
331
+ if dest.is_symlink():
332
+ dest.unlink()
333
+ if not (dest.exists() and dest.stat().st_size > 1000):
334
+ print(f"[*] Mengunduh {row['label']}...", flush=True)
335
+ _download_to_dest(row["repo"], row["file"], dest, token)
336
+
337
+ alt = row.get("alt_dest")
338
+ if alt is not None:
339
+ alt_path = pathlib.Path(alt)
340
+ if dest.exists() and dest.stat().st_size > 1000:
341
+ _link_or_copy(dest, alt_path)
342
+ print("[2/3] Semua model generator MiniMax-H3 telah siap.", flush=True)
343
+
344
+ def _init_comfy_nodes() -> None:
345
+ global _nodes_ready, server_instance
346
+ if _nodes_ready:
347
+ return
348
+
349
+ print("[3/3] Menginisialisasi Engine ComfyUI Generator (Full Standby)...", flush=True)
350
+ comfy_path = str(COMFY)
351
+ sys.path = [p for p in sys.path if p != comfy_path]
352
+ sys.path.insert(0, comfy_path)
353
+ for module_name in list(sys.modules):
354
+ if module_name in ("utils", "app") or module_name.startswith(("utils.", "app.")):
355
+ del sys.modules[module_name]
356
+
357
+ os.chdir(COMFY)
358
+
359
+ import execution
360
+ import nodes
361
+ import server
362
+
363
+ loop = asyncio.new_event_loop()
364
+ asyncio.set_event_loop(loop)
365
+
366
+ import inspect
367
+ sig = inspect.signature(server.PromptServer.__init__)
368
+ if "asset_manager" in sig.parameters:
369
+ try:
370
+ from app.assets.manager import default_asset_manager
371
+ asset_mgr = default_asset_manager()
372
+ except Exception:
373
+ class DummyAssetManager:
374
+ enabled = False
375
+ def startup(self): pass
376
+ def shutdown(self): pass
377
+ def register_routes(self, app, user_manager=None): pass
378
+ def ensure_scan_started(self): pass
379
+ def pause_background_scan(self): pass
380
+ def queue_output_scan(self): pass
381
+ def resume_background_scan(self): pass
382
+ def register_upload(self, *args, **kwargs): return None
383
+ def register_executed_output(self, *args, **kwargs): return None
384
+ def register_cached_output(self, *args, **kwargs): return None
385
+ def set_event_sink(self, sink): pass
386
+ asset_mgr = DummyAssetManager()
387
+ server_instance = server.PromptServer(loop, asset_mgr)
388
+ else:
389
+ server_instance = server.PromptServer(loop)
390
+
391
+ try:
392
+ execution.PromptQueue(server_instance)
393
+ except Exception:
394
+ pass
395
+
396
+ res = nodes.init_extra_nodes()
397
+ if asyncio.iscoroutine(res):
398
+ loop.run_until_complete(res)
399
+
400
+ _nodes_ready = True
401
+ print("[3/3] Engine ComfyUI Generator Siap & Berada dalam Mode Hot-Standby.", flush=True)
402
+
403
+ executor_instance = None
404
+
405
+ def _get_or_create_executor():
406
+ global executor_instance
407
+ if executor_instance is None:
408
+ import execution
409
+ executor_instance = execution.PromptExecutor(
410
+ server_instance,
411
+ cache_type=execution.CacheType.RAM_PRESSURE,
412
+ cache_args={"lru": 32, "ram": 60.0, "ram_inactive": 60.0},
413
+ )
414
+ return executor_instance
415
+
416
+ def _preload_models_to_ram():
417
+ """Me-load UNet INT8, LoRA TaoMate 3-Step, Video VAE INT8, dan Audio VAE ke RAM saat boot."""
418
+ print("[*] Pre-loading bobot model (UNet INT8 ~21GB, LoRA, VAE INT8) ke RAM...", flush=True)
419
+ try:
420
+ import nodes
421
+ unet_loader = nodes.UNETLoader()
422
+ vae_loader = nodes.VAELoader()
423
+ lora_loader = nodes.LoraLoaderModelOnly()
424
+
425
+ # 1. Load UNet INT8
426
+ print("[*] Pre-loading UNet INT8...", flush=True)
427
+ unet_res = unet_loader.load_unet("minimax_h3_fl2va_pruned_int8_convrot.safetensors", weight_dtype="default")
428
+ unet_model = unet_res[0]
429
+
430
+ # 2. Patch LoRA TaoMate 3-Step
431
+ print("[*] Pre-patching LoRA TaoMate 3-Step...", flush=True)
432
+ lora_loader.load_lora_model_only(unet_model, "minimax_h3_taomate_3step_lora_avg_rank_19_bf16.safetensors", 1.0)
433
+
434
+ # 3. Load Video VAE INT8 & Audio VAE
435
+ print("[*] Pre-loading Video VAE INT8 & Audio VAE...", flush=True)
436
+ vae_loader.load_vae("minimax_h3_video_vae_int8_convrot.safetensors")
437
+ vae_loader.load_vae("minimax_h3_audio_vae_fp32.safetensors")
438
+
439
+ print("[*] Pre-load model ke RAM berhasil! Semua model telah siap di memory (Zero Disk Reload).", flush=True)
440
+ except Exception as e:
441
+ print(f"[!] Warning saat pre-load model: {e}", flush=True)
442
+
443
+ # ============================================================
444
+ # 6. ROOT STARTUP PRE-WARMING (HOT-STANDBY OPTIMIZATION)
445
+ # ============================================================
446
+ def _startup_prewarm():
447
+ print("=" * 60, flush=True)
448
+ print("[startup] Memulai Pre-Warming Engine H3 Generator Hub...", flush=True)
449
+ _ensure_comfy()
450
+ _ensure_models()
451
+ _init_comfy_nodes()
452
+ _get_or_create_executor()
453
+ _preload_models_to_ram()
454
+ print("[startup] Pre-Warming Selesai. Siap Melayani Permintaan Instan.", flush=True)
455
+ print("=" * 60, flush=True)
456
+
457
+ _startup_prewarm()
458
+
459
+ # ============================================================
460
+ # 7. PERSISTENT LRU CLIENT POOL KE SPACE KONDISIONER
461
+ # ============================================================
462
+ @lru_cache(maxsize=32)
463
+ def _get_conditioner_client(space_id: str, ip_token: str | None) -> Client:
464
+ """Membuka dan mempertahankan sesi koneksi HTTP / WebSocket ke Space 1."""
465
+ headers = {"x-ip-token": ip_token} if ip_token else {}
466
+ print(f"[*] [LRU Pool] Membuka persistent connection ke Conditioner: {space_id}", flush=True)
467
+ return Client(space_id, headers=headers)
468
+
469
+ def fetch_remote_conditioning(
470
+ prompt: str,
471
+ first_frame_path: str,
472
+ last_frame_path: str | None,
473
+ duration: float,
474
+ ip_token: str | None,
475
+ ) -> str:
476
+ """Mengirim parameter ke Space 1 via Thin Wire API dan menerima file .safetensors."""
477
+ client = _get_conditioner_client(CONDITIONER_SPACE, ip_token)
478
+ print(f"[*] [Space 2] Meminta conditioning dari Space 1 ({CONDITIONER_SPACE})...", flush=True)
479
+
480
+ res = client.predict(
481
+ prompt=prompt,
482
+ first_frame_path=handle_file(first_frame_path),
483
+ last_frame_path=handle_file(last_frame_path) if last_frame_path else None,
484
+ duration=f"{duration:.0f}s",
485
+ api_name="/encode",
486
+ )
487
+
488
+ if isinstance(res, (tuple, list)):
489
+ remote_file = res[0]
490
+ elif isinstance(res, dict) and "path" in res:
491
+ remote_file = res["path"]
492
+ else:
493
+ remote_file = str(res)
494
+
495
+ size_kb = os.path.getsize(remote_file) / 1024.0 if os.path.exists(remote_file) else 0
496
+ print(f"[*] [Space 2] Berhasil menerima Thin Wire safetensors: {remote_file} ({size_kb:.2f} KB)", flush=True)
497
+ return remote_file
498
+
499
+ # ============================================================
500
+ # 8. LOGIKA RUNNER ZERO-OVERHEAD @SPACES.GPU (DYNAMIC DURATION)
501
+ # ============================================================
502
+ DURATION_GPU_MAP = {
503
+ 5.0: 55, # Waktu riil ~35s -> Minta 55s (ZeroGPU reserve: 82.5s)
504
+ 10.0: 75, # Waktu riil ~55s -> Minta 75s (ZeroGPU reserve: 112.5s)
505
+ 15.0: 100, # Waktu riil ~74.6s -> Minta 100s (ZeroGPU reserve: 150.0s)
506
+ }
507
+
508
+ def get_generator_duration(safetensors_path: str, video_duration: float | str = 5.0, seed: int = 0) -> int:
509
+ try:
510
+ dur = float(str(video_duration).replace("s", "").strip())
511
+ except Exception:
512
+ dur = 5.0
513
+ return DURATION_GPU_MAP.get(dur, int(min(max(50, 30 + dur * 4.5), 110)))
514
+
515
+ def _load_base_workflow() -> dict[str, Any]:
516
+ with open(WORKFLOW_FILE, "r", encoding="utf-8") as f:
517
+ return json.load(f)
518
+
519
+ def _run_comfy_generator_workflow(workflow: dict[str, Any]) -> str:
520
+ """Eksekusi workflow Denoising di Space 2."""
521
+ import execution
522
+
523
+ executor = _get_or_create_executor()
524
+ prompt_id = str(uuid.uuid4())
525
+
526
+ node_durations: list[tuple[str, str, float]] = []
527
+ orig_get_output_data = execution.get_output_data
528
+
529
+ def _profiling_get_output_data(obj, input_data_all, *args, **kwargs):
530
+ if isinstance(obj, str):
531
+ node_id = obj
532
+ node_info = workflow.get(node_id, {})
533
+ class_type = node_info.get("class_type", "UnknownNode")
534
+ node_title = node_info.get("_meta", {}).get("title", class_type)
535
+ label = f"[Node {node_id}: {node_title}]"
536
+ else:
537
+ node_id = "?"
538
+ class_type = obj.__class__.__name__
539
+ label = f"[{class_type}]"
540
+
541
+ t0 = time.time()
542
+ print(f"🚀 {label} Mulai dieksekusi...", flush=True)
543
+ try:
544
+ res = orig_get_output_data(obj, input_data_all, *args, **kwargs)
545
+ dur = time.time() - t0
546
+ node_durations.append((node_id, label, dur))
547
+ print(f"⏱️ {label} Selesai dalam: {dur:.2f}s", flush=True)
548
+ return res
549
+ except Exception as e:
550
+ dur = time.time() - t0
551
+ print(f"❌ {label} Gagal setelah: {dur:.2f}s ({e})", flush=True)
552
+ raise
553
+
554
+ execution.get_output_data = _profiling_get_output_data
555
+ t_workflow_start = time.time()
556
+
557
+ try:
558
+ executor.execute(
559
+ workflow,
560
+ prompt_id,
561
+ extra_data={},
562
+ execute_outputs=[NODE_OUTPUT_ID],
563
+ )
564
+ finally:
565
+ execution.get_output_data = orig_get_output_data
566
+ t_workflow_total = time.time() - t_workflow_start
567
+ print("\n" + "=" * 70, flush=True)
568
+ print("📊 REKAPITULASI PROFILING WAKTU SPACE 2 (PER NODE):", flush=True)
569
+ print("=" * 70, flush=True)
570
+ sorted_nodes = sorted(node_durations, key=lambda x: x[2], reverse=True)
571
+ for nid, label, dur in sorted_nodes:
572
+ pct = (dur / t_workflow_total * 100) if t_workflow_total > 0 else 0
573
+ bar = "█" * int(pct // 5)
574
+ print(f" {label:<45} : {dur:>6.2f}s ({pct:>5.1f}%) {bar}", flush=True)
575
+ print("-" * 70, flush=True)
576
+ print(f" ⏱️ TOTAL DURASI SAMPLING & DECODE : {t_workflow_total:.2f} detik", flush=True)
577
+ print("=" * 70 + "\n", flush=True)
578
+
579
+ if not executor.success:
580
+ err = (
581
+ executor.status_messages[-1]
582
+ if hasattr(executor, "status_messages") and executor.status_messages
583
+ else "ComfyUI execution gagal"
584
+ )
585
+ raise RuntimeError(str(err))
586
+
587
+ files = [
588
+ pathlib.Path(p)
589
+ for p in glob.glob(str(OUTPUT / "**" / "*.mp4"), recursive=True)
590
+ ]
591
+ if not files:
592
+ files = [
593
+ pathlib.Path(p)
594
+ for p in glob.glob(str(OUTPUT / "**" / "*.*"), recursive=True)
595
+ if p.endswith((".mp4", ".webm", ".mkv", ".mov"))
596
+ ]
597
+
598
+ if not files:
599
+ raise RuntimeError("Generation selesai tetapi file video output (.mp4) tidak ditemukan di folder output.")
600
+
601
+ latest_video = sorted(files, key=lambda p: p.stat().st_mtime, reverse=True)[0]
602
+ return str(latest_video)
603
+
604
+ @spaces.GPU(duration=get_generator_duration)
605
+ def run_generator_gpu(
606
+ safetensors_path: str,
607
+ video_duration: float | str = 5.0,
608
+ seed: int = 0,
609
+ ) -> str:
610
+ """Eksekusi murni GPU forward: Denoising TaoMate 3-Step LoRA + Video & Audio VAE Decode."""
611
+ wf = _load_base_workflow()
612
+
613
+ try:
614
+ dur_val = float(str(video_duration).replace("s", "").strip())
615
+ except Exception:
616
+ dur_val = 5.0
617
+
618
+ prefix = f"H3_vid_{uuid.uuid4().hex[:8]}"
619
+ wf["ext_h3_cond"]["inputs"]["safetensors_path"] = safetensors_path
620
+ wf["105_15"]["inputs"]["noise_seed"] = int(seed)
621
+ wf["105_9"]["inputs"]["steps"] = 3 # Kunci standar TaoMate 3-Step LoRA
622
+ wf["92"]["inputs"]["filename_prefix"] = f"video/{prefix}"
623
+
624
+ print(f"[*] [GPU Space 2] Menjalankan MiniMax-H3 Denoising (TaoMate 3-Step, Durasi={dur_val}s, Seed={seed})...", flush=True)
625
+
626
+ target_video = _run_comfy_generator_workflow(wf)
627
+ print(f"[*] [GPU Selesai] Video MP4 berhasil dibuat: {target_video}", flush=True)
628
+ return target_video
629
+
630
+ # ============================================================
631
+ # 9. PIPELINE END-TO-END DUAL-SPACE
632
+ # ============================================================
633
+ def generate_video_pipeline(
634
+ first_frame: str | None,
635
+ last_frame: str | None,
636
+ prompt: str,
637
+ duration_choice: str,
638
+ seed: float,
639
+ randomize_seed: bool,
640
+ request: gr.Request = None,
641
+ progress=gr.Progress(track_tqdm=True),
642
+ ) -> tuple[str | None, str, int]:
643
+ if not first_frame or not os.path.exists(first_frame):
644
+ raise gr.Error("First Frame (gambar keyframe awal) wajib diunggah untuk mode Image-to-Video!")
645
+
646
+ seed_int = int(seed)
647
+ if randomize_seed or seed_int == 0:
648
+ seed_int = random.randint(1, 1000000000000000)
649
+
650
+ try:
651
+ dur_val = float(str(duration_choice).replace("s", "").strip())
652
+ except Exception:
653
+ dur_val = 5.0
654
+
655
+ ip_token = None
656
+ if request and hasattr(request, "headers"):
657
+ ip_token = request.headers.get("x-ip-token")
658
+
659
+ t_total_start = time.time()
660
+
661
+ # Step 1: Conditioning via Space 1 (Thin Wire)
662
+ progress(0.1, desc="⚡ [Space 1] Menghubungi Qwen3-VL Conditioner & Keyframes Encoder...")
663
+ t_cond_start = time.time()
664
+ try:
665
+ safetensors_file = fetch_remote_conditioning(
666
+ prompt=prompt or "",
667
+ first_frame_path=first_frame,
668
+ last_frame_path=last_frame,
669
+ duration=dur_val,
670
+ ip_token=ip_token,
671
+ )
672
+ except Exception as e:
673
+ raise gr.Error(f"Gagal memproses conditioning di Space 1 ({CONDITIONER_SPACE}): {e}")
674
+ t_cond_elapsed = time.time() - t_cond_start
675
+
676
+ # Step 2: Denoising & Video Decode via Space 2 GPU
677
+ progress(0.4, desc=f"🎬 [Space 2] Denoising TaoMate 3-Step & Decoding Video ({dur_val:.0f}s)...")
678
+ t_gen_start = time.time()
679
+ try:
680
+ video_path = run_generator_gpu(
681
+ safetensors_path=safetensors_file,
682
+ video_duration=dur_val,
683
+ seed=seed_int,
684
+ )
685
+ except Exception as e:
686
+ raise gr.Error(f"Gagal saat proses denoising/video generation: {e}")
687
+ t_gen_elapsed = time.time() - t_gen_start
688
+ t_total_elapsed = time.time() - t_total_start
689
+
690
+ report = (
691
+ f"✅ Video Berhasil Dibuat!\n"
692
+ f"⏱️ Total Waktu: {t_total_elapsed:.2f}s | "
693
+ f"🧠 Space 1 (Conditioner): {t_cond_elapsed:.2f}s | "
694
+ f"🎬 Space 2 (Denoise & Decode): {t_gen_elapsed:.2f}s\n"
695
+ f"⚙️ Resolusi: Otomatis (0.4 MP) | Durasi: {dur_val:.0f}s | Steps: 3 (TaoMate LoRA) | Seed: {seed_int}"
696
+ )
697
+
698
+ return video_path, report, seed_int
699
+
700
+ # ============================================================
701
+ # 10. ASYNCHRONOUS REST TWIN-API (STANDAR INDUSTRI REPLICATE/FAL)
702
+ # ============================================================
703
+ def _cleanup_expired_jobs(ttl_seconds: int = 900) -> None:
704
+ now = time.time()
705
+ expired = [jid for jid, info in list(JOBS.items()) if now - info.get("created_at", 0) > ttl_seconds]
706
+ for jid in expired:
707
+ path = JOBS[jid].get("output_path")
708
+ if path and os.path.exists(path):
709
+ try:
710
+ os.remove(path)
711
+ except Exception:
712
+ pass
713
+ JOBS.pop(jid, None)
714
+
715
+ def process_media_value(val: str, prefix: str = "media") -> str:
716
+ """Download image dari URL atau decode dari base64, return absolute filepath."""
717
+ if not isinstance(val, str) or not val.strip():
718
+ return ""
719
+ val = val.strip()
720
+ if os.path.exists(val):
721
+ return val
722
+ if val.startswith("http://") or val.startswith("https://"):
723
+ ext = "png"
724
+ target_path = INPUT / f"{prefix}_{uuid.uuid4().hex[:8]}.{ext}"
725
+ r = http_requests.get(val, stream=True, timeout=60)
726
+ r.raise_for_status()
727
+ with open(target_path, "wb") as f:
728
+ for chunk in r.iter_content(chunk_size=8192):
729
+ f.write(chunk)
730
+ return str(target_path)
731
+ if val.startswith("data:"):
732
+ header, encoded = val.split(",", 1)
733
+ ext = "png"
734
+ if "jpeg" in header or "jpg" in header:
735
+ ext = "jpg"
736
+ elif "webp" in header:
737
+ ext = "webp"
738
+ target_path = INPUT / f"{prefix}_{uuid.uuid4().hex[:8]}.{ext}"
739
+ target_path.write_bytes(base64.b64decode(encoded))
740
+ return str(target_path)
741
+ return val
742
+
743
+ async def _job_worker(job_id: str, payload: dict[str, Any]) -> None:
744
+ JOBS[job_id]["status"] = "PROCESSING"
745
+ JOBS[job_id]["start_time"] = time.time()
746
+ try:
747
+ raw_first = payload.get("image") or payload.get("first_frame")
748
+ raw_last = payload.get("last_image") or payload.get("last_frame")
749
+ prompt = payload.get("prompt", "")
750
+ raw_dur = payload.get("duration", 5.0)
751
+ raw_seed = payload.get("seed", 0)
752
+
753
+ if not raw_first:
754
+ raise ValueError("Field 'image' (first frame) wajib disertakan.")
755
+
756
+ first_frame = await asyncio.to_thread(process_media_value, raw_first, "first_frame")
757
+ last_frame = await asyncio.to_thread(process_media_value, raw_last, "last_frame") if raw_last else None
758
+
759
+ try:
760
+ dur_val = float(str(raw_dur).replace("s", "").strip())
761
+ except Exception:
762
+ dur_val = 5.0
763
+
764
+ seed_val = int(raw_seed)
765
+ if seed_val <= 0:
766
+ seed_val = random.randint(1, 1000000000000000)
767
+
768
+ # 1. Hubungi Space 1 untuk conditioning
769
+ safetensors_file = await asyncio.to_thread(
770
+ fetch_remote_conditioning,
771
+ prompt=prompt,
772
+ first_frame_path=first_frame,
773
+ last_frame_path=last_frame,
774
+ duration=dur_val,
775
+ ip_token=None,
776
+ )
777
+
778
+ # 2. Eksekusi forward GPU Space 2
779
+ async with gpu_lock:
780
+ video_path = await asyncio.to_thread(
781
+ run_generator_gpu,
782
+ safetensors_path=safetensors_file,
783
+ video_duration=dur_val,
784
+ seed=seed_val,
785
+ )
786
+
787
+ if not video_path or not os.path.exists(video_path):
788
+ raise RuntimeError("Eksekusi selesai tetapi video output tidak ditemukan.")
789
+
790
+ JOBS[job_id]["status"] = "COMPLETED"
791
+ JOBS[job_id]["end_time"] = time.time()
792
+ JOBS[job_id]["output_path"] = str(video_path)
793
+ JOBS[job_id]["media_type"] = "video/mp4"
794
+ JOBS[job_id]["filename"] = pathlib.Path(video_path).name
795
+
796
+ except Exception as err:
797
+ JOBS[job_id]["status"] = "FAILED"
798
+ JOBS[job_id]["error"] = str(err)
799
+ JOBS[job_id]["end_time"] = time.time()
800
+
801
+ async def api_submit_job(request: Request):
802
+ try:
803
+ item = await request.json()
804
+ except Exception:
805
+ raise HTTPException(status_code=400, detail="Payload harus berupa JSON yang valid.")
806
+
807
+ _cleanup_expired_jobs()
808
+
809
+ job_id = f"job_{uuid.uuid4().hex[:12]}"
810
+ now = time.time()
811
+ JOBS[job_id] = {
812
+ "job_id": job_id,
813
+ "status": "QUEUED",
814
+ "created_at": now,
815
+ "start_time": None,
816
+ "end_time": None,
817
+ "output_path": None,
818
+ "media_type": None,
819
+ "error": None,
820
+ }
821
+
822
+ asyncio.create_task(_job_worker(job_id, item))
823
+
824
+ return Response(
825
+ content=json.dumps({
826
+ "status": "QUEUED",
827
+ "job_id": job_id,
828
+ "created_at": round(now, 3),
829
+ "check_url": f"/api/status/{job_id}",
830
+ }),
831
+ status_code=202,
832
+ media_type="application/json",
833
+ )
834
+
835
+ async def api_get_status(job_id: str):
836
+ _cleanup_expired_jobs()
837
+ job = JOBS.get(job_id)
838
+ if not job:
839
+ raise HTTPException(status_code=404, detail="Job ID tidak ditemukan atau telah kedaluwarsa.")
840
+
841
+ status = job["status"]
842
+ now = time.time()
843
+
844
+ if status == "QUEUED":
845
+ return {
846
+ "job_id": job_id,
847
+ "status": "QUEUED",
848
+ "elapsed_seconds": round(now - job["created_at"], 2),
849
+ }
850
+ elif status == "PROCESSING":
851
+ start = job["start_time"] or job["created_at"]
852
+ return {
853
+ "job_id": job_id,
854
+ "status": "PROCESSING",
855
+ "elapsed_seconds": round(now - start, 2),
856
+ }
857
+ elif status == "COMPLETED":
858
+ exec_time = round((job["end_time"] or now) - (job["start_time"] or job["created_at"]), 2)
859
+ return {
860
+ "job_id": job_id,
861
+ "status": "COMPLETED",
862
+ "execution_time_seconds": exec_time,
863
+ "media_type": job["media_type"],
864
+ "result_url": f"/api/result/{job_id}",
865
+ "error": None,
866
+ }
867
+ else: # FAILED
868
+ return {
869
+ "job_id": job_id,
870
+ "status": "FAILED",
871
+ "error": job.get("error"),
872
+ }
873
+
874
+ async def api_get_result(job_id: str):
875
+ job = JOBS.get(job_id)
876
+ if not job:
877
+ raise HTTPException(status_code=404, detail="Job ID tidak ditemukan atau telah kedaluwarsa.")
878
+
879
+ if job["status"] != "COMPLETED":
880
+ raise HTTPException(status_code=400, detail=f"Job belum selesai. Status saat ini: {job['status']}")
881
+
882
+ output_path = job.get("output_path")
883
+ if not output_path or not os.path.exists(output_path):
884
+ raise HTTPException(status_code=500, detail="File video tidak ditemukan di server.")
885
+
886
+ media_bytes = pathlib.Path(output_path).read_bytes()
887
+ exec_time = round((job["end_time"] or time.time()) - (job["start_time"] or job["created_at"]), 2)
888
+
889
+ headers = {
890
+ "Content-Disposition": f'inline; filename="{job.get("filename", "output.mp4")}"',
891
+ "X-Status": "success",
892
+ "X-Execution-Time": f"{exec_time:.2f}s",
893
+ }
894
+ return Response(content=media_bytes, media_type="video/mp4", headers=headers)
895
+
896
+ # ============================================================
897
+ # 11. ANTARMUKA GRADIO STUDIO MINIMALIS & ELEGAN
898
+ # ============================================================
899
+ custom_css = """
900
+ .container { max-width: 1200px; margin: auto; }
901
+ .generate-btn { font-size: 1.15rem !important; padding: 12px !important; font-weight: bold !important; }
902
+ """
903
+
904
+ with gr.Blocks(title="MiniMax-H3 Studio — Dual-Space ComfyUI", css=custom_css, theme=gr.themes.Soft()) as demo:
905
+ gr.Markdown(
906
+ """
907
+ # ⚡ MiniMax-H3 FL2VA — Dual-Space Video Studio & REST API
908
+ Aplikasi Video Generative canggih multimodal berbasis **MiniMax-H3 (FL2VA)** dengan arsitektur **Dual-Space ZeroGPU**.
909
+ - **Space 1 (`Conditioner`)**: Memproses Qwen3-VL 32B + Video VAE Keyframes via Thin Wire Protocol.
910
+ - **Space 2 (`Generator Hub`)**: Melakukan Denoising INT8 FL2VA + TaoMate 3-Step LoRA + Decode Video & Audio.
911
+ - **REST API Aktif**: Mendukung endpoint asinkron `/api/generate`, `/api/status/{job_id}`, dan `/api/result/{job_id}`.
912
+ """
913
+ )
914
+
915
+ with gr.Row():
916
+ with gr.Column(scale=5):
917
+ with gr.Row():
918
+ first_frame_input = gr.Image(type="filepath", label="First Frame (Keyframe Awal - Wajib)")
919
+ last_frame_input = gr.Image(type="filepath", label="Last Frame (Keyframe Akhir - Opsional)")
920
+
921
+ prompt_input = gr.Textbox(
922
+ label="Prompt Teks (Opsional)",
923
+ placeholder="Deskripsikan aksi atau suasana video sinematik...",
924
+ lines=2,
925
+ value="cinematic camera movement, natural motion, hyper-detailed",
926
+ )
927
+
928
+ with gr.Group():
929
+ duration_input = gr.Radio(
930
+ choices=["5s", "10s", "15s"],
931
+ value="5s",
932
+ label="Durasi Video",
933
+ )
934
+
935
+ with gr.Row():
936
+ seed_input = gr.Number(value=0, label="Seed (0 = Random)", precision=0)
937
+ randomize_seed_input = gr.Checkbox(label="🎲 Randomize Seed Setiap Generate", value=True)
938
+
939
+ btn_generate = gr.Button("🚀 Generate Video (Dual-Space I2V)", variant="primary", elem_classes=["generate-btn"])
940
+
941
+ with gr.Column(scale=5):
942
+ video_output = gr.Video(label="Hasil Video MiniMax-H3 (dengan Audio)", autoplay=True)
943
+ report_output = gr.Markdown(label="Status & Benchmark Log")
944
+
945
+ btn_generate.click(
946
+ fn=generate_video_pipeline,
947
+ inputs=[
948
+ first_frame_input,
949
+ last_frame_input,
950
+ prompt_input,
951
+ duration_input,
952
+ seed_input,
953
+ randomize_seed_input,
954
+ ],
955
+ outputs=[video_output, report_output, seed_input],
956
+ )
957
+
958
+ if __name__ == "__main__":
959
+ # 1. Jalankan Gradio dengan prevent_thread_lock=True
960
+ fastapi_app, _, _ = demo.queue(max_size=20).launch(show_error=True, prevent_thread_lock=True, ssr_mode=False)
961
+
962
+ # 2. Daftarkan 100% Asynchronous REST API
963
+ fastapi_app.post("/api/generate")(api_submit_job)
964
+ fastapi_app.post("/")(api_submit_job)
965
+ fastapi_app.get("/api/status/{job_id}")(api_get_status)
966
+ fastapi_app.get("/api/result/{job_id}")(api_get_result)
967
+
968
+ # 3. Kunci main thread agar melayani request
969
+ demo.block_thread()
custom_nodes/external_h3_conditioning/__init__.py ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ from .nodes import NODE_CLASS_MAPPINGS, NODE_DISPLAY_NAME_MAPPINGS
2
+
3
+ __all__ = ["NODE_CLASS_MAPPINGS", "NODE_DISPLAY_NAME_MAPPINGS"]
custom_nodes/external_h3_conditioning/nodes.py ADDED
@@ -0,0 +1,93 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ import os
2
+ import json
3
+ import torch
4
+ import safetensors.torch as st
5
+ from safetensors import safe_open
6
+
7
+ class ExternalH3ConditioningLoader:
8
+ """
9
+ Membaca tensor CONDITIONING multimodal dari file .safetensors yang dihasilkan Space 1.
10
+ Merekonstruksi cond_embed, minimax_token_tags, serta minimax_keyframes (latents keyframe).
11
+ Mengembalikan tuple: (conditioning, width, height, length)
12
+ """
13
+ @classmethod
14
+ def INPUT_TYPES(cls):
15
+ return {
16
+ "required": {
17
+ "safetensors_path": ("STRING", {"default": ""}),
18
+ }
19
+ }
20
+
21
+ RETURN_TYPES = ("CONDITIONING", "INT", "INT", "INT")
22
+ RETURN_NAMES = ("conditioning", "width", "height", "length")
23
+ FUNCTION = "load"
24
+ CATEGORY = "conditioning/minimax_h3"
25
+
26
+ def load(self, safetensors_path: str):
27
+ if not safetensors_path or not os.path.exists(safetensors_path):
28
+ raise FileNotFoundError(f"File conditioning safetensors tidak ditemukan di: {safetensors_path}")
29
+
30
+ tensors = st.load_file(safetensors_path)
31
+ if "cond_embed" not in tensors:
32
+ raise KeyError("Tensor 'cond_embed' tidak ditemukan di dalam file safetensors.")
33
+
34
+ meta = {}
35
+ with safe_open(safetensors_path, framework="pt") as handle:
36
+ meta = handle.metadata() or {}
37
+
38
+ cond_embed = tensors["cond_embed"]
39
+ # Pastikan format [batch, seq_len, dim]
40
+ if cond_embed.dim() == 2:
41
+ cond_embed = cond_embed.unsqueeze(0)
42
+
43
+ extra_dict = {}
44
+
45
+ # 1. Restore token tags
46
+ if "minimax_token_tags" in tensors:
47
+ extra_dict["minimax_token_tags"] = tensors["minimax_token_tags"]
48
+
49
+ # 2. Restore keyframes
50
+ keyframes_meta_str = meta.get("keyframes_meta", "")
51
+ if keyframes_meta_str:
52
+ try:
53
+ raw_entries = json.loads(keyframes_meta_str)
54
+ restored_keyframes = []
55
+ for entry in raw_entries:
56
+ kf_dict = {
57
+ "resolved_frame_index": int(entry.get("resolved_frame_index", 0))
58
+ }
59
+ tensor_key = entry.get("tensor_key", "")
60
+ if tensor_key and tensor_key in tensors:
61
+ kf_dict["latent"] = tensors[tensor_key]
62
+ restored_keyframes.append(kf_dict)
63
+ if restored_keyframes:
64
+ extra_dict["minimax_keyframes"] = restored_keyframes
65
+ print(f"[*] [Space 2] Berhasil memulihkan {len(restored_keyframes)} keyframe(s) dari conditioning", flush=True)
66
+ except Exception as e:
67
+ print(f"[!] Warning: Gagal memulihkan keyframes_meta: {e}", flush=True)
68
+
69
+ # 3. Restore pooled output jika ada
70
+ if "cond_pooled" in tensors:
71
+ cond_pooled = tensors["cond_pooled"]
72
+ if cond_pooled.dim() == 1:
73
+ cond_pooled = cond_pooled.unsqueeze(0)
74
+ extra_dict["pooled_output"] = cond_pooled
75
+
76
+ conditioning = [
77
+ (cond_embed, extra_dict)
78
+ ]
79
+
80
+ width = int(meta.get("width", 896))
81
+ height = int(meta.get("height", 504))
82
+ length = int(meta.get("length", 97))
83
+
84
+ print(f"[*] [Space 2] Conditioning siap: shape {list(cond_embed.shape)}, resolusi {width}x{height}, frames {length}", flush=True)
85
+ return (conditioning, width, height, length)
86
+
87
+
88
+ NODE_CLASS_MAPPINGS = {
89
+ "ExternalH3ConditioningLoader": ExternalH3ConditioningLoader,
90
+ }
91
+ NODE_DISPLAY_NAME_MAPPINGS = {
92
+ "ExternalH3ConditioningLoader": "External H3 Conditioning Loader",
93
+ }
requirements.txt ADDED
@@ -0,0 +1,29 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ --extra-index-url https://download.pytorch.org/whl/cu130
2
+ --extra-index-url https://download.pytorch.org/whl/cu128
3
+ torch
4
+ torchvision
5
+ torchaudio
6
+ torchsde
7
+ spaces
8
+ gradio>=5,<6
9
+ gradio_client>=1.0.0
10
+ huggingface_hub>=0.34.0
11
+ transformers>=4.48.0
12
+ accelerate>=0.26.0
13
+ safetensors
14
+ einops
15
+ scipy
16
+ numpy
17
+ pillow
18
+ psutil
19
+ websocket-client
20
+ spandrel
21
+ kornia
22
+ av
23
+ color-matcher
24
+ matplotlib
25
+ mss
26
+ opencv-python-headless
27
+ imageio
28
+ imageio-ffmpeg
29
+ requests
workflow_generator.json ADDED
@@ -0,0 +1,246 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "ext_h3_cond": {
3
+ "inputs": {
4
+ "safetensors_path": ""
5
+ },
6
+ "class_type": "ExternalH3ConditioningLoader",
7
+ "_meta": {
8
+ "title": "External H3 Conditioning Loader"
9
+ }
10
+ },
11
+ "canvas_init": {
12
+ "inputs": {
13
+ "width": [
14
+ "ext_h3_cond",
15
+ 1
16
+ ],
17
+ "height": [
18
+ "ext_h3_cond",
19
+ 2
20
+ ],
21
+ "length": [
22
+ "ext_h3_cond",
23
+ 3
24
+ ]
25
+ },
26
+ "class_type": "EmptyMiniMaxH3LatentAV",
27
+ "_meta": {
28
+ "title": "Empty MiniMax H3 AV Latent"
29
+ }
30
+ },
31
+ "105_6": {
32
+ "inputs": {
33
+ "unet_name": "minimax_h3_fl2va_pruned_int8_convrot.safetensors",
34
+ "weight_dtype": "default"
35
+ },
36
+ "class_type": "UNETLoader",
37
+ "_meta": {
38
+ "title": "Load Diffusion Model"
39
+ }
40
+ },
41
+ "105_121": {
42
+ "inputs": {
43
+ "lora_name": "minimax_h3_taomate_3step_lora_avg_rank_19_bf16.safetensors",
44
+ "strength_model": 1,
45
+ "model": [
46
+ "105_6",
47
+ 0
48
+ ]
49
+ },
50
+ "class_type": "LoraLoaderModelOnly",
51
+ "_meta": {
52
+ "title": "Load LoRA"
53
+ }
54
+ },
55
+ "105_132": {
56
+ "inputs": {
57
+ "attention": "comfy kitchen attention",
58
+ "model": [
59
+ "105_121",
60
+ 0
61
+ ]
62
+ },
63
+ "class_type": "ModelAttentionBackend",
64
+ "_meta": {
65
+ "title": "Model Attention Backend"
66
+ }
67
+ },
68
+ "105_129": {
69
+ "inputs": {
70
+ "shift_video": 12,
71
+ "shift_audio": 3,
72
+ "model": [
73
+ "105_132",
74
+ 0
75
+ ]
76
+ },
77
+ "class_type": "MiniMaxH3SigmaShift",
78
+ "_meta": {
79
+ "title": "ModelSamplingMiniMaxH3"
80
+ }
81
+ },
82
+ "105_17": {
83
+ "inputs": {
84
+ "sampler_name": "res_multistep"
85
+ },
86
+ "class_type": "KSamplerSelect",
87
+ "_meta": {
88
+ "title": "KSamplerSelect"
89
+ }
90
+ },
91
+ "105_15": {
92
+ "inputs": {
93
+ "noise_seed": 42
94
+ },
95
+ "class_type": "RandomNoise",
96
+ "_meta": {
97
+ "title": "RandomNoise"
98
+ }
99
+ },
100
+ "105_9": {
101
+ "inputs": {
102
+ "scheduler": "simple",
103
+ "steps": 3,
104
+ "denoise": 1,
105
+ "model": [
106
+ "105_129",
107
+ 0
108
+ ]
109
+ },
110
+ "class_type": "BasicScheduler",
111
+ "_meta": {
112
+ "title": "BasicScheduler"
113
+ }
114
+ },
115
+ "105_16": {
116
+ "inputs": {
117
+ "model": [
118
+ "105_129",
119
+ 0
120
+ ],
121
+ "conditioning": [
122
+ "ext_h3_cond",
123
+ 0
124
+ ]
125
+ },
126
+ "class_type": "BasicGuider",
127
+ "_meta": {
128
+ "title": "Basic Guider"
129
+ }
130
+ },
131
+ "105_14": {
132
+ "inputs": {
133
+ "noise": [
134
+ "105_15",
135
+ 0
136
+ ],
137
+ "guider": [
138
+ "105_16",
139
+ 0
140
+ ],
141
+ "sampler": [
142
+ "105_17",
143
+ 0
144
+ ],
145
+ "sigmas": [
146
+ "105_9",
147
+ 0
148
+ ],
149
+ "latent_image": [
150
+ "canvas_init",
151
+ 0
152
+ ]
153
+ },
154
+ "class_type": "SamplerCustomAdvanced",
155
+ "_meta": {
156
+ "title": "SamplerCustomAdvanced"
157
+ }
158
+ },
159
+ "105_11": {
160
+ "inputs": {
161
+ "vae_name": "minimax_h3_video_vae_int8_convrot.safetensors"
162
+ },
163
+ "class_type": "VAELoader",
164
+ "_meta": {
165
+ "title": "Load Video VAE"
166
+ }
167
+ },
168
+ "105_24": {
169
+ "inputs": {
170
+ "vae_name": "minimax_h3_audio_vae_fp32.safetensors"
171
+ },
172
+ "class_type": "VAELoader",
173
+ "_meta": {
174
+ "title": "Load Audio VAE"
175
+ }
176
+ },
177
+ "105_10": {
178
+ "inputs": {
179
+ "samples": [
180
+ "105_14",
181
+ 0
182
+ ],
183
+ "vae": [
184
+ "105_11",
185
+ 0
186
+ ]
187
+ },
188
+ "class_type": "VAEDecode",
189
+ "_meta": {
190
+ "title": "VAE Decode Video"
191
+ }
192
+ },
193
+ "105_23": {
194
+ "inputs": {
195
+ "samples": [
196
+ "105_14",
197
+ 0
198
+ ],
199
+ "vae": [
200
+ "105_24",
201
+ 0
202
+ ]
203
+ },
204
+ "class_type": "VAEDecodeAudio",
205
+ "_meta": {
206
+ "title": "VAE Decode Audio"
207
+ }
208
+ },
209
+ "105_91": {
210
+ "inputs": {
211
+ "fps": 24,
212
+ "bit_depth": 8,
213
+ "color_space": "sRGB",
214
+ "codec": "none",
215
+ "images": [
216
+ "105_10",
217
+ 0
218
+ ],
219
+ "audio": [
220
+ "105_23",
221
+ 0
222
+ ]
223
+ },
224
+ "class_type": "CreateVideo",
225
+ "_meta": {
226
+ "title": "Create Video"
227
+ }
228
+ },
229
+ "92": {
230
+ "inputs": {
231
+ "filename_prefix": "video/MiniMax_H3",
232
+ "format": "auto",
233
+ "format.codec": "auto",
234
+ "codec": "auto",
235
+ "video-preview": "",
236
+ "video": [
237
+ "105_91",
238
+ 0
239
+ ]
240
+ },
241
+ "class_type": "SaveVideo",
242
+ "_meta": {
243
+ "title": "Save Video"
244
+ }
245
+ }
246
+ }