akatz-ai commited on
Commit
5823b1d
·
verified ·
1 Parent(s): b1905ac

Publish only the final 1000-step character-swap LoRA

Browse files
LICENSE ADDED
@@ -0,0 +1,84 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ MiniMax H3 COMMUNITY LICENSE AGREEMENT
2
+ MiniMax H3 release date/License date: August 2, 2026.
3
+ The scope of this License Agreement (this “Agreement”) is expressly limited to the “Applicable Territory” as defined below.
4
+ By clicking to accept, or by using, reproducing, modifying, distributing, running, or displaying any portion or element of the MiniMax H3 Works (including through any Hosted Services) in any manner, you acknowledge and accept the terms of this Agreement, and this Agreement shall take immediate effect upon the occurrence of such act.
5
+ I. Definitions
6
+ 1. “Acceptable Use Policy” means the policy published by MiniMax in Exhibit A.
7
+ 2. “Agreement” means the terms and conditions set forth herein that govern the use, reproduction, distribution, modification, running, and display of the MiniMax H3 Works or any portion or element thereof.
8
+ 3. “Applicable Territory” means worldwide, excluding the Excluded Territories.
9
+ 4. “Documentation” means the specifications, manuals, and documentation concerning MiniMax H3 that are publicly released by MiniMax.
10
+ 5. “Excluded Territories” means the European Union, the United Kingdom, the Republic of Korea and the United States of America.
11
+ 6. “MiniMax H3” means the video generation model, together with its software and algorithms, including trained model weights, parameters (including optimizer states), machine-learning model code, inference-supporting code, and other elements thereof made publicly available by Us, as released at https://huggingface.co/MiniMaxAI/MiniMax-H3.
12
+ 7. “MiniMax H3 Works” means (i) the Materials, (ii) the Model Derivatives, and (iii) all derivatives thereof.
13
+ 8. “Hosted Services” means hosted services provided via application programming interfaces (APIs), web access, or any other electronic or remote means.
14
+ 9. “Licensee,” “you,” or “your” means the natural or legal person exercising rights and/or using the MiniMax H3 Works for any purpose in any field of use under this Agreement.
15
+ 10. “Materials” means, collectively, MiniMax H3 and the Documentation (and any portion thereof), in each case as made available by MiniMax under this Agreement and proprietary to MiniMax.
16
+ 11. “Model Derivatives” means all of the following: (i) any modification of MiniMax H3 or any Model Derivative thereof; (ii) any work based on MiniMax H3 or any Model Derivative thereof; or (iii) any other machine learning model created by transferring the patterns of the weights, parameters, operational patterns, or Outputs of MiniMax H3 or any Model Derivative thereof to another model, such that the latter model exhibits behavior similar to MiniMax H3 or its Model Derivatives, including by distillation methods, methods using intermediate data representations, or methods based on training using synthetic-data Outputs generated by MiniMax H3 or its Model Derivatives. For the avoidance of doubt, Outputs are not deemed Model Derivatives.
17
+ 12. “Output” means any result of operating or otherwise using MiniMax H3 or any Model Derivatives (including through Hosted Services).
18
+ 13. “Third Party” means any natural or legal person that is not under common control with us or with you.
19
+ 14. “Including” means “including but not limited to.”
20
+ 15. “We,” “Us” or “MiniMax” means Nanonoble Pte. Ltd..
21
+ II. Grant of Rights
22
+ Solely within the Applicable Territory, we grant you a non-exclusive, non-transferable, royalty-free, limited license to use, reproduce, distribute, create derivative works (including Model Derivatives), and modify the Materials in accordance with the terms of this Agreement and the Acceptable Use Policy, based on the intellectual property and other rights owned by MiniMax that are embodied in or used by the Materials. You shall not violate (or encourage or permit any person to violate) any term of this Agreement or the Acceptable Use Policy.
23
+ We will continuously evaluate the applicable laws, regulations and compliance requirements for the Excluded Territories. In the meantime, should any person in such Excluded Territories be interested in deploying our models, you are welcome to contact us about obtaining a license, which will be granted based on robust controls and guardrails for purposes of complying with the laws, regulations and compliance requirements of the Excluded Territories.
24
+ III. Distribution and Redistribution
25
+ Subject to and conditioned on your continuing compliance with this Agreement, including its territorial restrictions and the Acceptable Use Policy, and solely within the Applicable Territory, you may distribute or make available the MiniMax H3 Works to Third Parties within the Applicable Territory; provided, that all of the following conditions are met:
26
+ 1. You must provide a copy of this Agreement to all such Third Parties who receive the MiniMax H3 Works or use your products or services related thereto;
27
+ 2. You must cause any modified files to carry prominent notices stating that you have modified such files;
28
+ 3. You are encouraged to:
29
+ a. display a notice on any product or service developed using MiniMax H3 indicating that the product or service is “Powered by MiniMax H3”;
30
+ b. add an AI-generation identifier to files produced using generative AI models including MiniMax H3; and
31
+ c. publish at least one technical blog post or a public statement describing your experience using MiniMax H3 Works;
32
+ 4. All distributions to Third Parties (other than through Hosted Services) must be accompanied by a “NOTICE” text file containing the following notice:
33
+ “MiniMax H3 is licensed under the MiniMax H3 Community License Agreement, Copyright © 2026 MiniMax. All Rights Reserved.”
34
+ You may add your own copyright notices on your modifications; except as provided in this Section and in Section V, however, you may not impose additional or different terms and conditions on the use, reproduction, or distribution of your modifications or of any aggregate Model Derivatives, and your use, reproduction, modification, distribution, running, and display of the work must otherwise comply with the terms and conditions of this Agreement (including the provisions concerning the Applicable Territory). If you receive the MiniMax H3 Works from a Licensee as part of an integrated end-user product, the provisions of Section III of this Agreement do not apply to you, but Section V and Exhibit A remain applicable.
35
+ IV. Additional Commercial Terms
36
+ 1. You shall obtain a separate, prior written authorization from MiniMax by contacting api@minimax.io with the subject line “MiniMax H3 licensing - authorization request”, if your commercial products and services generate more than 20 million US dollars (or equivalent in other currencies) in yearly revenue.
37
+ 2. You shall prominently display “MiniMax H3”on the user interface of commercial product or service that uses MiniMax H3 or MiniMax H3 Works.
38
+ V. Use Restrictions
39
+ 1. Your use of the MiniMax H3 Works must comply with applicable laws and regulations (including trade-compliance laws and regulations) and must comply with the Acceptable Use Policy for the MiniMax H3 Works, which is incorporated into this Agreement by reference.
40
+ 2. Before providing access to the MiniMax H3 Works or any product, service, or Hosted Service incorporating them, you must bind each recipient or user to enforceable terms at least as protective as the use restrictions in this Section V and Exhibit A, and you must notify each recipient or user that those restrictions apply.
41
+ 3. You may not use the MiniMax H3 Works or any of their Outputs or results to improve any other artificial intelligence model (other than MiniMax H3 or its Model Derivatives).
42
+ 4. You may not use, reproduce, modify, distribute, or display the MiniMax H3 Works or any of their Outputs or results outside the Applicable Territory. Any such use outside the Applicable Territory is not authorized by this Agreement.
43
+ 5. If you provide or make available to any Third Party a product, service, or Hosted Service that permits the generation of Outputs using MiniMax H3 or any Model Derivative, you must, before making that product or service available and throughout its operation, implement, maintain, test, and periodically review reasonable and proportionate technical and organizational safeguards designed to prevent and mitigate access, uses, and Outputs that violate this Section V or Exhibit A, including uses or Outputs that infringe, misappropriate, or otherwise violate any Third Party’s intellectual-property or other rights. You must not knowingly disable, materially weaken, or permit the circumvention of those safeguards. You must maintain a reasonably accessible mechanism for reporting suspected violations. Upon receiving a good-faith report or otherwise obtaining actual knowledge of a violation, you must promptly investigate and take reasonable steps within your control to stop or mitigate the violation, including removing or disabling access to offending content or services and suspending or terminating repeat violators where appropriate. You are responsible for implementing and enforcing these requirements with respect to your products, services, systems, users, and downstream recipients.
44
+ VI. Intellectual Property
45
+ 1. Subject to MiniMax’s rights in the MiniMax H3 Works (and the intellectual property therein), and to your compliance with the terms and conditions of this Agreement, as between you and MiniMax, you will own the derivative works and modifications of the Materials that you have created or had created, as well as any Model Derivatives.
46
+ 2. Except for the limited license expressly granted in this paragraph, no trademark license is granted under this Agreement; with respect to MiniMax H3 Works, the Licensee may not use any name or mark owned by or associated with MiniMax or any of its affiliates, except as reasonably and customarily necessary to describe and distribute the MiniMax H3 Works. MiniMax hereby grants you a license to use the “MiniMax H3” mark (the “Mark”) within the Applicable Territory solely for the purpose of complying with Section III.3; provided, that you comply with all applicable trademark-protection laws. All goodwill arising from your use of the Mark shall inure to the benefit of MiniMax.
47
+ 3. If you bring or assert any suit or other legal proceeding (including a cross-claim or counterclaim in any action) against us or any other natural or legal person alleging that the Materials, any Output, or any portion of the foregoing infringes any intellectual property right or other right owned by you or for which you can obtain a license, all licenses granted to you under this Agreement will terminate as of the date such suit or proceeding is filed. You shall defend, indemnify, and hold us harmless against any Third-Party claim arising out of or related to the use or distribution of the MiniMax H3 Works by you or by any Third Party.
48
+ 4. MiniMax claims no rights over the Outputs you generate. You and your users are entirely responsible for the Outputs and any subsequent use thereof.
49
+ VII. Disclaimers and Limitations of Liability
50
+ 1. We have no obligation to support, update, provide training for, or develop any further version of the MiniMax H3 Works, or to grant any license with respect thereto.
51
+ 2. UNLESS AND ONLY TO THE EXTENT REQUIRED BY APPLICABLE LAW, THE MINIMAX H3 WORKS AND ANY OUTPUT AND RESULTS THEREFROM ARE PROVIDED “AS IS” WITHOUT ANY EXPRESS OR IMPLIED WARRANTIES OF ANY KIND INCLUDING ANY WARRANTIES OF TITLE, MERCHANTABILITY, NONINFRINGEMENT, COURSE OF DEALING, USAGE OF TRADE, OR FITNESS FOR A PARTICULAR PURPOSE. YOU ARE SOLELY RESPONSIBLE FOR DETERMINING THE APPROPRIATENESS OF USING, REPRODUCING, MODIFYING, PERFORMING, DISPLAYING OR DISTRIBUTING ANY OF THE MINIMAX H3 WORKS OR OUTPUTS AND ASSUME ANY AND ALL RISKS ASSOCIATED WITH YOUR OR A THIRD PARTY’S USE OR DISTRIBUTION OF ANY OF THE MINIMAX H3 WORKS OR OUTPUTS AND YOUR EXERCISE OF RIGHTS AND PERMISSIONS UNDER THIS AGREEMENT.
52
+ 3. TO THE FULLEST EXTENT PERMITTED BY APPLICABLE LAW, IN NO EVENT SHALL MINIMAX OR ITS AFFILIATES BE LIABLE UNDER ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, TORT, NEGLIGENCE, PRODUCTS LIABILITY, OR OTHERWISE, FOR ANY DAMAGES, INCLUDING ANY DIRECT, INDIRECT, SPECIAL, INCIDENTAL, EXEMPLARY, CONSEQUENTIAL OR PUNITIVE DAMAGES, OR LOST PROFITS OF ANY KIND ARISING FROM THIS AGREEMENT OR RELATED TO ANY OF THE MINIMAX H3 WORKS OR OUTPUTS, EVEN IF MINIMAX OR ITS AFFILIATES HAVE BEEN ADVISED OF THE POSSIBILITY OF ANY OF THE FOREGOING.
53
+ VIII. Term and Termination
54
+ 1. This Agreement is effective from the moment you accept this Agreement or begin accessing the Materials, and, subject to your compliance with its terms and conditions, will remain in effect until terminated as provided herein.
55
+ 2. If you breach any term or condition of this Agreement, we have the right to terminate this Agreement. Upon termination, you must immediately cease accessing, using, and distributing the MiniMax H3 Works; delete or destroy all copies within your possession or control; and notify each downstream recipient that your authorization has ended. The obligations in the preceding sentence and Sections VI.1, VI.3, VII, and IX survive termination.
56
+ IX. Governing Law and Jurisdiction
57
+ 1. This Agreement, and any dispute arising out of or related to this Agreement, shall be governed by the laws of the Hong Kong Special Administrative Region of the People’s Republic of China, without regard to its conflict-of-laws rules. The United Nations Convention on Contracts for the International Sale of Goods does not apply to this Agreement.
58
+ 2. Any dispute arising out of or related to this Agreement shall be subject to the exclusive jurisdiction of the courts of the Hong Kong Special Administrative Region of the People’s Republic of China with competent jurisdiction. Both MiniMax and the Licensee hereby consent to the exclusive jurisdiction of such courts for any such dispute.
59
+ Additional Note: Please note that the encoder of MiniMax H3 uses Qwen3-VL-32B, which is licensed under Apache 2.0 License: https://github.com/QwenLM/Qwen3-VL/blob/main/LICENSE.
60
+
61
+ Exhibit A — Acceptable Use Policy
62
+ MiniMax reserves the right to update this Acceptable Use Policy from time to time.
63
+ Last revised: August 2, 2026.
64
+ MiniMax is committed to promoting the safe and fair use of its tools and features, including MiniMax H3. You agree not to use MiniMax H3, any Model Derivatives, or any Output in any of the following ways:
65
+ 1. Use outside the Applicable Territory;
66
+ 2. Use in any manner that violates any applicable national, federal, state, local, or international law, regulation, or other legal requirement, or that infringes, misappropriates, or otherwise violates any Third Party’s intellectual-property or other proprietary rights, including through unauthorized reproduction, distribution, public display, public performance, or creation of derivative works;
67
+ 3. Use in any manner that may harm yourself or others;
68
+ 4. Use to repurpose or distribute the Outputs of MiniMax H3 or any Model Derivatives in order to harm yourself or others;
69
+ 5. Use to circumvent or bypass any safety guardrails or safeguards we have implemented;
70
+ 6. Use in any manner that exploits or harms, or intends to exploit or harm, minors;
71
+ 7. Use to generate or disseminate verifiably false information and/or content for the purpose of harming others or influencing elections;
72
+ 8. Use to manufacture or facilitate false online engagement, including fake reviews and other means of false online engagement;
73
+ 9. Use to intentionally defame, disparage, or otherwise harass others;
74
+ 10. Use to generate and/or disseminate malware (including ransomware) or any other content intended to damage electronic systems;
75
+ 11. Use to generate or disseminate personally identifiable information for the purpose of harming others;
76
+ 12. Use to generate or disseminate information (including images, code, posts, or articles) in or to any public environment (including via bot tweets or similar means) without clearly and prominently disclosing that such information and/or content is machine-generated;
77
+ 13. Use to impersonate another person without that person’s consent, authorization, or lawful right to do so;
78
+ 14. Use to make high-risk automated decisions in critical domains that affect individual safety, rights, or well-being (such as law enforcement, immigration, healthcare or medical services, critical-infrastructure management, product-safety components, essential services, credit, employment, housing, education, social scoring, or insurance);
79
+ 15. Use in any manner that violates or disregards the social, ethical, or moral standards of other countries or regions;
80
+ 16. Use to carry out, assist, threaten, incite, plan, advocate for, or encourage violent extremism or terrorism;
81
+ 17. Use for any purpose intended to discriminate against, or harm, individuals or groups based on protected characteristics or categories, online or offline social behavior, or known or predicted personality traits;
82
+ 18. Use to intentionally exploit the vulnerabilities of specific populations based on age, social, physical, or psychological characteristics, so as to materially distort the behavior of a member of that group in a manner that causes, or is likely to cause, physical or psychological harm to that person or to others;
83
+ 19. Use for military purposes;
84
+ 20. Use to engage in any unauthorized or unlicensed professional activity, including but not limited to financial, legal, medical or healthcare, or other professional practice.
NOTICE ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ MiniMax H3 is licensed under the MiniMax H3 Community License Agreement, Copyright © 2026 MiniMax. All Rights Reserved.
2
+
3
+ Modified by Akatz Labs: character-swap LoRA fine-tuning, 1,000 updates, September 2026. This repository contains adapter weights, not base-model weights.
README.md ADDED
@@ -0,0 +1,87 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ language:
3
+ - en
4
+ license: other
5
+ license_name: minimax-h3-community-license-agreement
6
+ license_link: LICENSE
7
+ base_model: Comfy-Org/MiniMax-H3
8
+ base_model_relation: adapter
9
+ pipeline_tag: video-to-video
10
+ library_name: diffusion-single-file
11
+ datasets:
12
+ - akatz-ai/H3-Character-Swap-v1
13
+ tags:
14
+ - minimax-h3
15
+ - lora
16
+ - character-swap
17
+ - video-editing
18
+ - ref2va
19
+ - comfyui
20
+ - ai-toolkit
21
+ ---
22
+
23
+ # MiniMax H3 Character Swap LoRA v1
24
+
25
+ An experimental character-replacement LoRA trained by Akatz Labs for **1,000 updates**. It aims to replace a selected character using an image reference while preserving the source scene. In our local comparisons it often preserved the background more closely than the base model, but motion timing, facial expressions, and hard cuts remain unreliable.
26
+
27
+ **Download:** [final 1,000-step LoRA](h3_character_swap_pro4500_1000.safetensors). Only the final 1,000-step checkpoint is published. Start with the final checkpoint at **strength 1.0**. This is an adapter, not a standalone model or a Turbo distillation LoRA.
28
+
29
+ **Dataset:** [H3 Character Swap v1](https://huggingface.co/datasets/akatz-ai/H3-Character-Swap-v1).
30
+
31
+ ## Use
32
+
33
+ 1. Use an H3 Ref2VA-capable runtime and obtain the base model and VAEs separately.
34
+ 2. Place the final `.safetensors` file in `ComfyUI/models/loras/` (or your configured shared LoRA directory), and apply it to the H3 model at strength 1.0 using a compatible model-only LoRA loader.
35
+ 3. Supply the source video as `<Video 1>` and replacement image or character sheet as `<Picture 1>`.
36
+ 4. Specify the target person in the prompt. No additional trigger word was trained.
37
+
38
+ Example:
39
+
40
+ > Replace only the man in the purple shirt in <Video 1> with the character in <Picture 1>. Keep the replacement character's identity, outfit, and art style from <Picture 1>. Preserve the source video's camera, background, lighting, objects, and all other people. Match the target person's position, scale, pose, and movement. Do not show the reference sheet or its background.
41
+
42
+ The training captions were shorter, for example `Swap the man in the purple shirt in <Video 1> with the character in <Picture 1>.` Preservation instructions helped some local evaluations, but stronger expression instructions sometimes suppressed the swap entirely. Prompt wording is not a guarantee of strict source alignment.
43
+
44
+ Short, continuous shots of roughly 4–5 seconds were more promising than our full 14-second tests. A precise maximum duration has **not** been established. Use 24 fps and your runtime's supported H3 frame grid. The character-swap LoRA does not require a Turbo LoRA, Spectrum, or Sol attention.
45
+
46
+ ## Base model and tested configurations
47
+
48
+ Training used `minimax_h3_ref2va_pruned_int8_convrot.safetensors` from [Comfy-Org/MiniMax-H3](https://huggingface.co/Comfy-Org/MiniMax-H3), plus the frozen [Ostris Ref2VA training assistant](https://huggingface.co/ostris/minimax_h3_training_adapter). The assistant and base weights are **not merged into this adapter** and are not distributed here. Model revisions and hashes are in `training/base-model-files.json`.
49
+
50
+ We evaluated on the Ref2VA base and a local FL2VA/Ref2VA hybrid (blocks 25–49, INT8). A later experimental configuration combined this character-swap adapter with a separate 768p Turbo 8-step LoRA, `res_multistep` / `simple`, and native Sol attention. Those are evaluation choices, not the training base or universal compatibility claims. The local hybrid is not bundled. The text-to-video-only FastH3 experiment did not demonstrate a useful replacement workflow.
51
+
52
+ ## Training record
53
+
54
+ | Setting | Recorded value |
55
+ | --- | --- |
56
+ | Updates / saves | 1,000 / every 250 updates |
57
+ | Hardware | RunPod RTX PRO 4500 Blackwell, 32 GB |
58
+ | LoRA rank / alpha | 16 / 16, excluding `adaln_proj` |
59
+ | Optimizer / learning rate | AdamW8bit / 5e-5 |
60
+ | Batch / accumulation | 1 / 1 |
61
+ | Precision | BF16, convrot8 transformer, NVFP4 text encoder |
62
+ | Edit target resolution | Area budget 1024; 1344×768 buckets |
63
+ | Video regularization resolution | Reduced area budget 384 |
64
+ | Regularization duration | 73 frames at 24 fps, approximately 3.04 seconds |
65
+ | Memory measures | Gradient checkpointing, layer offload, cached latents/text, chunked MLP |
66
+ | Sampling during training | Disabled |
67
+
68
+ The dataset has 94 synthetic image-edit triplets and 40 unchanged video/audio examples. Optimization used **76 edits and 32 regularization clips**; 18 edits and 8 clips were held out. The edit targets are single still images, with five-frame static source-video controls. This is **not** training on long moving character-swap targets. AI Toolkit interleaves regularization, so file-count ratios are not update ratios. The data include cross-style swaps and varied character sheets, but only one-character replacement targets.
69
+
70
+ The author estimates the overnight rental at around $11; this is not a measured cost benchmark. `training/train-1000-noeval.json` records the run configuration, with inactive local evaluation paths removed. The launcher and package snapshot are supplied for reference. The copied AI Toolkit source was not a Git checkout, so the embedded version label is not an exact source revision. See the dataset's runtime-contract fingerprints.
71
+
72
+ ## Findings and limitations
73
+
74
+ - Local side-by-side reviews suggest improved scene/background preservation relative to the base model. These are qualitative observations, not a benchmark score.
75
+ - Long windows can drift in framing, placement, or source timing. Hard cuts can become zooms or gradual repositioning.
76
+ - Close-up facial expressions may not match the original performance; added expression prompting was not consistently helpful.
77
+ - Two-character inference was tested, but multi-character replacements were not supervised in the training targets.
78
+ - Short-window continuation improved some joins but did not guarantee camera-cut timing or full source adherence.
79
+ - Generated audio was closer in some early LoRA comparisons but skipped or drifted in the continuation test. The later review used original source audio copied onto the generated video. That audio preservation is **postprocessing**, not proof of the LoRA's audio fidelity; remuxing does not repair lip-sync drift.
80
+
81
+ Next experiments include aligned moving swap targets, explicit hard-cut examples, expression supervision, and longer sequences. More VRAM alone is not a demonstrated quality fix.
82
+
83
+ ## License and attribution
84
+
85
+ This adapter is distributed under the [MiniMax H3 Community License Agreement](LICENSE), including its acceptable-use, distribution, commercial, and territorial provisions. **It is not Apache-2.0.** The agreement excludes the US, EU, UK, and Republic of Korea from its standard territorial grant and describes separate authorization. Read the complete upstream terms; this repository does not expand them. See NOTICE for attribution and the modification notice.
86
+
87
+ The separately published dataset uses Apache-2.0 for Akatz Labs contributions and preserves the upstream VidGen license/provenance. Its license does not replace this model's terms.
SHA256SUMS ADDED
@@ -0,0 +1,8 @@
 
 
 
 
 
 
 
 
 
1
+ 59b99642b95ea21630e311198ddbfffbfe05aadba0c2f5d884cbdf4efcc90f44 LICENSE
2
+ a48642bdfdf17dc2e22b08c737b6fdf37cd0d5eb3f0641a5cb30bd7fcbc88378 NOTICE
3
+ 54ae4176209198aa1d632a8195d0c5f6d2083b364c1047d60191f7cdd6e59468 README.md
4
+ 4b2a3f420ae804c0aa3422761ff84dbd1bf52eef6900ffab6d2e66df63cb4e79 h3_character_swap_pro4500_1000.safetensors
5
+ a4871c95fb7f74d88ff60c4cbb82cf05c65bc309f684efefb1572f87b25e4581 training/base-model-files.json
6
+ 79d4ebdbb1195e882629adbc7ec27640bd4a42bc24ae5a2961474667992b2e5d training/requirements-pinned.txt
7
+ 4aabe5d9758a926787eca28fe1bec142ca78160ba25b41d0dc5ab4b437c5e5ca training/train-1000-noeval.json
8
+ f574f99af8a96af93ddf8967f441495ccc7ea21135d89e3b9f80671cd4097a3d training/train-launcher.py
h3_character_swap_pro4500_1000.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:4b2a3f420ae804c0aa3422761ff84dbd1bf52eef6900ffab6d2e66df63cb4e79
3
+ size 155110320
training/base-model-files.json ADDED
@@ -0,0 +1,42 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ [
2
+ {
3
+ "repo": "Comfy-Org/MiniMax-H3",
4
+ "revision": "1c41cfca8ebba91d0af792a05c5761fb2c3a7975",
5
+ "filename": "diffusion_models/minimax_h3_ref2va_pruned_int8_convrot.safetensors",
6
+ "destination": "diffusion_models/minimax_h3_ref2va_pruned_int8_convrot.safetensors",
7
+ "size": 20970379616,
8
+ "sha256": "9255f52b6677845ad238f20dfaafa94727053694127ab7f255c048f0f9365779"
9
+ },
10
+ {
11
+ "repo": "Comfy-Org/MiniMax-H3",
12
+ "revision": "1c41cfca8ebba91d0af792a05c5761fb2c3a7975",
13
+ "filename": "text_encoders/qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors",
14
+ "destination": "text_encoders/qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors",
15
+ "size": 15687142551,
16
+ "sha256": "35a88d51044231fe332301d7a62aa81e3f2cba62febeb446e2c1e3e0ef76f2c6"
17
+ },
18
+ {
19
+ "repo": "Comfy-Org/MiniMax-H3",
20
+ "revision": "1c41cfca8ebba91d0af792a05c5761fb2c3a7975",
21
+ "filename": "vae/minimax_h3_video_vae_fp16.safetensors",
22
+ "destination": "vae/minimax_h3_video_vae_fp16.safetensors",
23
+ "size": 5207808496,
24
+ "sha256": "7c1f131492e7eddacaac9069a61b81bdd39de5cc96561e677c5eab1cdce5e522"
25
+ },
26
+ {
27
+ "repo": "Comfy-Org/MiniMax-H3",
28
+ "revision": "1c41cfca8ebba91d0af792a05c5761fb2c3a7975",
29
+ "filename": "vae/minimax_h3_audio_vae_fp32.safetensors",
30
+ "destination": "vae/minimax_h3_audio_vae_fp32.safetensors",
31
+ "size": 605254808,
32
+ "sha256": "8e505d95dd1561d47abd43d4238fd40d9bb1ae9e147ed0a4cba778d76ae4db48"
33
+ },
34
+ {
35
+ "repo": "ostris/minimax_h3_training_adapter",
36
+ "revision": "b2e664cb0bba5d96f3ab02fc142c825d4b2409e1",
37
+ "filename": "minimax_h3_ref2va_training_adapter_v1.safetensors",
38
+ "destination": "loras/minimax_h3_ref2va_training_adapter_v1.safetensors",
39
+ "size": 155110344,
40
+ "sha256": "79611559f85712b9ed3bd8a4009d07294dcf88b62ac566051128cfb65d87e0de"
41
+ }
42
+ ]
training/requirements-pinned.txt ADDED
@@ -0,0 +1,158 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ImageIO==2.37.4
2
+ Jinja2==3.1.6
3
+ Markdown==3.10.3
4
+ MarkupSafe==3.0.3
5
+ PyWavelets==1.10.0
6
+ PyYAML==6.0.3
7
+ Pygments==2.21.0
8
+ Werkzeug==3.1.8
9
+ absl-py==2.5.0
10
+ accelerate==1.15.0
11
+ albucore==0.0.16
12
+ albumentations==1.4.15
13
+ annotated-doc==0.0.5
14
+ annotated-types==0.8.0
15
+ antlr4-python3-runtime==4.9.3
16
+ anyio==4.15.1
17
+ audioread==3.1.0
18
+ av==16.0.1
19
+ bitsandbytes==0.50.2
20
+ brotli==1.2.0
21
+ certifi==2026.7.22
22
+ cffi==2.1.1
23
+ charset-normalizer==3.5.1
24
+ click==8.5.0
25
+ cloudpickle==3.1.2
26
+ contourpy==1.3.3
27
+ controlnet_aux==0.0.10
28
+ cuda-bindings==12.9.7
29
+ cuda-pathfinder==1.6.0
30
+ cuda-toolkit==12.8.1
31
+ cycler==0.12.1
32
+ decorator==5.3.1
33
+ einops==0.8.2
34
+ eval_type_backport==0.4.0
35
+ fastapi==0.141.1
36
+ filelock==3.32.3
37
+ flatten-json==0.1.14
38
+ fonttools==4.65.0
39
+ fsspec==2026.7.0
40
+ ftfy==6.3.1
41
+ gradio==6.27.0
42
+ gradio_client==2.7.0
43
+ groovy==0.1.2
44
+ grpcio==1.84.0
45
+ h11==0.16.0
46
+ hf-gradio==0.4.1
47
+ hf-xet==1.6.0
48
+ hf_transfer==0.1.9
49
+ httpcore==1.0.9
50
+ httpx==0.28.1
51
+ huggingface_hub==1.23.0
52
+ idna==3.19
53
+ importlib_metadata==9.0.1
54
+ invisible-watermark==0.2.0
55
+ joblib==1.6.0
56
+ kiwisolver==1.5.1
57
+ kornia==0.8.3
58
+ kornia_rs==0.1.14
59
+ lazy-loader==0.5
60
+ librosa==0.11.0
61
+ llvmlite==0.49.0
62
+ lpips==0.1.4
63
+ lycoris_lora==1.8.3
64
+ markdown-it-py==4.2.0
65
+ matplotlib==3.10.1
66
+ mdurl==0.1.2
67
+ mpmath==1.3.0
68
+ msgpack==1.2.2
69
+ mutagen==1.47.0
70
+ narwhals==2.26.0
71
+ networkx==3.6.1
72
+ ninja==1.13.2
73
+ numba==0.67.0
74
+ numpy==1.26.4
75
+ nvidia-cublas-cu12==12.8.4.1
76
+ nvidia-cuda-cupti-cu12==12.8.90
77
+ nvidia-cuda-nvrtc-cu12==12.8.93
78
+ nvidia-cuda-runtime-cu12==12.8.90
79
+ nvidia-cudnn-cu12==9.19.0.56
80
+ nvidia-cufft-cu12==11.3.3.83
81
+ nvidia-cufile-cu12==1.13.1.3
82
+ nvidia-curand-cu12==10.3.9.90
83
+ nvidia-cusolver-cu12==11.7.3.90
84
+ nvidia-cusparse-cu12==12.5.8.93
85
+ nvidia-cusparselt-cu12==0.7.1
86
+ nvidia-nccl-cu12==2.28.9
87
+ nvidia-nvjitlink-cu12==12.8.93
88
+ nvidia-nvshmem-cu12==3.4.5
89
+ nvidia-nvtx-cu12==12.8.90
90
+ omegaconf==2.3.1
91
+ open_clip_torch==3.3.0
92
+ opencv-python-headless==4.11.0.86
93
+ opencv-python==4.11.0.86
94
+ optimum-quanto==0.2.4
95
+ orjson==3.12.0
96
+ oyaml==1.0
97
+ packaging==26.3
98
+ pandas==3.0.5
99
+ peft==0.18.1
100
+ pillow==12.3.0
101
+ platformdirs==4.11.8
102
+ pooch==1.9.0
103
+ prodigyopt==1.1.2
104
+ protobuf==7.36.1
105
+ psutil==7.2.2
106
+ pycparser==3.0
107
+ pydantic==2.13.5
108
+ pydantic_core==2.46.5
109
+ pydub==0.25.1
110
+ pyparsing==3.3.2
111
+ python-dateutil==2.9.0.post0
112
+ python-dotenv==1.2.3
113
+ python-multipart==0.0.32
114
+ python-slugify==9.0.0
115
+ pytorch-fid==0.3.0
116
+ pytorch-wavelets==1.3.0
117
+ pytz==2026.3.post1
118
+ regex==2026.9.10
119
+ requests==2.34.2
120
+ rich==15.0.0
121
+ safehttpx==0.1.7
122
+ safetensors==0.9.0rc0
123
+ scikit-image==0.26.0
124
+ scikit-learn==1.9.1
125
+ scipy==1.12.0
126
+ semantic-version==2.10.0
127
+ sentencepiece==0.2.2
128
+ setuptools==78.1.0
129
+ shellingham==1.5.4
130
+ six==1.17.0
131
+ soundfile==0.14.0
132
+ soxr==1.1.0
133
+ starlette==1.6.0
134
+ sympy==1.14.0
135
+ tensorboard-data-server==0.7.2
136
+ tensorboard==2.21.0
137
+ text-unidecode==1.3
138
+ threadpoolctl==3.6.0
139
+ tifffile==2026.3.3
140
+ timm==1.0.22
141
+ tokenizers==0.22.2
142
+ toml==0.10.2
143
+ tomlkit==0.14.0
144
+ torch==2.11.0+cu128
145
+ torchao==0.10.0
146
+ torchaudio==2.11.0+cu128
147
+ torchcodec==0.11.0+cpu
148
+ torchvision==0.26.0+cu128
149
+ tqdm==4.70.1
150
+ transformers==5.5.3
151
+ triton==3.6.0
152
+ typer==0.27.2
153
+ typing-inspection==0.4.4
154
+ typing_extensions==4.16.0
155
+ urllib3==2.7.0
156
+ uvicorn==0.53.0
157
+ wcwidth==0.8.3
158
+ zipp==4.1.0
training/train-1000-noeval.json ADDED
@@ -0,0 +1,145 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "job": "extension",
3
+ "config": {
4
+ "name": "h3_character_swap_pro4500_1000",
5
+ "process": [
6
+ {
7
+ "type": "sd_trainer",
8
+ "training_folder": "${H3_OUTPUT_ROOT}",
9
+ "device": "cuda:0",
10
+ "network": {
11
+ "type": "lora",
12
+ "linear": 16,
13
+ "linear_alpha": 16,
14
+ "network_kwargs": {
15
+ "ignore_if_contains": [
16
+ "adaln_proj"
17
+ ]
18
+ }
19
+ },
20
+ "save": {
21
+ "dtype": "bf16",
22
+ "save_every": 250,
23
+ "max_step_saves_to_keep": 4,
24
+ "push_to_hub": false
25
+ },
26
+ "datasets": [
27
+ {
28
+ "caption_ext": "txt",
29
+ "default_caption": "",
30
+ "resolution": 1024,
31
+ "fps": 24,
32
+ "buckets": true,
33
+ "cache_latents_to_disk": true,
34
+ "caption_dropout_rate": 0,
35
+ "random_crop": false,
36
+ "flip_x": false,
37
+ "flip_y": false,
38
+ "num_workers": 0,
39
+ "cache_latents_num_workers": 1,
40
+ "full_size_control_images": true,
41
+ "shrink_video_to_frames": true,
42
+ "folder_path": "${H3_DATASET_ROOT}/data/edits/train/targets",
43
+ "control_path": [
44
+ "${H3_DATASET_ROOT}/data/edits/train/character_references",
45
+ "${H3_DATASET_ROOT}/data/edits/train/scene_videos"
46
+ ],
47
+ "num_frames": 1,
48
+ "auto_frame_count": true,
49
+ "trim_auto_frame_count_tail": true,
50
+ "do_audio": false
51
+ },
52
+ {
53
+ "caption_ext": "txt",
54
+ "default_caption": "",
55
+ "resolution": 384,
56
+ "fps": 24,
57
+ "buckets": true,
58
+ "cache_latents_to_disk": true,
59
+ "caption_dropout_rate": 0,
60
+ "random_crop": false,
61
+ "flip_x": false,
62
+ "flip_y": false,
63
+ "num_workers": 0,
64
+ "cache_latents_num_workers": 1,
65
+ "full_size_control_images": true,
66
+ "shrink_video_to_frames": true,
67
+ "folder_path": "${H3_DATASET_ROOT}/data/regularization/train/videos",
68
+ "control_path": "${H3_DATASET_ROOT}/data/regularization/train/videos",
69
+ "num_frames": 73,
70
+ "auto_frame_count": false,
71
+ "do_audio": true,
72
+ "is_reg": true,
73
+ "audio_normalize": false
74
+ }
75
+ ],
76
+ "train": {
77
+ "batch_size": 1,
78
+ "steps": 1000,
79
+ "gradient_accumulation_steps": 1,
80
+ "train_unet": true,
81
+ "train_text_encoder": false,
82
+ "cache_text_embeddings": true,
83
+ "skip_first_sample": true,
84
+ "gradient_checkpointing": true,
85
+ "noise_scheduler": "flowmatch",
86
+ "timestep_type": "shift",
87
+ "optimizer": "adamw8bit",
88
+ "optimizer_params": {
89
+ "weight_decay": 0.0001
90
+ },
91
+ "lr": 5e-05,
92
+ "weight_decay": 0.0001,
93
+ "dtype": "bf16",
94
+ "do_guidance_loss": true,
95
+ "guidance_loss_target": 3.5,
96
+ "disable_sampling": true
97
+ },
98
+ "model": {
99
+ "name_or_path": "Comfy-Org/MiniMax-H3",
100
+ "arch": "minimax_h3_ref2va",
101
+ "quantize": true,
102
+ "qtype": "convrot8",
103
+ "quantize_te": true,
104
+ "qtype_te": "nvfp4",
105
+ "low_vram": true,
106
+ "layer_offloading": true,
107
+ "layer_offloading_transformer_percent": 1.0,
108
+ "layer_offloading_text_encoder_percent": 1.0,
109
+ "assistant_lora_path": "${H3_MODEL_ROOT}/loras/minimax_h3_ref2va_training_adapter_v1.safetensors",
110
+ "model_kwargs": {
111
+ "image_refs_as_video": false,
112
+ "image_ref_video_frames": 5,
113
+ "partition": "ref2va_pruned",
114
+ "dit_ref2va_pruned_path": "${H3_MODEL_ROOT}/diffusion_models/minimax_h3_ref2va_pruned_int8_convrot.safetensors",
115
+ "text_encoder_path": "${H3_MODEL_ROOT}/text_encoders/qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors",
116
+ "video_vae_path": "${H3_MODEL_ROOT}/vae/minimax_h3_video_vae_fp16.safetensors",
117
+ "audio_vae_path": "${H3_MODEL_ROOT}/vae/minimax_h3_audio_vae_fp32.safetensors",
118
+ "sample_audio": false
119
+ }
120
+ },
121
+ "sample": {
122
+ "neg": "",
123
+ "sampler": "flowmatch",
124
+ "sample_every": 250,
125
+ "sample_start_step": 0,
126
+ "width": 1344,
127
+ "height": 768,
128
+ "num_frames": 73,
129
+ "fps": 24,
130
+ "guidance_scale": 1,
131
+ "sample_steps": 28,
132
+ "sample_audio": true,
133
+ "seed": 904231,
134
+ "samples": [],
135
+ "ext": "mp4",
136
+ "walk_seed": false
137
+ }
138
+ }
139
+ ]
140
+ },
141
+ "meta": {
142
+ "name": "h3-character-swap-v1",
143
+ "version": "1.0.0"
144
+ }
145
+ }
training/train-launcher.py ADDED
@@ -0,0 +1,71 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ import os,sys,runpy,traceback,json,time
2
+ config_path=os.path.abspath(sys.argv[1])
3
+ root=os.environ['AI_TOOLKIT_ROOT']
4
+ os.chdir(root);sys.path.insert(0,root)
5
+ from extensions_built_in.sd_trainer.SDTrainer import SDTrainer
6
+ import torch
7
+ from torch.utils.checkpoint import checkpoint
8
+ from extensions_built_in.diffusion_models.minimax_h3.src.transformer import MiniMaxH3Mlp
9
+ mlp_forward=MiniMaxH3Mlp.forward
10
+ def chunked_mlp(self,x):
11
+ if x.shape[-2] <= 1024:
12
+ return mlp_forward(self,x)
13
+ return torch.cat([checkpoint(mlp_forward,self,c,use_reentrant=False) if torch.is_grad_enabled() else mlp_forward(self,c) for c in x.split(1024,dim=-2)],dim=-2)
14
+ # Check values and gradients on the tokenwise operation before loading weights.
15
+ torch.manual_seed(123)
16
+ m=MiniMaxH3Mlp(16,32).double()
17
+ a=torch.randn(1,2051,16,dtype=torch.float64,requires_grad=True)
18
+ b=a.detach().clone().requires_grad_(True)
19
+ y=mlp_forward(m,a); y.square().sum().backward()
20
+ expected=[p.grad.clone() for p in m.parameters()]; m.zero_grad()
21
+ z=chunked_mlp(m,b); z.square().sum().backward()
22
+ assert torch.allclose(y,z,atol=1e-10,rtol=1e-10)
23
+ assert torch.allclose(a.grad,b.grad,atol=1e-10,rtol=1e-10)
24
+ assert all(torch.allclose(e,p.grad,atol=1e-9,rtol=1e-9) for e,p in zip(expected,m.parameters()))
25
+ print('CHUNK_PARITY_PASS: outputs, input gradients, parameter gradients',flush=True)
26
+ del m,a,b,y,z,expected
27
+ MiniMaxH3Mlp.forward=chunked_mlp
28
+
29
+ # Freeze evaluation conditioning with the full 73-frame video context.
30
+ # Upstream cache_sample_prompts constructs GenerateImageConfig without num_frames,
31
+ # which silently limits reference video conditioning to five frames.
32
+ from PIL import Image
33
+ from toolkit.config_modules import GenerateImageConfig
34
+ def cache_eval_prompts(self):
35
+ if self.train_config.disable_sampling:
36
+ return
37
+ vae_device=self.sd.vae.device
38
+ self.sd.vae.to("cpu")
39
+ torch.cuda.empty_cache()
40
+ self.sd.sample_prompts_cache=[]
41
+ for item in self.sample_config.samples:
42
+ gen=GenerateImageConfig(prompt=item.prompt,width=item.width,height=item.height,num_frames=item.num_frames,fps=item.fps,ctrl_img_1=item.ctrl_img_1,ctrl_img_2=item.ctrl_img_2,output_path='/tmp/h3-unused.mp4')
43
+ self.sd.prepare_sample_prompt_context(gen)
44
+ with torch.no_grad():
45
+ emb=self.sd.get_prompt_embeds(item.prompt,control_images=[Image.open(item.ctrl_img_1).convert('RGB'),item.ctrl_img_2]).to('cpu')
46
+ self.sd.sample_prompts_cache.append({'conditional':emb,'unconditional':emb})
47
+ self.sd.vae.to(vae_device)
48
+ print('EVAL_PROMPTS_CACHED: full 73-frame reference context',flush=True)
49
+ SDTrainer.cache_sample_prompts=cache_eval_prompts
50
+ torch.set_num_threads(16)
51
+ torch.set_num_interop_threads(4)
52
+
53
+ original=SDTrainer.hook_train_loop
54
+ def traced(self,batch):
55
+ bs=batch if isinstance(batch,list) else [batch]
56
+ info={'step':self.step_num,'regularization':[b.get_is_reg_list() for b in bs]}
57
+ print('SMOKE_STEP_START '+json.dumps(info),flush=True)
58
+ torch.cuda.synchronize(); started=time.monotonic()
59
+ torch.cuda.reset_peak_memory_stats()
60
+ try:
61
+ result=original(self,batch)
62
+ except Exception:
63
+ traceback.print_exc()
64
+ print('SMOKE_MEMORY '+json.dumps(dict(allocated=torch.cuda.max_memory_allocated(),reserved=torch.cuda.max_memory_reserved())),flush=True)
65
+ raise ValueError('Smoke batch failed; aborting rather than skipping')
66
+ torch.cuda.synchronize()
67
+ print('SMOKE_STEP_PASS '+json.dumps(dict(**info,elapsed_s=time.monotonic()-started,peak_allocated=torch.cuda.max_memory_allocated(),peak_reserved=torch.cuda.max_memory_reserved())),flush=True)
68
+ return result
69
+ SDTrainer.hook_train_loop=traced
70
+ sys.argv=[root+'/run.py',config_path]
71
+ runpy.run_path(root+'/run.py',run_name='__main__')