diff --git a/.gitattributes b/.gitattributes index a6344aac8c09253b3b630fb776ae94478aa0275b..c236849c1d4bfe9cf2467e424dec8c8883a7b860 100644 --- a/.gitattributes +++ b/.gitattributes @@ -33,3 +33,6 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text *.zip filter=lfs diff=lfs merge=lfs -text *.zst filter=lfs diff=lfs merge=lfs -text *tfevents* filter=lfs diff=lfs merge=lfs -text +processor/tokenizer.json filter=lfs diff=lfs merge=lfs -text +reproduction/samples/archive/2026-09-20-initial-poc/reference-clean-rgb.png filter=lfs diff=lfs merge=lfs -text +transformer/manifest.json filter=lfs diff=lfs merge=lfs -text diff --git a/LICENSE b/LICENSE new file mode 100644 index 0000000000000000000000000000000000000000..13ae08d5a5828f508cf0e852b1b316c0c0bbb9b3 --- /dev/null +++ b/LICENSE @@ -0,0 +1,55 @@ +Qwen RESEARCH LICENSE AGREEMENT + +Qwen RESEARCH LICENSE AGREEMENT Release Date: September 20, 2026 + +By clicking to agree or by using or distributing any portion or element of the Qwen Materials, you will be deemed to have recognized and accepted the content of this Agreement, which is effective immediately. + +1. Definitions + a. This Qwen RESEARCH LICENSE AGREEMENT (this "Agreement") shall mean the terms and conditions for use, reproduction, distribution and modification of the Materials as defined by this Agreement. + b. "We" (or "Us") shall mean Hangzhou Tongyi Laboratory Technology Co., Ltd. + c. "You" (or "Your") shall mean a natural person or legal entity exercising the rights granted by this Agreement and/or using the Materials for any purpose and in any field of use. + d. "Third Parties" shall mean individuals or legal entities that are not under common control with us or you. + e. "Qwen" shall mean the large language models, diffusion models, and software and algorithms, consisting of trained model weights, parameters (including optimizer states), machine-learning model code, inference-enabling code, training-enabling code, fine-tuning enabling code and other elements of the foregoing distributed by us. + f. "Materials" shall mean, collectively, our proprietary Qwen and Documentation (and any portion thereof) made available under this Agreement. + g. "Source" form shall mean the preferred form for making modifications, including but not limited to model source code, documentation source, and configuration files. + h. "Object" form shall mean any form resulting from mechanical transformation or translation of a Source form, including but not limited to compiled object code, generated documentation, and conversions to other media types. + i. "Non-Commercial" shall mean for research or evaluation purposes only. + +2. Grant of Rights + a. You are granted a non-exclusive, worldwide, non-transferable and royalty-free limited license under our intellectual property or other rights owned by us embodied in the Materials to use, reproduce, distribute, copy, create derivative works of, and make modifications to the Materials FOR NON-COMMERCIAL PURPOSES ONLY. + b. You shall not use the Materials for any commercial purpose without obtaining a separate commercial license from us. If you wish to use the Materials commercially, you shall request a license from us at model-business@notice.qwencloud.com. + +3. Redistribution +Subject to Section 2 (Grant of Rights), you may distribute copies or make the Materials, or derivative works thereof, available as part of a product or service that contains any of them, with or without modifications, and in Source or Object form, provided that you meet the following conditions: + a. You shall give any other recipients of the Materials or derivative works a copy of this Agreement; + b. You shall cause any modified files to carry prominent notices stating that you changed the files; + c. You shall retain in all copies of the Materials that you distribute the following attribution notices within a "Notice" text file distributed as a part of such copies: "Qwen is licensed under the Qwen RESEARCH LICENSE AGREEMENT, Copyright (c) 2026 Hangzhou Tongyi Laboratory Technology Co., Ltd. All Rights Reserved."; and + d. You may add your own copyright statement to your modifications and may provide additional or different license terms and conditions for use, reproduction, or distribution of your modifications, or for any such derivative works as a whole, provided your use, reproduction, and distribution of the work otherwise complies with the terms and conditions of this Agreement. + +4. Rules of use + a. The Materials may be subject to export controls or restrictions in China, the United States or other countries or regions. You shall comply with applicable laws and regulations in your use of the Materials. + b. If you use the Materials or any outputs or results therefrom to create, train, fine-tune, or improve an AI model that is distributed or made available, you shall prominently display “Built with Qwen” or “Improved using Qwen” in the related product documentation. + c. You shall not use "Qwen" as the primary name or identifier of any derivative works or products; reasonable descriptive use (e.g., "fine-tuned from Qwen Image") is permitted. + +5. Intellectual Property + a. We retain ownership of all intellectual property rights in and to the Materials and derivatives made by or for us. Conditioned upon compliance with the terms and conditions of this Agreement, with respect to any derivative works and modifications of the Materials that are made by you, you are and will be the owner of such derivative works and modifications. + b. No trademark license is granted to use the trade names, trademarks, service marks, or product names of us, except as required to fulfill notice requirements under this Agreement or as required for reasonable and customary use in describing and redistributing the Materials. + c. If you commence a lawsuit or other proceedings (including a cross-claim or counterclaim in a lawsuit) against us or any entity alleging that the Materials or any output therefrom, or any part of the foregoing, infringe any intellectual property or other right owned or licensable by you, then all licenses granted to you under this Agreement shall terminate as of the date such lawsuit or other proceeding is commenced or brought. + +6. Disclaimer of Warranty and Limitation of Liability + a. We are not obligated to support, update, provide training for, or develop any further version of the Qwen Materials or to grant any license thereto. + b. THE MATERIALS ARE PROVIDED "AS IS" WITHOUT ANY EXPRESS OR IMPLIED WARRANTY OF ANY KIND INCLUDING WARRANTIES OF MERCHANTABILITY, NONINFRINGEMENT, OR FITNESS FOR A PARTICULAR PURPOSE. WE MAKE NO WARRANTY AND ASSUME NO RESPONSIBILITY FOR THE SAFETY OR STABILITY OF THE MATERIALS AND ANY OUTPUT THEREFROM. + c. IN NO EVENT SHALL WE BE LIABLE TO YOU FOR ANY DAMAGES, INCLUDING, BUT NOT LIMITED TO ANY DIRECT, OR INDIRECT, SPECIAL OR CONSEQUENTIAL DAMAGES ARISING FROM YOUR USE OR INABILITY TO USE THE MATERIALS OR ANY OUTPUT OF IT, NO MATTER HOW IT’S CAUSED. + d. You will defend, indemnify and hold harmless us from and against any claim by any third party arising out of or related to your use or distribution of the Materials. + +7. Survival and Termination. + a. The term of this Agreement shall commence upon your acceptance of this Agreement or access to the Materials and will continue in full force and effect until terminated in accordance with the terms and conditions herein. + b. We may terminate this Agreement if you breach any of the terms or conditions of this Agreement. Upon termination of this Agreement, you must delete and cease use of the Materials. Sections 6 and 8 shall survive the termination of this Agreement. + +8. Governing Law and Jurisdiction. + a. This Agreement and any dispute arising out of or relating to it will be governed by the laws of China, without regard to conflict of law principles, and the UN Convention on Contracts for the International Sale of Goods does not apply to this Agreement. + b. The People's Courts in Hangzhou City shall have exclusive jurisdiction over any dispute arising out of this Agreement. + +9. Other Terms and Conditions. + a. Any arrangements, understandings, or agreements regarding the Material not stated herein are separate from and independent of the terms and conditions of this Agreement. You shall request a separate license from us, if you use the Materials in ways not expressly agreed to in this Agreement. + b. We shall not be bound by any additional or different terms or conditions communicated by you unless expressly agreed. diff --git a/NOTICE b/NOTICE new file mode 100644 index 0000000000000000000000000000000000000000..8bd7a1155eaca7e16f36a644a6a7d7cb3eec7141 --- /dev/null +++ b/NOTICE @@ -0,0 +1,9 @@ +Built with Qwen + +Qwen is licensed under the Qwen RESEARCH LICENSE AGREEMENT, Copyright (c) 2026 Hangzhou Tongyi Laboratory Technology Co., Ltd. All Rights Reserved. + +Mesmer Image 21 Nunchaku is a modified, independently calibrated quantization of Qwen Image 2.1. The transformer uses Nunchaku signed INT4 weights and activations with BF16 rank128 residual branches. The text encoder is serialized in bitsandbytes NF4; processor, scheduler and VAE originate from the pinned upstream release. These modifications are by MesmerTech, September 2026, and are not an official Qwen or Nunchaku release. + +The model and derivatives are for non-commercial research and evaluation under the accompanying LICENSE. Commercial use requires a separate license from Qwen. + +Runtime dependencies retain their respective licenses. The custom linear runtime uses MIT HAN Lab Nunchaku; the conversion packing adapter uses DeepCompressor. See THIRD_PARTY_NOTICES.md for pinned source and license links. diff --git a/README.md b/README.md new file mode 100644 index 0000000000000000000000000000000000000000..a31cb2870d9d68b1e29aa3a592b965b9765e382b --- /dev/null +++ b/README.md @@ -0,0 +1,65 @@ +--- +license: other +license_name: qwen-research +license_link: LICENSE +base_model: Qwen/Qwen-Image-2.1 +pipeline_tag: text-to-image +tags: +- image-editing +- nunchaku +- svdquant +- int4 +- research +--- +# Mesmer Image 21 Nunchaku + +**Built with Qwen.** My experimental rank128 INT4 quantization of Qwen Image 2.1, tested for generation and editing on a 16 GB RTX 4070 Ti SUPER. This is an independent implementation, not an official Qwen or Nunchaku release. + +**Research and evaluation only.** The [Qwen Research License](LICENSE) requires a separate commercial license. See [NOTICE](NOTICE) for modifications and attribution. + +[View my benchmark and full-size comparisons](https://mesmer.tools/benchmarks/qwen-image-2-1). + +I compared 21 scenarios at both 25 and 40 steps: 14 generation prompts, six single-image edits and one two-image edit. The selected version improves on my first INT4 attempt, but it is **not lossless**. Faces, hands, object counts and layouts can drift. Repeated runs with the same seed can also differ visibly, including with application caches disabled. + +| 1024 × 1024, RTX 4070 Ti SUPER | 25 steps | 40 steps | +|---|---:|---:| +| Generation, median of 14 | 15.50 s | 22.82 s | +| Single-reference editing, median of 6 | 19.48 s | 28.43 s | +| Two-reference editing, one case | 23.76 s | 34.35 s | + +Maximum observed GPU board memory: 11,994 MiB. These are local eager inference measurements, with CPU offload, built-in prefix KV reuse, and no prompt/reference LRU hits. They exclude cloud queue/startup/upload time. The comparison teacher uses a BF16 transformer **with an NF4 text encoder**, not a fully BF16 pipeline. Forty steps follows the upstream starting recommendation; 25 is a measured speed/quality option, not an equivalent-quality promise. + +## Contents + +- `transformer/`: 224 signed W4A4 block projections using upstream Nunchaku CUDA kernels, plus BF16 rank128 branches and unquantized boundary layers. Packed state is 4,655,177,728 bytes. +- `text_encoder/`: prequantized bitsandbytes NF4 Qwen3-VL with double quantization and BF16 compute. All 36 decoder layers are retained; the runtime removes only unused output components. +- `vae/`, `processor/`, `scheduler/`, `model_index.json`: pipeline components. +- `runtime/`: shared inference server and loader. The custom transformer needs this loader; stock pipeline loading alone will not recognize the packed format. +- `reproduction/`: calibration/export implementation, configuration and measured report. The obsolete first implementation is not shipped as the active runtime. + +Base revision: `b3179ad355be050328e483a9dfdd9e60cd62adfa`. Original selected transformer manifest SHA256: `cf69e83a646981d22ee6df8d5239b46a50df25d8eb73c9f0478feae87323e6cb`. The original manifest is retained unchanged, including calibration provenance. + +## Run + +Use Linux, CUDA 12.8, Torch 2.8.0 and the matching Nunchaku 1.2.1 wheel. This release was tested on Ada; do not use this INT4 build on RTX 5090/Blackwell. The pinned serving dependencies and Docker configuration are in `runtime/runpod/`. + +Download this repository to a local directory, review the runtime code, then: + +```bash +export QWEN_BACKEND=nunchaku +export QWEN_PREQUANT=/absolute/path/to/model +export QWEN_NUNCHAKU_CHECKPOINT="$QWEN_PREQUANT/transformer" +export HF_HUB_OFFLINE=1 +cd "$QWEN_PREQUANT/runtime" +uvicorn server:app --host 127.0.0.1 --port 8091 +``` + +```json +{"input":{"prompt":"A busy night market after rain. Two friends share an umbrella while a vendor hands them a paper bag. A handwritten sign reads FRESH BREAD.","width":1024,"height":1024,"steps":40,"CFGScale":1,"seed":42,"outputFormat":"PNG"}} +``` + +POST this body to `/runsync`. For edits add `referenceImages` with one or two public image URLs or data URIs. Optional `uploadUrl` accepts a presigned PUT URL; otherwise the response includes base64. The server serializes jobs. Current API limits are 256–1024 pixels per dimension in multiples of32, up to60 steps and up to2 references, even though the upstream model has broader capabilities. + +See `reproduction/REPORT.md` for comparison methodology and limitations. Per-layer calibration minimizes measured projection output error; this is not a reproduction of every official SVDQuant training/calibration choice, nor proof of end-to-end equivalence. + +For calibration reproduction, use the `reproduction/` directory as the POC workspace mounted at `/poc`, with the upstream BF16 model and a separate writable cache mounted at `/cache`. Copy `reproduction/artifacts/qwen21-calibration` to `/cache/qwen21-calibration`. The calibration edit reference is included at its original relative path. Follow `reproduction/nunchaku_backend/README.md`; randomized SVD/CUDA reductions mean byte-identical regenerated checkpoints are not promised. diff --git a/SHA256SUMS b/SHA256SUMS new file mode 100644 index 0000000000000000000000000000000000000000..8d2819fe6732a663a78244ac7e83902b9be0fe3b --- /dev/null +++ b/SHA256SUMS @@ -0,0 +1,338 @@ +8dc973f024ff95966bea25866efa443fd16776dcb1001e681e3d467ea572b28d LICENSE +0bed3a6a716e52829e0d78110339d6c028bb4d3893c7221cf9efe2ba485f8bed NOTICE +b07d6247562cfe168dbc98df3435a3b1ffe85f9a8148d9c0bf20e97cef7be38c README.md +affa01b1499e11cc0fd209fcc6775e659bcf773959e708e268e27c69be808114 THIRD_PARTY_NOTICES.md +11ce832ce35b332259dcaa4bcc370bb391694959ce3f74409a321a6b32919a57 model_index.json +3636d0f0bd6bef02654cdffdc447b79cb2cef8ab02cc75267345946291a489e4 processor/chat_template.jinja +8993ae056f017c142d0dd91c3950b1bcfb4043e118d2f24436ec8e67272113b8 processor/processor_config.json +be75606093db2094d7cd20f3c2f385c212750648bd6ea4fb2bf507a6a4c55506 processor/tokenizer.json +09bbc8d235d72ce43f6ab3c1d15e55319f52b672780f977f40441345048c2443 processor/tokenizer_config.json +604692671915a90ee81430fd7e1097ec626a05659d9a0f80412f19da50ddcc2d reproduction/Dockerfile +8dc973f024ff95966bea25866efa443fd16776dcb1001e681e3d467ea572b28d reproduction/LICENSE +0bed3a6a716e52829e0d78110339d6c028bb4d3893c7221cf9efe2ba485f8bed reproduction/NOTICE +7ad48b6b89d0abebde961696b5fd33862d14e7e2d91de3fad74264943102a0bd reproduction/REPORT.md +affa01b1499e11cc0fd209fcc6775e659bcf773959e708e268e27c69be808114 reproduction/THIRD_PARTY_NOTICES.md +7f032d0930881f8352309e10084b3e5e4418ac81fb1c69a9bc7457f5fa731018 reproduction/artifacts/qwen21-calibration/activation_stats.safetensors +2efdac15b2ad81492b394604e807921fa4268d7e4d410f25261fe6fb78e0d833 reproduction/artifacts/qwen21-calibration/calibration.json +5050e49d34b4dcac35f75fcd25638cc401105ffb7595339a7551e67a4d8abcae reproduction/environment.freeze.txt +f748f62cb3daf7aae90c01d455889a4a6472716bfe806596f88411695469c998 reproduction/experiments/fidelity-v3/README.md +7ad48b6b89d0abebde961696b5fd33862d14e7e2d91de3fad74264943102a0bd reproduction/experiments/fidelity-v3/REPORT.md +191c65872a09c37b0672df820b8db7fa3ae639c9bc2207bbb3e437124bf7f68b reproduction/experiments/fidelity-v3/calibration-jobs.json +5759b11e2380f888ada234e5b125a1aa6d5f2a5694a35522e0908f50bd95a23b reproduction/experiments/fidelity-v3/candidate-remaining.json +4d0f6d5d7c4cc3f83676d1d72826546955c1b15446b8b9e2c5fef7c0395f0dfa reproduction/experiments/fidelity-v3/expanded-jobs.json +f1fdc8b4d7cda630c480f794f0da53e0caf39e5cf78736f28af9c4fc4e5d9fb5 reproduction/experiments/fidelity-v3/expanded-rubric.md +a3a1ea160e2f2c0fdbb731b6f21d7d388176abc626ca2f6600024255c746951f reproduction/experiments/fidelity-v3/original-18-jobs.json +ba9ae69d9b203bc5e5c1c998e726ca697da49fddd869ebea423b235fe5cd1441 reproduction/experiments/fidelity-v3/performance.md +8993e784dfc3dca10f508bc4f5658bb900b5b15305082dac1d019975624e0e7a reproduction/experiments/fidelity-v3/quality-jobs-all.json +03492ea5750ca7ae9b62d2e8e0dc03ab266af846e078a83363ded3f15dbcb32a reproduction/experiments/fidelity-v3/quality-jobs.json +f416d01bbb950380ec181d6a9431dc8c106270deb7fc38625deb967e5dbe09ee reproduction/experiments/fidelity-v3/selection.json +f5571496730809e5d3edfabf087f9dfc6efcd1f90fd538a83fa50aded1ecce09 reproduction/experiments/fidelity-v3/teacher-priority-bf16.json +b584a8116aaa3584d0268704fa0df28d4db632f3de42b27694ad372218e312b7 reproduction/experiments/fidelity-v3/teacher-priority.json +7f52529a6eed5bca8d2f327cff6617e4121e2d80cce4f86d6346482de51c8fd7 reproduction/experiments/fidelity-v3/teacher-remaining.json +2787ad3ea400a2df3f96b45643cc8c480976e787bb17d2e9e58b54e7872195d8 reproduction/lean_encoder.py +a94c85703a3ae47eea0ca7edacfc4d9dfe5b96235a9cf8a4a61df69056251630 reproduction/nunchaku_backend/README.md +888e726eec3fbf637073e7861a546b2c11cff6a00b5cad60cbe905b3a2782069 reproduction/nunchaku_backend/__init__.py +bd864b154571b93825f4bd260e66f745e6d1d0fd1767ca8918f759b8127f20fc reproduction/nunchaku_backend/baseline_candidate.py +36242e9d2b06d1c77542b030061def860ec8263458d4e57cfa1495f6df51c7f7 reproduction/nunchaku_backend/baseline_candidate_test.py +d5c0409c0804a66c99549d16ea741e36f2e2fb029ea3662c2592bc54916926a8 reproduction/nunchaku_backend/cache_repeat_diagnostic.py +dcbc86bf6f6b2cd20317a9aa26c30b6986fed6aab1b231c507f3c88d44fee939 reproduction/nunchaku_backend/checkpoint_io.py +5680d0a4cd7d941d69e0df0e436d8e6794c8331ad31c9d61fd95ac1cdfa67bfa reproduction/nunchaku_backend/cleanup_test.py +81f1a59b36b15834a1ba81d1ecad627bc23b19fda96f26d2ae4cbf80bd71d765 reproduction/nunchaku_backend/collect_v3.py +2f6bdeb29564fbfcf46162df876ab56fbd913b31db8884345285caae3d3caf2b reproduction/nunchaku_backend/compare_iterations_v3.py +e170904ae03cdf47551c763d182376218e02b28b47be47e524b18877614ebed1 reproduction/nunchaku_backend/convert_activation_reference.py +765bd94af8fd02957b5b1b9cd18b7c4432eb8011c3e376fa6dec52380c76ef69 reproduction/nunchaku_backend/convert_recipe_v3.md +37e5a426e20416d516a8ba4a1e7ebd55c5366fea042467961852951a3c9fad5e reproduction/nunchaku_backend/convert_reference.py +3f0b095b08c3d9022559e9c1dc6d9412b42b444e31897e437040f2c3347b8051 reproduction/nunchaku_backend/convert_reference_notes.md +f6d5ec714ad98077e0bea90cc0af3e08c417e6a0ef4ea135863939294d3be046 reproduction/nunchaku_backend/convert_reference_test.py +bac0f1bc6eb5feeae974154363fc2ad978852045b7182165e9c3b8297b7aa7fc reproduction/nunchaku_backend/denoiser_probe_v3.py +84bd1d8b1993062686a7406f994e2cb2d597912f2a2364c2f7916718a6b10823 reproduction/nunchaku_backend/diagnose_kernel.py +7fcb10f22c5191eabb82222e160ba7434f02099d644c057d74aa25f75cd227fa reproduction/nunchaku_backend/export_v3.py +0d43c9c3435dcfd3f821c7e4ce8a7a2cfdc5fc9a329dfca0340756e0550d1471 reproduction/nunchaku_backend/export_v3_test.py +7aea9e2f87be048d43a43f9c83a42e14cd313d2d53588612533d44fbd840cf9b reproduction/nunchaku_backend/gptq_v3.py +a4c48c1b1f6d4abd4d4a8236719049e53827d6da1613c32e3eb2820bf317949b reproduction/nunchaku_backend/gptq_v3_test.py +26965ca14b8e4a8f9332bf93a3859f0f22e386313de6fab61995cef95e0995b3 reproduction/nunchaku_backend/hybrid_v3.py +73141bec2b023b8342459ef3a07f95f7bb1e74dc03f29cacb956c8dfd791b3bf reproduction/nunchaku_backend/iterations100_probe.md +e70362903fe14b3a71bafadcb89d11d5b0e80866f296f19f6eca0c529957d28b reproduction/nunchaku_backend/kernel_probe.py +ac4a8ecc41312dbd675ea309e1532c53fd2d267ac6e60f5c4f9f4fbc4d2db073 reproduction/nunchaku_backend/layout.py +cc0bfa981c0b615f6e7548f7c583d72cea834f7677cc5167225c655b20c5b4b1 reproduction/nunchaku_backend/mlp_parent_probe_v3.py +097acd9b5e792155af27d33930fb3663ec4337d6cd4dba192b28eb09082125a8 reproduction/nunchaku_backend/mlp_parent_probe_v3_test.py +770de03be86c5a83b1376f2e612a7459dac2e2e7a99aba43528a986d069a9d5b reproduction/nunchaku_backend/optimize_v3.py +6b21f991b863ed8500b2808a709a79fa09ea1c99f19a072dbfcd4b0e2f48b9d0 reproduction/nunchaku_backend/packing.py +8e9c2ce022a6569800ea575a8968c7bea413271e1bd83a3183c71574998b55a1 reproduction/nunchaku_backend/probe_iterations_v3.sh +f7dc00d8f6d643b9a2983acd164f2b0c03abd9379f84421356a582dd46b85dd7 reproduction/nunchaku_backend/rank_probe_v3.md +d7a04a4d5a573fe7c74d6e0e409cfaabdf414ccfc22c305e8c7fc2dfc2d4d737 reproduction/nunchaku_backend/rank_probe_v3.py +d42a9be112a8d47e91cd6ef2f940137bb5f670240e54bf793ef4f2291a280746 reproduction/nunchaku_backend/rank_probe_v3_test.py +1cee4d9da46125e7f362ad8253ac22921e142209464e41276fcaec84c3f066e7 reproduction/nunchaku_backend/rank_upgrade_helpers.py +aef04a267f40dbf75003689fda7d274c25081aee27da7eb3e896e0592e256139 reproduction/nunchaku_backend/rank_upgrade_helpers_test.py +c35a12c28cb982e81c1c5600782db379d4541fedcc54c34ba960472c32b46a94 reproduction/nunchaku_backend/refine_checkpoint_v3.py +dac9996d5df928a22c719e572256233f1d6331a0d1f93faa994c3930d29852c3 reproduction/nunchaku_backend/refine_v3_test.py +67d239d98e03e0fcb8b6d2b6031978b35f0ab39676633898d58e881d2c418ae6 reproduction/nunchaku_backend/refinement100_ready.md +8fec8d2e829561c9372f92105fed9c384d1e25fa969f741172f492d74e4a20a8 reproduction/nunchaku_backend/role_sweep_v3.py +19f4697e8f8bf20c0c528a82d6ff8159ae2527527c768b84753705593f44eb10 reproduction/nunchaku_backend/runner_hybrid_test.py +bb0cad2bfec1ea2b7ea5962db4831c10bbed9f86aa6ef0b04bf4cd30bf03168e reproduction/nunchaku_backend/runtime.py +d7c5c968984e10a0c3ba8e861e8d6a4dc1e84d23a8898c799374c923cbc4579d reproduction/nunchaku_backend/summarize_v3.py +e1d8c87c9bc867b4dfb0ac581d995abe1824202040e94c471f7077590bd11527 reproduction/nunchaku_backend/teacher_v3.py +91562daf3bdb60217ff6fa234969cd6e2b2652ddcd29e0039ca6e5e52e047dcb reproduction/nunchaku_backend/upgrade_mlp_rank_v3.md +db5b42d8949857e2ae77f8cd3bf741a5c51e7d26d0236658ac6fdc93fb392c0d reproduction/nunchaku_backend/upgrade_mlp_rank_v3.py +ae7b0cd2d1511aaf699192e5760313409866f0b0e48ca8ac211863c1819af458 reproduction/nunchaku_backend/upgrade_mlp_rank_v3_test.py +671985da2939b5565ad0b60891ebca8ad8e778421a711ae021e35153710df8c3 reproduction/nunchaku_backend/upstream/deepcompressor-tree.json +032e43f8335089bd1ffe112797478c6b7a46e1b204c048d75776289301b0553f reproduction/nunchaku_backend/upstream/nunchaku-tree.json +245cc971a635403851fe77986213a59190e244383c2c75741dc0789facf6ed4d reproduction/nunchaku_backend/upstream/nunchaku__models__linear.py +a3c209cc785afc3d77f6d78b81b24975671344ff324284c899771cc10be6c833 reproduction/nunchaku_backend/upstream/nunchaku__models__transformers__transformer_qwenimage.py +3d865dcb286e8296d1a6e3876f2576675136fdba860406139acdd6e086d1cdcc reproduction/nunchaku_backend/upstream/nunchaku__ops__quantize.py +a369e7db138a950ae9ae4f80dd1d20e37029603d2c5025c371e0a77f2b29e73c reproduction/nunchaku_backend/v3_notes.md +c939a0dc977f0bcc21a202888996b2d33cf6b8e351848869ca8416450a93873b reproduction/nunchaku_backend/v3_test.py +04fdc67a215261ad700c1356d937087ce8bd6d2b1debd1eb7e3f783f84a55d15 reproduction/nunchaku_backend/validate_rank_upgrade_cpu.py +27996383f5ae4568d5bec0512e128ff800edeb6abaa4e5ee2c64c0b48b1ebf26 reproduction/nunchaku_backend/validate_v3_cpu.py +6f7e1926881a78b0e342667447d9110642715153e00bb3f677a776004c497ef0 reproduction/repeatability.md +c62923ee98c8a7876e99b0dfeaa7e6553b09b93a8d6004ff1b8b96fe859f23d3 reproduction/runner.py +a191a4896478811560e3cf679c23a33b9530252c6d6a72e6136e10418d51c803 reproduction/samples/archive/2026-09-20-initial-poc/reference-clean-rgb.png +4e231258338bcd8fc2a4fdcd0b5c9b6092f069170fde3137b1c862c1cd95872e reproduction/server.py +5ece0c1ef602659ad3a71eb01a05bc951848430cd688141685673fc793f7be6b runtime/API.md +8dc973f024ff95966bea25866efa443fd16776dcb1001e681e3d467ea572b28d runtime/LICENSE +0bed3a6a716e52829e0d78110339d6c028bb4d3893c7221cf9efe2ba485f8bed runtime/NOTICE +affa01b1499e11cc0fd209fcc6775e659bcf773959e708e268e27c69be808114 runtime/THIRD_PARTY_NOTICES.md +d26107b7df25fea9792afdce862b1f5ac7a5d3512fdbfe3575c9876a17666ab1 runtime/lean_encoder.py +a48779c4bf66738834d01d31da0fbaf817be9a614a5e79d7b1b4ca2341fa2481 runtime/nunchaku_backend/__init__.py +f0242706d5b5439aee5641b6e16cb576825992d12f3c85037928e783aa7f4b82 runtime/nunchaku_backend/layout.py +5c15b3b3459c26e4c614fe3efdda5e1a1220cbb4f63e340ff726b8f28881dc6d runtime/nunchaku_backend/runtime.py +1f63fa88b385ca4747043ea6d71645aff5d03a3c943700b4163f79b4dba77498 runtime/runner.py +8e0c34d3bafc0ffaef52095d640da340ad90534fb6ab121a03ee22fc6538510f runtime/runpod/Dockerfile +c799f3ee5bb8d176018dbe96d8f07db0c6e1bad6a79b5f72fe08a20154734268 runtime/runpod/download_models.py +ba3be5f3ae7945c15aba7acfbc7d41065d6d88993ccf89ae4b41f6abacd25d90 runtime/runpod/handler.py +d6f05ef501106f7a082f8585e028c8eb9955ae02c4cbcccf6824ac62f1b51204 runtime/runpod/requirements.txt +c21172fa44d6865969fe4de8f24e709c954a73a5b8714db64024cdb578ee18a9 runtime/server.py +b73113f657b5ea5aa7be0b6c7bffa671e3e91fb565991ce875cd9ebb7a9737d7 scheduler/scheduler_config.json +c0bc48f2d1a50def0a306a328236b7e1a260b13da8282ec1673405c12dbd1fca text_encoder/config.json +81207a236dbd29cd0df5208c43225ff6d0155e3fe63a40ab2f7331fcf04f4834 text_encoder/generation_config.json +61c75987e552c49b9bb804b6665b70e6df10d808397d7dcfe2e6ce94e7292210 text_encoder/model.safetensors +3e64be674ff9cf15412ebaa996e7bcbdb1132f0bb3e6e0b6bacbe1dfd78397d3 transformer/config.json +cf69e83a646981d22ee6df8d5239b46a50df25d8eb73c9f0478feae87323e6cb transformer/manifest.json +91cedb20bfd9f986b3dfd666508f7ef59bf2f4eba98784e15071715b9082aba6 transformer/model-boundary.safetensors +be954118c35254d2512ab5a469765dcc834b9cd9ce23c7f0206de1a63f0d04e5 transformer/model-layer-000.safetensors +805910c5b758ae10aa3bd41bd35a0226a0a7ffda8dd5249ef5dc8642b17533f3 transformer/model-layer-001.safetensors +dee0003b353df79eb67c4fe949e484999b194158c4cbbd144e8d4f7898e57569 transformer/model-layer-002.safetensors +3b38108461d9afd2248223c28f208b175822cfdc03d6cac27bd9be8b58ce6b4a transformer/model-layer-003.safetensors +bf2dedbba5139a3ef770ed583bc6b5391598448fdf81e977f62210830e470493 transformer/model-layer-004.safetensors +0d8b6b13096092574df713cfa59333c80d275a8eb983d041cb45b7ef510999c0 transformer/model-layer-005.safetensors +43e266de81565062f7e07e8b283ae1678deb0bf3124643e8a1746428a3816098 transformer/model-layer-006.safetensors +a385b1e5c2bf279906832f96fd028b35939ed4d5e8468ef8a7403695eafb010b transformer/model-layer-007.safetensors +438e4c30185fc9f196777196eeafeeca6fbcd0f738b9bbec1d7fc37dd1c51e21 transformer/model-layer-008.safetensors +f83cbe6924cec25cccd785e226af722ecc1e68a800e57bdfcedbac3d730cb651 transformer/model-layer-009.safetensors +db6a92ce00cda731450c7b79d5d39fd7df6ce3ffdda6bf58297fd47431fc3bf7 transformer/model-layer-010.safetensors +4dc2ec89217357598e6829b46fc4f60c85c46567efbd0f0696cbef036ad11091 transformer/model-layer-011.safetensors +54d032380b5660ec2c2b9b1f73ef125fdd05f234164597eb2a6fcc13923945db transformer/model-layer-012.safetensors +124ae09e4b97a1667808651d8f445a84b7d8da42bbd84f4d7379b41c6c9d86d2 transformer/model-layer-013.safetensors +a166d614f6a6941fb7a7508d2a611610abc7bf7f6d1527f52f9c7d7db435f688 transformer/model-layer-014.safetensors +4cf90f6c54659fbb2f259b1454da2b9d457f4b12bba9834568fd43bdf8e6acdb transformer/model-layer-015.safetensors +4fe7a02c345da061343ac51feea997dfbbee408588d1894ce52ca45804ee7554 transformer/model-layer-016.safetensors +403f07d5311ae7da8ed97ace3ece9d3d56284d0a764cee35d9419df1ce79a515 transformer/model-layer-017.safetensors +fa11beb0cfdb4eaa571a6968872be1f75e25d06e00efef6e44179fc72e04a021 transformer/model-layer-018.safetensors +4adcdb2c60bda8d4ec3c96c715cad036dd1c43da11d0ab225dfe72d09ca0dacb transformer/model-layer-019.safetensors +95593972d08bd06bba59cc6a7ec9ee2a61a78d6b53ef5e0d01de34719b18cb9e transformer/model-layer-020.safetensors +e68b95ecfc1df57b3b91dc813d57740acbf40302fa5e5f6989634ea44bd80d26 transformer/model-layer-021.safetensors +6c13a089c4784a26862cfaa59312f09fc71f95b72c4d26b8065ee54b6bc7c421 transformer/model-layer-022.safetensors +121526e8517c376058d032eecbc2445df4e26faa26fa04361473d83fb94ec960 transformer/model-layer-023.safetensors +238667fe2e13404a8adaf981a3283f2593ae859096d26bce454bf574d0a61434 transformer/model-layer-024.safetensors +e3f3f259fb31dc98060e1de725cad827ca1b0eae0bf5bfa3ab4bc50ff0d11f55 transformer/model-layer-025.safetensors +896efca07b1bcd0272af0cd3d580c64762591e738c9b9bff2eec81feda9b6cf2 transformer/model-layer-026.safetensors +1feff8f4d8d07777e3c914c0152443d643257e9155a364829fe14fa52eaca563 transformer/model-layer-027.safetensors +562ed669c40ca3bd447109beb66df73ad21c00aa82b7b91d50c5d34475fc88c8 transformer/model-layer-028.safetensors +eb27895adb5ff1402342f530c9244b02344edcaf5cdb7571b55170dc87aa8b66 transformer/model-layer-029.safetensors +7ead55cbc062e1a32f48ca2e5bc17f311f0a80e6125e6bad1d4bd30c762434f3 transformer/model-layer-030.safetensors +b01bbd72216709fb7888e201f1b25a70505740e34ab23d399c17465e324c75e9 transformer/model-layer-031.safetensors +49db7476493465a03fce9fc9085e27e596440246b7a0d182074e569114d52f06 transformer/model-layer-032.safetensors +c67faa03d75949ce4a209210de02cdf58192877d445a0264b0030c55f7c374b9 transformer/model-layer-033.safetensors +38b273218108066350b01c2006a6ab667092868dbbc144b390dddf8550be2988 transformer/model-layer-034.safetensors +046ac87ebda321a370208bbc879d62f62d439463b297e05a11bbec12990edb7a transformer/model-layer-035.safetensors +e1e8afb5ccbb88266124e88a233e7545e27591e9e0ac07c8ca9fb45783569eaa transformer/model-layer-036.safetensors +d3628d58f6e3394ebf0e790c8dd44685b76ee027154c6d5fff863308b5758ab6 transformer/model-layer-037.safetensors +e73d297573fb3fb87d909441a81d82ed56eed0facd8a38d42836ed8cc61fbb84 transformer/model-layer-038.safetensors +6604928bf9b57f4ec5189fd534249e064758d2c3c319db2231b608cabd968a4e transformer/model-layer-039.safetensors +2c07c7123bd68cb52e5f69abdee4e2d4cc945dc8c8436662b6ff004422f821dd transformer/model-layer-040.safetensors +4177a378ff702ee8e079a980bd47cc759a8f8e25a57b63d6746110ee9ed5e257 transformer/model-layer-041.safetensors +a200b9c9383eb3c969ea124941df9fb157821c327e4a2991a1073ca0b274943b transformer/model-layer-042.safetensors +681f1584a3b9200fdea7ca0f55fc263ca01f1d8f5d80326fc8a00b09f73bbd9d transformer/model-layer-043.safetensors +241e6e0ba34c77390c9173c5b512d4544b74c27c3f71e1834d14d900af1f4cd2 transformer/model-layer-044.safetensors +5276d95df07f1e59f18388a1f2d210e1c5cfe95f4136df5ea6aa27d2ce112a3f transformer/model-layer-045.safetensors +18b52c52508cfa67e73a8c74db48dee8d9b2453a3234557d97e1f3b2ca1dfb9b transformer/model-layer-046.safetensors +543bc08d4fc332aa6a290dc026b62ebfc22d05e181f16346ad36003ae1e8566e transformer/model-layer-047.safetensors +8717d187ac5dc06762b8c578b791d932221b2adffc85d499d4b8e3a9103af03c transformer/model-layer-048.safetensors +56962839346c87c9dc42328b8af42f56a4d678c7bfbea57d36bc7b3dfc7ec9b1 transformer/model-layer-049.safetensors +ed43ccd63f133a9a191ac6e2037b43f4effeee1e44887f377ddc6459c9d1ef1f transformer/model-layer-050.safetensors +e3597b0ff8f3a2b798d0873873b9627f0c7e0ef5ceeac560251fdc66fe8f6195 transformer/model-layer-051.safetensors +fce9a7a76138781ec49ba8054f7d2f8b770aa7b9bcf78787d3db7d8cde159ff4 transformer/model-layer-052.safetensors +22b532bffe7dd1ded582cb37b970fa483fe6eb197acdaa8778d6b700a740393a transformer/model-layer-053.safetensors +ae24df341390707a611a12cd08763386397a6c271e7b1e0dbf062f5116e56af5 transformer/model-layer-054.safetensors +089382d22679fe93818b79bd53a52b413c3984d366cdd1331bde0d4f959d4782 transformer/model-layer-055.safetensors +12a9d61e10e9623df25d571e42a1b1006aaf0e3eec573e89f53cf6c21a4bf120 transformer/model-layer-056.safetensors +6ca13e70f2ed1616e2c9c5604aff6daeebb226d21b04effeeb9b1de1de1ea99f transformer/model-layer-057.safetensors +d01cfd242427e7bbe69d73245ec95e1d22e43f62bf45f53595019a703135968d transformer/model-layer-058.safetensors +378f3c3fcdb57101f5760e988aff466923f5bf0332584860aedf24c3352b8617 transformer/model-layer-059.safetensors +662886be033ee5f2ee3541ed7187d3b752dd8176c4688dbb72890d5aaa891532 transformer/model-layer-060.safetensors +64375e1fe523f359aca412649e1c30ab1496269f0737e230fade32adcfbf6347 transformer/model-layer-061.safetensors +37658199e6de2fb0e049ae0fe30cb16407db44922274cbc911245cec8fa3332b transformer/model-layer-062.safetensors +702e1a78ebfd91fe2b6efbef78fefd1257d0509320636e21e557bd86f5e2153c transformer/model-layer-063.safetensors +1c3f65152c3afba62069cda1c22df5798184a29c869de87d835e311798906530 transformer/model-layer-064.safetensors +610ba7bd5f30ad75be63f2368b5e76320007ded4d3551873c0366d9602d681ca transformer/model-layer-065.safetensors +f910131d9dbdc89949e03f087ac979b01fbfd1f43eed7ae1db583107d89de328 transformer/model-layer-066.safetensors +bdeee3b356428a7e75a19982522273328c20708cc572f12fe442f1dde57a3d15 transformer/model-layer-067.safetensors +3fc30d0a256df124ca9d254f2c677e01c114b729239a589848b3039289de947c transformer/model-layer-068.safetensors +9b6ca9886ac5cf675000ad407d008857f46c0b50617f926d53af1a028a00e278 transformer/model-layer-069.safetensors +40ad6cdbd4a624bba006875b65055bfcb78dbf8a3dd22e719f2a5afafdde36d3 transformer/model-layer-070.safetensors +333c2026b5f2c6f2746241824a1ef7560fac6a2bb9930c160cca9218b5c0e49d transformer/model-layer-071.safetensors +852177773dcff10bae0917265572f7658cd6cb06512a197b6cbba770f269c703 transformer/model-layer-072.safetensors +7ec3d0b4a46c5683f3dedd572298dab524cf3f6c22d2db4483786f2f4f1b7f4c transformer/model-layer-073.safetensors +e042abf43b93871fe6b0a0b3a6374d1d2435308643df8fe720d7e9899f25b8f5 transformer/model-layer-074.safetensors +ad29ca50a411ef8070919e4217520d1de0f41ca4666c6a2a3b475595a9e0e9ae transformer/model-layer-075.safetensors +04471891428d103172959be5afc598594537ce52929fb8ef4e007cd869726a81 transformer/model-layer-076.safetensors +2b94f5d716a82f220d3f846b70402346610671e48617084843ec161e558155ea transformer/model-layer-077.safetensors +8f686e842a4abea031616d9be089dc679f2b953b2b3bf575a1e109df7eeedbaf transformer/model-layer-078.safetensors +f080bed05b77c27fa76a2e4f5b7394f59a3a7732c853e358788f95d564ab62f3 transformer/model-layer-079.safetensors +60cdb41cdfecd42bcc2313e0c30c03109deaf88349472234a1f065ebe6165c96 transformer/model-layer-080.safetensors +c2fff5071e36336544a28ea5bab59f34a739b1887b201b1c0c77e4af4112d9ab transformer/model-layer-081.safetensors +622a1e8c1e6ebb3acff58b676d822cd3ab7cf849da5d692090f6dbf18e123fbd transformer/model-layer-082.safetensors +b213327da54a43dd807d91785c035be6b5f1ff651b0e523e8a7d4215231517c8 transformer/model-layer-083.safetensors +ca594ed744f93b6c247582ce8b238f3f7741f92b11b8de175cd2f898c3fd1201 transformer/model-layer-084.safetensors +6178d8df840e6cd804a85a93400ab524a723aa1779109278ed40c1915bf58176 transformer/model-layer-085.safetensors +cc0fbfa403743d2b45586649e5842c25993c46e69e41842a4c0e95494791b4d5 transformer/model-layer-086.safetensors +5aaea51447d8f892bed089a2ed4bd30cffe8a017430423493a2e9fdeaa806fd8 transformer/model-layer-087.safetensors +1fb2952585109eb0e33e7fb280f950850b685207c694603f0e241ba19a0efc2c transformer/model-layer-088.safetensors +3595b8921f6308f9df9a8bc73c68d28b0628ee3346c2b755a9b0e152ad50dbff transformer/model-layer-089.safetensors +779cba9de87994d34841be62fb9fb3098edfea5f8a32719a192820a95b04b570 transformer/model-layer-090.safetensors +1d8d72051cf8d046e9d3804ad00c0feb887e8fa89adcd3c74190ebdb510b6241 transformer/model-layer-091.safetensors +c83ff768a390c18f08ff89120787f05661ba2646a88eb128a80cdefad310aa13 transformer/model-layer-092.safetensors +b39f0766395452bcd08186a57659bdf8fde0e14302ed6d486913794281ff9cc5 transformer/model-layer-093.safetensors +7ca843ab6f7c7c503f31370980a6509e092081518df1c8a5a190bafe761b46d9 transformer/model-layer-094.safetensors +4780e7617421b97b54d3fbd997a15410009298cd8ddf901a32dc3662fb6c0c45 transformer/model-layer-095.safetensors +ef31e5488181a2e5a8efffaa15ab993ea1fbd698aa282a0aff12e787d702aa58 transformer/model-layer-096.safetensors +fe65aa51d7ae66bd32c7647a32a9e6a6f8e205247da5ecedee6b820b42b457af transformer/model-layer-097.safetensors +d9530418ef182c32baf1ebbf6597d4f395795e9597cc0ae3480f5a7bc3ae589f transformer/model-layer-098.safetensors +f061a9bd99376abf7896896f5e3fa39cc75b5dfee8dd3112dc30fd76cd6c99df transformer/model-layer-099.safetensors +ef0d94bf7d96ddcf977127d461593fd1b1fe92287ebec8b1c039b72337ffbc49 transformer/model-layer-100.safetensors +eb9e524affd0d6fff922656c540c695f09d5d0ef0dc3bda142c8155b2e03a33c transformer/model-layer-101.safetensors +17d356995ea24b6b3f78e421ef5dae1b3e4614f5688c761e0ac44782f97b3587 transformer/model-layer-102.safetensors +e943711c2e0f0c182336c85bbf3cea24a5c0e736498b54dd87e8f3b56fa7850f transformer/model-layer-103.safetensors +7b89be9f61d558b436aa50cb76791fa27acee5ec2b4d5b39b9f81b52e40d2ba9 transformer/model-layer-104.safetensors +997415569ddeae044e68348e149237ce72997afe5807eeed9f153b8cb90ea9f6 transformer/model-layer-105.safetensors +cb839042bf02639fb0f8d7f30245a644ab718865f8219b069eb11ef52a4b3700 transformer/model-layer-106.safetensors +37a4279e3632da1979e7e228bb56fa251f2af33b32423685139680ae37e3ed79 transformer/model-layer-107.safetensors +a72a4ca1452fde96faeb4913cabeb3990afd41fc7a59881fadad9b449e707869 transformer/model-layer-108.safetensors +cdb26f30293cf565b39dcc48432cb34a1127680ddcd98d5a7843d4529026e583 transformer/model-layer-109.safetensors +1640282696d83ecd960d9e53e494a2a0c258c572c0baf40a5d7293e171bf9d70 transformer/model-layer-110.safetensors +c9fdb01cbb69e977e15d19fd5c676a978bd8de8e77e4d87b128fe4e112f1443b transformer/model-layer-111.safetensors +bc8406984a60b35a95cc3d3df9fa4043598fd4c4a7369ba46701c1c82e295329 transformer/model-layer-112.safetensors +153830f28286160f9c24d532bc060ff5a88f66377ee5599a88fa0b75d18206b8 transformer/model-layer-113.safetensors +7c716c2b9efde0599f0d6d5a9c8b3fa372a3ccd35caf1c2f7bae9d59895ce810 transformer/model-layer-114.safetensors +1dc07c10fe6e9ae35c6fb80e6a1db2074a170a149129924908845e610c949414 transformer/model-layer-115.safetensors +a8a6a19494147a2b0b54d1b0c1a478ccb7124da5c85589bfa432f2dd2a79d705 transformer/model-layer-116.safetensors +47073f8778699492cd7c3efd5f518cf43d594f42a5e2bb73163eb5562a0df56d transformer/model-layer-117.safetensors +1e8a2eeee518d9b399cc6d462d86950fd82e2dcec2a1782e49b4f800365b10a5 transformer/model-layer-118.safetensors +9ab2de6d358176842de9cb233f23eb72c979e3d4b930d6d1d1c810ed163b6b47 transformer/model-layer-119.safetensors +1b70e6dcd33f38c84f407d528e1244e392af653b9b20320b7e1b08426d7126d0 transformer/model-layer-120.safetensors +7c25bcdec9d6343adccc80d1d8736e41513043592df3c7d51b843ab9f4f42a56 transformer/model-layer-121.safetensors +09e0aede5e349c62df8cb67a4fe4f384394cc632dabe683c6d60bd011b691b51 transformer/model-layer-122.safetensors +7406a84e2cb4384b5f91f7d096da029c90f05f61c27c433444cd6ab7eb29f0f5 transformer/model-layer-123.safetensors +7f554df5ac153938fe5f19c8da9b386e4e7d0c302cc41cb8c00c9bed86356809 transformer/model-layer-124.safetensors +0d379ea34ac4cb43b0395ccadc262c8d4f3c2467d8705358fbc2ee6ef80abc7e transformer/model-layer-125.safetensors +f6e6bdfd817ab781c7a897d7a1278baf418990d129c5e2aa9fa53da7c6b54445 transformer/model-layer-126.safetensors +6a053d3ce96af8761049b45ba3e9973260c589241075fb6c7aad97af85261d19 transformer/model-layer-127.safetensors +b3ee8d8f7c634f20269b03f2ed4f9b07aa66322b3a1778551ad4cc895eac2f64 transformer/model-layer-128.safetensors +fc5d83671933ce3687969b89fa4d3cbe417d87791f6a8eff115df5b5d5d69507 transformer/model-layer-129.safetensors +bbd42cd7cf806884036cc47f3eb8ce63a39225074069dd937f358df36375f6f7 transformer/model-layer-130.safetensors +35399d61dbfef25d553d2ad4a9ab9aaddaf09d34c7a8c1fb26dd81d0359a3b32 transformer/model-layer-131.safetensors +76891b27ccbe864d5dc81287f1efcc62f24cf9bf2cad8a95873d836ab3df1362 transformer/model-layer-132.safetensors +f652df6c51d36d2f84ac3fb2c64f31c90854b61fa8bf9f1316f541533e52c4b9 transformer/model-layer-133.safetensors +8092cc3d159db04987ec29ea2dc66f311e1dab723891a98be65c28564dd119b1 transformer/model-layer-134.safetensors +86a53245dd56b67d3c6d0f82ead2b958499224c8cc4d4413856380ef5e7c59aa transformer/model-layer-135.safetensors +d7bd44c6263219fd40557e7ee70e0b6f32798933adb6c2a9f6a610905fb228c8 transformer/model-layer-136.safetensors +8d1fa8656256b8c3de2e20012d8c417922d23c2dc2364d52c8e799bc61ef23c9 transformer/model-layer-137.safetensors +1dabfe4025ba54ab98fbadc3186358c3e8a2d8f3f35e88f16fbc8a740345e284 transformer/model-layer-138.safetensors +b5b778bef0930e115b05f6fdafa5f15af628dc0893b9e123d0057ad0c7354dcb transformer/model-layer-139.safetensors +9a2caab7233b4b3e6cba26bdd5cad20e481068d7711edd80ebda3a1818a02bbd transformer/model-layer-140.safetensors +6f185e3ca0aecd04909b919c4b2f27828ca38142707a2db68d37b874a5ae3325 transformer/model-layer-141.safetensors +b8216d42ab6e5a2a8c9acc2ef8e5792c3033aff240b7c7f1c5852a2657f9328c transformer/model-layer-142.safetensors +83a99ec434c8d942ffcdb01095c57eea23c69d936e505c758987a2cb8f83802e transformer/model-layer-143.safetensors +1f2f38a7b782158b024e780af5061c0a925f85217655457302507392d4d7fa1c transformer/model-layer-144.safetensors +55deabfd19590e3c31d09c0289341c28e10a3022f7664cb5b3027beb5b173dc9 transformer/model-layer-145.safetensors +fbdec697030590d2918d16a57cb4df0bd0ed985f86b2cc6e9fc8ecf3d6ec9b74 transformer/model-layer-146.safetensors +79d19289ab817171111ad3ebf21f0d899828d33d7d8294d18acde940b457278f transformer/model-layer-147.safetensors +5bff9dcdbc570a1a708d1e885916384c5250567a6ff9485cbfc5a1acb6ba4e92 transformer/model-layer-148.safetensors +a07bb1a5c12fb027651d81e97be8bb5d1d5f9c0941a293e11aaf352416465feb transformer/model-layer-149.safetensors +595b454b5adf39d47d578425d0565ee58b99b7c15e0ee8460a930718857b2adc transformer/model-layer-150.safetensors +ebb02d3899c2c549bf745ee9ce4bd5a0399c7dcfcece4d5ec095bc97476d0611 transformer/model-layer-151.safetensors +3dde89f1f23a5f29ff717be2a2a466a1fba52770eac4708e4d5c1fc60b0ea6bd transformer/model-layer-152.safetensors +f4c08a3a7f9c4db79f222e1a0bd4292570b1639a4c92077cc3a67731b35fabb0 transformer/model-layer-153.safetensors +7aecccd706d0a48e73ed91c47bf48241a1dab1dcdc35c1211bf7662745e3ab77 transformer/model-layer-154.safetensors +a6d713ad190ddf60d32ccbcce641fa2a4624a34fb7721e4ff7f06c5a591edce7 transformer/model-layer-155.safetensors +aca79eaf47f044ece19a2700c253ed460d14af7f010ecffda533f0f39e21a1af transformer/model-layer-156.safetensors +1a50e43104053ceb8044f7eb23429eb65617aecc4dab6630493256dbedf9f90a transformer/model-layer-157.safetensors +0d6e0b4a265ff0e5b74efc3d2c1531033fdd88215d7f98e12353ed6dcf53bf65 transformer/model-layer-158.safetensors +4ceae84b1d40e869902c0389cdb4cdda6c0e7f25c655df54639a903b5a1dcb23 transformer/model-layer-159.safetensors +f5480155d2336e66accfdc3ff68bdd7add76e2ee69a52d4b79f84008122f4775 transformer/model-layer-160.safetensors +d90c15eda925b9e4078346c3d6f8ab244acc22f7564284ceecaa8da534c792d6 transformer/model-layer-161.safetensors +937689f2e2b07dd3eb5c676de842613df063c09d919512a5a0dd4527595d1a73 transformer/model-layer-162.safetensors +0c0510d9bb9133bbc158ef839b613c405d2c48527efdc4233d40d72bf12729e7 transformer/model-layer-163.safetensors +2304466eb2d328c41b90ea9ccf36a0ce9f8c6c546c61b16b927d2d31b0fa6f3f transformer/model-layer-164.safetensors +2d4c60764ea4134625ba6a53bf680916401dd2d542726773eb23ae98d675d939 transformer/model-layer-165.safetensors +2af0f011d764cb17330079edc8bcd81579008b435ccfa09535d2d617f2540e40 transformer/model-layer-166.safetensors +bfb70cc86e8346d1e8ec99b0f1169c7f09045c6dcafc9f4c83e37e792d8ccc1a transformer/model-layer-167.safetensors +6b62b906fd2e95b94cb49ff99a30ccab6de09f43ab90d822073f06065ef34848 transformer/model-layer-168.safetensors +bd2c6280cb61f2df44303881c2f03cca3b8144085f3103abef7cde079b6ba61c transformer/model-layer-169.safetensors +f6e436975ca7f716bb2cef12cb6f580c04ace3407c9ac93fdcca38dec253fad2 transformer/model-layer-170.safetensors +63d8cdeac345b59db7c7d13ef181db297f722f482ba403221087d88586ca51ae transformer/model-layer-171.safetensors +076ffdc562304bb46f13a63153c276fed0c5e7a016869bb8abff4aa96bc8d73a transformer/model-layer-172.safetensors +92560f36cb2caf425f0f582a59d9cb881dbed4cb9b24b7e269295756e14857b7 transformer/model-layer-173.safetensors +96f258672dea3940c766f99102d2e7875eefdb122bb470c9917b03544a71a7f8 transformer/model-layer-174.safetensors +07a4006bfdf83fa3edf1d76a91d32569d7d01ab6d6df8094f7610695fb12eb5c transformer/model-layer-175.safetensors +a7a469067f641186b3ee98d3599018146123476ea8cf17f34f564fd2b1be7331 transformer/model-layer-176.safetensors +1068a02cb088b6ff81828bfe41ddc29e5ed39d5110fcd0765329eb9eceef13de transformer/model-layer-177.safetensors +232d48d785782768233dabd00869e5fc830e1ed3afdbdfc1584c4e4998e269e6 transformer/model-layer-178.safetensors +65dfba28d0eb981c4511519315124786d17b7dda06a862b950c42e64e257bef3 transformer/model-layer-179.safetensors +154299e21c12fc6f6ffe0c925b680b037e9770ba1f4c290b92ae7e9b6d82ad07 transformer/model-layer-180.safetensors +f8c8f0218cddb0766c81c8e53ed75a381c2dbe6caa51c37a648ecd72566ec9c5 transformer/model-layer-181.safetensors +569cfa0dc65e1f0962ecadd00014ffeafdca746a962b6125fe0a3b695ceecb7b transformer/model-layer-182.safetensors +9a1e032e18799849da0cdcfd12212dd8c64f0d30cfd845c01210b58ca65eda4b transformer/model-layer-183.safetensors +a6eadd8851c292ed4720e86fb1843b4a6aac42c6eb28de8d1655a6c3986337db transformer/model-layer-184.safetensors +538fe7cd6635b778c45ca08f7325345f5ca8011394d22ecd14552fbcdae1b156 transformer/model-layer-185.safetensors +aaea8e82f725e47bfaae1c203e43bb72726925a0f6c7723c71fe9324d96e49dc transformer/model-layer-186.safetensors +65491dfc924613b72091a7023eda6a1bfcfddc935f55fdbb8e9167ae9e926c01 transformer/model-layer-187.safetensors +ee2a07daae8304502a47b24814e4acf020c1db8447d11d9d3b915f1df736f071 transformer/model-layer-188.safetensors +097fd99046f597a1c420a5631c329c570d28d43d9e39246022a04391606aacee transformer/model-layer-189.safetensors +e455979028643775c982bcca099c238373eb398890470844bd752a92935570d1 transformer/model-layer-190.safetensors +f80d4b0a029bf21df4fcb2d1f5dbe64bdca08b944f4f855330b6bed858f62b45 transformer/model-layer-191.safetensors +fc8620ef733372e04139aa28cb66e79521d822f62d0d5a9fae0b8331cc6058f7 transformer/model-layer-192.safetensors +040ceabbe19d1bb7ba6d3ffa003911c04fb9ae058d4904a94eef63214c972ed3 transformer/model-layer-193.safetensors +67d90b1f8937a77cef64eb98f7d4524b7bfb63ca8091e37b5e3d5b9cb7de983f transformer/model-layer-194.safetensors +aa44539d208f33cec16ff224233ed529e099dc3c3ce55171ffe0471df564ab74 transformer/model-layer-195.safetensors +530a027205a2ececbbb5fc4d1d7b583203c02d52cc1510be952d97b688d4a588 transformer/model-layer-196.safetensors +b22addafc782d0026fe5c8a38fe86cde27142db717f5d3814bc21eb5e7c96f72 transformer/model-layer-197.safetensors +d35be6bf6784806af8c1fec6f42fb0c334d26dd7f3f6b0d5eda436ba6eee2f12 transformer/model-layer-198.safetensors +02e7643b75d7b32f660fe066cc27fb25a15a4bbd2de7f7b34893cea28d39c83c transformer/model-layer-199.safetensors +5e130350f6946bfde73715a0772685f89172fc06ffc7887d4e913edd3ca31a64 transformer/model-layer-200.safetensors +4e530d8d190724ff383d8c64672bf4c6c641f3e5690d6078e231d882285c6401 transformer/model-layer-201.safetensors +cd314577248b0ad6f33f2b73e783fd7190298d9d130a9963c91559264f5135ac transformer/model-layer-202.safetensors +7ebaae021b537a48f8d5a819920a6c3ecc5b2313c9f172d53d8db89c90693178 transformer/model-layer-203.safetensors +c3c9287c6a83dcde1b0e60e08276bbd6ea7524c5a663b1d495044a2a6f80354b transformer/model-layer-204.safetensors +1270ede15fe5c4591d2fd57abce23b2a7fa7541bb772e344a1c49ec93d0472c3 transformer/model-layer-205.safetensors +e8cf69b2c8ddf028a77f16d64063431ec46a208f6245e0a769412757e4724d65 transformer/model-layer-206.safetensors +72b08fa0e31cf3bd99e6a6236df6ad282093046df19a01498c78d551ac86251a transformer/model-layer-207.safetensors +e7576cdccbe9919168cd79afc6b123669b103db4a22f85aee43669f0208b6add transformer/model-layer-208.safetensors +406764ed05b591c51c6de957ba08c14b70c3005d7ed7f51b43d0669558c46d32 transformer/model-layer-209.safetensors +cc0fdc115f94600e85fb67716296156b8e961749d5611ea4d20c08f2c53c3e9f transformer/model-layer-210.safetensors +261f7f9c98592065b59de6ab65436755f6e51cb89a1225a9d2617f89dbca90c2 transformer/model-layer-211.safetensors +b0e4e8d429ed025f7861f0d7e24778db961e3e457afdbed0b3b1dd9359a00606 transformer/model-layer-212.safetensors +958b6bc3ed202a331ac2c7356b36a44d15133dabb6dfd9df9a00bedd973e4b37 transformer/model-layer-213.safetensors +4b11f13bd803c408fccf70880038ed91b2bdfdb7dbb9af55e44ffd86d948665b transformer/model-layer-214.safetensors +85d5055784d4dd7b742d080672e1fd48d7de3ad2281c85f7b59136aaaa9ec886 transformer/model-layer-215.safetensors +9d3f778c492229cfffbe3f89fbfb81d70c482289081f102b981f22415eba35ed transformer/model-layer-216.safetensors +b67a198870946e16bf78c3437ef8944224a300da262efacf565742f3b95bd8fa transformer/model-layer-217.safetensors +68c03ece759f1c35bafbadf5aae72a12d747825cef792b335d58f1fa8a029cac transformer/model-layer-218.safetensors +203bdb319411779aebb2bf6d2a897d2b8bc25d5139dc854c7d3e269e9ed61d17 transformer/model-layer-219.safetensors +145609d723d1d92fa9caaec598bb1a4dda4da2fe597d28b01a1e715dee387619 transformer/model-layer-220.safetensors +197f0cac2cc9128fe788c1d97293c1f9bd01166cda4503dc976ba2f322f8a7af transformer/model-layer-221.safetensors +27fe4be4ba4b522fc21c32cb74c93b39217fcb7af95eabb3d9788c54dcd456ce transformer/model-layer-222.safetensors +ba5619d3690bcb7b364efe3868a9e96125a2ca283e3a50ef370b524973fb19dd transformer/model-layer-223.safetensors +9a00de7da981c4e7b52dad8af812688bcf02ad5d21533d5a412b95e9a0815a4b transformer/model.safetensors.index.json +15c66897aa5b7d34158680d5cb8e0710ae8a40c7e06fdec6fc0174bfafbe81b0 vae/config.json +71879ffd5321e6d10c3c87513e2b474b1252efa7f3dec2969214a9bf06a6dd5c vae/diffusion_pytorch_model.safetensors diff --git a/THIRD_PARTY_NOTICES.md b/THIRD_PARTY_NOTICES.md new file mode 100644 index 0000000000000000000000000000000000000000..f9a44a036fb0343f436b6a5db1b7bb059178cbe6 --- /dev/null +++ b/THIRD_PARTY_NOTICES.md @@ -0,0 +1,10 @@ +# Third-party components + +- Qwen Image 2.1, revision `b3179ad355be050328e483a9dfdd9e60cd62adfa`: [Qwen Research License](https://huggingface.co/Qwen/Qwen-Image-2.1/blob/b3179ad355be050328e483a9dfdd9e60cd62adfa/LICENSE). Full local copy: LICENSE. +- Nunchaku 1.2.1: [upstream source and license](https://github.com/nunchaku-tech/nunchaku/tree/v1.2.1). Generic SVDQW4A4Linear runtime; no claim of official Qwen Image 2.1 support. +- DeepCompressor `69f3473f5e1c1504bae35cc50c7858ef900a9b17`: [source and license](https://github.com/mit-han-lab/deepcompressor/tree/69f3473f5e1c1504bae35cc50c7858ef900a9b17). Conversion only. +- Diffusers `80c7ed262aeffbeb43ef13ae04baeb9b84515a69`: [Apache-2.0 source](https://github.com/huggingface/diffusers/tree/80c7ed262aeffbeb43ef13ae04baeb9b84515a69). +- Transformers 5.17.0: [Apache-2.0 source](https://github.com/huggingface/transformers). +- bitsandbytes 0.50.2: [MIT source](https://github.com/bitsandbytes-foundation/bitsandbytes). + +Dependency packages in the container include their own license metadata. These software licenses do not replace the model's research-only terms. diff --git a/model_index.json b/model_index.json new file mode 100644 index 0000000000000000000000000000000000000000..f123b34c9af96af94399de834ff8a04fe4aa2175 --- /dev/null +++ b/model_index.json @@ -0,0 +1,25 @@ +{ + "_class_name": "QwenImage21Pipeline", + "_diffusers_version": "0.41.0.dev0", + "_name_or_path": "Qwen/Qwen-Image-2.1", + "processor": [ + "transformers", + "Qwen3VLProcessor" + ], + "scheduler": [ + "diffusers", + "FlowMatchEulerDiscreteScheduler" + ], + "text_encoder": [ + "transformers", + "Qwen3VLForConditionalGeneration" + ], + "transformer": [ + "diffusers", + "QwenImage21Transformer2DModel" + ], + "vae": [ + "diffusers", + "AutoencoderKLQwenImage21" + ] +} diff --git a/processor/chat_template.jinja b/processor/chat_template.jinja new file mode 100644 index 0000000000000000000000000000000000000000..124386803f142761528f710e77ae483f5f8c4fc4 --- /dev/null +++ b/processor/chat_template.jinja @@ -0,0 +1,120 @@ +{%- if tools %} + {{- '<|im_start|>system\n' }} + {%- if messages[0].role == 'system' %} + {%- if messages[0].content is string %} + {{- messages[0].content }} + {%- else %} + {%- for content in messages[0].content %} + {%- if 'text' in content %} + {{- content.text }} + {%- endif %} + {%- endfor %} + {%- endif %} + {{- '\n\n' }} + {%- endif %} + {{- "# Tools\n\nYou may call one or more functions to assist with the user query.\n\nYou are provided with function signatures within XML tags:\n" }} + {%- for tool in tools %} + {{- "\n" }} + {{- tool | tojson }} + {%- endfor %} + {{- "\n\n\nFor each function call, return a json object with function name and arguments within XML tags:\n\n{\"name\": , \"arguments\": }\n<|im_end|>\n" }} +{%- else %} + {%- if messages[0].role == 'system' %} + {{- '<|im_start|>system\n' }} + {%- if messages[0].content is string %} + {{- messages[0].content }} + {%- else %} + {%- for content in messages[0].content %} + {%- if 'text' in content %} + {{- content.text }} + {%- endif %} + {%- endfor %} + {%- endif %} + {{- '<|im_end|>\n' }} + {%- endif %} +{%- endif %} +{%- set image_count = namespace(value=0) %} +{%- set video_count = namespace(value=0) %} +{%- for message in messages %} + {%- if message.role == "user" %} + {{- '<|im_start|>' + message.role + '\n' }} + {%- if message.content is string %} + {{- message.content }} + {%- else %} + {%- for content in message.content %} + {%- if content.type == 'image' or 'image' in content or 'image_url' in content %} + {%- set image_count.value = image_count.value + 1 %} + {%- if add_vision_id %}Picture {{ image_count.value }}: {% endif -%} + <|vision_start|><|image_pad|><|vision_end|> + {%- elif content.type == 'video' or 'video' in content %} + {%- set video_count.value = video_count.value + 1 %} + {%- if add_vision_id %}Video {{ video_count.value }}: {% endif -%} + <|vision_start|><|video_pad|><|vision_end|> + {%- elif 'text' in content %} + {{- content.text }} + {%- endif %} + {%- endfor %} + {%- endif %} + {{- '<|im_end|>\n' }} + {%- elif message.role == "assistant" %} + {{- '<|im_start|>' + message.role + '\n' }} + {%- if message.content is string %} + {{- message.content }} + {%- else %} + {%- for content_item in message.content %} + {%- if 'text' in content_item %} + {{- content_item.text }} + {%- endif %} + {%- endfor %} + {%- endif %} + {%- if message.tool_calls %} + {%- for tool_call in message.tool_calls %} + {%- if (loop.first and message.content) or (not loop.first) %} + {{- '\n' }} + {%- endif %} + {%- if tool_call.function %} + {%- set tool_call = tool_call.function %} + {%- endif %} + {{- '\n{"name": "' }} + {{- tool_call.name }} + {{- '", "arguments": ' }} + {%- if tool_call.arguments is string %} + {{- tool_call.arguments }} + {%- else %} + {{- tool_call.arguments | tojson }} + {%- endif %} + {{- '}\n' }} + {%- endfor %} + {%- endif %} + {{- '<|im_end|>\n' }} + {%- elif message.role == "tool" %} + {%- if loop.first or (messages[loop.index0 - 1].role != "tool") %} + {{- '<|im_start|>user' }} + {%- endif %} + {{- '\n\n' }} + {%- if message.content is string %} + {{- message.content }} + {%- else %} + {%- for content in message.content %} + {%- if content.type == 'image' or 'image' in content or 'image_url' in content %} + {%- set image_count.value = image_count.value + 1 %} + {%- if add_vision_id %}Picture {{ image_count.value }}: {% endif -%} + <|vision_start|><|image_pad|><|vision_end|> + {%- elif content.type == 'video' or 'video' in content %} + {%- set video_count.value = video_count.value + 1 %} + {%- if add_vision_id %}Video {{ video_count.value }}: {% endif -%} + <|vision_start|><|video_pad|><|vision_end|> + {%- elif 'text' in content %} + {{- content.text }} + {%- endif %} + {%- endfor %} + {%- endif %} + {{- '\n' }} + {%- if loop.last or (messages[loop.index0 + 1].role != "tool") %} + {{- '<|im_end|>\n' }} + {%- endif %} + {%- endif %} +{%- endfor %} +{%- if add_generation_prompt %} + {{- '<|im_start|>assistant\n' }} +{%- endif %} diff --git a/processor/processor_config.json b/processor/processor_config.json new file mode 100644 index 0000000000000000000000000000000000000000..ed4741ad96488761e3fcc3fc63b830bb3b3c87e1 --- /dev/null +++ b/processor/processor_config.json @@ -0,0 +1,65 @@ +{ + "image_processor": { + "data_format": "channels_first", + "default_to_square": true, + "do_convert_rgb": true, + "do_normalize": true, + "do_rescale": true, + "do_resize": true, + "image_mean": [ + 0.5, + 0.5, + 0.5 + ], + "image_processor_type": "Qwen2VLImageProcessor", + "image_std": [ + 0.5, + 0.5, + 0.5 + ], + "merge_size": 2, + "patch_size": 16, + "resample": 3, + "rescale_factor": 0.00392156862745098, + "size": { + "longest_edge": 16777216, + "shortest_edge": 65536 + }, + "temporal_patch_size": 2 + }, + "processor_class": "Qwen3VLProcessor", + "video_processor": { + "data_format": "channels_first", + "default_to_square": true, + "do_convert_rgb": true, + "do_normalize": true, + "do_rescale": true, + "do_resize": true, + "do_sample_frames": true, + "fps": 2, + "image_mean": [ + 0.5, + 0.5, + 0.5 + ], + "image_std": [ + 0.5, + 0.5, + 0.5 + ], + "max_frames": 768, + "max_video_tokens": 768, + "merge_size": 2, + "min_frames": 4, + "patch_size": 16, + "resample": 3, + "rescale_factor": 0.00392156862745098, + "return_metadata": false, + "size": { + "longest_edge": 25165824, + "shortest_edge": 4096 + }, + "temporal_patch_size": 2, + "video_processor_type": "Qwen3VLVideoProcessor" + } +} diff --git a/processor/tokenizer.json b/processor/tokenizer.json new file mode 100644 index 0000000000000000000000000000000000000000..c7afbed2efcdf019f88ab0572ec29d3bf595dfe2 --- /dev/null +++ b/processor/tokenizer.json @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:be75606093db2094d7cd20f3c2f385c212750648bd6ea4fb2bf507a6a4c55506 +size 11422650 diff --git a/processor/tokenizer_config.json b/processor/tokenizer_config.json new file mode 100644 index 0000000000000000000000000000000000000000..e103cbdf14329084db5c9e610e9877b2be7efa9b --- /dev/null +++ b/processor/tokenizer_config.json @@ -0,0 +1,16 @@ +{ + "add_prefix_space": false, + "backend": "tokenizers", + "bos_token": null, + "clean_up_tokenization_spaces": false, + "eos_token": "<|im_end|>", + "errors": "replace", + "is_local": true, + "local_files_only": false, + "model_max_length": 262144, + "pad_token": "<|endoftext|>", + "processor_class": "Qwen3VLProcessor", + "split_special_tokens": false, + "tokenizer_class": "Qwen2Tokenizer", + "unk_token": null +} diff --git a/reproduction/Dockerfile b/reproduction/Dockerfile new file mode 100644 index 0000000000000000000000000000000000000000..e0b60d8a29d39258e3cdb72efafe328fa0cf7cbe --- /dev/null +++ b/reproduction/Dockerfile @@ -0,0 +1,17 @@ +# Cached home-PC base; immutable digest preserves torch 2.8.0 / CUDA 12.8. +FROM mesmerlord/flux2-klein-runpod@sha256:c6622a7307520c3001afe7969bdb44858f3b27d100a4066285a5f54e14a4af0e +USER root +WORKDIR /poc +RUN python -m pip install --no-cache-dir transformers==5.17.0 bitsandbytes==0.50.2 accelerate==1.12.0 huggingface-hub==1.32.0 tokenizers==0.23.2 safetensors==0.8.0 pydantic==2.13.4 httpx==0.28.1 pillow==12.2.0 starlette==0.47.3 fastapi==0.116.1 uvicorn==0.35.0 \ + 'diffusers @ https://github.com/huggingface/diffusers/archive/80c7ed262aeffbeb43ef13ae04baeb9b84515a69.zip' +RUN python -m pip install --no-cache-dir --no-deps torchvision==0.23.0 --index-url https://download.pytorch.org/whl/cu128 +ADD https://github.com/mit-han-lab/deepcompressor/archive/69f3473f5e1c1504bae35cc50c7858ef900a9b17.tar.gz /tmp/deepcompressor.tar.gz +RUN mkdir -p /opt/deepcompressor && tar -xzf /tmp/deepcompressor.tar.gz --strip-components=1 -C /opt/deepcompressor && rm /tmp/deepcompressor.tar.gz +COPY runner.py lean_encoder.py server.py /poc/ +COPY nunchaku_backend/ /poc/nunchaku_backend/ +COPY scripts/download.py /poc/scripts/ +ENV KLEIN_LIGHT_IMPORTS=0 HF_HUB_OFFLINE=0 TRANSFORMERS_OFFLINE=0 HF_HOME=/cache/huggingface TORCHINDUCTOR_CACHE_DIR=/cache/torchinductor TRITON_CACHE_DIR=/cache/triton PYTHONUNBUFFERED=1 +ENV PYTHONPATH=/opt/deepcompressor:/poc +ENTRYPOINT [] +HEALTHCHECK --interval=30s --timeout=5s --start-period=120s --retries=3 CMD python -c "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8091/readyz',timeout=3)" +CMD ["uvicorn", "server:app", "--host", "0.0.0.0", "--port", "8091"] diff --git a/reproduction/LICENSE b/reproduction/LICENSE new file mode 100644 index 0000000000000000000000000000000000000000..13ae08d5a5828f508cf0e852b1b316c0c0bbb9b3 --- /dev/null +++ b/reproduction/LICENSE @@ -0,0 +1,55 @@ +Qwen RESEARCH LICENSE AGREEMENT + +Qwen RESEARCH LICENSE AGREEMENT Release Date: September 20, 2026 + +By clicking to agree or by using or distributing any portion or element of the Qwen Materials, you will be deemed to have recognized and accepted the content of this Agreement, which is effective immediately. + +1. Definitions + a. This Qwen RESEARCH LICENSE AGREEMENT (this "Agreement") shall mean the terms and conditions for use, reproduction, distribution and modification of the Materials as defined by this Agreement. + b. "We" (or "Us") shall mean Hangzhou Tongyi Laboratory Technology Co., Ltd. + c. "You" (or "Your") shall mean a natural person or legal entity exercising the rights granted by this Agreement and/or using the Materials for any purpose and in any field of use. + d. "Third Parties" shall mean individuals or legal entities that are not under common control with us or you. + e. "Qwen" shall mean the large language models, diffusion models, and software and algorithms, consisting of trained model weights, parameters (including optimizer states), machine-learning model code, inference-enabling code, training-enabling code, fine-tuning enabling code and other elements of the foregoing distributed by us. + f. "Materials" shall mean, collectively, our proprietary Qwen and Documentation (and any portion thereof) made available under this Agreement. + g. "Source" form shall mean the preferred form for making modifications, including but not limited to model source code, documentation source, and configuration files. + h. "Object" form shall mean any form resulting from mechanical transformation or translation of a Source form, including but not limited to compiled object code, generated documentation, and conversions to other media types. + i. "Non-Commercial" shall mean for research or evaluation purposes only. + +2. Grant of Rights + a. You are granted a non-exclusive, worldwide, non-transferable and royalty-free limited license under our intellectual property or other rights owned by us embodied in the Materials to use, reproduce, distribute, copy, create derivative works of, and make modifications to the Materials FOR NON-COMMERCIAL PURPOSES ONLY. + b. You shall not use the Materials for any commercial purpose without obtaining a separate commercial license from us. If you wish to use the Materials commercially, you shall request a license from us at model-business@notice.qwencloud.com. + +3. Redistribution +Subject to Section 2 (Grant of Rights), you may distribute copies or make the Materials, or derivative works thereof, available as part of a product or service that contains any of them, with or without modifications, and in Source or Object form, provided that you meet the following conditions: + a. You shall give any other recipients of the Materials or derivative works a copy of this Agreement; + b. You shall cause any modified files to carry prominent notices stating that you changed the files; + c. You shall retain in all copies of the Materials that you distribute the following attribution notices within a "Notice" text file distributed as a part of such copies: "Qwen is licensed under the Qwen RESEARCH LICENSE AGREEMENT, Copyright (c) 2026 Hangzhou Tongyi Laboratory Technology Co., Ltd. All Rights Reserved."; and + d. You may add your own copyright statement to your modifications and may provide additional or different license terms and conditions for use, reproduction, or distribution of your modifications, or for any such derivative works as a whole, provided your use, reproduction, and distribution of the work otherwise complies with the terms and conditions of this Agreement. + +4. Rules of use + a. The Materials may be subject to export controls or restrictions in China, the United States or other countries or regions. You shall comply with applicable laws and regulations in your use of the Materials. + b. If you use the Materials or any outputs or results therefrom to create, train, fine-tune, or improve an AI model that is distributed or made available, you shall prominently display “Built with Qwen” or “Improved using Qwen” in the related product documentation. + c. You shall not use "Qwen" as the primary name or identifier of any derivative works or products; reasonable descriptive use (e.g., "fine-tuned from Qwen Image") is permitted. + +5. Intellectual Property + a. We retain ownership of all intellectual property rights in and to the Materials and derivatives made by or for us. Conditioned upon compliance with the terms and conditions of this Agreement, with respect to any derivative works and modifications of the Materials that are made by you, you are and will be the owner of such derivative works and modifications. + b. No trademark license is granted to use the trade names, trademarks, service marks, or product names of us, except as required to fulfill notice requirements under this Agreement or as required for reasonable and customary use in describing and redistributing the Materials. + c. If you commence a lawsuit or other proceedings (including a cross-claim or counterclaim in a lawsuit) against us or any entity alleging that the Materials or any output therefrom, or any part of the foregoing, infringe any intellectual property or other right owned or licensable by you, then all licenses granted to you under this Agreement shall terminate as of the date such lawsuit or other proceeding is commenced or brought. + +6. Disclaimer of Warranty and Limitation of Liability + a. We are not obligated to support, update, provide training for, or develop any further version of the Qwen Materials or to grant any license thereto. + b. THE MATERIALS ARE PROVIDED "AS IS" WITHOUT ANY EXPRESS OR IMPLIED WARRANTY OF ANY KIND INCLUDING WARRANTIES OF MERCHANTABILITY, NONINFRINGEMENT, OR FITNESS FOR A PARTICULAR PURPOSE. WE MAKE NO WARRANTY AND ASSUME NO RESPONSIBILITY FOR THE SAFETY OR STABILITY OF THE MATERIALS AND ANY OUTPUT THEREFROM. + c. IN NO EVENT SHALL WE BE LIABLE TO YOU FOR ANY DAMAGES, INCLUDING, BUT NOT LIMITED TO ANY DIRECT, OR INDIRECT, SPECIAL OR CONSEQUENTIAL DAMAGES ARISING FROM YOUR USE OR INABILITY TO USE THE MATERIALS OR ANY OUTPUT OF IT, NO MATTER HOW IT’S CAUSED. + d. You will defend, indemnify and hold harmless us from and against any claim by any third party arising out of or related to your use or distribution of the Materials. + +7. Survival and Termination. + a. The term of this Agreement shall commence upon your acceptance of this Agreement or access to the Materials and will continue in full force and effect until terminated in accordance with the terms and conditions herein. + b. We may terminate this Agreement if you breach any of the terms or conditions of this Agreement. Upon termination of this Agreement, you must delete and cease use of the Materials. Sections 6 and 8 shall survive the termination of this Agreement. + +8. Governing Law and Jurisdiction. + a. This Agreement and any dispute arising out of or relating to it will be governed by the laws of China, without regard to conflict of law principles, and the UN Convention on Contracts for the International Sale of Goods does not apply to this Agreement. + b. The People's Courts in Hangzhou City shall have exclusive jurisdiction over any dispute arising out of this Agreement. + +9. Other Terms and Conditions. + a. Any arrangements, understandings, or agreements regarding the Material not stated herein are separate from and independent of the terms and conditions of this Agreement. You shall request a separate license from us, if you use the Materials in ways not expressly agreed to in this Agreement. + b. We shall not be bound by any additional or different terms or conditions communicated by you unless expressly agreed. diff --git a/reproduction/NOTICE b/reproduction/NOTICE new file mode 100644 index 0000000000000000000000000000000000000000..8bd7a1155eaca7e16f36a644a6a7d7cb3eec7141 --- /dev/null +++ b/reproduction/NOTICE @@ -0,0 +1,9 @@ +Built with Qwen + +Qwen is licensed under the Qwen RESEARCH LICENSE AGREEMENT, Copyright (c) 2026 Hangzhou Tongyi Laboratory Technology Co., Ltd. All Rights Reserved. + +Mesmer Image 21 Nunchaku is a modified, independently calibrated quantization of Qwen Image 2.1. The transformer uses Nunchaku signed INT4 weights and activations with BF16 rank128 residual branches. The text encoder is serialized in bitsandbytes NF4; processor, scheduler and VAE originate from the pinned upstream release. These modifications are by MesmerTech, September 2026, and are not an official Qwen or Nunchaku release. + +The model and derivatives are for non-commercial research and evaluation under the accompanying LICENSE. Commercial use requires a separate license from Qwen. + +Runtime dependencies retain their respective licenses. The custom linear runtime uses MIT HAN Lab Nunchaku; the conversion packing adapter uses DeepCompressor. See THIRD_PARTY_NOTICES.md for pinned source and license links. diff --git a/reproduction/REPORT.md b/reproduction/REPORT.md new file mode 100644 index 0000000000000000000000000000000000000000..7d25dc36bfd83be51b87deefedf21b67928279b7 --- /dev/null +++ b/reproduction/REPORT.md @@ -0,0 +1,91 @@ +# Qwen Image 2.1 fidelity revision — September 21, 2026 + +The new evaluation contains 21 scenarios at both 25 and 40 steps. It includes interacting groups with specific roles, a wheelchair award ceremony, a family reunion, a radio interview, dense bilingual text, a three-panel comic, museum geometry, product layouts, local preservation edits and two-reference composition. Initial samples remain archived; all new comparisons are in `samples/fidelity-v3/`. + +The serving choice is calibrated rank 128. The larger selective rank 512 candidate is retained as an experiment: despite better numerical error, direct image inspection found extra objects and geometry regressions. All 126 comparison images (42 each for the BF16-transformer teacher, calibrated rank 128 and selective rank 512) are complete, locally verified and covered by direct visual reviews. These are 42 matched jobs, not 126 independent prompts. The selected API is running and its functional generation/edit/cache checks are complete; strict same-seed repeatability failed, as documented below. + +## What changed in the converter + +The previous one-pass rank 32 approximation was replaced by calibration against actual Nunchaku CUDA output. The new search selects smoothing and low-rank residual fits using validation output error, checks held-out rows after selection, and accepts optional GPTQ correction only when validation improves. Rank 128 is used across 224 quantized linear layers. Each linear remains W4A4 plus a BF16 low-rank branch; original normalization, positional encoding, attention and KV-cache behavior remain in the Diffusers implementation. + +The SVDQuant paper and official Nunchaku/DeepCompressor implementation motivated the calibration and decomposition work. This is a custom Qwen Image 2.1 adapter and bounded calibration implementation, not a reproduction of every official calibration setting. The exact shipped Klein calibration recipe is not public. See the [SVDQuant source review](../../research/svdquant-paper-v3.md) and [converter audit](../../research/nunchaku-v3-audit.md). + +A follow-up sensitivity sweep found that the MLP projection and output layers contribute substantially to denoiser error. Temporarily restoring both roles to BF16 reduced the three diagnostic relative-L2 errors to 5.61%, 2.74% and 9.28%, but adds roughly 4.15 GiB of model state. This is a diagnostic, not a measured image-quality or safe two-reference serving configuration. + +The practical candidate instead raises only the 32 MLP projection low-rank branches from 128 to 512. All 32 selected 512 by validation. Other 192 quantized linears and 193 untouched shards remain byte-identical. It retains all 224 INT4 kernels and adds 384 MiB of model tensors. The complete checkpoint has 5,057,830,912 bytes (4.710 GiB) of model tensors. Independent CPU loading verified 1,417 finite state tensors, every output/source shard hash and all 193 untouched shard identities; load time was 3.48 seconds. Across the 32 upgraded layers, held-out MSE ratios to actual rank 128 range from 0.531 to 0.656, with median 0.642. Model state is not peak VRAM. + +## Numerical evidence and rejected changes + +The first rank 128 converter improved held-out linear MSE against a regenerated one-pass rank 32 recipe in all 224 linears; median ratio was 0.6064. That numerical baseline is not the exact archived v2 checkpoint. Actual v2 images and whole-denoiser checks use its real saved checkpoint. + +Three fixed-recipe runs allowing 100 fitting iterations did not justify a full conversion: one control stayed unchanged, one slightly better validation result worsened held-out error, and another slightly regressed. More iterations were rejected instead of presumed beneficial. + +The rank 512 probe improved held-out projection MSE by 35.5–38.6% in three representative layers. Through the complete MLP, improvement was smaller: 7.9–12.2%. Block 11 rank 256 improved raw projection error while slightly worsening the parent MLP's validation error. These checks demonstrate why linear error alone cannot establish image fidelity. + +| Actual checkpoint | First-step denoiser relative L2 | Middle step | Final step | +|---|---:|---:|---:| +| Archived v2 |12.48%|6.54%|21.87%| +| Calibrated rank 128 |8.85%|5.08%|19.43%| +| Selective MLP rank 512 |8.46%|4.72%|17.93%| + +This is one disjoint development editing prompt at steps 0/20/39 of 40. Every checkpoint builds its own prefix KV cache; no teacher cache is reused. These are conditional predictions on the teacher trajectory, not free-running image scores. + +## Image review and step count + +Forty steps is the [official starting recommendation](https://huggingface.co/docs/diffusers/main/api/pipelines/qwenimage21). Both 25 and 40 are evaluated here at 1024×1024, with fixed seeds, CFG 1 and prefix KV caching. More steps do not reliably fix incorrect counts or relationships. + +The rank 128 review covers all 42 teacher/candidate pairs. It restores the fantasy telescope and improves the old pancake-count error, retains all 14 bilingual poster strings and six comic dialogue lines, and performs the local coat/poster edits with strong preservation. It still changes some faces, hand positions, poses, object scale and design details. Its 25-step chess banner loses a digit in 2026; 40 steps restores it. Its museum drawings add tables and confuse routes at both step counts. The product campaign improves requested cup placement relative to the teacher, demonstrating that visual similarity and prompt compliance are different measurements. + +The BF16 teacher also has real limitations: its 25-step chess title reads CIT CHES FINAL, the comic does not stage keys under the chair, the product relocation edit leaves its cup on the pedestal, and the two-reference portrait looks off-camera. These are shared model/instruction failures, not evidence that every candidate difference is caused by quantization. The teacher uses the same NF4 encoder, so it isolates transformer approximation rather than representing a fully BF16 pipeline. + +Direct observations and image hashes are recorded in the [rank 128 scorecard](../../research/quality-v3-scorecard.json), with root versus independent reviewer attribution, and the [complete selective rank 512 review index](../../research/quality-v3-rank512-review-index.md). Similarity metrics compare the same labels/seeds and unaligned 1024×1024 images. SSIM and PSNR measure resemblance, not a percentage of semantic quality. + +The **matched original 18-job cohort** is available for all four candidates. Each is compared with the corresponding BF16-transformer image; these medians pool nine scenarios at both step counts. + +| Candidate | Images | Median SSIM | Median PSNR (dB) | +|---|---:|---:|---:| +| Archived Nunchaku v2 | 18 | 0.7608 | 18.52 | +| Calibrated rank 128 | 18 | 0.8508 | 20.57 | +| Selective rank 512 | 18 | 0.8287 | 19.96 | +| NF4 transformer | 18 | 0.8237 | 20.84 | + +The **full matched 42-job cohort** is available for the two new Nunchaku candidates. It includes the original 18 jobs plus 24 expanded jobs. + +| Candidate | Images | Median SSIM | Median PSNR (dB) | +|---|---:|---:|---:| +| Calibrated rank 128 | 42 | 0.8269 | 19.99 | +| Selective rank 512 | 42 | 0.8207 | 19.62 | + +Rank 128 has higher median SSIM than NF4 on the original 18 jobs, while NF4 has higher median PSNR. Selective rank 512 trails rank 128 on both medians in both matched cohorts despite its better denoiser probe. NF4 was not measured on the full 42-job cohort. Comparisons across the two tables would mix different prompt sets. Exact labels and per-image metrics are in the [original 18-job results](../../results/compare-fidelity-v3-original18.json) and [full 42-job results](../../results/compare-fidelity-v3-full42.json). + +## Selective rank 512 decision + +The larger branch is not promoted. All 42 selective outputs now have direct review coverage, including fixed-input checks for all 14 edits. In the 12 original generation jobs it improves the thin perfume peel and restores a more teacher-like towel grip at 40 steps, but produces seven blueberries and two knives in the 40-step breakfast, two cats in the 25-step fantasy scene, and a distorted horn-like telescope at 40. Botanical wording and lemon counts remain correct. + +The 16 expanded generation jobs show further mixed effects. Selective rank 512 restores the chess year and camera-at-eye action, but adds six kitchen rolls at 25 steps and six/five cafe tables at 25/40. The 40-step museum also gains at least three shelf groups instead of two. Bilingual text and comic dialogue remain readable, while the comic still fails to stage the keys under the chair. Product cups sit correctly on the tabletop at 25 but return to the pedestal at 40, matching a teacher error and losing rank 128's adherence improvement. Forty steps trades failures rather than consistently repairing them. + +Localized towel, coat and poster edits retain strong visual preservation, with small texture and typography changes and no decisive broad improvement over rank 128. The 40-step two-reference portrait restores the unobscured bottle emblem and supporting grip more closely to the teacher; off-camera gaze and rendering drift remain. Product-relocation edits still leave the cup on the pedestal in all compared backends. These observations support retaining the smaller checkpoint for this POC, not a claim that rank 128 is universally better or near-lossless. + +## Timing, memory and reproducibility + +For the selected rank 128 checkpoint, median generation takes **15.50 s at 25 steps / 22.82 s at 40**, versus **27.51 s / 42.76 s** for the streamed BF16 transformer. These medians use the same 14 generation scenarios per step count. Median one-reference edit inference takes **19.48 s / 28.43 s** across six matched scenarios per step count. The single two-reference portrait takes **23.76 s / 34.35 s**; these are individual runs, not repeated-run medians. Sampled board occupancy reaches **11,994 MiB (11.71 GiB)** in that two-reference case. Selective rank 512 is slightly slower (**16.06 s / 23.56 s** median generation) and reaches **12,374 MiB** with two references. + +See [performance.md](performance.md) and the [per-job measurement summary](../../results/fidelity-v3-summary.json) for verified per-job measurements. CLI comparisons disable prompt/reference LRU reuse, retain per-request prefix KV, use an untiled BF16 VAE and release dead KV before decoding. Timings cover engine inference, excluding process startup and image file loading/saving. Board occupancy is sampled separately from allocator peaks; initial teacher jobs without board telemetry are explicitly missing those values. + +GPU 0 is an RTX 4070 Ti SUPER with 16,376 MiB. Klein remains on its existing GPU 1. GPU jobs run serially. The implementation stages the large NF4 encoder on CPU between requests; it retains all 36 decoder layers and removes only the unused LM-head/final-normalization path after exact feature-parity verification. Persistent compilation caches and optional compilation remain available; this Nunchaku quality revision uses eager inference. + +Reproduction commands are in the experiment README, including calibration, checkpoint conversion, selective rank upgrade and the frozen 42-job image suite. Docker/model/library revisions are pinned. Saved source state, service rollback information and historical failures remain preserved. + +## Service verification + +The calibrated rank128 API is healthy on PC 2 loopback port 8091. Four real 25-step requests completed: bilingual generation at 18.22 s, cached repeat 14.81 s, two-reference portrait 26.42 s, and cached repeat 22.93 s. These are client wall times from one run each. The prompt cache hit on both repeats and the reference-latent cache hit twice on the two-reference repeat. All four responses were COMPLETED with valid 1024×1024 PNGs. Engine startup measured 8.78 s from the saved checkpoint, including ML imports; this is a cached-disk process start, not machine boot. Maximum sampled API board occupancy was 11,986 MiB. See [final API measurements](../../results/api-fidelity-v3-final.json). + +**Strict byte-identical repeatability failed.** The failure is preserved, not converted into a tolerance pass. A controlled four-run diagnostic compared two uncached runs and a cache miss/hit: all prompt tensors and first-transformer inputs were exactly equal in values, shapes, strides and dtypes, and stored cache tensors remained unchanged. Nevertheless the first transformer outputs differed even without caching. Uncached repeats had RGB MAE 5.06/255; cache miss/hit had 6.48/255. This isolates the observed divergence to the transformer forward, but does not identify the exact kernel or operation. Official low-rank atomic reductions are a hypothesis, not a confirmed sole cause. See [diagnosis](../../research/api-cache-v3-diagnosis.md) and [input fingerprints](../../results/repeatability-diagnostic/diagnostic.json). + +The final API verification used an explicit record-differences mode so it could preserve and inspect every returned image while retaining `strict_repeatability_passed=false`. Both posters retained the requested wording; the two-reference repeats changed hand placement, bench and suitcase position while retaining the person/product attributes. Same-seed variation can affect composition, not merely file bytes. The image-suite comparisons therefore describe individual recorded realizations; they do not establish repeatable equivalence or isolate every visual difference solely to quantization. + +Klein remains healthy with the same container and original start time on GPU1; Qwen alone occupies GPU0 for this POC. Higgs and GPU0 Comfy remain stopped under the user's authorization, GPU1 Comfy and connectivity remain intact, and experiment telemetry is stopped. [Final service state](../../results/final-state-v3.json) and `scripts/rollback.sh` preserve the recovery path. + +## Limits + +This is a 1024×1024 POC below the release's recommended native 2K resolution. It does not establish near-lossless quality, superiority to Klein, full-BF16 encoder fidelity, ten-reference operation, native 2K speed or a production-router rollout. Calibration is small and edit splits share one source photo; prefix/reference-token coverage is limited. The visual suite was reused after earlier candidate results informed further work, so these are diagnostic comparisons rather than a fresh unseen acceptance set. Joint attention and full-block optimization are not implemented by this converter. The model uses the Qwen Research License; the downloaded license is retained with the release research. diff --git a/reproduction/THIRD_PARTY_NOTICES.md b/reproduction/THIRD_PARTY_NOTICES.md new file mode 100644 index 0000000000000000000000000000000000000000..f9a44a036fb0343f436b6a5db1b7bb059178cbe6 --- /dev/null +++ b/reproduction/THIRD_PARTY_NOTICES.md @@ -0,0 +1,10 @@ +# Third-party components + +- Qwen Image 2.1, revision `b3179ad355be050328e483a9dfdd9e60cd62adfa`: [Qwen Research License](https://huggingface.co/Qwen/Qwen-Image-2.1/blob/b3179ad355be050328e483a9dfdd9e60cd62adfa/LICENSE). Full local copy: LICENSE. +- Nunchaku 1.2.1: [upstream source and license](https://github.com/nunchaku-tech/nunchaku/tree/v1.2.1). Generic SVDQW4A4Linear runtime; no claim of official Qwen Image 2.1 support. +- DeepCompressor `69f3473f5e1c1504bae35cc50c7858ef900a9b17`: [source and license](https://github.com/mit-han-lab/deepcompressor/tree/69f3473f5e1c1504bae35cc50c7858ef900a9b17). Conversion only. +- Diffusers `80c7ed262aeffbeb43ef13ae04baeb9b84515a69`: [Apache-2.0 source](https://github.com/huggingface/diffusers/tree/80c7ed262aeffbeb43ef13ae04baeb9b84515a69). +- Transformers 5.17.0: [Apache-2.0 source](https://github.com/huggingface/transformers). +- bitsandbytes 0.50.2: [MIT source](https://github.com/bitsandbytes-foundation/bitsandbytes). + +Dependency packages in the container include their own license metadata. These software licenses do not replace the model's research-only terms. diff --git a/reproduction/artifacts/qwen21-calibration/activation_stats.safetensors b/reproduction/artifacts/qwen21-calibration/activation_stats.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..72879ea6de39c33950ab5094e38d57cbac47a5d7 --- /dev/null +++ b/reproduction/artifacts/qwen21-calibration/activation_stats.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:7f032d0930881f8352309e10084b3e5e4418ac81fb1c69a9bc7457f5fa731018 +size 4743896 diff --git a/reproduction/artifacts/qwen21-calibration/calibration.json b/reproduction/artifacts/qwen21-calibration/calibration.json new file mode 100644 index 0000000000000000000000000000000000000000..720137c2fcef8be50abae0dbc38029068081cd62 --- /dev/null +++ b/reproduction/artifacts/qwen21-calibration/calibration.json @@ -0,0 +1,293 @@ +{ + "format": "qwen21-input-absmax-v1", + "source": "nf4-trajectories-approximate", + "observations": { + "transformer_blocks.0.attn.to_q": 82130, + "transformer_blocks.0.attn.to_k": 82130, + "transformer_blocks.0.attn.to_v": 82130, + "transformer_blocks.0.attn.to_out.0": 82130, + "transformer_blocks.0.img_mlp.gate_layer": 82130, + "transformer_blocks.0.img_mlp.proj": 82130, + "transformer_blocks.0.img_mlp.out": 82130, + "transformer_blocks.1.attn.to_q": 82130, + "transformer_blocks.1.attn.to_k": 82130, + "transformer_blocks.1.attn.to_v": 82130, + "transformer_blocks.1.attn.to_out.0": 82130, + "transformer_blocks.1.img_mlp.gate_layer": 82130, + "transformer_blocks.1.img_mlp.proj": 82130, + "transformer_blocks.1.img_mlp.out": 82130, + "transformer_blocks.2.attn.to_q": 82130, + "transformer_blocks.2.attn.to_k": 82130, + "transformer_blocks.2.attn.to_v": 82130, + "transformer_blocks.2.attn.to_out.0": 82130, + "transformer_blocks.2.img_mlp.gate_layer": 82130, + "transformer_blocks.2.img_mlp.proj": 82130, + "transformer_blocks.2.img_mlp.out": 82130, + "transformer_blocks.3.attn.to_q": 82130, + "transformer_blocks.3.attn.to_k": 82130, + "transformer_blocks.3.attn.to_v": 82130, + "transformer_blocks.3.attn.to_out.0": 82130, + "transformer_blocks.3.img_mlp.gate_layer": 82130, + "transformer_blocks.3.img_mlp.proj": 82130, + "transformer_blocks.3.img_mlp.out": 82130, + "transformer_blocks.4.attn.to_q": 82130, + "transformer_blocks.4.attn.to_k": 82130, + "transformer_blocks.4.attn.to_v": 82130, + "transformer_blocks.4.attn.to_out.0": 82130, + "transformer_blocks.4.img_mlp.gate_layer": 82130, + "transformer_blocks.4.img_mlp.proj": 82130, + "transformer_blocks.4.img_mlp.out": 82130, + "transformer_blocks.5.attn.to_q": 82130, + "transformer_blocks.5.attn.to_k": 82130, + "transformer_blocks.5.attn.to_v": 82130, + "transformer_blocks.5.attn.to_out.0": 82130, + "transformer_blocks.5.img_mlp.gate_layer": 82130, + "transformer_blocks.5.img_mlp.proj": 82130, + "transformer_blocks.5.img_mlp.out": 82130, + "transformer_blocks.6.attn.to_q": 82130, + "transformer_blocks.6.attn.to_k": 82130, + "transformer_blocks.6.attn.to_v": 82130, + "transformer_blocks.6.attn.to_out.0": 82130, + "transformer_blocks.6.img_mlp.gate_layer": 82130, + "transformer_blocks.6.img_mlp.proj": 82130, + "transformer_blocks.6.img_mlp.out": 82130, + "transformer_blocks.7.attn.to_q": 82130, + "transformer_blocks.7.attn.to_k": 82130, + "transformer_blocks.7.attn.to_v": 82130, + "transformer_blocks.7.attn.to_out.0": 82130, + "transformer_blocks.7.img_mlp.gate_layer": 82130, + "transformer_blocks.7.img_mlp.proj": 82130, + "transformer_blocks.7.img_mlp.out": 82130, + "transformer_blocks.8.attn.to_q": 82130, + "transformer_blocks.8.attn.to_k": 82130, + "transformer_blocks.8.attn.to_v": 82130, + "transformer_blocks.8.attn.to_out.0": 82130, + "transformer_blocks.8.img_mlp.gate_layer": 82130, + "transformer_blocks.8.img_mlp.proj": 82130, + "transformer_blocks.8.img_mlp.out": 82130, + "transformer_blocks.9.attn.to_q": 82130, + "transformer_blocks.9.attn.to_k": 82130, + "transformer_blocks.9.attn.to_v": 82130, + "transformer_blocks.9.attn.to_out.0": 82130, + "transformer_blocks.9.img_mlp.gate_layer": 82130, + "transformer_blocks.9.img_mlp.proj": 82130, + "transformer_blocks.9.img_mlp.out": 82130, + "transformer_blocks.10.attn.to_q": 82130, + "transformer_blocks.10.attn.to_k": 82130, + "transformer_blocks.10.attn.to_v": 82130, + "transformer_blocks.10.attn.to_out.0": 82130, + "transformer_blocks.10.img_mlp.gate_layer": 82130, + "transformer_blocks.10.img_mlp.proj": 82130, + "transformer_blocks.10.img_mlp.out": 82130, + "transformer_blocks.11.attn.to_q": 82130, + "transformer_blocks.11.attn.to_k": 82130, + "transformer_blocks.11.attn.to_v": 82130, + "transformer_blocks.11.attn.to_out.0": 82130, + "transformer_blocks.11.img_mlp.gate_layer": 82130, + "transformer_blocks.11.img_mlp.proj": 82130, + "transformer_blocks.11.img_mlp.out": 82130, + "transformer_blocks.12.attn.to_q": 82130, + "transformer_blocks.12.attn.to_k": 82130, + "transformer_blocks.12.attn.to_v": 82130, + "transformer_blocks.12.attn.to_out.0": 82130, + "transformer_blocks.12.img_mlp.gate_layer": 82130, + "transformer_blocks.12.img_mlp.proj": 82130, + "transformer_blocks.12.img_mlp.out": 82130, + "transformer_blocks.13.attn.to_q": 82130, + "transformer_blocks.13.attn.to_k": 82130, + "transformer_blocks.13.attn.to_v": 82130, + "transformer_blocks.13.attn.to_out.0": 82130, + "transformer_blocks.13.img_mlp.gate_layer": 82130, + "transformer_blocks.13.img_mlp.proj": 82130, + "transformer_blocks.13.img_mlp.out": 82130, + "transformer_blocks.14.attn.to_q": 82130, + "transformer_blocks.14.attn.to_k": 82130, + "transformer_blocks.14.attn.to_v": 82130, + "transformer_blocks.14.attn.to_out.0": 82130, + "transformer_blocks.14.img_mlp.gate_layer": 82130, + "transformer_blocks.14.img_mlp.proj": 82130, + "transformer_blocks.14.img_mlp.out": 82130, + "transformer_blocks.15.attn.to_q": 82130, + "transformer_blocks.15.attn.to_k": 82130, + "transformer_blocks.15.attn.to_v": 82130, + "transformer_blocks.15.attn.to_out.0": 82130, + "transformer_blocks.15.img_mlp.gate_layer": 82130, + "transformer_blocks.15.img_mlp.proj": 82130, + "transformer_blocks.15.img_mlp.out": 82130, + "transformer_blocks.16.attn.to_q": 82130, + "transformer_blocks.16.attn.to_k": 82130, + "transformer_blocks.16.attn.to_v": 82130, + "transformer_blocks.16.attn.to_out.0": 82130, + "transformer_blocks.16.img_mlp.gate_layer": 82130, + "transformer_blocks.16.img_mlp.proj": 82130, + "transformer_blocks.16.img_mlp.out": 82130, + "transformer_blocks.17.attn.to_q": 82130, + "transformer_blocks.17.attn.to_k": 82130, + "transformer_blocks.17.attn.to_v": 82130, + "transformer_blocks.17.attn.to_out.0": 82130, + "transformer_blocks.17.img_mlp.gate_layer": 82130, + "transformer_blocks.17.img_mlp.proj": 82130, + "transformer_blocks.17.img_mlp.out": 82130, + "transformer_blocks.18.attn.to_q": 82130, + "transformer_blocks.18.attn.to_k": 82130, + "transformer_blocks.18.attn.to_v": 82130, + "transformer_blocks.18.attn.to_out.0": 82130, + "transformer_blocks.18.img_mlp.gate_layer": 82130, + "transformer_blocks.18.img_mlp.proj": 82130, + "transformer_blocks.18.img_mlp.out": 82130, + "transformer_blocks.19.attn.to_q": 82130, + "transformer_blocks.19.attn.to_k": 82130, + "transformer_blocks.19.attn.to_v": 82130, + "transformer_blocks.19.attn.to_out.0": 82130, + "transformer_blocks.19.img_mlp.gate_layer": 82130, + "transformer_blocks.19.img_mlp.proj": 82130, + "transformer_blocks.19.img_mlp.out": 82130, + "transformer_blocks.20.attn.to_q": 82130, + "transformer_blocks.20.attn.to_k": 82130, + "transformer_blocks.20.attn.to_v": 82130, + "transformer_blocks.20.attn.to_out.0": 82130, + "transformer_blocks.20.img_mlp.gate_layer": 82130, + "transformer_blocks.20.img_mlp.proj": 82130, + "transformer_blocks.20.img_mlp.out": 82130, + "transformer_blocks.21.attn.to_q": 82130, + "transformer_blocks.21.attn.to_k": 82130, + "transformer_blocks.21.attn.to_v": 82130, + "transformer_blocks.21.attn.to_out.0": 82130, + "transformer_blocks.21.img_mlp.gate_layer": 82130, + "transformer_blocks.21.img_mlp.proj": 82130, + "transformer_blocks.21.img_mlp.out": 82130, + "transformer_blocks.22.attn.to_q": 82130, + "transformer_blocks.22.attn.to_k": 82130, + "transformer_blocks.22.attn.to_v": 82130, + "transformer_blocks.22.attn.to_out.0": 82130, + "transformer_blocks.22.img_mlp.gate_layer": 82130, + "transformer_blocks.22.img_mlp.proj": 82130, + "transformer_blocks.22.img_mlp.out": 82130, + "transformer_blocks.23.attn.to_q": 82130, + "transformer_blocks.23.attn.to_k": 82130, + "transformer_blocks.23.attn.to_v": 82130, + "transformer_blocks.23.attn.to_out.0": 82130, + "transformer_blocks.23.img_mlp.gate_layer": 82130, + "transformer_blocks.23.img_mlp.proj": 82130, + "transformer_blocks.23.img_mlp.out": 82130, + "transformer_blocks.24.attn.to_q": 82130, + "transformer_blocks.24.attn.to_k": 82130, + "transformer_blocks.24.attn.to_v": 82130, + "transformer_blocks.24.attn.to_out.0": 82130, + "transformer_blocks.24.img_mlp.gate_layer": 82130, + "transformer_blocks.24.img_mlp.proj": 82130, + "transformer_blocks.24.img_mlp.out": 82130, + "transformer_blocks.25.attn.to_q": 82130, + "transformer_blocks.25.attn.to_k": 82130, + "transformer_blocks.25.attn.to_v": 82130, + "transformer_blocks.25.attn.to_out.0": 82130, + "transformer_blocks.25.img_mlp.gate_layer": 82130, + "transformer_blocks.25.img_mlp.proj": 82130, + "transformer_blocks.25.img_mlp.out": 82130, + "transformer_blocks.26.attn.to_q": 82130, + "transformer_blocks.26.attn.to_k": 82130, + "transformer_blocks.26.attn.to_v": 82130, + "transformer_blocks.26.attn.to_out.0": 82130, + "transformer_blocks.26.img_mlp.gate_layer": 82130, + "transformer_blocks.26.img_mlp.proj": 82130, + "transformer_blocks.26.img_mlp.out": 82130, + "transformer_blocks.27.attn.to_q": 82130, + "transformer_blocks.27.attn.to_k": 82130, + "transformer_blocks.27.attn.to_v": 82130, + "transformer_blocks.27.attn.to_out.0": 82130, + "transformer_blocks.27.img_mlp.gate_layer": 82130, + "transformer_blocks.27.img_mlp.proj": 82130, + "transformer_blocks.27.img_mlp.out": 82130, + "transformer_blocks.28.attn.to_q": 82130, + "transformer_blocks.28.attn.to_k": 82130, + "transformer_blocks.28.attn.to_v": 82130, + "transformer_blocks.28.attn.to_out.0": 82130, + "transformer_blocks.28.img_mlp.gate_layer": 82130, + "transformer_blocks.28.img_mlp.proj": 82130, + "transformer_blocks.28.img_mlp.out": 82130, + "transformer_blocks.29.attn.to_q": 82130, + "transformer_blocks.29.attn.to_k": 82130, + "transformer_blocks.29.attn.to_v": 82130, + "transformer_blocks.29.attn.to_out.0": 82130, + "transformer_blocks.29.img_mlp.gate_layer": 82130, + "transformer_blocks.29.img_mlp.proj": 82130, + "transformer_blocks.29.img_mlp.out": 82130, + "transformer_blocks.30.attn.to_q": 82130, + "transformer_blocks.30.attn.to_k": 82130, + "transformer_blocks.30.attn.to_v": 82130, + "transformer_blocks.30.attn.to_out.0": 82130, + "transformer_blocks.30.img_mlp.gate_layer": 82130, + "transformer_blocks.30.img_mlp.proj": 82130, + "transformer_blocks.30.img_mlp.out": 82130, + "transformer_blocks.31.attn.to_q": 82130, + "transformer_blocks.31.attn.to_k": 82130, + "transformer_blocks.31.attn.to_v": 82130, + "transformer_blocks.31.attn.to_out.0": 82130, + "transformer_blocks.31.img_mlp.gate_layer": 82130, + "transformer_blocks.31.img_mlp.proj": 82130, + "transformer_blocks.31.img_mlp.out": 82130 + }, + "metadata": { + "jobs_file": "/poc/nunchaku_backend/calibration_jobs.json", + "jobs": [ + { + "label": "calibration-landscape", + "prompt": "Photorealistic mountain valley at sunrise, a winding river and pine forest, clouds behind snowy peaks, crisp natural colors.", + "width": 1024, + "height": 1024, + "seed": 731 + }, + { + "label": "calibration-chess-portrait", + "prompt": "Editorial photograph of an elderly chess player concentrating over a wooden chessboard in a quiet cafe, soft side lighting, natural facial texture.", + "width": 1024, + "height": 1024, + "seed": 732 + }, + { + "label": "calibration-subway", + "prompt": "A candid documentary photograph of a modern subway platform with commuters, silver train arriving, overhead fluorescent lighting and deep perspective.", + "width": 1024, + "height": 1024, + "seed": 733 + }, + { + "label": "calibration-fruit", + "prompt": "A still life of oranges, green pears and a cut pomegranate in a blue ceramic bowl on a linen cloth, realistic daylight and detailed fruit texture.", + "width": 1024, + "height": 1024, + "seed": 734 + }, + { + "label": "calibration-mug-white", + "prompt": "Change the mug to matte white ceramic. Preserve the composition, spoon, table and lighting.", + "images": [ + "/poc/samples/archive/2026-09-20-initial-poc/reference-clean-rgb.png" + ], + "width": 1024, + "height": 1024, + "seed": 735 + }, + { + "label": "calibration-mug-background", + "prompt": "Add a small green houseplant in a simple pot behind the mug near the wall. Preserve the mug, spoon and table.", + "images": [ + "/poc/samples/archive/2026-09-20-initial-poc/reference-clean-rgb.png" + ], + "width": 1024, + "height": 1024, + "seed": 736 + } + ], + "steps": 8, + "recorded_steps": [ + 0, + 4, + 7 + ], + "teacher": "NF4 approximate trajectories", + "evaluation_prompts_and_references_held_out": true + }, + "instrumented_latency": true, + "layers": 224 +} diff --git a/reproduction/environment.freeze.txt b/reproduction/environment.freeze.txt new file mode 100644 index 0000000000000000000000000000000000000000..fffcb18a03471bc67200b6a61b61cc2edacfa513 --- /dev/null +++ b/reproduction/environment.freeze.txt @@ -0,0 +1,193 @@ +accelerate==1.12.0 +aiodns==4.0.4 +aiohappyeyeballs==2.6.2 +aiohttp==3.14.1 +aiohttp-retry==2.9.1 +aiosignal==1.4.0 +annotated-doc==0.0.4 +annotated-types==0.7.0 +anyio==4.14.0 +archspec @ file:///home/conda/feedstock_root/build_artifacts/archspec_1737352602016/work +asttokens @ file:///home/conda/feedstock_root/build_artifacts/asttokens_1733250440834/work +astunparse==1.6.3 +attrs @ file:///home/conda/feedstock_root/build_artifacts/attrs_1741918516150/work +backoff==2.2.1 +backports.zstd==1.6.0 +bcrypt==5.0.0 +beautifulsoup4 @ file:///home/conda/feedstock_root/build_artifacts/beautifulsoup4_1744783198182/work +bitsandbytes==0.50.2 +boltons @ file:///home/conda/feedstock_root/build_artifacts/boltons_1749686179973/work +boto3==1.43.34 +botocore==1.43.34 +brotli==1.2.0 +certifi @ file:///home/conda/feedstock_root/build_artifacts/certifi_1754231422783/work/certifi +cffi==2.0.0 +chardet @ file:///home/conda/feedstock_root/build_artifacts/chardet_1741797914774/work +charset-normalizer @ file:///home/conda/feedstock_root/build_artifacts/charset-normalizer_1746214863626/work +click==8.5.0 +cmake==4.0.3 +colorama @ file:///home/conda/feedstock_root/build_artifacts/colorama_1733218098505/work +conda @ file:///home/conda/feedstock_root/build_artifacts/conda_1754405241914/work/conda-src +conda-build @ file:///home/conda/feedstock_root/build_artifacts/conda-build_1754316273870/work +conda-libmamba-solver @ file:///home/conda/feedstock_root/build_artifacts/conda-libmamba-solver_1742219570693/work/src +conda-package-handling @ file:///home/conda/feedstock_root/build_artifacts/conda-package-handling_1736345463896/work +conda_index @ file:///home/conda/feedstock_root/build_artifacts/conda-index_1748375757308/work +conda_package_streaming @ file:///home/conda/feedstock_root/build_artifacts/conda-package-streaming_1751548120229/work +cryptography==46.0.7 +decorator @ file:///home/conda/feedstock_root/build_artifacts/decorator_1740384970518/work +detect-installer==0.1.0 +diffusers @ https://github.com/huggingface/diffusers/archive/80c7ed262aeffbeb43ef13ae04baeb9b84515a69.zip#sha256=7b0d00b5b44f5ca1745d477ae4a0c2da125d58b89cff8470c422d3181a797a09 +distro @ file:///home/conda/feedstock_root/build_artifacts/distro_1734729835256/work +dnspython==2.7.0 +einops==0.8.2 +email-validator==2.3.0 +evalidate @ file:///home/conda/feedstock_root/build_artifacts/bld/rattler-build_evalidate_1746793833/work +exceptiongroup @ file:///home/conda/feedstock_root/build_artifacts/exceptiongroup_1746947292760/work +executing @ file:///home/conda/feedstock_root/build_artifacts/executing_1745502089858/work +expecttest==0.3.0 +fastapi==0.116.1 +fastapi-cli==0.0.27 +fastapi-cloud-cli==0.20.0 +fastar==0.11.0 +filelock @ file:///home/conda/feedstock_root/build_artifacts/filelock_1741969488311/work +frozendict @ file:///home/conda/feedstock_root/build_artifacts/frozendict_1728841334936/work +frozenlist==1.8.0 +fsspec==2025.7.0 +h11==0.16.0 +h2 @ file:///home/conda/feedstock_root/build_artifacts/h2_1738578511449/work +hf-xet==1.6.0 +hpack @ file:///home/conda/feedstock_root/build_artifacts/hpack_1737618293087/work +httpcore==1.0.9 +httptools==0.8.0 +httpx==0.28.1 +huggingface_hub==1.32.0 +hyperframe @ file:///home/conda/feedstock_root/build_artifacts/hyperframe_1737618333194/work +hypothesis==6.137.1 +idna @ file:///home/conda/feedstock_root/build_artifacts/idna_1733211830134/work +importlib_metadata==9.0.0 +inquirerpy==0.3.4 +invoke==3.0.3 +ipython @ file:///home/conda/feedstock_root/build_artifacts/bld/rattler-build_ipython_1751465044/work +ipython_pygments_lexers @ file:///home/conda/feedstock_root/build_artifacts/ipython_pygments_lexers_1737123620466/work +itsdangerous==2.2.0 +jedi @ file:///home/conda/feedstock_root/build_artifacts/jedi_1733300866624/work +Jinja2 @ file:///home/conda/feedstock_root/build_artifacts/jinja2_1741263328855/work +jmespath==1.1.0 +jsonpatch @ file:///home/conda/feedstock_root/build_artifacts/jsonpatch_1733814567314/work +jsonpointer @ file:///home/conda/feedstock_root/build_artifacts/jsonpointer_1725302941992/work +jsonschema @ file:///home/conda/feedstock_root/build_artifacts/bld/rattler-build_jsonschema_1752925388/work +jsonschema-specifications @ file:///tmp/tmpuvkyqc9y/src +libarchive-c @ file:///home/conda/feedstock_root/build_artifacts/bld/rattler-build_python-libarchive-c_1747927321/work +libmambapy @ file:///home/conda/feedstock_root/build_artifacts/mamba-split_1746515836725/work/libmambapy +lief @ file:///home/conda/feedstock_root/build_artifacts/lief_1750151383011/work/api/python +lintrunner==0.12.7 +markdown-it-py==4.2.0 +MarkupSafe @ file:///home/conda/feedstock_root/build_artifacts/markupsafe_1733219680183/work +matplotlib-inline @ file:///home/conda/feedstock_root/build_artifacts/matplotlib-inline_1733416936468/work +mdurl==0.1.2 +menuinst @ file:///home/conda/feedstock_root/build_artifacts/menuinst_1753546279984/work +mpmath==1.3.0 +msgpack @ file:///home/conda/feedstock_root/build_artifacts/msgpack-python_1749813202382/work +multidict==6.7.1 +networkx==3.5 +ninja==1.11.1.4 +numpy @ file:///home/conda/feedstock_root/build_artifacts/bld/rattler-build_numpy_1753401560/work/dist/numpy-2.3.2-cp311-cp311-linux_x86_64.whl#sha256=469440884ff65cdd5601d2ff9e1c5b9bcc2d8176c37e37ec78a0666703a54157 +nunchaku @ https://github.com/nunchaku-tech/nunchaku/releases/download/v1.2.1/nunchaku-1.2.1+cu12.8torch2.8-cp311-cp311-linux_x86_64.whl#sha256=77dab1a3abdff16d5cbff70e26a5700e6fccb77b4207c3d53127d5f0224bd82d +nvidia-cublas-cu12==12.8.4.1 +nvidia-cuda-cupti-cu12==12.8.90 +nvidia-cuda-nvrtc-cu12==12.8.93 +nvidia-cuda-runtime-cu12==12.8.90 +nvidia-cudnn-cu12==9.10.2.21 +nvidia-cufft-cu12==11.3.3.83 +nvidia-cufile-cu12==1.13.1.3 +nvidia-curand-cu12==10.3.9.90 +nvidia-cusolver-cu12==11.7.3.90 +nvidia-cusparse-cu12==12.5.8.93 +nvidia-cusparselt-cu12==0.7.1 +nvidia-nccl-cu12==2.27.3 +nvidia-nvjitlink-cu12==12.8.93 +nvidia-nvtx-cu12==12.8.90 +optree==0.17.0 +packaging @ file:///home/conda/feedstock_root/build_artifacts/bld/rattler-build_packaging_1745345660/work +paramiko==5.0.0 +parso @ file:///home/conda/feedstock_root/build_artifacts/parso_1733271261340/work +peft==0.19.1 +pexpect @ file:///home/conda/feedstock_root/build_artifacts/pexpect_1733301927746/work +pfzy==0.3.4 +pickleshare @ file:///home/conda/feedstock_root/build_artifacts/pickleshare_1733327343728/work +pillow==12.2.0 +pkginfo @ file:///home/conda/feedstock_root/build_artifacts/pkginfo_1739984581450/work +platformdirs @ file:///home/conda/feedstock_root/build_artifacts/bld/rattler-build_platformdirs_1746710438/work +pluggy @ file:///home/conda/feedstock_root/build_artifacts/pluggy_1747339660894/work +prettytable==3.17.0 +prompt_toolkit @ file:///home/conda/feedstock_root/build_artifacts/prompt-toolkit_1744724089886/work +propcache==0.5.2 +protobuf==7.35.1 +psutil @ file:///home/conda/feedstock_root/build_artifacts/psutil_1740663149797/work +ptyprocess @ file:///home/conda/feedstock_root/build_artifacts/ptyprocess_1733302279685/work/dist/ptyprocess-0.7.0-py2.py3-none-any.whl#sha256=92c32ff62b5fd8cf325bec5ab90d7be3d2a8ca8c8a3813ff487a8d2002630d1f +pure_eval @ file:///home/conda/feedstock_root/build_artifacts/pure_eval_1733569405015/work +py-cpuinfo==9.0.0 +pycares==5.0.1 +pycosat @ file:///home/conda/feedstock_root/build_artifacts/pycosat_1732588400443/work +pycparser @ file:///home/conda/feedstock_root/build_artifacts/bld/rattler-build_pycparser_1733195786/work +pydantic==2.13.4 +pydantic-extra-types==2.11.1 +pydantic-settings==2.14.2 +pydantic_core==2.46.4 +Pygments @ file:///home/conda/feedstock_root/build_artifacts/pygments_1750615794071/work +PyNaCl==1.6.2 +PySocks @ file:///home/conda/feedstock_root/build_artifacts/pysocks_1733217236728/work +python-dateutil==2.9.0.post0 +python-dotenv==1.2.2 +python-etcd==0.4.5 +python-multipart==0.0.32 +pytz @ file:///home/conda/feedstock_root/build_artifacts/pytz_1742920838005/work +PyYAML @ file:///home/conda/feedstock_root/build_artifacts/pyyaml_1737454647378/work +referencing @ file:///home/conda/feedstock_root/build_artifacts/bld/rattler-build_referencing_1737836872/work +regex==2026.5.9 +requests==2.34.2 +rich==15.0.0 +rich-toolkit==0.20.1 +rignore==0.7.6 +rpds-py @ file:///home/conda/feedstock_root/build_artifacts/bld/rattler-build_rpds-py_1751468291/work +ruamel.yaml @ file:///home/conda/feedstock_root/build_artifacts/ruamel.yaml_1749479918291/work +ruamel.yaml.clib @ file:///home/conda/feedstock_root/build_artifacts/ruamel.yaml.clib_1728724459810/work +runpod==1.9.1 +s3transfer==0.19.0 +safetensors==0.8.0 +sentencepiece==0.2.1 +sentry-sdk==2.63.0 +shellingham==1.5.4 +six==1.17.0 +sortedcontainers==2.4.0 +soupsieve @ file:///home/conda/feedstock_root/build_artifacts/soupsieve_1746563585861/work +stack_data @ file:///home/conda/feedstock_root/build_artifacts/stack_data_1733569443808/work +starlette==0.47.3 +sympy==1.14.0 +tokenizers==0.23.2 +tomli==2.4.1 +tomlkit==0.15.0 +torch==2.8.0+cu128 +torchaudio==2.8.0+cu128 +torchelastic==0.2.2 +torchvision==0.23.0+cu128 +tqdm @ file:///home/conda/feedstock_root/build_artifacts/tqdm_1735661334605/work +tqdm-loggable==0.4.1 +traitlets @ file:///home/conda/feedstock_root/build_artifacts/traitlets_1733367359838/work +transformers==5.17.0 +triton==3.4.0 +truststore @ file:///home/conda/feedstock_root/build_artifacts/bld/rattler-build_truststore_1739009763/work +typer==0.25.1 +types-dataclasses==0.6.6 +typing-inspection==0.4.2 +typing_extensions @ file:///home/conda/feedstock_root/build_artifacts/bld/rattler-build_typing_extensions_1751643513/work +urllib3 @ file:///home/conda/feedstock_root/build_artifacts/urllib3_1750271362675/work +uvicorn==0.35.0 +uvloop==0.22.1 +watchdog==6.0.0 +watchfiles==1.2.0 +wcwidth @ file:///home/conda/feedstock_root/build_artifacts/wcwidth_1733231326287/work +websockets==16.0 +yarl==1.24.2 +zipp==4.1.0 +zstandard==0.23.0 diff --git a/reproduction/experiments/fidelity-v3/README.md b/reproduction/experiments/fidelity-v3/README.md new file mode 100644 index 0000000000000000000000000000000000000000..caba9d3bec05c24a40b67d5ab44ae41593953e18 --- /dev/null +++ b/reproduction/experiments/fidelity-v3/README.md @@ -0,0 +1,82 @@ +# Qwen Image 2.1 fidelity revision + +This experiment improves the custom Nunchaku converter against the original BF16 transformer. The encoder stays NF4 and the VAE stays BF16 in every comparison, so the reference isolates transformer quantization; it is not a fully BF16 pipeline. + +The previous results remain under `samples/varied-v2/`. Initial POC samples remain archived under `samples/archive/2026-09-20-initial-poc/`. New outputs go only under `samples/fidelity-v3/`. + +## Calibration and evaluation separation + +`calibration-jobs.json` contains four training, two validation and two held-out calibration jobs. Each split contains generation and editing. Editing instructions differ, but these calibration splits share one source photograph; this is a limitation. None of the final evaluation prompts or reference photographs is used for calibration. + +The initial three-layer probe used 16-step BF16 trajectories. The final collection uses 40 steps and samples invocations 0, 10, 20, 30 and 39. First-step sampling explicitly covers text, reference-image and target-image tokens. Later sampling covers target tokens. Per-layer files retain 512 training, 256 validation and 256 held-out rows, plus full-token training channel maxima. The manifest records actual jobs, sampling, model revision and timestamps. + +`quality-jobs-all.json` freezes 21 scenarios at both 25 and 40 steps: the original six generation prompts and three edits, plus eight new generation and four new editing scenarios in `expanded-jobs.json`. There are 42 jobs per evaluated backend. The new cases cover role-specific groups, coordinated hand interactions, a wheelchair award presentation, a family reunion, dense English/Chinese type, dialogue continuity, diagram geometry, a product campaign, a newsroom, local preservation and two-reference identity/product composition. `expanded-rubric.md` defines their checks. `quality-jobs.json` remains the original nine-job 40-step subset. + +The expanded edits use fixed BF16 reference images generated by earlier jobs. Their file contents, including any source-model imperfections, must be held identical across candidates. Inspect counting, anatomy, geometry, lettering, materials and edit preservation separately from image-similarity scores. + +## Reference integrity + +`teacher_v3.py` uses streamed block offloading to fit the original BF16 transformer on GPU 0. A resident-versus-offloaded 512-pixel, four-step test produced byte-identical PNGs; the measured result is `results/teacher-v3-parity.json`. That is an implementation parity check, not a broad image-quality result. + +## Reproduction + +Use the pinned Dockerfile and the same GPU-isolated Docker invocation as the main POC. Mount the project at `/poc` and its cache at `/cache`. Only physical GPU UUID `GPU-65bcac75-a854-7804-3e12-aca221530033` is exposed. Klein remains on physical GPU 1. Run GPU experiments serially. + +Inside the container: + +```bash +python -m nunchaku_backend.collect_v3 \ + --jobs /poc/experiments/fidelity-v3/calibration-jobs.json \ + --out /cache/qwen21-activation-v3-40 --steps 40 + +python -m nunchaku_backend.teacher_v3 \ + --jobs /poc/experiments/fidelity-v3/quality-jobs-all.json \ + --sample-dir /poc/samples/fidelity-v3/bf16-dit + +python -m nunchaku_backend.export_v3 \ + --model-path /cache/huggingface/hub/models--Qwen--Qwen-Image-2.1/snapshots/b3179ad355be050328e483a9dfdd9e60cd62adfa \ + --activations /cache/qwen21-activation-v3-40 \ + --baseline-calibration /cache/qwen21-calibration \ + --out /cache/qwen-nunchaku-v3-r128 --device cuda:0 \ + --ranks 128 --alphas .25 .5 .75 \ + --families activation_only smoothquant --weighting none \ + --iterations 16 --final-gptq --factorization balanced + +python /poc/runner.py --backend nunchaku \ + --nunchaku-checkpoint /cache/qwen-nunchaku-v3-r128 \ + --prequant /cache/qwen-nf4 --no-cache \ + --jobs /poc/experiments/fidelity-v3/quality-jobs-all.json \ + --sample-dir /poc/samples/fidelity-v3/nunchaku-v3-r128 +``` + +The subsequent selective upgrade keeps the other 192 quantized projections and BF16 boundary tensors byte-identical, while choosing rank128, 256 or 512 for each of the 32 MLP projection layers using validation kernel output error. All 32 selected rank512 in this run. This adds 384 MiB of model tensors; it is not a peak-VRAM estimate. + +```bash +python -m nunchaku_backend.upgrade_mlp_rank_v3 \ + --source-checkpoint /cache/qwen-nunchaku-v3-r128 \ + --model-path /cache/huggingface/hub/models--Qwen--Qwen-Image-2.1/snapshots/b3179ad355be050328e483a9dfdd9e60cd62adfa \ + --activations /cache/qwen21-activation-v3-40 \ + --baseline-calibration /cache/qwen21-calibration \ + --out /cache/qwen-nunchaku-v3-mlpproj-upgrade \ + --device cuda:0 --ranks 256 512 --iterations 16 + +python /poc/runner.py --backend nunchaku \ + --nunchaku-checkpoint /cache/qwen-nunchaku-v3-mlpproj-upgrade \ + --prequant /cache/qwen-nf4 --no-cache \ + --jobs /poc/experiments/fidelity-v3/quality-jobs-all.json \ + --sample-dir /poc/samples/fidelity-v3/nunchaku-v3-mlpproj512 +``` + +The checkpoint exporter records its exact search settings, activation-file hashes, per-layer candidate errors and selection decisions in the checkpoint manifest. Selection uses validation outputs from the actual Nunchaku kernel; held-out rows are measured only after selection. GPTQ is retained only when validation improves. The original one-pass recipe remains a candidate, although rerunning randomized SVD is not necessarily a byte-for-byte reproduction of the archived v2 checkpoint. + +After generation, run `scripts/build_fidelity_gallery.py` locally to rebuild the offline comparison. Results must distinguish measured numerical approximation, visible image quality, latency and VRAM. No near-lossless claim follows from per-layer error alone. + +## Sources + +See `research/svdquant-paper-v3.md` for the paper and official implementation audit, and `research/nunchaku-v3-audit.md` for the shipped Klein checkpoint inspection. The new converter uses a bounded search and randomized SVD; RMS weighting and ridge correction are explicitly experimental extensions, and its GPTQ implementation retains original channel order. It does not reproduce every official calibration detail, especially joint attention/block-level selection. + +## Selected serving configuration + +Use `docker compose -f compose.yaml -f compose.nunchaku-v3.yaml up -d` for calibrated rank128. The selective512 override remains experimental and is not selected. Both candidates complete42/42 jobs; the comparison gallery contains126 new quality images plus the archived comparison controls. + +FinalAPI functional checks are complete, while strict same-seed repeatability failed. Reproduce the full diagnostic recording with `python /poc/scripts/api_fidelity_smoke.py --record-repeat-differences --sample-dir samples/fidelity-v3/api-new-run --report results/api-new-run.json`; use a fresh output directory. Omitting the flag preserves the strict assertion after all images/differences are saved. This is not a tolerance-based repeatability pass. See REPORT.md for measured limitations. diff --git a/reproduction/experiments/fidelity-v3/REPORT.md b/reproduction/experiments/fidelity-v3/REPORT.md new file mode 100644 index 0000000000000000000000000000000000000000..7d25dc36bfd83be51b87deefedf21b67928279b7 --- /dev/null +++ b/reproduction/experiments/fidelity-v3/REPORT.md @@ -0,0 +1,91 @@ +# Qwen Image 2.1 fidelity revision — September 21, 2026 + +The new evaluation contains 21 scenarios at both 25 and 40 steps. It includes interacting groups with specific roles, a wheelchair award ceremony, a family reunion, a radio interview, dense bilingual text, a three-panel comic, museum geometry, product layouts, local preservation edits and two-reference composition. Initial samples remain archived; all new comparisons are in `samples/fidelity-v3/`. + +The serving choice is calibrated rank 128. The larger selective rank 512 candidate is retained as an experiment: despite better numerical error, direct image inspection found extra objects and geometry regressions. All 126 comparison images (42 each for the BF16-transformer teacher, calibrated rank 128 and selective rank 512) are complete, locally verified and covered by direct visual reviews. These are 42 matched jobs, not 126 independent prompts. The selected API is running and its functional generation/edit/cache checks are complete; strict same-seed repeatability failed, as documented below. + +## What changed in the converter + +The previous one-pass rank 32 approximation was replaced by calibration against actual Nunchaku CUDA output. The new search selects smoothing and low-rank residual fits using validation output error, checks held-out rows after selection, and accepts optional GPTQ correction only when validation improves. Rank 128 is used across 224 quantized linear layers. Each linear remains W4A4 plus a BF16 low-rank branch; original normalization, positional encoding, attention and KV-cache behavior remain in the Diffusers implementation. + +The SVDQuant paper and official Nunchaku/DeepCompressor implementation motivated the calibration and decomposition work. This is a custom Qwen Image 2.1 adapter and bounded calibration implementation, not a reproduction of every official calibration setting. The exact shipped Klein calibration recipe is not public. See the [SVDQuant source review](../../research/svdquant-paper-v3.md) and [converter audit](../../research/nunchaku-v3-audit.md). + +A follow-up sensitivity sweep found that the MLP projection and output layers contribute substantially to denoiser error. Temporarily restoring both roles to BF16 reduced the three diagnostic relative-L2 errors to 5.61%, 2.74% and 9.28%, but adds roughly 4.15 GiB of model state. This is a diagnostic, not a measured image-quality or safe two-reference serving configuration. + +The practical candidate instead raises only the 32 MLP projection low-rank branches from 128 to 512. All 32 selected 512 by validation. Other 192 quantized linears and 193 untouched shards remain byte-identical. It retains all 224 INT4 kernels and adds 384 MiB of model tensors. The complete checkpoint has 5,057,830,912 bytes (4.710 GiB) of model tensors. Independent CPU loading verified 1,417 finite state tensors, every output/source shard hash and all 193 untouched shard identities; load time was 3.48 seconds. Across the 32 upgraded layers, held-out MSE ratios to actual rank 128 range from 0.531 to 0.656, with median 0.642. Model state is not peak VRAM. + +## Numerical evidence and rejected changes + +The first rank 128 converter improved held-out linear MSE against a regenerated one-pass rank 32 recipe in all 224 linears; median ratio was 0.6064. That numerical baseline is not the exact archived v2 checkpoint. Actual v2 images and whole-denoiser checks use its real saved checkpoint. + +Three fixed-recipe runs allowing 100 fitting iterations did not justify a full conversion: one control stayed unchanged, one slightly better validation result worsened held-out error, and another slightly regressed. More iterations were rejected instead of presumed beneficial. + +The rank 512 probe improved held-out projection MSE by 35.5–38.6% in three representative layers. Through the complete MLP, improvement was smaller: 7.9–12.2%. Block 11 rank 256 improved raw projection error while slightly worsening the parent MLP's validation error. These checks demonstrate why linear error alone cannot establish image fidelity. + +| Actual checkpoint | First-step denoiser relative L2 | Middle step | Final step | +|---|---:|---:|---:| +| Archived v2 |12.48%|6.54%|21.87%| +| Calibrated rank 128 |8.85%|5.08%|19.43%| +| Selective MLP rank 512 |8.46%|4.72%|17.93%| + +This is one disjoint development editing prompt at steps 0/20/39 of 40. Every checkpoint builds its own prefix KV cache; no teacher cache is reused. These are conditional predictions on the teacher trajectory, not free-running image scores. + +## Image review and step count + +Forty steps is the [official starting recommendation](https://huggingface.co/docs/diffusers/main/api/pipelines/qwenimage21). Both 25 and 40 are evaluated here at 1024×1024, with fixed seeds, CFG 1 and prefix KV caching. More steps do not reliably fix incorrect counts or relationships. + +The rank 128 review covers all 42 teacher/candidate pairs. It restores the fantasy telescope and improves the old pancake-count error, retains all 14 bilingual poster strings and six comic dialogue lines, and performs the local coat/poster edits with strong preservation. It still changes some faces, hand positions, poses, object scale and design details. Its 25-step chess banner loses a digit in 2026; 40 steps restores it. Its museum drawings add tables and confuse routes at both step counts. The product campaign improves requested cup placement relative to the teacher, demonstrating that visual similarity and prompt compliance are different measurements. + +The BF16 teacher also has real limitations: its 25-step chess title reads CIT CHES FINAL, the comic does not stage keys under the chair, the product relocation edit leaves its cup on the pedestal, and the two-reference portrait looks off-camera. These are shared model/instruction failures, not evidence that every candidate difference is caused by quantization. The teacher uses the same NF4 encoder, so it isolates transformer approximation rather than representing a fully BF16 pipeline. + +Direct observations and image hashes are recorded in the [rank 128 scorecard](../../research/quality-v3-scorecard.json), with root versus independent reviewer attribution, and the [complete selective rank 512 review index](../../research/quality-v3-rank512-review-index.md). Similarity metrics compare the same labels/seeds and unaligned 1024×1024 images. SSIM and PSNR measure resemblance, not a percentage of semantic quality. + +The **matched original 18-job cohort** is available for all four candidates. Each is compared with the corresponding BF16-transformer image; these medians pool nine scenarios at both step counts. + +| Candidate | Images | Median SSIM | Median PSNR (dB) | +|---|---:|---:|---:| +| Archived Nunchaku v2 | 18 | 0.7608 | 18.52 | +| Calibrated rank 128 | 18 | 0.8508 | 20.57 | +| Selective rank 512 | 18 | 0.8287 | 19.96 | +| NF4 transformer | 18 | 0.8237 | 20.84 | + +The **full matched 42-job cohort** is available for the two new Nunchaku candidates. It includes the original 18 jobs plus 24 expanded jobs. + +| Candidate | Images | Median SSIM | Median PSNR (dB) | +|---|---:|---:|---:| +| Calibrated rank 128 | 42 | 0.8269 | 19.99 | +| Selective rank 512 | 42 | 0.8207 | 19.62 | + +Rank 128 has higher median SSIM than NF4 on the original 18 jobs, while NF4 has higher median PSNR. Selective rank 512 trails rank 128 on both medians in both matched cohorts despite its better denoiser probe. NF4 was not measured on the full 42-job cohort. Comparisons across the two tables would mix different prompt sets. Exact labels and per-image metrics are in the [original 18-job results](../../results/compare-fidelity-v3-original18.json) and [full 42-job results](../../results/compare-fidelity-v3-full42.json). + +## Selective rank 512 decision + +The larger branch is not promoted. All 42 selective outputs now have direct review coverage, including fixed-input checks for all 14 edits. In the 12 original generation jobs it improves the thin perfume peel and restores a more teacher-like towel grip at 40 steps, but produces seven blueberries and two knives in the 40-step breakfast, two cats in the 25-step fantasy scene, and a distorted horn-like telescope at 40. Botanical wording and lemon counts remain correct. + +The 16 expanded generation jobs show further mixed effects. Selective rank 512 restores the chess year and camera-at-eye action, but adds six kitchen rolls at 25 steps and six/five cafe tables at 25/40. The 40-step museum also gains at least three shelf groups instead of two. Bilingual text and comic dialogue remain readable, while the comic still fails to stage the keys under the chair. Product cups sit correctly on the tabletop at 25 but return to the pedestal at 40, matching a teacher error and losing rank 128's adherence improvement. Forty steps trades failures rather than consistently repairing them. + +Localized towel, coat and poster edits retain strong visual preservation, with small texture and typography changes and no decisive broad improvement over rank 128. The 40-step two-reference portrait restores the unobscured bottle emblem and supporting grip more closely to the teacher; off-camera gaze and rendering drift remain. Product-relocation edits still leave the cup on the pedestal in all compared backends. These observations support retaining the smaller checkpoint for this POC, not a claim that rank 128 is universally better or near-lossless. + +## Timing, memory and reproducibility + +For the selected rank 128 checkpoint, median generation takes **15.50 s at 25 steps / 22.82 s at 40**, versus **27.51 s / 42.76 s** for the streamed BF16 transformer. These medians use the same 14 generation scenarios per step count. Median one-reference edit inference takes **19.48 s / 28.43 s** across six matched scenarios per step count. The single two-reference portrait takes **23.76 s / 34.35 s**; these are individual runs, not repeated-run medians. Sampled board occupancy reaches **11,994 MiB (11.71 GiB)** in that two-reference case. Selective rank 512 is slightly slower (**16.06 s / 23.56 s** median generation) and reaches **12,374 MiB** with two references. + +See [performance.md](performance.md) and the [per-job measurement summary](../../results/fidelity-v3-summary.json) for verified per-job measurements. CLI comparisons disable prompt/reference LRU reuse, retain per-request prefix KV, use an untiled BF16 VAE and release dead KV before decoding. Timings cover engine inference, excluding process startup and image file loading/saving. Board occupancy is sampled separately from allocator peaks; initial teacher jobs without board telemetry are explicitly missing those values. + +GPU 0 is an RTX 4070 Ti SUPER with 16,376 MiB. Klein remains on its existing GPU 1. GPU jobs run serially. The implementation stages the large NF4 encoder on CPU between requests; it retains all 36 decoder layers and removes only the unused LM-head/final-normalization path after exact feature-parity verification. Persistent compilation caches and optional compilation remain available; this Nunchaku quality revision uses eager inference. + +Reproduction commands are in the experiment README, including calibration, checkpoint conversion, selective rank upgrade and the frozen 42-job image suite. Docker/model/library revisions are pinned. Saved source state, service rollback information and historical failures remain preserved. + +## Service verification + +The calibrated rank128 API is healthy on PC 2 loopback port 8091. Four real 25-step requests completed: bilingual generation at 18.22 s, cached repeat 14.81 s, two-reference portrait 26.42 s, and cached repeat 22.93 s. These are client wall times from one run each. The prompt cache hit on both repeats and the reference-latent cache hit twice on the two-reference repeat. All four responses were COMPLETED with valid 1024×1024 PNGs. Engine startup measured 8.78 s from the saved checkpoint, including ML imports; this is a cached-disk process start, not machine boot. Maximum sampled API board occupancy was 11,986 MiB. See [final API measurements](../../results/api-fidelity-v3-final.json). + +**Strict byte-identical repeatability failed.** The failure is preserved, not converted into a tolerance pass. A controlled four-run diagnostic compared two uncached runs and a cache miss/hit: all prompt tensors and first-transformer inputs were exactly equal in values, shapes, strides and dtypes, and stored cache tensors remained unchanged. Nevertheless the first transformer outputs differed even without caching. Uncached repeats had RGB MAE 5.06/255; cache miss/hit had 6.48/255. This isolates the observed divergence to the transformer forward, but does not identify the exact kernel or operation. Official low-rank atomic reductions are a hypothesis, not a confirmed sole cause. See [diagnosis](../../research/api-cache-v3-diagnosis.md) and [input fingerprints](../../results/repeatability-diagnostic/diagnostic.json). + +The final API verification used an explicit record-differences mode so it could preserve and inspect every returned image while retaining `strict_repeatability_passed=false`. Both posters retained the requested wording; the two-reference repeats changed hand placement, bench and suitcase position while retaining the person/product attributes. Same-seed variation can affect composition, not merely file bytes. The image-suite comparisons therefore describe individual recorded realizations; they do not establish repeatable equivalence or isolate every visual difference solely to quantization. + +Klein remains healthy with the same container and original start time on GPU1; Qwen alone occupies GPU0 for this POC. Higgs and GPU0 Comfy remain stopped under the user's authorization, GPU1 Comfy and connectivity remain intact, and experiment telemetry is stopped. [Final service state](../../results/final-state-v3.json) and `scripts/rollback.sh` preserve the recovery path. + +## Limits + +This is a 1024×1024 POC below the release's recommended native 2K resolution. It does not establish near-lossless quality, superiority to Klein, full-BF16 encoder fidelity, ten-reference operation, native 2K speed or a production-router rollout. Calibration is small and edit splits share one source photo; prefix/reference-token coverage is limited. The visual suite was reused after earlier candidate results informed further work, so these are diagnostic comparisons rather than a fresh unseen acceptance set. Joint attention and full-block optimization are not implemented by this converter. The model uses the Qwen Research License; the downloaded license is retained with the release research. diff --git a/reproduction/experiments/fidelity-v3/calibration-jobs.json b/reproduction/experiments/fidelity-v3/calibration-jobs.json new file mode 100644 index 0000000000000000000000000000000000000000..10edd71aff8ab9b41b38d57d5ff4bcca4e240e4c --- /dev/null +++ b/reproduction/experiments/fidelity-v3/calibration-jobs.json @@ -0,0 +1,75 @@ +[ + { + "label": "calibration-landscape", + "prompt": "Photorealistic mountain valley at sunrise, a winding river and pine forest, clouds behind snowy peaks, crisp natural colors.", + "width": 1024, + "height": 1024, + "seed": 731, + "split": "train" + }, + { + "label": "calibration-chess-portrait", + "prompt": "Editorial photograph of an elderly chess player concentrating over a wooden chessboard in a quiet cafe, soft side lighting, natural facial texture.", + "width": 1024, + "height": 1024, + "seed": 732, + "split": "train" + }, + { + "label": "calibration-subway", + "prompt": "A candid documentary photograph of a modern subway platform with commuters, silver train arriving, overhead fluorescent lighting and deep perspective.", + "width": 1024, + "height": 1024, + "seed": 733, + "split": "train" + }, + { + "label": "calibration-mug-white", + "prompt": "Change the mug to matte white ceramic. Preserve the composition, spoon, table and lighting.", + "images": [ + "/poc/samples/archive/2026-09-20-initial-poc/reference-clean-rgb.png" + ], + "width": 1024, + "height": 1024, + "seed": 735, + "split": "train" + }, + { + "label": "validation-market", + "prompt": "Documentary photograph of an outdoor farmers market with vegetables, baskets and shoppers beneath striped awnings in soft morning light.", + "seed": 2401, + "split": "validation", + "width": 1024, + "height": 1024 + }, + { + "label": "validation-mug-edit", + "prompt": "Place a folded linen cloth under the mug. Preserve the mug, table, spoon and background.", + "seed": 2402, + "images": [ + "/poc/samples/archive/2026-09-20-initial-poc/reference-clean-rgb.png" + ], + "split": "validation", + "width": 1024, + "height": 1024 + }, + { + "label": "heldout-sailboat", + "prompt": "A watercolor illustration of two small sailboats moored beside a wooden jetty in a quiet lake, distant mountains and pine trees, overcast sky.", + "seed": 3401, + "split": "heldout", + "width": 1024, + "height": 1024 + }, + { + "label": "heldout-mug-edit", + "prompt": "Change the tabletop to polished dark walnut wood while retaining the mug and spoon and wall with the same framing and lighting.", + "seed": 3402, + "images": [ + "/poc/samples/archive/2026-09-20-initial-poc/reference-clean-rgb.png" + ], + "split": "heldout", + "width": 1024, + "height": 1024 + } +] diff --git a/reproduction/experiments/fidelity-v3/candidate-remaining.json b/reproduction/experiments/fidelity-v3/candidate-remaining.json new file mode 100644 index 0000000000000000000000000000000000000000..6488390a2a376af5b19810caa319a17762236029 --- /dev/null +++ b/reproduction/experiments/fidelity-v3/candidate-remaining.json @@ -0,0 +1,326 @@ +[ + { + "label": "editorial-25", + "prompt": "Candid editorial photograph in a sunlit pottery studio. Exactly two adults: an older woman with short silver hair on the left demonstrates shaping a clay bowl on a pottery wheel; a younger man with curly dark hair on the right watches and holds a folded blue towel with both hands. Their hands are anatomically natural and clearly visible, with clay on the woman's fingertips. Wooden shelves hold irregular ceramic vases, one trailing green plant hangs high in the back right, and soft morning light enters from a large window on the left. Waist-up environmental composition, documentary photography, natural skin texture, restrained warm colors, realistic depth of field. No lettering.", + "steps": 25, + "seed": 118, + "width": 1024, + "height": 1024 + }, + { + "label": "rainy-city-40", + "prompt": "Cinematic street-level architectural photograph of a narrow Tokyo side street at blue hour after rain. A tiny warmly lit ramen restaurant occupies the left foreground, with a striped navy awning and a vertical sign that reads \"RAMEN\". Exactly three bicycles are parked along the right wall. A person holding a transparent umbrella walks away at center, wearing a mustard yellow raincoat. Overhead wires cross between weathered three-story buildings, red lanterns glow farther down the street, and puddles reflect the signs. Strong one-point perspective, realistic glass and wet asphalt, fine distant details, balanced blue and amber lighting.", + "steps": 40, + "seed": 227, + "width": 1024, + "height": 1024 + }, + { + "label": "rainy-city-25", + "prompt": "Cinematic street-level architectural photograph of a narrow Tokyo side street at blue hour after rain. A tiny warmly lit ramen restaurant occupies the left foreground, with a striped navy awning and a vertical sign that reads \"RAMEN\". Exactly three bicycles are parked along the right wall. A person holding a transparent umbrella walks away at center, wearing a mustard yellow raincoat. Overhead wires cross between weathered three-story buildings, red lanterns glow farther down the street, and puddles reflect the signs. Strong one-point perspective, realistic glass and wet asphalt, fine distant details, balanced blue and amber lighting.", + "steps": 25, + "seed": 227, + "width": 1024, + "height": 1024 + }, + { + "label": "botanical-poster-25", + "prompt": "Sophisticated botanical exhibition poster on warm ivory paper, flat front-on view, crisp print design. At the top the large dark green serif headline reads exactly \"THE SECRET GARDEN\" on two centered lines. Under it a smaller line reads exactly \"BOTANICAL STUDIES / 2026\". The center contains a delicate detailed watercolor illustration of a lemon branch with exactly two yellow lemons, white blossoms and dark green leaves, surrounded by generous negative space. Along the bottom three evenly spaced text blocks read \"12 OCTOBER\", \"10 AM - 6 PM\", and \"GLASSHOUSE No. 4\". A thin rectangular green border surrounds the entire design. Elegant typography, no other text, no mockup, no shadows outside the paper.", + "steps": 25, + "seed": 336, + "width": 1024, + "height": 1024 + }, + { + "label": "botanical-poster-40", + "prompt": "Sophisticated botanical exhibition poster on warm ivory paper, flat front-on view, crisp print design. At the top the large dark green serif headline reads exactly \"THE SECRET GARDEN\" on two centered lines. Under it a smaller line reads exactly \"BOTANICAL STUDIES / 2026\". The center contains a delicate detailed watercolor illustration of a lemon branch with exactly two yellow lemons, white blossoms and dark green leaves, surrounded by generous negative space. Along the bottom three evenly spaced text blocks read \"12 OCTOBER\", \"10 AM - 6 PM\", and \"GLASSHOUSE No. 4\". A thin rectangular green border surrounds the entire design. Elegant typography, no other text, no mockup, no shadows outside the paper.", + "steps": 40, + "seed": 336, + "width": 1024, + "height": 1024 + }, + { + "label": "food-spatial-25", + "prompt": "Overhead high-end food photograph of a neatly arranged breakfast on a pale stone table. A large round white plate is centered. On the plate, exactly three pancakes form a vertical stack, topped by exactly five blueberries and a small square of butter. To the left of the plate is a silver fork, to the right is a silver knife with its blade facing the plate. A clear glass of orange juice sits at the upper right, a small white bowl containing sliced strawberries sits at the upper left, and a folded rust-colored linen napkin lies beneath the fork. A little maple syrup runs down only the right edge of the pancake stack. Soft natural light from the upper left, realistic food texture, all objects fully inside frame.", + "steps": 25, + "seed": 445, + "width": 1024, + "height": 1024 + }, + { + "label": "fantasy-cutaway-25", + "prompt": "Intricate isometric cutaway illustration of a cozy three-level library built inside an enormous ancient tree, in a hand-painted storybook style. Ground floor: a round green entrance door, a circular rug, and a sleeping orange cat beside a small wood stove. Middle floor: curved bookshelves filled with colorful books, a writing desk by an oval window, and a brass telescope pointing outside. Top floor: a glass-domed reading nook with two red armchairs and a hanging lantern. A single wooden spiral staircase visibly connects all three floors. Exposed roots wrap around mossy rocks, tiny mushrooms cluster near the door, and the leafy crown surrounds the dome without hiding it. Consistent isometric perspective, warm interior lighting, cool twilight forest background, detailed but readable, no text.", + "steps": 25, + "seed": 554, + "width": 1024, + "height": 1024 + }, + { + "label": "glass-product-40", + "prompt": "Luxury studio still life photograph on a polished dark green marble plinth. A rectangular clear glass perfume bottle half filled with pale amber liquid stands in the center, with a brushed gold cylindrical cap and a small cream label reading exactly \"LUMEN\". A thin curved strip of orange peel lies in front of the bottle. Behind and to the left is a translucent ribbed glass sphere; behind and to the right are two glossy dark green leaves. Hard afternoon sunlight from the upper left creates realistic caustics, transparent overlapping shadows, precise refraction through the bottle and a soft reflection on the marble. Deep emerald background, controlled commercial composition, realistic materials, no additional text.", + "steps": 40, + "seed": 663, + "width": 1024, + "height": 1024 + }, + { + "label": "glass-product-25", + "prompt": "Luxury studio still life photograph on a polished dark green marble plinth. A rectangular clear glass perfume bottle half filled with pale amber liquid stands in the center, with a brushed gold cylindrical cap and a small cream label reading exactly \"LUMEN\". A thin curved strip of orange peel lies in front of the bottle. Behind and to the left is a translucent ribbed glass sphere; behind and to the right are two glossy dark green leaves. Hard afternoon sunlight from the upper left creates realistic caustics, transparent overlapping shadows, precise refraction through the bottle and a soft reflection on the marble. Deep emerald background, controlled commercial composition, realistic materials, no additional text.", + "steps": 25, + "seed": 663, + "width": 1024, + "height": 1024 + }, + { + "label": "editorial-edit-25", + "prompt": "Replace only the blue towel held by the younger man with a bright red towel of the same size and folded shape. Preserve both people's identities, facial expressions, hands, clothing, clay bowl, studio shelves, window light, framing and photographic style.", + "steps": 25, + "seed": 774, + "width": 1024, + "height": 1024, + "images": [ + "/poc/samples/varied-v2/nf4/editorial-40.png" + ] + }, + { + "label": "editorial-edit-40", + "prompt": "Replace only the blue towel held by the younger man with a bright red towel of the same size and folded shape. Preserve both people's identities, facial expressions, hands, clothing, clay bowl, studio shelves, window light, framing and photographic style.", + "steps": 40, + "seed": 774, + "width": 1024, + "height": 1024, + "images": [ + "/poc/samples/varied-v2/nf4/editorial-40.png" + ] + }, + { + "label": "poster-edit-25", + "prompt": "Edit only the text at the top: replace \"THE SECRET GARDEN\" with \"THE LEMON HOUSE\". Keep the same dark green serif type style and centered placement. Preserve all other text exactly, the lemon branch illustration, paper background, border, spacing and overall design.", + "steps": 25, + "seed": 885, + "width": 1024, + "height": 1024, + "images": [ + "/poc/samples/varied-v2/nf4/botanical-poster-40.png" + ] + }, + { + "label": "poster-edit-40", + "prompt": "Edit only the text at the top: replace \"THE SECRET GARDEN\" with \"THE LEMON HOUSE\". Keep the same dark green serif type style and centered placement. Preserve all other text exactly, the lemon branch illustration, paper background, border, spacing and overall design.", + "steps": 40, + "seed": 885, + "width": 1024, + "height": 1024, + "images": [ + "/poc/samples/varied-v2/nf4/botanical-poster-40.png" + ] + }, + { + "label": "city-edit-25", + "prompt": "Transform this rainy Tokyo street photograph into a richly textured hand-painted watercolor illustration. Preserve the layout of the buildings, the ramen restaurant on the left, all three bicycles on the right, the central person with the transparent umbrella and yellow raincoat, the overhead wires, blue-hour lighting, and puddle reflections. Keep the RAMEN sign legible.", + "steps": 25, + "seed": 996, + "width": 1024, + "height": 1024, + "images": [ + "/poc/samples/varied-v2/nf4/rainy-city-40.png" + ] + }, + { + "label": "city-edit-40", + "prompt": "Transform this rainy Tokyo street photograph into a richly textured hand-painted watercolor illustration. Preserve the layout of the buildings, the ramen restaurant on the left, all three bicycles on the right, the central person with the transparent umbrella and yellow raincoat, the overhead wires, blue-hour lighting, and puddle reflections. Keep the RAMEN sign legible.", + "steps": 40, + "seed": 996, + "width": 1024, + "height": 1024, + "images": [ + "/poc/samples/varied-v2/nf4/rainy-city-40.png" + ] + }, + { + "label": "expanded-community-kitchen-25", + "prompt": "Documentary photograph for a community cooking program, wide eye-level composition in a bright, practical teaching kitchen. Exactly six adults, all visible from at least the waist up. At the left counter, a woman with short black hair in a green apron chops carrots on a wooden board, with one hand holding the knife and the other safely curled on the carrot. At center, a silver-haired man in a white apron holds a ladle over a large soup pot. Immediately to his right, a woman in a mustard cardigan passes a rectangular tray of four bread rolls to a younger man in a navy shirt, who receives it with both hands; their hands meet opposite edges of the same tray. At the far right, a man in a red apron rinses a colander of leafy greens at the sink. Behind the central counter, a woman wearing glasses and a purple sweater reads a paper recipe card. Natural varied expressions, clear individual faces, plausible arms and hands attached to their owners, no extra people, realistic morning window light and stainless-steel reflections. No readable signage. The image should feel like a real coordinated cooking session, not a posed group portrait.", + "steps": 25, + "seed": 21011, + "width": 1024, + "height": 1024 + }, + { + "label": "expanded-chess-awards-25", + "prompt": "Editorial event photograph of a small chess championship award presentation on a low indoor stage. Exactly five adults. At center, a smiling woman with a short dark bob sits in a clearly recognizable manual wheelchair; she and the presenter standing immediately to her left jointly hold one gold trophy between them. The presenter is an older man wearing a charcoal suit. Immediately to the winner's right, the runner-up, a young man in a cream sweater, wears one silver medal and applauds. At the far left, a female photographer kneels with a camera held to her eye and aims toward the winner. At the far right, a host in a burgundy suit holds a microphone and a small cue card. All five people fit fully in frame, with believable hands, wheelchair wheels and footrests. Behind them, a clean white stage banner reads exactly \"CITY CHESS FINAL\" with \"2026\" centered on the line below. A small chess table with a board stands at the rear left. Warm indoor event lighting, realistic fabric and skin, no audience or additional figures.", + "steps": 25, + "seed": 21022, + "width": 1024, + "height": 1024 + }, + { + "label": "expanded-station-reunion-25", + "prompt": "Natural editorial photograph of a family reunion beside a stationary passenger train on a quiet outdoor railway platform. Exactly four people, full bodies in frame, no other passengers. On the left is a woman in her thirties with an oval face, dark curly hair tied in a low bun, small round gold earrings, a navy knee-length coat and a mustard yellow scarf. She holds the left hand of a young girl standing at center; the girl wears a red raincoat and carries a small stuffed rabbit in her free hand. On the right, a man with close-cropped brown hair, a light gray jacket and a red backpack bends slightly to embrace an older gray-haired woman wearing a teal coat; the older woman faces him with one arm around his shoulder. The mother and girl watch them with happy, understated expressions. One upright green suitcase stands beside the mother's left leg. Anatomically plausible hands and embraces with clear ownership of each arm, natural eye contact, soft overcast daylight, authentic train windows and platform paving. Do not crop feet or add duplicate people. No readable logos or signs.", + "steps": 25, + "seed": 21033, + "width": 1024, + "height": 1024 + }, + { + "label": "expanded-station-reunion-40", + "prompt": "Natural editorial photograph of a family reunion beside a stationary passenger train on a quiet outdoor railway platform. Exactly four people, full bodies in frame, no other passengers. On the left is a woman in her thirties with an oval face, dark curly hair tied in a low bun, small round gold earrings, a navy knee-length coat and a mustard yellow scarf. She holds the left hand of a young girl standing at center; the girl wears a red raincoat and carries a small stuffed rabbit in her free hand. On the right, a man with close-cropped brown hair, a light gray jacket and a red backpack bends slightly to embrace an older gray-haired woman wearing a teal coat; the older woman faces him with one arm around his shoulder. The mother and girl watch them with happy, understated expressions. One upright green suitcase stands beside the mother's left leg. Anatomically plausible hands and embraces with clear ownership of each arm, natural eye contact, soft overcast daylight, authentic train windows and platform paving. Do not crop feet or add duplicate people. No readable logos or signs.", + "steps": 40, + "seed": 21033, + "width": 1024, + "height": 1024 + }, + { + "label": "expanded-bilingual-festival-25", + "prompt": "A polished bilingual neighborhood film festival poster, square format, flat front-on print artwork with generous margins, warm cream paper and deep navy typography accented by vermilion. The large centered English title at the top reads exactly \"RIVERLIGHT FILM WEEK\". Directly beneath it, an equally clear Chinese title reads exactly \"河畔电影周\". The next centered line reads exactly \"18–20 SEPTEMBER 2026\". Below a thin vermilion rule, a tidy two-column program occupies the middle: the left column contains times, and the right column contains one English film title with its Chinese title directly underneath. Row one: \"18:00\", \"A Quiet River\", \"静静的河\". Row two: \"19:30\", \"Night Market\", \"夜市\". Row three: \"21:00\", \"Home Again\", \"再次回家\". Maintain aligned columns and clearly separated rows. At the bottom, centered on separate lines, print exactly \"RIVERSIDE CINEMA / 河畔影院\" and \"FREE ENTRY / 免费入场\". A small restrained illustration of a red bridge over two blue river lines sits between the program and footer without touching any letters. Sharp readable Latin and Chinese lettering, consistent type hierarchy, no additional text, no mockup perspective.", + "steps": 25, + "seed": 21044, + "width": 1024, + "height": 1024 + }, + { + "label": "expanded-library-comic-25", + "prompt": "A polished, friendly three-panel comic strip in a square canvas, with exactly three equal vertical panels arranged left to right, clear white gutters and a thin dark border around each panel. Keep the same two adult characters and library reading-room setting in all panels: Maya has a chin-length black bob, round glasses and a red sweater; Leo has short curly brown hair and a blue button-up shirt. Each panel contains exactly two white speech bubbles, with clear tails pointing to the correct speaker. Panel 1: Maya stands beside a wooden reading chair, looks worried and searches her bag. Maya says exactly \"Where are my keys?\" Leo, standing to her right, says exactly \"Check your pocket.\" Panel 2: Maya turns out an empty coat pocket and says exactly \"Not here!\" Leo points down at a small key ring visible under that same chair and says exactly \"Under the chair.\" Panel 3: Maya stands upright holding that key ring up, smiles and says exactly \"Found them. Thanks!\" Leo smiles and says exactly \"You're welcome.\" Crisp ink lines, restrained warm colors, expressive natural poses, consistent faces and clothing, believable hands, a coherent sequence with the chair in the same room. Readable dialogue with no narration captions and no additional panels or characters.", + "steps": 25, + "seed": 21055, + "width": 1024, + "height": 1024 + }, + { + "label": "expanded-museum-plan-25", + "prompt": "A clear visitor floor-plan diagram for a small museum, square format, flat orthographic top-down vector design on white, with thick charcoal exterior walls, thinner interior walls and clearly visible doorway gaps. Title centered above the plan: \"MUSEUM VISITOR MAP\". Inside one rectangular building, arrange five labeled spaces: \"GALLERY A\" in a large upper-left room, \"GALLERY B\" in a large upper-right room, \"MAIN HALL\" in a wide horizontal central space spanning the building, \"CAFE\" in the lower-left room, and \"SHOP\" in the lower-right room. A narrow unlabeled central entrance aisle runs from a doorway at the bottom edge into the main hall, separating the cafe and shop. Label that bottom doorway \"ENTRANCE\" just outside the building. Each of the four corner rooms has one doorway directly into the main hall. Give Gallery A a pale blue fill, Gallery B pale green, the cafe pale orange and the shop pale violet; keep the hall and entrance aisle white. Put exactly two simple rectangular display plinth symbols in each gallery, exactly three round table symbols in the cafe, and exactly two parallel shelf symbols in the shop. A single continuous blue visitor-route line with arrowheads begins outside the entrance, goes through the entrance aisle and main hall, and ends inside Gallery A, passing through actual doorway openings without crossing any walls. Small bottom legend: a blue arrow icon followed by exactly \"SUGGESTED ROUTE\". Spacious legible labeling, precise geometry, no perspective, no people and no extra rooms.", + "steps": 25, + "seed": 21066, + "width": 1024, + "height": 1024 + }, + { + "label": "expanded-museum-plan-40", + "prompt": "A clear visitor floor-plan diagram for a small museum, square format, flat orthographic top-down vector design on white, with thick charcoal exterior walls, thinner interior walls and clearly visible doorway gaps. Title centered above the plan: \"MUSEUM VISITOR MAP\". Inside one rectangular building, arrange five labeled spaces: \"GALLERY A\" in a large upper-left room, \"GALLERY B\" in a large upper-right room, \"MAIN HALL\" in a wide horizontal central space spanning the building, \"CAFE\" in the lower-left room, and \"SHOP\" in the lower-right room. A narrow unlabeled central entrance aisle runs from a doorway at the bottom edge into the main hall, separating the cafe and shop. Label that bottom doorway \"ENTRANCE\" just outside the building. Each of the four corner rooms has one doorway directly into the main hall. Give Gallery A a pale blue fill, Gallery B pale green, the cafe pale orange and the shop pale violet; keep the hall and entrance aisle white. Put exactly two simple rectangular display plinth symbols in each gallery, exactly three round table symbols in the cafe, and exactly two parallel shelf symbols in the shop. A single continuous blue visitor-route line with arrowheads begins outside the entrance, goes through the entrance aisle and main hall, and ends inside Gallery A, passing through actual doorway openings without crossing any walls. Small bottom legend: a blue arrow icon followed by exactly \"SUGGESTED ROUTE\". Spacious legible labeling, precise geometry, no perspective, no people and no extra rooms.", + "steps": 40, + "seed": 21066, + "width": 1024, + "height": 1024 + }, + { + "label": "expanded-ridgeline-campaign-25", + "prompt": "Premium studio product campaign for an insulated water bottle, square print advertisement. Center one tall matte cobalt-blue metal bottle on a low pale stone rectangular pedestal. The bottle has a black screw cap, one narrow orange vertical stripe on its front, and a small white mountain-triangle emblem near its base. On the tabletop to the left of the pedestal is one short yellow enamel cup; to the right is one short turquoise enamel cup. There are exactly three products in total: the bottle and the two cups, all fully visible and separate. Clean warm gray seamless backdrop, soft directional daylight from the upper left, realistic brushed metal and enamel, restrained shadows. At the top, large elegant dark navy sans-serif lettering reads exactly \"TAKE THE LONG WAY\". Directly underneath it, smaller lettering reads exactly \"RIDGELINE / EVERYDAY ADVENTURE\". At the bottom center, small text reads exactly \"750 mL · BUILT TO REUSE\". Plenty of negative space around the type, crisp product edges, no additional props, hands or extra words. Front three-quarter product view with the orange stripe and mountain emblem clearly visible.", + "steps": 25, + "seed": 21077, + "width": 1024, + "height": 1024 + }, + { + "label": "expanded-ridgeline-campaign-40", + "prompt": "Premium studio product campaign for an insulated water bottle, square print advertisement. Center one tall matte cobalt-blue metal bottle on a low pale stone rectangular pedestal. The bottle has a black screw cap, one narrow orange vertical stripe on its front, and a small white mountain-triangle emblem near its base. On the tabletop to the left of the pedestal is one short yellow enamel cup; to the right is one short turquoise enamel cup. There are exactly three products in total: the bottle and the two cups, all fully visible and separate. Clean warm gray seamless backdrop, soft directional daylight from the upper left, realistic brushed metal and enamel, restrained shadows. At the top, large elegant dark navy sans-serif lettering reads exactly \"TAKE THE LONG WAY\". Directly underneath it, smaller lettering reads exactly \"RIDGELINE / EVERYDAY ADVENTURE\". At the bottom center, small text reads exactly \"750 mL · BUILT TO REUSE\". Plenty of negative space around the type, crisp product edges, no additional props, hands or extra words. Front three-quarter product view with the orange stripe and mountain emblem clearly visible.", + "steps": 40, + "seed": 21077, + "width": 1024, + "height": 1024 + }, + { + "label": "expanded-newsroom-interview-25", + "prompt": "Behind-the-scenes editorial photograph of a radio newsroom recording an interview, wide eye-level square composition. Exactly four adults with clearly separated faces and bodies. On the left, a woman journalist with a dark ponytail, white shirt and black headphones sits at a small table and holds an open notebook in her left hand while gesturing toward the guest with her right. On the right, a middle-aged man with a trimmed beard, brown jacket and silver headphones sits opposite her, speaking toward one black microphone on an articulated desk boom; a second matching microphone points toward the journalist. Behind a glass partition at center, a sound engineer in a blue sweater sits at a mixing console with one hand on a fader. Next to the engineer, a producer wearing a green shirt stands holding up a white card that reads exactly \"2 MIN\". Above the studio window, a red illuminated sign reads exactly \"ON AIR\". The two seated people look at each other, the engineer watches the controls, and the producer faces the studio. Plausible microphone arms and cables, convincing glass with subtle reflections that do not duplicate faces, believable fingers, soft warm studio lighting, documentary realism and no other people.", + "steps": 25, + "seed": 21088, + "width": 1024, + "height": 1024 + }, + { + "label": "expanded-newsroom-interview-40", + "prompt": "Behind-the-scenes editorial photograph of a radio newsroom recording an interview, wide eye-level square composition. Exactly four adults with clearly separated faces and bodies. On the left, a woman journalist with a dark ponytail, white shirt and black headphones sits at a small table and holds an open notebook in her left hand while gesturing toward the guest with her right. On the right, a middle-aged man with a trimmed beard, brown jacket and silver headphones sits opposite her, speaking toward one black microphone on an articulated desk boom; a second matching microphone points toward the journalist. Behind a glass partition at center, a sound engineer in a blue sweater sits at a mixing console with one hand on a fader. Next to the engineer, a producer wearing a green shirt stands holding up a white card that reads exactly \"2 MIN\". Above the studio window, a red illuminated sign reads exactly \"ON AIR\". The two seated people look at each other, the engineer watches the controls, and the producer faces the studio. Plausible microphone arms and cables, convincing glass with subtle reflections that do not duplicate faces, believable fingers, soft warm studio lighting, documentary realism and no other people.", + "steps": 40, + "seed": 21088, + "width": 1024, + "height": 1024 + }, + { + "label": "expanded-station-coat-edit-25", + "prompt": "Change only the mother's navy knee-length coat to a rich burgundy red fabric of the same cut, length, folds and texture. She is the woman on the left with dark curly hair in a low bun, gold earrings and a mustard yellow scarf. Preserve her exact face and identity, expression, gaze, body pose and handholding with the girl. Preserve the yellow scarf, all other clothing, the girl and stuffed rabbit, the father and grandmother's embrace, every person's hands and face, the green suitcase, train, platform, lighting, crop and photographic style. This is a localized garment recoloring, with the rest of the photograph retained.", + "steps": 25, + "seed": 21101, + "width": 1024, + "height": 1024, + "images": [ + "/poc/samples/fidelity-v3/bf16-dit/expanded-station-reunion-40.png" + ] + }, + { + "label": "expanded-station-coat-edit-40", + "prompt": "Change only the mother's navy knee-length coat to a rich burgundy red fabric of the same cut, length, folds and texture. She is the woman on the left with dark curly hair in a low bun, gold earrings and a mustard yellow scarf. Preserve her exact face and identity, expression, gaze, body pose and handholding with the girl. Preserve the yellow scarf, all other clothing, the girl and stuffed rabbit, the father and grandmother's embrace, every person's hands and face, the green suitcase, train, platform, lighting, crop and photographic style. This is a localized garment recoloring, with the rest of the photograph retained.", + "steps": 40, + "seed": 21101, + "width": 1024, + "height": 1024, + "images": [ + "/poc/samples/fidelity-v3/bf16-dit/expanded-station-reunion-40.png" + ] + }, + { + "label": "expanded-festival-type-edit-25", + "prompt": "Update only two text lines in this festival poster. Replace the top English headline \"RIVERLIGHT FILM WEEK\" with exactly \"RIVERLIGHT FILM NIGHTS\", keeping it centered in the same headline area with the same navy font style and a sensible size adjustment if needed. Replace the date line \"18–20 SEPTEMBER 2026\" with exactly \"25–27 SEPTEMBER 2026\". Preserve the Chinese headline \"河畔电影周\" exactly, every program time and English/Chinese film title, both footer lines, the two-column alignment, row spacing, bridge illustration, rules, colors, margins and paper background. Do not translate or reword any other text, and do not redesign the poster.", + "steps": 25, + "seed": 21112, + "width": 1024, + "height": 1024, + "images": [ + "/poc/samples/fidelity-v3/bf16-dit/expanded-bilingual-festival-40.png" + ] + }, + { + "label": "expanded-festival-type-edit-40", + "prompt": "Update only two text lines in this festival poster. Replace the top English headline \"RIVERLIGHT FILM WEEK\" with exactly \"RIVERLIGHT FILM NIGHTS\", keeping it centered in the same headline area with the same navy font style and a sensible size adjustment if needed. Replace the date line \"18–20 SEPTEMBER 2026\" with exactly \"25–27 SEPTEMBER 2026\". Preserve the Chinese headline \"河畔电影周\" exactly, every program time and English/Chinese film title, both footer lines, the two-column alignment, row spacing, bridge illustration, rules, colors, margins and paper background. Do not translate or reword any other text, and do not redesign the poster.", + "steps": 40, + "seed": 21112, + "width": 1024, + "height": 1024, + "images": [ + "/poc/samples/fidelity-v3/bf16-dit/expanded-bilingual-festival-40.png" + ] + }, + { + "label": "expanded-multiref-portrait-25", + "prompt": "Create a natural lifestyle campaign photograph using both references. From the first image, use only the mother on the left: preserve her recognizable oval face, dark curly hair tied in a low bun, small round gold earrings, navy coat and mustard yellow scarf. From the second image, use the central cobalt-blue insulated bottle: preserve its tall shape, black screw cap, single narrow orange front stripe and small white mountain-triangle emblem near the base. Show this same woman seated alone on a wooden railway-platform bench, waist-up, smiling gently toward the camera while holding that bottle upright with both hands around its lower half; the cap, orange stripe and emblem should remain visible. Put one green suitcase on the ground beside the bench, partly visible in frame. Use soft overcast daylight and a stationary passenger train softly out of focus behind her. Realistic hands and fingers, recognizable identity, coherent scale and contact shadows. Do not include the other people from the first reference, the cups or pedestal from the second reference, or any poster lettering. The intended result is one coherent photograph with exactly one woman and one bottle.", + "steps": 25, + "seed": 21123, + "width": 1024, + "height": 1024, + "images": [ + "/poc/samples/fidelity-v3/bf16-dit/expanded-station-reunion-40.png", + "/poc/samples/fidelity-v3/bf16-dit/expanded-ridgeline-campaign-40.png" + ] + }, + { + "label": "expanded-multiref-portrait-40", + "prompt": "Create a natural lifestyle campaign photograph using both references. From the first image, use only the mother on the left: preserve her recognizable oval face, dark curly hair tied in a low bun, small round gold earrings, navy coat and mustard yellow scarf. From the second image, use the central cobalt-blue insulated bottle: preserve its tall shape, black screw cap, single narrow orange front stripe and small white mountain-triangle emblem near the base. Show this same woman seated alone on a wooden railway-platform bench, waist-up, smiling gently toward the camera while holding that bottle upright with both hands around its lower half; the cap, orange stripe and emblem should remain visible. Put one green suitcase on the ground beside the bench, partly visible in frame. Use soft overcast daylight and a stationary passenger train softly out of focus behind her. Realistic hands and fingers, recognizable identity, coherent scale and contact shadows. Do not include the other people from the first reference, the cups or pedestal from the second reference, or any poster lettering. The intended result is one coherent photograph with exactly one woman and one bottle.", + "steps": 40, + "seed": 21123, + "width": 1024, + "height": 1024, + "images": [ + "/poc/samples/fidelity-v3/bf16-dit/expanded-station-reunion-40.png", + "/poc/samples/fidelity-v3/bf16-dit/expanded-ridgeline-campaign-40.png" + ] + }, + { + "label": "expanded-product-relocation-edit-25", + "prompt": "Edit the arrangement of the small cups in this product campaign while preserving the central bottle and all typography. Remove the yellow cup that is currently on the left, restoring the tabletop and its lighting naturally. Move the turquoise cup from the right side to the far left of the pedestal, fully visible with a small clear gap from the pedestal; leave the former right-hand position empty. The final image must contain exactly two products: the unchanged cobalt-blue bottle on its original pedestal and the one turquoise cup on the left. Preserve the bottle's shape, cap, orange stripe, white mountain emblem, location and scale; preserve the pedestal, backdrop, framing, light direction, and the exact text \"TAKE THE LONG WAY\", \"RIDGELINE / EVERYDAY ADVENTURE\", and \"750 mL · BUILT TO REUSE\". Give the moved cup a physically consistent contact shadow. Do not introduce any extra objects or redesign the campaign.", + "steps": 25, + "seed": 21134, + "width": 1024, + "height": 1024, + "images": [ + "/poc/samples/fidelity-v3/bf16-dit/expanded-ridgeline-campaign-40.png" + ] + }, + { + "label": "expanded-product-relocation-edit-40", + "prompt": "Edit the arrangement of the small cups in this product campaign while preserving the central bottle and all typography. Remove the yellow cup that is currently on the left, restoring the tabletop and its lighting naturally. Move the turquoise cup from the right side to the far left of the pedestal, fully visible with a small clear gap from the pedestal; leave the former right-hand position empty. The final image must contain exactly two products: the unchanged cobalt-blue bottle on its original pedestal and the one turquoise cup on the left. Preserve the bottle's shape, cap, orange stripe, white mountain emblem, location and scale; preserve the pedestal, backdrop, framing, light direction, and the exact text \"TAKE THE LONG WAY\", \"RIDGELINE / EVERYDAY ADVENTURE\", and \"750 mL · BUILT TO REUSE\". Give the moved cup a physically consistent contact shadow. Do not introduce any extra objects or redesign the campaign.", + "steps": 40, + "seed": 21134, + "width": 1024, + "height": 1024, + "images": [ + "/poc/samples/fidelity-v3/bf16-dit/expanded-ridgeline-campaign-40.png" + ] + } +] diff --git a/reproduction/experiments/fidelity-v3/expanded-jobs.json b/reproduction/experiments/fidelity-v3/expanded-jobs.json new file mode 100644 index 0000000000000000000000000000000000000000..7aa1232b5cf878fb7cc94337b31f6ee7d801a9e0 --- /dev/null +++ b/reproduction/experiments/fidelity-v3/expanded-jobs.json @@ -0,0 +1,220 @@ +[ + { + "label": "expanded-community-kitchen-25", + "prompt": "Documentary photograph for a community cooking program, wide eye-level composition in a bright, practical teaching kitchen. Exactly six adults, all visible from at least the waist up. At the left counter, a woman with short black hair in a green apron chops carrots on a wooden board, with one hand holding the knife and the other safely curled on the carrot. At center, a silver-haired man in a white apron holds a ladle over a large soup pot. Immediately to his right, a woman in a mustard cardigan passes a rectangular tray of four bread rolls to a younger man in a navy shirt, who receives it with both hands; their hands meet opposite edges of the same tray. At the far right, a man in a red apron rinses a colander of leafy greens at the sink. Behind the central counter, a woman wearing glasses and a purple sweater reads a paper recipe card. Natural varied expressions, clear individual faces, plausible arms and hands attached to their owners, no extra people, realistic morning window light and stainless-steel reflections. No readable signage. The image should feel like a real coordinated cooking session, not a posed group portrait.", + "steps": 25, + "seed": 21011, + "width": 1024, + "height": 1024 + }, + { + "label": "expanded-community-kitchen-40", + "prompt": "Documentary photograph for a community cooking program, wide eye-level composition in a bright, practical teaching kitchen. Exactly six adults, all visible from at least the waist up. At the left counter, a woman with short black hair in a green apron chops carrots on a wooden board, with one hand holding the knife and the other safely curled on the carrot. At center, a silver-haired man in a white apron holds a ladle over a large soup pot. Immediately to his right, a woman in a mustard cardigan passes a rectangular tray of four bread rolls to a younger man in a navy shirt, who receives it with both hands; their hands meet opposite edges of the same tray. At the far right, a man in a red apron rinses a colander of leafy greens at the sink. Behind the central counter, a woman wearing glasses and a purple sweater reads a paper recipe card. Natural varied expressions, clear individual faces, plausible arms and hands attached to their owners, no extra people, realistic morning window light and stainless-steel reflections. No readable signage. The image should feel like a real coordinated cooking session, not a posed group portrait.", + "steps": 40, + "seed": 21011, + "width": 1024, + "height": 1024 + }, + { + "label": "expanded-chess-awards-25", + "prompt": "Editorial event photograph of a small chess championship award presentation on a low indoor stage. Exactly five adults. At center, a smiling woman with a short dark bob sits in a clearly recognizable manual wheelchair; she and the presenter standing immediately to her left jointly hold one gold trophy between them. The presenter is an older man wearing a charcoal suit. Immediately to the winner's right, the runner-up, a young man in a cream sweater, wears one silver medal and applauds. At the far left, a female photographer kneels with a camera held to her eye and aims toward the winner. At the far right, a host in a burgundy suit holds a microphone and a small cue card. All five people fit fully in frame, with believable hands, wheelchair wheels and footrests. Behind them, a clean white stage banner reads exactly \"CITY CHESS FINAL\" with \"2026\" centered on the line below. A small chess table with a board stands at the rear left. Warm indoor event lighting, realistic fabric and skin, no audience or additional figures.", + "steps": 25, + "seed": 21022, + "width": 1024, + "height": 1024 + }, + { + "label": "expanded-chess-awards-40", + "prompt": "Editorial event photograph of a small chess championship award presentation on a low indoor stage. Exactly five adults. At center, a smiling woman with a short dark bob sits in a clearly recognizable manual wheelchair; she and the presenter standing immediately to her left jointly hold one gold trophy between them. The presenter is an older man wearing a charcoal suit. Immediately to the winner's right, the runner-up, a young man in a cream sweater, wears one silver medal and applauds. At the far left, a female photographer kneels with a camera held to her eye and aims toward the winner. At the far right, a host in a burgundy suit holds a microphone and a small cue card. All five people fit fully in frame, with believable hands, wheelchair wheels and footrests. Behind them, a clean white stage banner reads exactly \"CITY CHESS FINAL\" with \"2026\" centered on the line below. A small chess table with a board stands at the rear left. Warm indoor event lighting, realistic fabric and skin, no audience or additional figures.", + "steps": 40, + "seed": 21022, + "width": 1024, + "height": 1024 + }, + { + "label": "expanded-station-reunion-25", + "prompt": "Natural editorial photograph of a family reunion beside a stationary passenger train on a quiet outdoor railway platform. Exactly four people, full bodies in frame, no other passengers. On the left is a woman in her thirties with an oval face, dark curly hair tied in a low bun, small round gold earrings, a navy knee-length coat and a mustard yellow scarf. She holds the left hand of a young girl standing at center; the girl wears a red raincoat and carries a small stuffed rabbit in her free hand. On the right, a man with close-cropped brown hair, a light gray jacket and a red backpack bends slightly to embrace an older gray-haired woman wearing a teal coat; the older woman faces him with one arm around his shoulder. The mother and girl watch them with happy, understated expressions. One upright green suitcase stands beside the mother's left leg. Anatomically plausible hands and embraces with clear ownership of each arm, natural eye contact, soft overcast daylight, authentic train windows and platform paving. Do not crop feet or add duplicate people. No readable logos or signs.", + "steps": 25, + "seed": 21033, + "width": 1024, + "height": 1024 + }, + { + "label": "expanded-station-reunion-40", + "prompt": "Natural editorial photograph of a family reunion beside a stationary passenger train on a quiet outdoor railway platform. Exactly four people, full bodies in frame, no other passengers. On the left is a woman in her thirties with an oval face, dark curly hair tied in a low bun, small round gold earrings, a navy knee-length coat and a mustard yellow scarf. She holds the left hand of a young girl standing at center; the girl wears a red raincoat and carries a small stuffed rabbit in her free hand. On the right, a man with close-cropped brown hair, a light gray jacket and a red backpack bends slightly to embrace an older gray-haired woman wearing a teal coat; the older woman faces him with one arm around his shoulder. The mother and girl watch them with happy, understated expressions. One upright green suitcase stands beside the mother's left leg. Anatomically plausible hands and embraces with clear ownership of each arm, natural eye contact, soft overcast daylight, authentic train windows and platform paving. Do not crop feet or add duplicate people. No readable logos or signs.", + "steps": 40, + "seed": 21033, + "width": 1024, + "height": 1024 + }, + { + "label": "expanded-bilingual-festival-25", + "prompt": "A polished bilingual neighborhood film festival poster, square format, flat front-on print artwork with generous margins, warm cream paper and deep navy typography accented by vermilion. The large centered English title at the top reads exactly \"RIVERLIGHT FILM WEEK\". Directly beneath it, an equally clear Chinese title reads exactly \"河畔电影周\". The next centered line reads exactly \"18–20 SEPTEMBER 2026\". Below a thin vermilion rule, a tidy two-column program occupies the middle: the left column contains times, and the right column contains one English film title with its Chinese title directly underneath. Row one: \"18:00\", \"A Quiet River\", \"静静的河\". Row two: \"19:30\", \"Night Market\", \"夜市\". Row three: \"21:00\", \"Home Again\", \"再次回家\". Maintain aligned columns and clearly separated rows. At the bottom, centered on separate lines, print exactly \"RIVERSIDE CINEMA / 河畔影院\" and \"FREE ENTRY / 免费入场\". A small restrained illustration of a red bridge over two blue river lines sits between the program and footer without touching any letters. Sharp readable Latin and Chinese lettering, consistent type hierarchy, no additional text, no mockup perspective.", + "steps": 25, + "seed": 21044, + "width": 1024, + "height": 1024 + }, + { + "label": "expanded-bilingual-festival-40", + "prompt": "A polished bilingual neighborhood film festival poster, square format, flat front-on print artwork with generous margins, warm cream paper and deep navy typography accented by vermilion. The large centered English title at the top reads exactly \"RIVERLIGHT FILM WEEK\". Directly beneath it, an equally clear Chinese title reads exactly \"河畔电影周\". The next centered line reads exactly \"18–20 SEPTEMBER 2026\". Below a thin vermilion rule, a tidy two-column program occupies the middle: the left column contains times, and the right column contains one English film title with its Chinese title directly underneath. Row one: \"18:00\", \"A Quiet River\", \"静静的河\". Row two: \"19:30\", \"Night Market\", \"夜市\". Row three: \"21:00\", \"Home Again\", \"再次回家\". Maintain aligned columns and clearly separated rows. At the bottom, centered on separate lines, print exactly \"RIVERSIDE CINEMA / 河畔影院\" and \"FREE ENTRY / 免费入场\". A small restrained illustration of a red bridge over two blue river lines sits between the program and footer without touching any letters. Sharp readable Latin and Chinese lettering, consistent type hierarchy, no additional text, no mockup perspective.", + "steps": 40, + "seed": 21044, + "width": 1024, + "height": 1024 + }, + { + "label": "expanded-library-comic-25", + "prompt": "A polished, friendly three-panel comic strip in a square canvas, with exactly three equal vertical panels arranged left to right, clear white gutters and a thin dark border around each panel. Keep the same two adult characters and library reading-room setting in all panels: Maya has a chin-length black bob, round glasses and a red sweater; Leo has short curly brown hair and a blue button-up shirt. Each panel contains exactly two white speech bubbles, with clear tails pointing to the correct speaker. Panel 1: Maya stands beside a wooden reading chair, looks worried and searches her bag. Maya says exactly \"Where are my keys?\" Leo, standing to her right, says exactly \"Check your pocket.\" Panel 2: Maya turns out an empty coat pocket and says exactly \"Not here!\" Leo points down at a small key ring visible under that same chair and says exactly \"Under the chair.\" Panel 3: Maya stands upright holding that key ring up, smiles and says exactly \"Found them. Thanks!\" Leo smiles and says exactly \"You're welcome.\" Crisp ink lines, restrained warm colors, expressive natural poses, consistent faces and clothing, believable hands, a coherent sequence with the chair in the same room. Readable dialogue with no narration captions and no additional panels or characters.", + "steps": 25, + "seed": 21055, + "width": 1024, + "height": 1024 + }, + { + "label": "expanded-library-comic-40", + "prompt": "A polished, friendly three-panel comic strip in a square canvas, with exactly three equal vertical panels arranged left to right, clear white gutters and a thin dark border around each panel. Keep the same two adult characters and library reading-room setting in all panels: Maya has a chin-length black bob, round glasses and a red sweater; Leo has short curly brown hair and a blue button-up shirt. Each panel contains exactly two white speech bubbles, with clear tails pointing to the correct speaker. Panel 1: Maya stands beside a wooden reading chair, looks worried and searches her bag. Maya says exactly \"Where are my keys?\" Leo, standing to her right, says exactly \"Check your pocket.\" Panel 2: Maya turns out an empty coat pocket and says exactly \"Not here!\" Leo points down at a small key ring visible under that same chair and says exactly \"Under the chair.\" Panel 3: Maya stands upright holding that key ring up, smiles and says exactly \"Found them. Thanks!\" Leo smiles and says exactly \"You're welcome.\" Crisp ink lines, restrained warm colors, expressive natural poses, consistent faces and clothing, believable hands, a coherent sequence with the chair in the same room. Readable dialogue with no narration captions and no additional panels or characters.", + "steps": 40, + "seed": 21055, + "width": 1024, + "height": 1024 + }, + { + "label": "expanded-museum-plan-25", + "prompt": "A clear visitor floor-plan diagram for a small museum, square format, flat orthographic top-down vector design on white, with thick charcoal exterior walls, thinner interior walls and clearly visible doorway gaps. Title centered above the plan: \"MUSEUM VISITOR MAP\". Inside one rectangular building, arrange five labeled spaces: \"GALLERY A\" in a large upper-left room, \"GALLERY B\" in a large upper-right room, \"MAIN HALL\" in a wide horizontal central space spanning the building, \"CAFE\" in the lower-left room, and \"SHOP\" in the lower-right room. A narrow unlabeled central entrance aisle runs from a doorway at the bottom edge into the main hall, separating the cafe and shop. Label that bottom doorway \"ENTRANCE\" just outside the building. Each of the four corner rooms has one doorway directly into the main hall. Give Gallery A a pale blue fill, Gallery B pale green, the cafe pale orange and the shop pale violet; keep the hall and entrance aisle white. Put exactly two simple rectangular display plinth symbols in each gallery, exactly three round table symbols in the cafe, and exactly two parallel shelf symbols in the shop. A single continuous blue visitor-route line with arrowheads begins outside the entrance, goes through the entrance aisle and main hall, and ends inside Gallery A, passing through actual doorway openings without crossing any walls. Small bottom legend: a blue arrow icon followed by exactly \"SUGGESTED ROUTE\". Spacious legible labeling, precise geometry, no perspective, no people and no extra rooms.", + "steps": 25, + "seed": 21066, + "width": 1024, + "height": 1024 + }, + { + "label": "expanded-museum-plan-40", + "prompt": "A clear visitor floor-plan diagram for a small museum, square format, flat orthographic top-down vector design on white, with thick charcoal exterior walls, thinner interior walls and clearly visible doorway gaps. Title centered above the plan: \"MUSEUM VISITOR MAP\". Inside one rectangular building, arrange five labeled spaces: \"GALLERY A\" in a large upper-left room, \"GALLERY B\" in a large upper-right room, \"MAIN HALL\" in a wide horizontal central space spanning the building, \"CAFE\" in the lower-left room, and \"SHOP\" in the lower-right room. A narrow unlabeled central entrance aisle runs from a doorway at the bottom edge into the main hall, separating the cafe and shop. Label that bottom doorway \"ENTRANCE\" just outside the building. Each of the four corner rooms has one doorway directly into the main hall. Give Gallery A a pale blue fill, Gallery B pale green, the cafe pale orange and the shop pale violet; keep the hall and entrance aisle white. Put exactly two simple rectangular display plinth symbols in each gallery, exactly three round table symbols in the cafe, and exactly two parallel shelf symbols in the shop. A single continuous blue visitor-route line with arrowheads begins outside the entrance, goes through the entrance aisle and main hall, and ends inside Gallery A, passing through actual doorway openings without crossing any walls. Small bottom legend: a blue arrow icon followed by exactly \"SUGGESTED ROUTE\". Spacious legible labeling, precise geometry, no perspective, no people and no extra rooms.", + "steps": 40, + "seed": 21066, + "width": 1024, + "height": 1024 + }, + { + "label": "expanded-ridgeline-campaign-25", + "prompt": "Premium studio product campaign for an insulated water bottle, square print advertisement. Center one tall matte cobalt-blue metal bottle on a low pale stone rectangular pedestal. The bottle has a black screw cap, one narrow orange vertical stripe on its front, and a small white mountain-triangle emblem near its base. On the tabletop to the left of the pedestal is one short yellow enamel cup; to the right is one short turquoise enamel cup. There are exactly three products in total: the bottle and the two cups, all fully visible and separate. Clean warm gray seamless backdrop, soft directional daylight from the upper left, realistic brushed metal and enamel, restrained shadows. At the top, large elegant dark navy sans-serif lettering reads exactly \"TAKE THE LONG WAY\". Directly underneath it, smaller lettering reads exactly \"RIDGELINE / EVERYDAY ADVENTURE\". At the bottom center, small text reads exactly \"750 mL · BUILT TO REUSE\". Plenty of negative space around the type, crisp product edges, no additional props, hands or extra words. Front three-quarter product view with the orange stripe and mountain emblem clearly visible.", + "steps": 25, + "seed": 21077, + "width": 1024, + "height": 1024 + }, + { + "label": "expanded-ridgeline-campaign-40", + "prompt": "Premium studio product campaign for an insulated water bottle, square print advertisement. Center one tall matte cobalt-blue metal bottle on a low pale stone rectangular pedestal. The bottle has a black screw cap, one narrow orange vertical stripe on its front, and a small white mountain-triangle emblem near its base. On the tabletop to the left of the pedestal is one short yellow enamel cup; to the right is one short turquoise enamel cup. There are exactly three products in total: the bottle and the two cups, all fully visible and separate. Clean warm gray seamless backdrop, soft directional daylight from the upper left, realistic brushed metal and enamel, restrained shadows. At the top, large elegant dark navy sans-serif lettering reads exactly \"TAKE THE LONG WAY\". Directly underneath it, smaller lettering reads exactly \"RIDGELINE / EVERYDAY ADVENTURE\". At the bottom center, small text reads exactly \"750 mL · BUILT TO REUSE\". Plenty of negative space around the type, crisp product edges, no additional props, hands or extra words. Front three-quarter product view with the orange stripe and mountain emblem clearly visible.", + "steps": 40, + "seed": 21077, + "width": 1024, + "height": 1024 + }, + { + "label": "expanded-newsroom-interview-25", + "prompt": "Behind-the-scenes editorial photograph of a radio newsroom recording an interview, wide eye-level square composition. Exactly four adults with clearly separated faces and bodies. On the left, a woman journalist with a dark ponytail, white shirt and black headphones sits at a small table and holds an open notebook in her left hand while gesturing toward the guest with her right. On the right, a middle-aged man with a trimmed beard, brown jacket and silver headphones sits opposite her, speaking toward one black microphone on an articulated desk boom; a second matching microphone points toward the journalist. Behind a glass partition at center, a sound engineer in a blue sweater sits at a mixing console with one hand on a fader. Next to the engineer, a producer wearing a green shirt stands holding up a white card that reads exactly \"2 MIN\". Above the studio window, a red illuminated sign reads exactly \"ON AIR\". The two seated people look at each other, the engineer watches the controls, and the producer faces the studio. Plausible microphone arms and cables, convincing glass with subtle reflections that do not duplicate faces, believable fingers, soft warm studio lighting, documentary realism and no other people.", + "steps": 25, + "seed": 21088, + "width": 1024, + "height": 1024 + }, + { + "label": "expanded-newsroom-interview-40", + "prompt": "Behind-the-scenes editorial photograph of a radio newsroom recording an interview, wide eye-level square composition. Exactly four adults with clearly separated faces and bodies. On the left, a woman journalist with a dark ponytail, white shirt and black headphones sits at a small table and holds an open notebook in her left hand while gesturing toward the guest with her right. On the right, a middle-aged man with a trimmed beard, brown jacket and silver headphones sits opposite her, speaking toward one black microphone on an articulated desk boom; a second matching microphone points toward the journalist. Behind a glass partition at center, a sound engineer in a blue sweater sits at a mixing console with one hand on a fader. Next to the engineer, a producer wearing a green shirt stands holding up a white card that reads exactly \"2 MIN\". Above the studio window, a red illuminated sign reads exactly \"ON AIR\". The two seated people look at each other, the engineer watches the controls, and the producer faces the studio. Plausible microphone arms and cables, convincing glass with subtle reflections that do not duplicate faces, believable fingers, soft warm studio lighting, documentary realism and no other people.", + "steps": 40, + "seed": 21088, + "width": 1024, + "height": 1024 + }, + { + "label": "expanded-station-coat-edit-25", + "prompt": "Change only the mother's navy knee-length coat to a rich burgundy red fabric of the same cut, length, folds and texture. She is the woman on the left with dark curly hair in a low bun, gold earrings and a mustard yellow scarf. Preserve her exact face and identity, expression, gaze, body pose and handholding with the girl. Preserve the yellow scarf, all other clothing, the girl and stuffed rabbit, the father and grandmother's embrace, every person's hands and face, the green suitcase, train, platform, lighting, crop and photographic style. This is a localized garment recoloring, with the rest of the photograph retained.", + "steps": 25, + "seed": 21101, + "width": 1024, + "height": 1024, + "images": [ + "/poc/samples/fidelity-v3/bf16-dit/expanded-station-reunion-40.png" + ] + }, + { + "label": "expanded-station-coat-edit-40", + "prompt": "Change only the mother's navy knee-length coat to a rich burgundy red fabric of the same cut, length, folds and texture. She is the woman on the left with dark curly hair in a low bun, gold earrings and a mustard yellow scarf. Preserve her exact face and identity, expression, gaze, body pose and handholding with the girl. Preserve the yellow scarf, all other clothing, the girl and stuffed rabbit, the father and grandmother's embrace, every person's hands and face, the green suitcase, train, platform, lighting, crop and photographic style. This is a localized garment recoloring, with the rest of the photograph retained.", + "steps": 40, + "seed": 21101, + "width": 1024, + "height": 1024, + "images": [ + "/poc/samples/fidelity-v3/bf16-dit/expanded-station-reunion-40.png" + ] + }, + { + "label": "expanded-festival-type-edit-25", + "prompt": "Update only two text lines in this festival poster. Replace the top English headline \"RIVERLIGHT FILM WEEK\" with exactly \"RIVERLIGHT FILM NIGHTS\", keeping it centered in the same headline area with the same navy font style and a sensible size adjustment if needed. Replace the date line \"18–20 SEPTEMBER 2026\" with exactly \"25–27 SEPTEMBER 2026\". Preserve the Chinese headline \"河畔电影周\" exactly, every program time and English/Chinese film title, both footer lines, the two-column alignment, row spacing, bridge illustration, rules, colors, margins and paper background. Do not translate or reword any other text, and do not redesign the poster.", + "steps": 25, + "seed": 21112, + "width": 1024, + "height": 1024, + "images": [ + "/poc/samples/fidelity-v3/bf16-dit/expanded-bilingual-festival-40.png" + ] + }, + { + "label": "expanded-festival-type-edit-40", + "prompt": "Update only two text lines in this festival poster. Replace the top English headline \"RIVERLIGHT FILM WEEK\" with exactly \"RIVERLIGHT FILM NIGHTS\", keeping it centered in the same headline area with the same navy font style and a sensible size adjustment if needed. Replace the date line \"18–20 SEPTEMBER 2026\" with exactly \"25–27 SEPTEMBER 2026\". Preserve the Chinese headline \"河畔电影周\" exactly, every program time and English/Chinese film title, both footer lines, the two-column alignment, row spacing, bridge illustration, rules, colors, margins and paper background. Do not translate or reword any other text, and do not redesign the poster.", + "steps": 40, + "seed": 21112, + "width": 1024, + "height": 1024, + "images": [ + "/poc/samples/fidelity-v3/bf16-dit/expanded-bilingual-festival-40.png" + ] + }, + { + "label": "expanded-multiref-portrait-25", + "prompt": "Create a natural lifestyle campaign photograph using both references. From the first image, use only the mother on the left: preserve her recognizable oval face, dark curly hair tied in a low bun, small round gold earrings, navy coat and mustard yellow scarf. From the second image, use the central cobalt-blue insulated bottle: preserve its tall shape, black screw cap, single narrow orange front stripe and small white mountain-triangle emblem near the base. Show this same woman seated alone on a wooden railway-platform bench, waist-up, smiling gently toward the camera while holding that bottle upright with both hands around its lower half; the cap, orange stripe and emblem should remain visible. Put one green suitcase on the ground beside the bench, partly visible in frame. Use soft overcast daylight and a stationary passenger train softly out of focus behind her. Realistic hands and fingers, recognizable identity, coherent scale and contact shadows. Do not include the other people from the first reference, the cups or pedestal from the second reference, or any poster lettering. The intended result is one coherent photograph with exactly one woman and one bottle.", + "steps": 25, + "seed": 21123, + "width": 1024, + "height": 1024, + "images": [ + "/poc/samples/fidelity-v3/bf16-dit/expanded-station-reunion-40.png", + "/poc/samples/fidelity-v3/bf16-dit/expanded-ridgeline-campaign-40.png" + ] + }, + { + "label": "expanded-multiref-portrait-40", + "prompt": "Create a natural lifestyle campaign photograph using both references. From the first image, use only the mother on the left: preserve her recognizable oval face, dark curly hair tied in a low bun, small round gold earrings, navy coat and mustard yellow scarf. From the second image, use the central cobalt-blue insulated bottle: preserve its tall shape, black screw cap, single narrow orange front stripe and small white mountain-triangle emblem near the base. Show this same woman seated alone on a wooden railway-platform bench, waist-up, smiling gently toward the camera while holding that bottle upright with both hands around its lower half; the cap, orange stripe and emblem should remain visible. Put one green suitcase on the ground beside the bench, partly visible in frame. Use soft overcast daylight and a stationary passenger train softly out of focus behind her. Realistic hands and fingers, recognizable identity, coherent scale and contact shadows. Do not include the other people from the first reference, the cups or pedestal from the second reference, or any poster lettering. The intended result is one coherent photograph with exactly one woman and one bottle.", + "steps": 40, + "seed": 21123, + "width": 1024, + "height": 1024, + "images": [ + "/poc/samples/fidelity-v3/bf16-dit/expanded-station-reunion-40.png", + "/poc/samples/fidelity-v3/bf16-dit/expanded-ridgeline-campaign-40.png" + ] + }, + { + "label": "expanded-product-relocation-edit-25", + "prompt": "Edit the arrangement of the small cups in this product campaign while preserving the central bottle and all typography. Remove the yellow cup that is currently on the left, restoring the tabletop and its lighting naturally. Move the turquoise cup from the right side to the far left of the pedestal, fully visible with a small clear gap from the pedestal; leave the former right-hand position empty. The final image must contain exactly two products: the unchanged cobalt-blue bottle on its original pedestal and the one turquoise cup on the left. Preserve the bottle's shape, cap, orange stripe, white mountain emblem, location and scale; preserve the pedestal, backdrop, framing, light direction, and the exact text \"TAKE THE LONG WAY\", \"RIDGELINE / EVERYDAY ADVENTURE\", and \"750 mL · BUILT TO REUSE\". Give the moved cup a physically consistent contact shadow. Do not introduce any extra objects or redesign the campaign.", + "steps": 25, + "seed": 21134, + "width": 1024, + "height": 1024, + "images": [ + "/poc/samples/fidelity-v3/bf16-dit/expanded-ridgeline-campaign-40.png" + ] + }, + { + "label": "expanded-product-relocation-edit-40", + "prompt": "Edit the arrangement of the small cups in this product campaign while preserving the central bottle and all typography. Remove the yellow cup that is currently on the left, restoring the tabletop and its lighting naturally. Move the turquoise cup from the right side to the far left of the pedestal, fully visible with a small clear gap from the pedestal; leave the former right-hand position empty. The final image must contain exactly two products: the unchanged cobalt-blue bottle on its original pedestal and the one turquoise cup on the left. Preserve the bottle's shape, cap, orange stripe, white mountain emblem, location and scale; preserve the pedestal, backdrop, framing, light direction, and the exact text \"TAKE THE LONG WAY\", \"RIDGELINE / EVERYDAY ADVENTURE\", and \"750 mL · BUILT TO REUSE\". Give the moved cup a physically consistent contact shadow. Do not introduce any extra objects or redesign the campaign.", + "steps": 40, + "seed": 21134, + "width": 1024, + "height": 1024, + "images": [ + "/poc/samples/fidelity-v3/bf16-dit/expanded-ridgeline-campaign-40.png" + ] + } +] diff --git a/reproduction/experiments/fidelity-v3/expanded-rubric.md b/reproduction/experiments/fidelity-v3/expanded-rubric.md new file mode 100644 index 0000000000000000000000000000000000000000..9d4916db1cfff6972ba350495c9de53190e1812f --- /dev/null +++ b/reproduction/experiments/fidelity-v3/expanded-rubric.md @@ -0,0 +1,92 @@ +# Expanded Qwen fidelity evaluation + +This is an **evaluation-only** set. Do not use its prompts, seeds, teacher images, editing references or captured activations to fit, calibrate or select a quantization configuration. Freeze candidate settings before running it. If these results motivate another change, report that reuse and obtain a fresh independent evaluation set for the next acceptance claim. + +`expanded-jobs.json` contains 12 scenarios, each at 25 and 40 steps, for 24 jobs at 1024×1024: eight text-to-image scenarios followed by four editing scenarios. Each pair has identical prompt, seed, size and references; only the step count and label differ. Seeds are new and fixed. The manifest uses the existing runner schema. + +## Order, controls and evidence + +1. Generate the BF16-transformer teacher for the eight generation scenarios first. The three required 40-step reference images are `expanded-station-reunion-40.png`, `expanded-bilingual-festival-40.png`, and `expanded-ridgeline-campaign-40.png`, all under `/poc/samples/fidelity-v3/bf16-dit/`. +2. Use these exact reference bytes for every teacher and candidate editing job, including both step counts. Never substitute a candidate's own generated reference. Record reference SHA256 values with the run. The multi-reference job receives the station image first and the product image second. +3. Keep seed, scheduler, KV mode, text encoder, VAE, resolution and preprocessing fixed within teacher/candidate comparisons. Record any exception. This teacher isolates transformer precision; it is not an entirely BF16 pipeline if its encoder is quantized. +4. Inspect the actual teacher references before scoring edits. If a required source feature is absent or unreadable, mark that edit criterion **source-limited** and describe what is visible. Do not claim preservation of an identity, text or product feature the reference never established. +5. Review at full resolution. Record adherence, visual defects and drift from the teacher separately. A teacher can fail the prompt; a candidate can fix a prompt error while becoming less pixel-similar. SSIM or PSNR cannot adjudicate that distinction. + +For each criterion use **pass / partial / fail / source-limited**, with one sentence of visible evidence. Partial means the intended result is recognizable but incomplete or ambiguous; identify the exact ambiguity. Record failures independently rather than averaging unrelated properties into an unexplained score. Review 25 and 40 steps separately, then state whether the additional steps visibly helped, hurt or made no material difference for that case. + +## Generation scenarios + +| Scenario and seed | Critical requested content | Composition and physical plausibility | Secondary review | +|---|---|---|---| +| `expanded-community-kitchen`, 21011 | Exactly six adults with six distinct roles: green-apron woman chopping carrots at left; silver-haired man ladling soup at center; mustard-cardigan woman passing a tray to navy-shirt man at center-right; red-apron man rinsing greens at right; glasses/purple-sweater woman reading a recipe behind center. The one tray contains four bread rolls. | Passing and receiving hands touch opposite edges of the same tray; arms attach to the right people. Knife/holding hand and ladle/pot contact are plausible. All six are visible at least waist-up and engaged rather than posed. | Face separation, anatomy, kitchen tools, reflections, food texture, coordinated documentary appearance. Note any background figure/reflection counted as a spurious person. | +| `expanded-chess-awards`, 21022 | Exactly five adults: seated woman winner at center, older male presenter immediately left, cream-sweater male runner-up immediately right, kneeling female photographer far left, burgundy-suited host far right. Winner and presenter jointly hold one gold trophy; runner-up wears one silver medal and applauds. Banner is `CITY CHESS FINAL` with `2026` below. | Winner sits in a recognizable manual wheelchair with coherent wheels/footrests; all five fit fully in frame. Trophy hand contact, applauding hands, camera held to eye and microphone/cue-card grips are plausible. | Small rear-left chess table/board; actual stage presentation; no additional audience. Do not penalize assistive-device use or natural bodily variation; score incoherent rendering and missed requested geometry. | +| `expanded-station-reunion`, 21033 | Exactly four people: mother left with dark low bun, oval face, gold earrings, navy coat/yellow scarf; red-coated girl at center holding mother's hand and carrying a stuffed rabbit; gray-jacket/red-backpack father at right embracing teal-coated gray-haired grandmother. One green suitcase by mother's left leg. | Clear handholding and embrace, correct arm ownership, believable eye contact. Full bodies/feet in frame. Mother and girl watch the reunion; father and grandmother face each other. | Preserve a usable mother identity for later edits. Train/platform realism, coat/scarf details, coherent suitcase, no duplicate people or cropped feet. | +| `expanded-bilingual-festival`, 21044 | All 14 specified text strings below are readable and correct. Three program rows have aligned time/title columns, each English title directly above its corresponding Chinese title. | Header/date, thin rule, program, bridge illustration and two footer lines form a clear hierarchy without collisions. Cream/navy/vermilion flat square poster, generous margins. | Chinese glyph integrity, Latin spelling, punctuation and typography. Check letters at native resolution; decorative pseudo-text is a failure even when the poster looks polished. | +| `expanded-library-comic`, 21055 | Exactly three equal vertical panels left-to-right; the same Maya (black bob/glasses/red sweater) and Leo (curly brown hair/blue shirt) in each. Exactly two speech bubbles per panel with the six exact lines below and correct speakers. Sequence: searching bag; empty pocket plus keys under chair; Maya holding recovered keys. | Bubble tails identify their speakers; reading order is clear. Same reading-room/chair continuity. Hands, pointing, bag, pocket and key ring form a coherent causal sequence. | Expressions change appropriately. No extra panels, characters, captions or missing key ring. A plausible illustration with scrambled dialogue or event order fails those criteria. | +| `expanded-museum-plan`, 21066 | Exactly five labeled spaces: Gallery A upper left, Gallery B upper right, Main Hall central full width, Cafe lower left, Shop lower right. Two display plinths per gallery, three round cafe tables, two shop shelves. All eight label strings below correct. | Bottom-center entrance aisle separates Cafe/Shop and joins the hall; each corner room has one doorway into hall. One continuous blue route goes entrance → aisle → hall → Gallery A through doorways, crossing no walls. | Orthographic readable geometry, distinct requested room colors, wall/door consistency, arrow direction, no extra rooms or perspective. Do not count entrance aisle as a sixth room. | +| `expanded-ridgeline-campaign`, 21077 | Exactly three products: central cobalt bottle on pedestal, yellow cup left, turquoise cup right. Bottle has black cap, one narrow orange stripe and a small white mountain-triangle emblem near its base. All three exact text lines below correct. | Separate fully visible objects, coherent cap/bottle form, plausible contact shadows, stable front three-quarter view, clear type hierarchy and negative space. | Material rendering, stripe/emblem clarity, usable reference product identity, consistent studio lighting. Additional props/products or illegible bottom text are explicit failures. | +| `expanded-newsroom-interview`, 21088 | Exactly four adults: ponytailed woman journalist left with notebook and gesture; bearded male guest right; blue-sweater engineer at console behind glass; standing green-shirt producer beside engineer holding `2 MIN`. Red sign reads `ON AIR`. Exactly two desk microphones, one aimed at each seated speaker. | Journalist/guest look at each other; engineer's hand reaches a fader; producer faces studio. Coherent microphone booms/cables, notebook grip, cue-card grip, glass and reflections without duplicate people. | Studio realism, faces/hands, desk objects, seated/staff depth separation and readable signs. | + +### Exact text checks + +For each string, record correct / incorrect / absent. Record wrong characters or punctuation explicitly. Wrapping and harmless whitespace changes may be noted separately from actual spelling errors; do not silently normalize wording, numeral changes or Chinese glyph substitutions. + +**Bilingual festival: 14 strings** + +| Position | Exact text | +|---|---| +| English headline | `RIVERLIGHT FILM WEEK` | +| Chinese headline | `河畔电影周` | +| Date | `18–20 SEPTEMBER 2026` | +| Row 1 time | `18:00` | +| Row 1 English | `A Quiet River` | +| Row 1 Chinese | `静静的河` | +| Row 2 time | `19:30` | +| Row 2 English | `Night Market` | +| Row 2 Chinese | `夜市` | +| Row 3 time | `21:00` | +| Row 3 English | `Home Again` | +| Row 3 Chinese | `再次回家` | +| Venue footer | `RIVERSIDE CINEMA / 河畔影院` | +| Admission footer | `FREE ENTRY / 免费入场` | + +**Comic: six bubbles, attribution matters** + +| Panel | Maya | Leo | +|---|---|---| +| 1 | `Where are my keys?` | `Check your pocket.` | +| 2 | `Not here!` | `Under the chair.` | +| 3 | `Found them. Thanks!` | `You're welcome.` | + +**Museum: eight strings:** `MUSEUM VISITOR MAP`, `GALLERY A`, `GALLERY B`, `MAIN HALL`, `CAFE`, `SHOP`, `ENTRANCE`, `SUGGESTED ROUTE`. + +**Product campaign: three strings:** `TAKE THE LONG WAY`, `RIDGELINE / EVERYDAY ADVENTURE`, `750 mL · BUILT TO REUSE`. + +## Editing scenarios + +| Scenario and seed | Fixed teacher reference(s) | Required change | Protected evidence / failure modes | +|---|---|---|---| +| `expanded-station-coat-edit`, 21101 | `expanded-station-reunion-40.png` | Only mother's navy coat becomes burgundy, preserving cut, length, folds and fabric. | Mother identity/face/gaze/pose; yellow scarf; exact handholding; all four people and their hands/clothing; father/grandmother embrace; rabbit; suitcase; train/platform/crop/light. Report recoloring spill, reconstructed fingers/faces, changed pose, unintended wardrobe changes and background redraw. Use original reference as the preservation baseline. | +| `expanded-festival-type-edit`, 21112 | `expanded-bilingual-festival-40.png` | Headline becomes `RIVERLIGHT FILM NIGHTS`; date becomes `25–27 SEPTEMBER 2026`. Headline may resize sensibly within its original area. | Chinese headline, all nine program strings, both footer lines and their alignment; illustration, rules, margins, background/colors. Check every original string that was actually legible. Fail undesired translation, missing rows, font/style redesign or collateral Chinese corruption separately from the two requested replacements. | +| `expanded-multiref-portrait`, 21123 | First `expanded-station-reunion-40.png`; second `expanded-ridgeline-campaign-40.png` | One seated mother from reference 1 holds one bottle from reference 2 in both hands in a waist-up railway-platform portrait. She smiles toward the camera; one green suitcase is partly visible by the bench. | Recognizable mother face, low bun, gold earrings, navy coat/yellow scarf; recognizable bottle shape, black cap, orange stripe/emblem. New pose/background are intentional. Hands must grip plausibly without hiding cap/stripe/emblem. Check coherent scale/light, exactly one person/bottle, no imported other people/cups/pedestal/poster text. Judge identity and product fidelity separately; whole-frame pixel similarity is not the target here. | +| `expanded-product-relocation-edit`, 21134 | `expanded-ridgeline-campaign-40.png` | Remove yellow cup; relocate turquoise cup from right to far left of pedestal with a visible gap. Final products: unchanged bottle plus one turquoise cup, exactly two. Former cup locations should be empty/naturally restored. | Bottle location/scale/shape/cap/stripe/emblem, pedestal, all three text strings, crop/light/background. Check no ghost cup, extra handle, duplicate cup, right-side residual object or typography redraw. Moved cup must contact tabletop with consistent shadow. | + +## Review record + +Use one row per backend/scenario/step count. Add per-string or per-role details when the summary would hide a failure. + +| Backend/checkpoint | Scenario / steps | Adherence findings | Anatomical / physical defects | Drift from teacher | Edit protected-region findings | Source-limited criteria | Overall evidence, no unsupported ranking | +|---|---|---|---|---|---|---|---| +| Pending | — | — | — | — | — | — | — | + +Record visually meaningful candidate improvements as well as regressions. Keep text accuracy, people interactions, diagram logic, editing preservation and image fidelity distinct in the final report. A count of passed criteria may summarize this finite set, but cannot establish general model quality or statistical significance. + +To measure image fidelity after matching teacher/candidate images exist, pass this manifest explicitly to the CPU comparison tool: + +```sh +/tmp/qwen-api-test-py311/bin/python scripts/compare_fidelity_v3.py \ + --jobs experiments/fidelity-v3/expanded-jobs.json \ + --output results/compare-fidelity-v3-expanded.json +``` + +The prior v2 folders have no outputs for these new labels and should remain explicitly missing. For a direct new-candidate comparison, specify only actual candidate folders with repeated `--candidate NAME=FOLDER`; do not mix different available-case cohorts into a ranking. No measurements or image-quality claims have been made in this rubric. diff --git a/reproduction/experiments/fidelity-v3/original-18-jobs.json b/reproduction/experiments/fidelity-v3/original-18-jobs.json new file mode 100644 index 0000000000000000000000000000000000000000..98bc9b51f65bd51f662aa6a38e35b41a06fbd478 --- /dev/null +++ b/reproduction/experiments/fidelity-v3/original-18-jobs.json @@ -0,0 +1,164 @@ +[ + { + "label": "editorial-25", + "prompt": "Candid editorial photograph in a sunlit pottery studio. Exactly two adults: an older woman with short silver hair on the left demonstrates shaping a clay bowl on a pottery wheel; a younger man with curly dark hair on the right watches and holds a folded blue towel with both hands. Their hands are anatomically natural and clearly visible, with clay on the woman's fingertips. Wooden shelves hold irregular ceramic vases, one trailing green plant hangs high in the back right, and soft morning light enters from a large window on the left. Waist-up environmental composition, documentary photography, natural skin texture, restrained warm colors, realistic depth of field. No lettering.", + "steps": 25, + "seed": 118, + "width": 1024, + "height": 1024 + }, + { + "label": "editorial-40", + "prompt": "Candid editorial photograph in a sunlit pottery studio. Exactly two adults: an older woman with short silver hair on the left demonstrates shaping a clay bowl on a pottery wheel; a younger man with curly dark hair on the right watches and holds a folded blue towel with both hands. Their hands are anatomically natural and clearly visible, with clay on the woman's fingertips. Wooden shelves hold irregular ceramic vases, one trailing green plant hangs high in the back right, and soft morning light enters from a large window on the left. Waist-up environmental composition, documentary photography, natural skin texture, restrained warm colors, realistic depth of field. No lettering.", + "steps": 40, + "seed": 118, + "width": 1024, + "height": 1024 + }, + { + "label": "rainy-city-40", + "prompt": "Cinematic street-level architectural photograph of a narrow Tokyo side street at blue hour after rain. A tiny warmly lit ramen restaurant occupies the left foreground, with a striped navy awning and a vertical sign that reads \"RAMEN\". Exactly three bicycles are parked along the right wall. A person holding a transparent umbrella walks away at center, wearing a mustard yellow raincoat. Overhead wires cross between weathered three-story buildings, red lanterns glow farther down the street, and puddles reflect the signs. Strong one-point perspective, realistic glass and wet asphalt, fine distant details, balanced blue and amber lighting.", + "steps": 40, + "seed": 227, + "width": 1024, + "height": 1024 + }, + { + "label": "rainy-city-25", + "prompt": "Cinematic street-level architectural photograph of a narrow Tokyo side street at blue hour after rain. A tiny warmly lit ramen restaurant occupies the left foreground, with a striped navy awning and a vertical sign that reads \"RAMEN\". Exactly three bicycles are parked along the right wall. A person holding a transparent umbrella walks away at center, wearing a mustard yellow raincoat. Overhead wires cross between weathered three-story buildings, red lanterns glow farther down the street, and puddles reflect the signs. Strong one-point perspective, realistic glass and wet asphalt, fine distant details, balanced blue and amber lighting.", + "steps": 25, + "seed": 227, + "width": 1024, + "height": 1024 + }, + { + "label": "botanical-poster-25", + "prompt": "Sophisticated botanical exhibition poster on warm ivory paper, flat front-on view, crisp print design. At the top the large dark green serif headline reads exactly \"THE SECRET GARDEN\" on two centered lines. Under it a smaller line reads exactly \"BOTANICAL STUDIES / 2026\". The center contains a delicate detailed watercolor illustration of a lemon branch with exactly two yellow lemons, white blossoms and dark green leaves, surrounded by generous negative space. Along the bottom three evenly spaced text blocks read \"12 OCTOBER\", \"10 AM - 6 PM\", and \"GLASSHOUSE No. 4\". A thin rectangular green border surrounds the entire design. Elegant typography, no other text, no mockup, no shadows outside the paper.", + "steps": 25, + "seed": 336, + "width": 1024, + "height": 1024 + }, + { + "label": "botanical-poster-40", + "prompt": "Sophisticated botanical exhibition poster on warm ivory paper, flat front-on view, crisp print design. At the top the large dark green serif headline reads exactly \"THE SECRET GARDEN\" on two centered lines. Under it a smaller line reads exactly \"BOTANICAL STUDIES / 2026\". The center contains a delicate detailed watercolor illustration of a lemon branch with exactly two yellow lemons, white blossoms and dark green leaves, surrounded by generous negative space. Along the bottom three evenly spaced text blocks read \"12 OCTOBER\", \"10 AM - 6 PM\", and \"GLASSHOUSE No. 4\". A thin rectangular green border surrounds the entire design. Elegant typography, no other text, no mockup, no shadows outside the paper.", + "steps": 40, + "seed": 336, + "width": 1024, + "height": 1024 + }, + { + "label": "food-spatial-40", + "prompt": "Overhead high-end food photograph of a neatly arranged breakfast on a pale stone table. A large round white plate is centered. On the plate, exactly three pancakes form a vertical stack, topped by exactly five blueberries and a small square of butter. To the left of the plate is a silver fork, to the right is a silver knife with its blade facing the plate. A clear glass of orange juice sits at the upper right, a small white bowl containing sliced strawberries sits at the upper left, and a folded rust-colored linen napkin lies beneath the fork. A little maple syrup runs down only the right edge of the pancake stack. Soft natural light from the upper left, realistic food texture, all objects fully inside frame.", + "steps": 40, + "seed": 445, + "width": 1024, + "height": 1024 + }, + { + "label": "food-spatial-25", + "prompt": "Overhead high-end food photograph of a neatly arranged breakfast on a pale stone table. A large round white plate is centered. On the plate, exactly three pancakes form a vertical stack, topped by exactly five blueberries and a small square of butter. To the left of the plate is a silver fork, to the right is a silver knife with its blade facing the plate. A clear glass of orange juice sits at the upper right, a small white bowl containing sliced strawberries sits at the upper left, and a folded rust-colored linen napkin lies beneath the fork. A little maple syrup runs down only the right edge of the pancake stack. Soft natural light from the upper left, realistic food texture, all objects fully inside frame.", + "steps": 25, + "seed": 445, + "width": 1024, + "height": 1024 + }, + { + "label": "fantasy-cutaway-25", + "prompt": "Intricate isometric cutaway illustration of a cozy three-level library built inside an enormous ancient tree, in a hand-painted storybook style. Ground floor: a round green entrance door, a circular rug, and a sleeping orange cat beside a small wood stove. Middle floor: curved bookshelves filled with colorful books, a writing desk by an oval window, and a brass telescope pointing outside. Top floor: a glass-domed reading nook with two red armchairs and a hanging lantern. A single wooden spiral staircase visibly connects all three floors. Exposed roots wrap around mossy rocks, tiny mushrooms cluster near the door, and the leafy crown surrounds the dome without hiding it. Consistent isometric perspective, warm interior lighting, cool twilight forest background, detailed but readable, no text.", + "steps": 25, + "seed": 554, + "width": 1024, + "height": 1024 + }, + { + "label": "fantasy-cutaway-40", + "prompt": "Intricate isometric cutaway illustration of a cozy three-level library built inside an enormous ancient tree, in a hand-painted storybook style. Ground floor: a round green entrance door, a circular rug, and a sleeping orange cat beside a small wood stove. Middle floor: curved bookshelves filled with colorful books, a writing desk by an oval window, and a brass telescope pointing outside. Top floor: a glass-domed reading nook with two red armchairs and a hanging lantern. A single wooden spiral staircase visibly connects all three floors. Exposed roots wrap around mossy rocks, tiny mushrooms cluster near the door, and the leafy crown surrounds the dome without hiding it. Consistent isometric perspective, warm interior lighting, cool twilight forest background, detailed but readable, no text.", + "steps": 40, + "seed": 554, + "width": 1024, + "height": 1024 + }, + { + "label": "glass-product-40", + "prompt": "Luxury studio still life photograph on a polished dark green marble plinth. A rectangular clear glass perfume bottle half filled with pale amber liquid stands in the center, with a brushed gold cylindrical cap and a small cream label reading exactly \"LUMEN\". A thin curved strip of orange peel lies in front of the bottle. Behind and to the left is a translucent ribbed glass sphere; behind and to the right are two glossy dark green leaves. Hard afternoon sunlight from the upper left creates realistic caustics, transparent overlapping shadows, precise refraction through the bottle and a soft reflection on the marble. Deep emerald background, controlled commercial composition, realistic materials, no additional text.", + "steps": 40, + "seed": 663, + "width": 1024, + "height": 1024 + }, + { + "label": "glass-product-25", + "prompt": "Luxury studio still life photograph on a polished dark green marble plinth. A rectangular clear glass perfume bottle half filled with pale amber liquid stands in the center, with a brushed gold cylindrical cap and a small cream label reading exactly \"LUMEN\". A thin curved strip of orange peel lies in front of the bottle. Behind and to the left is a translucent ribbed glass sphere; behind and to the right are two glossy dark green leaves. Hard afternoon sunlight from the upper left creates realistic caustics, transparent overlapping shadows, precise refraction through the bottle and a soft reflection on the marble. Deep emerald background, controlled commercial composition, realistic materials, no additional text.", + "steps": 25, + "seed": 663, + "width": 1024, + "height": 1024 + }, + { + "label": "editorial-edit-25", + "prompt": "Replace only the blue towel held by the younger man with a bright red towel of the same size and folded shape. Preserve both people's identities, facial expressions, hands, clothing, clay bowl, studio shelves, window light, framing and photographic style.", + "steps": 25, + "seed": 774, + "width": 1024, + "height": 1024, + "images": [ + "/poc/samples/varied-v2/nf4/editorial-40.png" + ] + }, + { + "label": "editorial-edit-40", + "prompt": "Replace only the blue towel held by the younger man with a bright red towel of the same size and folded shape. Preserve both people's identities, facial expressions, hands, clothing, clay bowl, studio shelves, window light, framing and photographic style.", + "steps": 40, + "seed": 774, + "width": 1024, + "height": 1024, + "images": [ + "/poc/samples/varied-v2/nf4/editorial-40.png" + ] + }, + { + "label": "poster-edit-25", + "prompt": "Edit only the text at the top: replace \"THE SECRET GARDEN\" with \"THE LEMON HOUSE\". Keep the same dark green serif type style and centered placement. Preserve all other text exactly, the lemon branch illustration, paper background, border, spacing and overall design.", + "steps": 25, + "seed": 885, + "width": 1024, + "height": 1024, + "images": [ + "/poc/samples/varied-v2/nf4/botanical-poster-40.png" + ] + }, + { + "label": "poster-edit-40", + "prompt": "Edit only the text at the top: replace \"THE SECRET GARDEN\" with \"THE LEMON HOUSE\". Keep the same dark green serif type style and centered placement. Preserve all other text exactly, the lemon branch illustration, paper background, border, spacing and overall design.", + "steps": 40, + "seed": 885, + "width": 1024, + "height": 1024, + "images": [ + "/poc/samples/varied-v2/nf4/botanical-poster-40.png" + ] + }, + { + "label": "city-edit-25", + "prompt": "Transform this rainy Tokyo street photograph into a richly textured hand-painted watercolor illustration. Preserve the layout of the buildings, the ramen restaurant on the left, all three bicycles on the right, the central person with the transparent umbrella and yellow raincoat, the overhead wires, blue-hour lighting, and puddle reflections. Keep the RAMEN sign legible.", + "steps": 25, + "seed": 996, + "width": 1024, + "height": 1024, + "images": [ + "/poc/samples/varied-v2/nf4/rainy-city-40.png" + ] + }, + { + "label": "city-edit-40", + "prompt": "Transform this rainy Tokyo street photograph into a richly textured hand-painted watercolor illustration. Preserve the layout of the buildings, the ramen restaurant on the left, all three bicycles on the right, the central person with the transparent umbrella and yellow raincoat, the overhead wires, blue-hour lighting, and puddle reflections. Keep the RAMEN sign legible.", + "steps": 40, + "seed": 996, + "width": 1024, + "height": 1024, + "images": [ + "/poc/samples/varied-v2/nf4/rainy-city-40.png" + ] + } +] diff --git a/reproduction/experiments/fidelity-v3/performance.md b/reproduction/experiments/fidelity-v3/performance.md new file mode 100644 index 0000000000000000000000000000000000000000..7ac71fbc232fec271e341e41e85a2a5b4ea06d7f --- /dev/null +++ b/reproduction/experiments/fidelity-v3/performance.md @@ -0,0 +1,30 @@ +# Fidelity experiment performance + +Serial engine inference, excludes image-file loading/saving and process startup; no prompt/reference LRU reuse. Per-request prefix KV remains enabled. + +Original BF16 DiT; NF4 text/vision encoder and BF16 untiled VAE fixed across candidates. + +| Backend | Reference images | Steps | Images | Median inference (s) | Median transformer (s) | Max allocated (MiB) | Sampled board max (MiB) | +|---|---:|---:|---:|---:|---:|---:|---:| +| bf16-dit | 0 | 25 | 14 | 27.51 | 25.31 | 7387 | 9346 | +| bf16-dit | 0 | 40 | 14 | 42.76 | 40.54 | 7387 | 9348 | +| bf16-dit | 1 | 25 | 6 | 31.98 | 29.01 | 7391 | 12030 | +| bf16-dit | 1 | 40 | 6 | 48.70 | 45.74 | 7391 | 12030 | +| bf16-dit | 2 | 25 | 1 | 36.79 | 33.14 | 7398 | 15120 | +| bf16-dit | 2 | 40 | 1 | 55.28 | 51.62 | 7400 | 14826 | +| nunchaku-v3-mlpproj512 | 0 | 25 | 14 | 16.06 | 13.21 | 7388 | 9344 | +| nunchaku-v3-mlpproj512 | 0 | 40 | 14 | 23.56 | 20.68 | 7388 | 9346 | +| nunchaku-v3-mlpproj512 | 1 | 25 | 6 | 20.02 | 16.49 | 8242 | 9488 | +| nunchaku-v3-mlpproj512 | 1 | 40 | 6 | 29.23 | 25.72 | 8242 | 9488 | +| nunchaku-v3-mlpproj512 | 2 | 25 | 1 | 24.41 | 20.08 | 10911 | 12374 | +| nunchaku-v3-mlpproj512 | 2 | 40 | 1 | 35.35 | 31.04 | 10911 | 12374 | +| nunchaku-v3-r128 | 0 | 25 | 14 | 15.50 | 12.71 | 7388 | 9344 | +| nunchaku-v3-r128 | 0 | 40 | 14 | 22.82 | 19.93 | 7388 | 9629 | +| nunchaku-v3-r128 | 1 | 25 | 6 | 19.48 | 16.04 | 7862 | 9346 | +| nunchaku-v3-r128 | 1 | 40 | 6 | 28.43 | 24.92 | 7862 | 9344 | +| nunchaku-v3-r128 | 2 | 25 | 1 | 23.76 | 19.54 | 10531 | 11994 | +| nunchaku-v3-r128 | 2 | 40 | 1 | 34.35 | 30.19 | 10531 | 11994 | + +Complete paired run: True. See the JSON report for missing jobs, validation errors, exact settings, timestamps and image/reference hashes. + +These timings describe this fixed 1024-pixel suite on physical GPU 0. The original v2 timings used different LRU-cache states and must not be compared as controlled end-to-end speed measurements. Model-state byte costs do not substitute for measured peak VRAM. diff --git a/reproduction/experiments/fidelity-v3/quality-jobs-all.json b/reproduction/experiments/fidelity-v3/quality-jobs-all.json new file mode 100644 index 0000000000000000000000000000000000000000..474610f792657975273ed282ab82437ceb4e9d8e --- /dev/null +++ b/reproduction/experiments/fidelity-v3/quality-jobs-all.json @@ -0,0 +1,382 @@ +[ + { + "label": "editorial-25", + "prompt": "Candid editorial photograph in a sunlit pottery studio. Exactly two adults: an older woman with short silver hair on the left demonstrates shaping a clay bowl on a pottery wheel; a younger man with curly dark hair on the right watches and holds a folded blue towel with both hands. Their hands are anatomically natural and clearly visible, with clay on the woman's fingertips. Wooden shelves hold irregular ceramic vases, one trailing green plant hangs high in the back right, and soft morning light enters from a large window on the left. Waist-up environmental composition, documentary photography, natural skin texture, restrained warm colors, realistic depth of field. No lettering.", + "steps": 25, + "seed": 118, + "width": 1024, + "height": 1024 + }, + { + "label": "editorial-40", + "prompt": "Candid editorial photograph in a sunlit pottery studio. Exactly two adults: an older woman with short silver hair on the left demonstrates shaping a clay bowl on a pottery wheel; a younger man with curly dark hair on the right watches and holds a folded blue towel with both hands. Their hands are anatomically natural and clearly visible, with clay on the woman's fingertips. Wooden shelves hold irregular ceramic vases, one trailing green plant hangs high in the back right, and soft morning light enters from a large window on the left. Waist-up environmental composition, documentary photography, natural skin texture, restrained warm colors, realistic depth of field. No lettering.", + "steps": 40, + "seed": 118, + "width": 1024, + "height": 1024 + }, + { + "label": "rainy-city-40", + "prompt": "Cinematic street-level architectural photograph of a narrow Tokyo side street at blue hour after rain. A tiny warmly lit ramen restaurant occupies the left foreground, with a striped navy awning and a vertical sign that reads \"RAMEN\". Exactly three bicycles are parked along the right wall. A person holding a transparent umbrella walks away at center, wearing a mustard yellow raincoat. Overhead wires cross between weathered three-story buildings, red lanterns glow farther down the street, and puddles reflect the signs. Strong one-point perspective, realistic glass and wet asphalt, fine distant details, balanced blue and amber lighting.", + "steps": 40, + "seed": 227, + "width": 1024, + "height": 1024 + }, + { + "label": "rainy-city-25", + "prompt": "Cinematic street-level architectural photograph of a narrow Tokyo side street at blue hour after rain. A tiny warmly lit ramen restaurant occupies the left foreground, with a striped navy awning and a vertical sign that reads \"RAMEN\". Exactly three bicycles are parked along the right wall. A person holding a transparent umbrella walks away at center, wearing a mustard yellow raincoat. Overhead wires cross between weathered three-story buildings, red lanterns glow farther down the street, and puddles reflect the signs. Strong one-point perspective, realistic glass and wet asphalt, fine distant details, balanced blue and amber lighting.", + "steps": 25, + "seed": 227, + "width": 1024, + "height": 1024 + }, + { + "label": "botanical-poster-25", + "prompt": "Sophisticated botanical exhibition poster on warm ivory paper, flat front-on view, crisp print design. At the top the large dark green serif headline reads exactly \"THE SECRET GARDEN\" on two centered lines. Under it a smaller line reads exactly \"BOTANICAL STUDIES / 2026\". The center contains a delicate detailed watercolor illustration of a lemon branch with exactly two yellow lemons, white blossoms and dark green leaves, surrounded by generous negative space. Along the bottom three evenly spaced text blocks read \"12 OCTOBER\", \"10 AM - 6 PM\", and \"GLASSHOUSE No. 4\". A thin rectangular green border surrounds the entire design. Elegant typography, no other text, no mockup, no shadows outside the paper.", + "steps": 25, + "seed": 336, + "width": 1024, + "height": 1024 + }, + { + "label": "botanical-poster-40", + "prompt": "Sophisticated botanical exhibition poster on warm ivory paper, flat front-on view, crisp print design. At the top the large dark green serif headline reads exactly \"THE SECRET GARDEN\" on two centered lines. Under it a smaller line reads exactly \"BOTANICAL STUDIES / 2026\". The center contains a delicate detailed watercolor illustration of a lemon branch with exactly two yellow lemons, white blossoms and dark green leaves, surrounded by generous negative space. Along the bottom three evenly spaced text blocks read \"12 OCTOBER\", \"10 AM - 6 PM\", and \"GLASSHOUSE No. 4\". A thin rectangular green border surrounds the entire design. Elegant typography, no other text, no mockup, no shadows outside the paper.", + "steps": 40, + "seed": 336, + "width": 1024, + "height": 1024 + }, + { + "label": "food-spatial-40", + "prompt": "Overhead high-end food photograph of a neatly arranged breakfast on a pale stone table. A large round white plate is centered. On the plate, exactly three pancakes form a vertical stack, topped by exactly five blueberries and a small square of butter. To the left of the plate is a silver fork, to the right is a silver knife with its blade facing the plate. A clear glass of orange juice sits at the upper right, a small white bowl containing sliced strawberries sits at the upper left, and a folded rust-colored linen napkin lies beneath the fork. A little maple syrup runs down only the right edge of the pancake stack. Soft natural light from the upper left, realistic food texture, all objects fully inside frame.", + "steps": 40, + "seed": 445, + "width": 1024, + "height": 1024 + }, + { + "label": "food-spatial-25", + "prompt": "Overhead high-end food photograph of a neatly arranged breakfast on a pale stone table. A large round white plate is centered. On the plate, exactly three pancakes form a vertical stack, topped by exactly five blueberries and a small square of butter. To the left of the plate is a silver fork, to the right is a silver knife with its blade facing the plate. A clear glass of orange juice sits at the upper right, a small white bowl containing sliced strawberries sits at the upper left, and a folded rust-colored linen napkin lies beneath the fork. A little maple syrup runs down only the right edge of the pancake stack. Soft natural light from the upper left, realistic food texture, all objects fully inside frame.", + "steps": 25, + "seed": 445, + "width": 1024, + "height": 1024 + }, + { + "label": "fantasy-cutaway-25", + "prompt": "Intricate isometric cutaway illustration of a cozy three-level library built inside an enormous ancient tree, in a hand-painted storybook style. Ground floor: a round green entrance door, a circular rug, and a sleeping orange cat beside a small wood stove. Middle floor: curved bookshelves filled with colorful books, a writing desk by an oval window, and a brass telescope pointing outside. Top floor: a glass-domed reading nook with two red armchairs and a hanging lantern. A single wooden spiral staircase visibly connects all three floors. Exposed roots wrap around mossy rocks, tiny mushrooms cluster near the door, and the leafy crown surrounds the dome without hiding it. Consistent isometric perspective, warm interior lighting, cool twilight forest background, detailed but readable, no text.", + "steps": 25, + "seed": 554, + "width": 1024, + "height": 1024 + }, + { + "label": "fantasy-cutaway-40", + "prompt": "Intricate isometric cutaway illustration of a cozy three-level library built inside an enormous ancient tree, in a hand-painted storybook style. Ground floor: a round green entrance door, a circular rug, and a sleeping orange cat beside a small wood stove. Middle floor: curved bookshelves filled with colorful books, a writing desk by an oval window, and a brass telescope pointing outside. Top floor: a glass-domed reading nook with two red armchairs and a hanging lantern. A single wooden spiral staircase visibly connects all three floors. Exposed roots wrap around mossy rocks, tiny mushrooms cluster near the door, and the leafy crown surrounds the dome without hiding it. Consistent isometric perspective, warm interior lighting, cool twilight forest background, detailed but readable, no text.", + "steps": 40, + "seed": 554, + "width": 1024, + "height": 1024 + }, + { + "label": "glass-product-40", + "prompt": "Luxury studio still life photograph on a polished dark green marble plinth. A rectangular clear glass perfume bottle half filled with pale amber liquid stands in the center, with a brushed gold cylindrical cap and a small cream label reading exactly \"LUMEN\". A thin curved strip of orange peel lies in front of the bottle. Behind and to the left is a translucent ribbed glass sphere; behind and to the right are two glossy dark green leaves. Hard afternoon sunlight from the upper left creates realistic caustics, transparent overlapping shadows, precise refraction through the bottle and a soft reflection on the marble. Deep emerald background, controlled commercial composition, realistic materials, no additional text.", + "steps": 40, + "seed": 663, + "width": 1024, + "height": 1024 + }, + { + "label": "glass-product-25", + "prompt": "Luxury studio still life photograph on a polished dark green marble plinth. A rectangular clear glass perfume bottle half filled with pale amber liquid stands in the center, with a brushed gold cylindrical cap and a small cream label reading exactly \"LUMEN\". A thin curved strip of orange peel lies in front of the bottle. Behind and to the left is a translucent ribbed glass sphere; behind and to the right are two glossy dark green leaves. Hard afternoon sunlight from the upper left creates realistic caustics, transparent overlapping shadows, precise refraction through the bottle and a soft reflection on the marble. Deep emerald background, controlled commercial composition, realistic materials, no additional text.", + "steps": 25, + "seed": 663, + "width": 1024, + "height": 1024 + }, + { + "label": "editorial-edit-25", + "prompt": "Replace only the blue towel held by the younger man with a bright red towel of the same size and folded shape. Preserve both people's identities, facial expressions, hands, clothing, clay bowl, studio shelves, window light, framing and photographic style.", + "steps": 25, + "seed": 774, + "width": 1024, + "height": 1024, + "images": [ + "/poc/samples/varied-v2/nf4/editorial-40.png" + ] + }, + { + "label": "editorial-edit-40", + "prompt": "Replace only the blue towel held by the younger man with a bright red towel of the same size and folded shape. Preserve both people's identities, facial expressions, hands, clothing, clay bowl, studio shelves, window light, framing and photographic style.", + "steps": 40, + "seed": 774, + "width": 1024, + "height": 1024, + "images": [ + "/poc/samples/varied-v2/nf4/editorial-40.png" + ] + }, + { + "label": "poster-edit-25", + "prompt": "Edit only the text at the top: replace \"THE SECRET GARDEN\" with \"THE LEMON HOUSE\". Keep the same dark green serif type style and centered placement. Preserve all other text exactly, the lemon branch illustration, paper background, border, spacing and overall design.", + "steps": 25, + "seed": 885, + "width": 1024, + "height": 1024, + "images": [ + "/poc/samples/varied-v2/nf4/botanical-poster-40.png" + ] + }, + { + "label": "poster-edit-40", + "prompt": "Edit only the text at the top: replace \"THE SECRET GARDEN\" with \"THE LEMON HOUSE\". Keep the same dark green serif type style and centered placement. Preserve all other text exactly, the lemon branch illustration, paper background, border, spacing and overall design.", + "steps": 40, + "seed": 885, + "width": 1024, + "height": 1024, + "images": [ + "/poc/samples/varied-v2/nf4/botanical-poster-40.png" + ] + }, + { + "label": "city-edit-25", + "prompt": "Transform this rainy Tokyo street photograph into a richly textured hand-painted watercolor illustration. Preserve the layout of the buildings, the ramen restaurant on the left, all three bicycles on the right, the central person with the transparent umbrella and yellow raincoat, the overhead wires, blue-hour lighting, and puddle reflections. Keep the RAMEN sign legible.", + "steps": 25, + "seed": 996, + "width": 1024, + "height": 1024, + "images": [ + "/poc/samples/varied-v2/nf4/rainy-city-40.png" + ] + }, + { + "label": "city-edit-40", + "prompt": "Transform this rainy Tokyo street photograph into a richly textured hand-painted watercolor illustration. Preserve the layout of the buildings, the ramen restaurant on the left, all three bicycles on the right, the central person with the transparent umbrella and yellow raincoat, the overhead wires, blue-hour lighting, and puddle reflections. Keep the RAMEN sign legible.", + "steps": 40, + "seed": 996, + "width": 1024, + "height": 1024, + "images": [ + "/poc/samples/varied-v2/nf4/rainy-city-40.png" + ] + }, + { + "label": "expanded-community-kitchen-25", + "prompt": "Documentary photograph for a community cooking program, wide eye-level composition in a bright, practical teaching kitchen. Exactly six adults, all visible from at least the waist up. At the left counter, a woman with short black hair in a green apron chops carrots on a wooden board, with one hand holding the knife and the other safely curled on the carrot. At center, a silver-haired man in a white apron holds a ladle over a large soup pot. Immediately to his right, a woman in a mustard cardigan passes a rectangular tray of four bread rolls to a younger man in a navy shirt, who receives it with both hands; their hands meet opposite edges of the same tray. At the far right, a man in a red apron rinses a colander of leafy greens at the sink. Behind the central counter, a woman wearing glasses and a purple sweater reads a paper recipe card. Natural varied expressions, clear individual faces, plausible arms and hands attached to their owners, no extra people, realistic morning window light and stainless-steel reflections. No readable signage. The image should feel like a real coordinated cooking session, not a posed group portrait.", + "steps": 25, + "seed": 21011, + "width": 1024, + "height": 1024 + }, + { + "label": "expanded-community-kitchen-40", + "prompt": "Documentary photograph for a community cooking program, wide eye-level composition in a bright, practical teaching kitchen. Exactly six adults, all visible from at least the waist up. At the left counter, a woman with short black hair in a green apron chops carrots on a wooden board, with one hand holding the knife and the other safely curled on the carrot. At center, a silver-haired man in a white apron holds a ladle over a large soup pot. Immediately to his right, a woman in a mustard cardigan passes a rectangular tray of four bread rolls to a younger man in a navy shirt, who receives it with both hands; their hands meet opposite edges of the same tray. At the far right, a man in a red apron rinses a colander of leafy greens at the sink. Behind the central counter, a woman wearing glasses and a purple sweater reads a paper recipe card. Natural varied expressions, clear individual faces, plausible arms and hands attached to their owners, no extra people, realistic morning window light and stainless-steel reflections. No readable signage. The image should feel like a real coordinated cooking session, not a posed group portrait.", + "steps": 40, + "seed": 21011, + "width": 1024, + "height": 1024 + }, + { + "label": "expanded-chess-awards-25", + "prompt": "Editorial event photograph of a small chess championship award presentation on a low indoor stage. Exactly five adults. At center, a smiling woman with a short dark bob sits in a clearly recognizable manual wheelchair; she and the presenter standing immediately to her left jointly hold one gold trophy between them. The presenter is an older man wearing a charcoal suit. Immediately to the winner's right, the runner-up, a young man in a cream sweater, wears one silver medal and applauds. At the far left, a female photographer kneels with a camera held to her eye and aims toward the winner. At the far right, a host in a burgundy suit holds a microphone and a small cue card. All five people fit fully in frame, with believable hands, wheelchair wheels and footrests. Behind them, a clean white stage banner reads exactly \"CITY CHESS FINAL\" with \"2026\" centered on the line below. A small chess table with a board stands at the rear left. Warm indoor event lighting, realistic fabric and skin, no audience or additional figures.", + "steps": 25, + "seed": 21022, + "width": 1024, + "height": 1024 + }, + { + "label": "expanded-chess-awards-40", + "prompt": "Editorial event photograph of a small chess championship award presentation on a low indoor stage. Exactly five adults. At center, a smiling woman with a short dark bob sits in a clearly recognizable manual wheelchair; she and the presenter standing immediately to her left jointly hold one gold trophy between them. The presenter is an older man wearing a charcoal suit. Immediately to the winner's right, the runner-up, a young man in a cream sweater, wears one silver medal and applauds. At the far left, a female photographer kneels with a camera held to her eye and aims toward the winner. At the far right, a host in a burgundy suit holds a microphone and a small cue card. All five people fit fully in frame, with believable hands, wheelchair wheels and footrests. Behind them, a clean white stage banner reads exactly \"CITY CHESS FINAL\" with \"2026\" centered on the line below. A small chess table with a board stands at the rear left. Warm indoor event lighting, realistic fabric and skin, no audience or additional figures.", + "steps": 40, + "seed": 21022, + "width": 1024, + "height": 1024 + }, + { + "label": "expanded-station-reunion-25", + "prompt": "Natural editorial photograph of a family reunion beside a stationary passenger train on a quiet outdoor railway platform. Exactly four people, full bodies in frame, no other passengers. On the left is a woman in her thirties with an oval face, dark curly hair tied in a low bun, small round gold earrings, a navy knee-length coat and a mustard yellow scarf. She holds the left hand of a young girl standing at center; the girl wears a red raincoat and carries a small stuffed rabbit in her free hand. On the right, a man with close-cropped brown hair, a light gray jacket and a red backpack bends slightly to embrace an older gray-haired woman wearing a teal coat; the older woman faces him with one arm around his shoulder. The mother and girl watch them with happy, understated expressions. One upright green suitcase stands beside the mother's left leg. Anatomically plausible hands and embraces with clear ownership of each arm, natural eye contact, soft overcast daylight, authentic train windows and platform paving. Do not crop feet or add duplicate people. No readable logos or signs.", + "steps": 25, + "seed": 21033, + "width": 1024, + "height": 1024 + }, + { + "label": "expanded-station-reunion-40", + "prompt": "Natural editorial photograph of a family reunion beside a stationary passenger train on a quiet outdoor railway platform. Exactly four people, full bodies in frame, no other passengers. On the left is a woman in her thirties with an oval face, dark curly hair tied in a low bun, small round gold earrings, a navy knee-length coat and a mustard yellow scarf. She holds the left hand of a young girl standing at center; the girl wears a red raincoat and carries a small stuffed rabbit in her free hand. On the right, a man with close-cropped brown hair, a light gray jacket and a red backpack bends slightly to embrace an older gray-haired woman wearing a teal coat; the older woman faces him with one arm around his shoulder. The mother and girl watch them with happy, understated expressions. One upright green suitcase stands beside the mother's left leg. Anatomically plausible hands and embraces with clear ownership of each arm, natural eye contact, soft overcast daylight, authentic train windows and platform paving. Do not crop feet or add duplicate people. No readable logos or signs.", + "steps": 40, + "seed": 21033, + "width": 1024, + "height": 1024 + }, + { + "label": "expanded-bilingual-festival-25", + "prompt": "A polished bilingual neighborhood film festival poster, square format, flat front-on print artwork with generous margins, warm cream paper and deep navy typography accented by vermilion. The large centered English title at the top reads exactly \"RIVERLIGHT FILM WEEK\". Directly beneath it, an equally clear Chinese title reads exactly \"河畔电影周\". The next centered line reads exactly \"18–20 SEPTEMBER 2026\". Below a thin vermilion rule, a tidy two-column program occupies the middle: the left column contains times, and the right column contains one English film title with its Chinese title directly underneath. Row one: \"18:00\", \"A Quiet River\", \"静静的河\". Row two: \"19:30\", \"Night Market\", \"夜市\". Row three: \"21:00\", \"Home Again\", \"再次回家\". Maintain aligned columns and clearly separated rows. At the bottom, centered on separate lines, print exactly \"RIVERSIDE CINEMA / 河畔影院\" and \"FREE ENTRY / 免费入场\". A small restrained illustration of a red bridge over two blue river lines sits between the program and footer without touching any letters. Sharp readable Latin and Chinese lettering, consistent type hierarchy, no additional text, no mockup perspective.", + "steps": 25, + "seed": 21044, + "width": 1024, + "height": 1024 + }, + { + "label": "expanded-bilingual-festival-40", + "prompt": "A polished bilingual neighborhood film festival poster, square format, flat front-on print artwork with generous margins, warm cream paper and deep navy typography accented by vermilion. The large centered English title at the top reads exactly \"RIVERLIGHT FILM WEEK\". Directly beneath it, an equally clear Chinese title reads exactly \"河畔电影周\". The next centered line reads exactly \"18–20 SEPTEMBER 2026\". Below a thin vermilion rule, a tidy two-column program occupies the middle: the left column contains times, and the right column contains one English film title with its Chinese title directly underneath. Row one: \"18:00\", \"A Quiet River\", \"静静的河\". Row two: \"19:30\", \"Night Market\", \"夜市\". Row three: \"21:00\", \"Home Again\", \"再次回家\". Maintain aligned columns and clearly separated rows. At the bottom, centered on separate lines, print exactly \"RIVERSIDE CINEMA / 河畔影院\" and \"FREE ENTRY / 免费入场\". A small restrained illustration of a red bridge over two blue river lines sits between the program and footer without touching any letters. Sharp readable Latin and Chinese lettering, consistent type hierarchy, no additional text, no mockup perspective.", + "steps": 40, + "seed": 21044, + "width": 1024, + "height": 1024 + }, + { + "label": "expanded-library-comic-25", + "prompt": "A polished, friendly three-panel comic strip in a square canvas, with exactly three equal vertical panels arranged left to right, clear white gutters and a thin dark border around each panel. Keep the same two adult characters and library reading-room setting in all panels: Maya has a chin-length black bob, round glasses and a red sweater; Leo has short curly brown hair and a blue button-up shirt. Each panel contains exactly two white speech bubbles, with clear tails pointing to the correct speaker. Panel 1: Maya stands beside a wooden reading chair, looks worried and searches her bag. Maya says exactly \"Where are my keys?\" Leo, standing to her right, says exactly \"Check your pocket.\" Panel 2: Maya turns out an empty coat pocket and says exactly \"Not here!\" Leo points down at a small key ring visible under that same chair and says exactly \"Under the chair.\" Panel 3: Maya stands upright holding that key ring up, smiles and says exactly \"Found them. Thanks!\" Leo smiles and says exactly \"You're welcome.\" Crisp ink lines, restrained warm colors, expressive natural poses, consistent faces and clothing, believable hands, a coherent sequence with the chair in the same room. Readable dialogue with no narration captions and no additional panels or characters.", + "steps": 25, + "seed": 21055, + "width": 1024, + "height": 1024 + }, + { + "label": "expanded-library-comic-40", + "prompt": "A polished, friendly three-panel comic strip in a square canvas, with exactly three equal vertical panels arranged left to right, clear white gutters and a thin dark border around each panel. Keep the same two adult characters and library reading-room setting in all panels: Maya has a chin-length black bob, round glasses and a red sweater; Leo has short curly brown hair and a blue button-up shirt. Each panel contains exactly two white speech bubbles, with clear tails pointing to the correct speaker. Panel 1: Maya stands beside a wooden reading chair, looks worried and searches her bag. Maya says exactly \"Where are my keys?\" Leo, standing to her right, says exactly \"Check your pocket.\" Panel 2: Maya turns out an empty coat pocket and says exactly \"Not here!\" Leo points down at a small key ring visible under that same chair and says exactly \"Under the chair.\" Panel 3: Maya stands upright holding that key ring up, smiles and says exactly \"Found them. Thanks!\" Leo smiles and says exactly \"You're welcome.\" Crisp ink lines, restrained warm colors, expressive natural poses, consistent faces and clothing, believable hands, a coherent sequence with the chair in the same room. Readable dialogue with no narration captions and no additional panels or characters.", + "steps": 40, + "seed": 21055, + "width": 1024, + "height": 1024 + }, + { + "label": "expanded-museum-plan-25", + "prompt": "A clear visitor floor-plan diagram for a small museum, square format, flat orthographic top-down vector design on white, with thick charcoal exterior walls, thinner interior walls and clearly visible doorway gaps. Title centered above the plan: \"MUSEUM VISITOR MAP\". Inside one rectangular building, arrange five labeled spaces: \"GALLERY A\" in a large upper-left room, \"GALLERY B\" in a large upper-right room, \"MAIN HALL\" in a wide horizontal central space spanning the building, \"CAFE\" in the lower-left room, and \"SHOP\" in the lower-right room. A narrow unlabeled central entrance aisle runs from a doorway at the bottom edge into the main hall, separating the cafe and shop. Label that bottom doorway \"ENTRANCE\" just outside the building. Each of the four corner rooms has one doorway directly into the main hall. Give Gallery A a pale blue fill, Gallery B pale green, the cafe pale orange and the shop pale violet; keep the hall and entrance aisle white. Put exactly two simple rectangular display plinth symbols in each gallery, exactly three round table symbols in the cafe, and exactly two parallel shelf symbols in the shop. A single continuous blue visitor-route line with arrowheads begins outside the entrance, goes through the entrance aisle and main hall, and ends inside Gallery A, passing through actual doorway openings without crossing any walls. Small bottom legend: a blue arrow icon followed by exactly \"SUGGESTED ROUTE\". Spacious legible labeling, precise geometry, no perspective, no people and no extra rooms.", + "steps": 25, + "seed": 21066, + "width": 1024, + "height": 1024 + }, + { + "label": "expanded-museum-plan-40", + "prompt": "A clear visitor floor-plan diagram for a small museum, square format, flat orthographic top-down vector design on white, with thick charcoal exterior walls, thinner interior walls and clearly visible doorway gaps. Title centered above the plan: \"MUSEUM VISITOR MAP\". Inside one rectangular building, arrange five labeled spaces: \"GALLERY A\" in a large upper-left room, \"GALLERY B\" in a large upper-right room, \"MAIN HALL\" in a wide horizontal central space spanning the building, \"CAFE\" in the lower-left room, and \"SHOP\" in the lower-right room. A narrow unlabeled central entrance aisle runs from a doorway at the bottom edge into the main hall, separating the cafe and shop. Label that bottom doorway \"ENTRANCE\" just outside the building. Each of the four corner rooms has one doorway directly into the main hall. Give Gallery A a pale blue fill, Gallery B pale green, the cafe pale orange and the shop pale violet; keep the hall and entrance aisle white. Put exactly two simple rectangular display plinth symbols in each gallery, exactly three round table symbols in the cafe, and exactly two parallel shelf symbols in the shop. A single continuous blue visitor-route line with arrowheads begins outside the entrance, goes through the entrance aisle and main hall, and ends inside Gallery A, passing through actual doorway openings without crossing any walls. Small bottom legend: a blue arrow icon followed by exactly \"SUGGESTED ROUTE\". Spacious legible labeling, precise geometry, no perspective, no people and no extra rooms.", + "steps": 40, + "seed": 21066, + "width": 1024, + "height": 1024 + }, + { + "label": "expanded-ridgeline-campaign-25", + "prompt": "Premium studio product campaign for an insulated water bottle, square print advertisement. Center one tall matte cobalt-blue metal bottle on a low pale stone rectangular pedestal. The bottle has a black screw cap, one narrow orange vertical stripe on its front, and a small white mountain-triangle emblem near its base. On the tabletop to the left of the pedestal is one short yellow enamel cup; to the right is one short turquoise enamel cup. There are exactly three products in total: the bottle and the two cups, all fully visible and separate. Clean warm gray seamless backdrop, soft directional daylight from the upper left, realistic brushed metal and enamel, restrained shadows. At the top, large elegant dark navy sans-serif lettering reads exactly \"TAKE THE LONG WAY\". Directly underneath it, smaller lettering reads exactly \"RIDGELINE / EVERYDAY ADVENTURE\". At the bottom center, small text reads exactly \"750 mL · BUILT TO REUSE\". Plenty of negative space around the type, crisp product edges, no additional props, hands or extra words. Front three-quarter product view with the orange stripe and mountain emblem clearly visible.", + "steps": 25, + "seed": 21077, + "width": 1024, + "height": 1024 + }, + { + "label": "expanded-ridgeline-campaign-40", + "prompt": "Premium studio product campaign for an insulated water bottle, square print advertisement. Center one tall matte cobalt-blue metal bottle on a low pale stone rectangular pedestal. The bottle has a black screw cap, one narrow orange vertical stripe on its front, and a small white mountain-triangle emblem near its base. On the tabletop to the left of the pedestal is one short yellow enamel cup; to the right is one short turquoise enamel cup. There are exactly three products in total: the bottle and the two cups, all fully visible and separate. Clean warm gray seamless backdrop, soft directional daylight from the upper left, realistic brushed metal and enamel, restrained shadows. At the top, large elegant dark navy sans-serif lettering reads exactly \"TAKE THE LONG WAY\". Directly underneath it, smaller lettering reads exactly \"RIDGELINE / EVERYDAY ADVENTURE\". At the bottom center, small text reads exactly \"750 mL · BUILT TO REUSE\". Plenty of negative space around the type, crisp product edges, no additional props, hands or extra words. Front three-quarter product view with the orange stripe and mountain emblem clearly visible.", + "steps": 40, + "seed": 21077, + "width": 1024, + "height": 1024 + }, + { + "label": "expanded-newsroom-interview-25", + "prompt": "Behind-the-scenes editorial photograph of a radio newsroom recording an interview, wide eye-level square composition. Exactly four adults with clearly separated faces and bodies. On the left, a woman journalist with a dark ponytail, white shirt and black headphones sits at a small table and holds an open notebook in her left hand while gesturing toward the guest with her right. On the right, a middle-aged man with a trimmed beard, brown jacket and silver headphones sits opposite her, speaking toward one black microphone on an articulated desk boom; a second matching microphone points toward the journalist. Behind a glass partition at center, a sound engineer in a blue sweater sits at a mixing console with one hand on a fader. Next to the engineer, a producer wearing a green shirt stands holding up a white card that reads exactly \"2 MIN\". Above the studio window, a red illuminated sign reads exactly \"ON AIR\". The two seated people look at each other, the engineer watches the controls, and the producer faces the studio. Plausible microphone arms and cables, convincing glass with subtle reflections that do not duplicate faces, believable fingers, soft warm studio lighting, documentary realism and no other people.", + "steps": 25, + "seed": 21088, + "width": 1024, + "height": 1024 + }, + { + "label": "expanded-newsroom-interview-40", + "prompt": "Behind-the-scenes editorial photograph of a radio newsroom recording an interview, wide eye-level square composition. Exactly four adults with clearly separated faces and bodies. On the left, a woman journalist with a dark ponytail, white shirt and black headphones sits at a small table and holds an open notebook in her left hand while gesturing toward the guest with her right. On the right, a middle-aged man with a trimmed beard, brown jacket and silver headphones sits opposite her, speaking toward one black microphone on an articulated desk boom; a second matching microphone points toward the journalist. Behind a glass partition at center, a sound engineer in a blue sweater sits at a mixing console with one hand on a fader. Next to the engineer, a producer wearing a green shirt stands holding up a white card that reads exactly \"2 MIN\". Above the studio window, a red illuminated sign reads exactly \"ON AIR\". The two seated people look at each other, the engineer watches the controls, and the producer faces the studio. Plausible microphone arms and cables, convincing glass with subtle reflections that do not duplicate faces, believable fingers, soft warm studio lighting, documentary realism and no other people.", + "steps": 40, + "seed": 21088, + "width": 1024, + "height": 1024 + }, + { + "label": "expanded-station-coat-edit-25", + "prompt": "Change only the mother's navy knee-length coat to a rich burgundy red fabric of the same cut, length, folds and texture. She is the woman on the left with dark curly hair in a low bun, gold earrings and a mustard yellow scarf. Preserve her exact face and identity, expression, gaze, body pose and handholding with the girl. Preserve the yellow scarf, all other clothing, the girl and stuffed rabbit, the father and grandmother's embrace, every person's hands and face, the green suitcase, train, platform, lighting, crop and photographic style. This is a localized garment recoloring, with the rest of the photograph retained.", + "steps": 25, + "seed": 21101, + "width": 1024, + "height": 1024, + "images": [ + "/poc/samples/fidelity-v3/bf16-dit/expanded-station-reunion-40.png" + ] + }, + { + "label": "expanded-station-coat-edit-40", + "prompt": "Change only the mother's navy knee-length coat to a rich burgundy red fabric of the same cut, length, folds and texture. She is the woman on the left with dark curly hair in a low bun, gold earrings and a mustard yellow scarf. Preserve her exact face and identity, expression, gaze, body pose and handholding with the girl. Preserve the yellow scarf, all other clothing, the girl and stuffed rabbit, the father and grandmother's embrace, every person's hands and face, the green suitcase, train, platform, lighting, crop and photographic style. This is a localized garment recoloring, with the rest of the photograph retained.", + "steps": 40, + "seed": 21101, + "width": 1024, + "height": 1024, + "images": [ + "/poc/samples/fidelity-v3/bf16-dit/expanded-station-reunion-40.png" + ] + }, + { + "label": "expanded-festival-type-edit-25", + "prompt": "Update only two text lines in this festival poster. Replace the top English headline \"RIVERLIGHT FILM WEEK\" with exactly \"RIVERLIGHT FILM NIGHTS\", keeping it centered in the same headline area with the same navy font style and a sensible size adjustment if needed. Replace the date line \"18–20 SEPTEMBER 2026\" with exactly \"25–27 SEPTEMBER 2026\". Preserve the Chinese headline \"河畔电影周\" exactly, every program time and English/Chinese film title, both footer lines, the two-column alignment, row spacing, bridge illustration, rules, colors, margins and paper background. Do not translate or reword any other text, and do not redesign the poster.", + "steps": 25, + "seed": 21112, + "width": 1024, + "height": 1024, + "images": [ + "/poc/samples/fidelity-v3/bf16-dit/expanded-bilingual-festival-40.png" + ] + }, + { + "label": "expanded-festival-type-edit-40", + "prompt": "Update only two text lines in this festival poster. Replace the top English headline \"RIVERLIGHT FILM WEEK\" with exactly \"RIVERLIGHT FILM NIGHTS\", keeping it centered in the same headline area with the same navy font style and a sensible size adjustment if needed. Replace the date line \"18–20 SEPTEMBER 2026\" with exactly \"25–27 SEPTEMBER 2026\". Preserve the Chinese headline \"河畔电影周\" exactly, every program time and English/Chinese film title, both footer lines, the two-column alignment, row spacing, bridge illustration, rules, colors, margins and paper background. Do not translate or reword any other text, and do not redesign the poster.", + "steps": 40, + "seed": 21112, + "width": 1024, + "height": 1024, + "images": [ + "/poc/samples/fidelity-v3/bf16-dit/expanded-bilingual-festival-40.png" + ] + }, + { + "label": "expanded-multiref-portrait-25", + "prompt": "Create a natural lifestyle campaign photograph using both references. From the first image, use only the mother on the left: preserve her recognizable oval face, dark curly hair tied in a low bun, small round gold earrings, navy coat and mustard yellow scarf. From the second image, use the central cobalt-blue insulated bottle: preserve its tall shape, black screw cap, single narrow orange front stripe and small white mountain-triangle emblem near the base. Show this same woman seated alone on a wooden railway-platform bench, waist-up, smiling gently toward the camera while holding that bottle upright with both hands around its lower half; the cap, orange stripe and emblem should remain visible. Put one green suitcase on the ground beside the bench, partly visible in frame. Use soft overcast daylight and a stationary passenger train softly out of focus behind her. Realistic hands and fingers, recognizable identity, coherent scale and contact shadows. Do not include the other people from the first reference, the cups or pedestal from the second reference, or any poster lettering. The intended result is one coherent photograph with exactly one woman and one bottle.", + "steps": 25, + "seed": 21123, + "width": 1024, + "height": 1024, + "images": [ + "/poc/samples/fidelity-v3/bf16-dit/expanded-station-reunion-40.png", + "/poc/samples/fidelity-v3/bf16-dit/expanded-ridgeline-campaign-40.png" + ] + }, + { + "label": "expanded-multiref-portrait-40", + "prompt": "Create a natural lifestyle campaign photograph using both references. From the first image, use only the mother on the left: preserve her recognizable oval face, dark curly hair tied in a low bun, small round gold earrings, navy coat and mustard yellow scarf. From the second image, use the central cobalt-blue insulated bottle: preserve its tall shape, black screw cap, single narrow orange front stripe and small white mountain-triangle emblem near the base. Show this same woman seated alone on a wooden railway-platform bench, waist-up, smiling gently toward the camera while holding that bottle upright with both hands around its lower half; the cap, orange stripe and emblem should remain visible. Put one green suitcase on the ground beside the bench, partly visible in frame. Use soft overcast daylight and a stationary passenger train softly out of focus behind her. Realistic hands and fingers, recognizable identity, coherent scale and contact shadows. Do not include the other people from the first reference, the cups or pedestal from the second reference, or any poster lettering. The intended result is one coherent photograph with exactly one woman and one bottle.", + "steps": 40, + "seed": 21123, + "width": 1024, + "height": 1024, + "images": [ + "/poc/samples/fidelity-v3/bf16-dit/expanded-station-reunion-40.png", + "/poc/samples/fidelity-v3/bf16-dit/expanded-ridgeline-campaign-40.png" + ] + }, + { + "label": "expanded-product-relocation-edit-25", + "prompt": "Edit the arrangement of the small cups in this product campaign while preserving the central bottle and all typography. Remove the yellow cup that is currently on the left, restoring the tabletop and its lighting naturally. Move the turquoise cup from the right side to the far left of the pedestal, fully visible with a small clear gap from the pedestal; leave the former right-hand position empty. The final image must contain exactly two products: the unchanged cobalt-blue bottle on its original pedestal and the one turquoise cup on the left. Preserve the bottle's shape, cap, orange stripe, white mountain emblem, location and scale; preserve the pedestal, backdrop, framing, light direction, and the exact text \"TAKE THE LONG WAY\", \"RIDGELINE / EVERYDAY ADVENTURE\", and \"750 mL · BUILT TO REUSE\". Give the moved cup a physically consistent contact shadow. Do not introduce any extra objects or redesign the campaign.", + "steps": 25, + "seed": 21134, + "width": 1024, + "height": 1024, + "images": [ + "/poc/samples/fidelity-v3/bf16-dit/expanded-ridgeline-campaign-40.png" + ] + }, + { + "label": "expanded-product-relocation-edit-40", + "prompt": "Edit the arrangement of the small cups in this product campaign while preserving the central bottle and all typography. Remove the yellow cup that is currently on the left, restoring the tabletop and its lighting naturally. Move the turquoise cup from the right side to the far left of the pedestal, fully visible with a small clear gap from the pedestal; leave the former right-hand position empty. The final image must contain exactly two products: the unchanged cobalt-blue bottle on its original pedestal and the one turquoise cup on the left. Preserve the bottle's shape, cap, orange stripe, white mountain emblem, location and scale; preserve the pedestal, backdrop, framing, light direction, and the exact text \"TAKE THE LONG WAY\", \"RIDGELINE / EVERYDAY ADVENTURE\", and \"750 mL · BUILT TO REUSE\". Give the moved cup a physically consistent contact shadow. Do not introduce any extra objects or redesign the campaign.", + "steps": 40, + "seed": 21134, + "width": 1024, + "height": 1024, + "images": [ + "/poc/samples/fidelity-v3/bf16-dit/expanded-ridgeline-campaign-40.png" + ] + } +] diff --git a/reproduction/experiments/fidelity-v3/quality-jobs.json b/reproduction/experiments/fidelity-v3/quality-jobs.json new file mode 100644 index 0000000000000000000000000000000000000000..0a0cbccbde9ced67c56094cdb24864459938d20d --- /dev/null +++ b/reproduction/experiments/fidelity-v3/quality-jobs.json @@ -0,0 +1,83 @@ +[ + { + "label": "editorial-40", + "prompt": "Candid editorial photograph in a sunlit pottery studio. Exactly two adults: an older woman with short silver hair on the left demonstrates shaping a clay bowl on a pottery wheel; a younger man with curly dark hair on the right watches and holds a folded blue towel with both hands. Their hands are anatomically natural and clearly visible, with clay on the woman's fingertips. Wooden shelves hold irregular ceramic vases, one trailing green plant hangs high in the back right, and soft morning light enters from a large window on the left. Waist-up environmental composition, documentary photography, natural skin texture, restrained warm colors, realistic depth of field. No lettering.", + "steps": 40, + "seed": 118, + "width": 1024, + "height": 1024 + }, + { + "label": "rainy-city-40", + "prompt": "Cinematic street-level architectural photograph of a narrow Tokyo side street at blue hour after rain. A tiny warmly lit ramen restaurant occupies the left foreground, with a striped navy awning and a vertical sign that reads \"RAMEN\". Exactly three bicycles are parked along the right wall. A person holding a transparent umbrella walks away at center, wearing a mustard yellow raincoat. Overhead wires cross between weathered three-story buildings, red lanterns glow farther down the street, and puddles reflect the signs. Strong one-point perspective, realistic glass and wet asphalt, fine distant details, balanced blue and amber lighting.", + "steps": 40, + "seed": 227, + "width": 1024, + "height": 1024 + }, + { + "label": "botanical-poster-40", + "prompt": "Sophisticated botanical exhibition poster on warm ivory paper, flat front-on view, crisp print design. At the top the large dark green serif headline reads exactly \"THE SECRET GARDEN\" on two centered lines. Under it a smaller line reads exactly \"BOTANICAL STUDIES / 2026\". The center contains a delicate detailed watercolor illustration of a lemon branch with exactly two yellow lemons, white blossoms and dark green leaves, surrounded by generous negative space. Along the bottom three evenly spaced text blocks read \"12 OCTOBER\", \"10 AM - 6 PM\", and \"GLASSHOUSE No. 4\". A thin rectangular green border surrounds the entire design. Elegant typography, no other text, no mockup, no shadows outside the paper.", + "steps": 40, + "seed": 336, + "width": 1024, + "height": 1024 + }, + { + "label": "food-spatial-40", + "prompt": "Overhead high-end food photograph of a neatly arranged breakfast on a pale stone table. A large round white plate is centered. On the plate, exactly three pancakes form a vertical stack, topped by exactly five blueberries and a small square of butter. To the left of the plate is a silver fork, to the right is a silver knife with its blade facing the plate. A clear glass of orange juice sits at the upper right, a small white bowl containing sliced strawberries sits at the upper left, and a folded rust-colored linen napkin lies beneath the fork. A little maple syrup runs down only the right edge of the pancake stack. Soft natural light from the upper left, realistic food texture, all objects fully inside frame.", + "steps": 40, + "seed": 445, + "width": 1024, + "height": 1024 + }, + { + "label": "fantasy-cutaway-40", + "prompt": "Intricate isometric cutaway illustration of a cozy three-level library built inside an enormous ancient tree, in a hand-painted storybook style. Ground floor: a round green entrance door, a circular rug, and a sleeping orange cat beside a small wood stove. Middle floor: curved bookshelves filled with colorful books, a writing desk by an oval window, and a brass telescope pointing outside. Top floor: a glass-domed reading nook with two red armchairs and a hanging lantern. A single wooden spiral staircase visibly connects all three floors. Exposed roots wrap around mossy rocks, tiny mushrooms cluster near the door, and the leafy crown surrounds the dome without hiding it. Consistent isometric perspective, warm interior lighting, cool twilight forest background, detailed but readable, no text.", + "steps": 40, + "seed": 554, + "width": 1024, + "height": 1024 + }, + { + "label": "glass-product-40", + "prompt": "Luxury studio still life photograph on a polished dark green marble plinth. A rectangular clear glass perfume bottle half filled with pale amber liquid stands in the center, with a brushed gold cylindrical cap and a small cream label reading exactly \"LUMEN\". A thin curved strip of orange peel lies in front of the bottle. Behind and to the left is a translucent ribbed glass sphere; behind and to the right are two glossy dark green leaves. Hard afternoon sunlight from the upper left creates realistic caustics, transparent overlapping shadows, precise refraction through the bottle and a soft reflection on the marble. Deep emerald background, controlled commercial composition, realistic materials, no additional text.", + "steps": 40, + "seed": 663, + "width": 1024, + "height": 1024 + }, + { + "label": "editorial-edit-40", + "prompt": "Replace only the blue towel held by the younger man with a bright red towel of the same size and folded shape. Preserve both people's identities, facial expressions, hands, clothing, clay bowl, studio shelves, window light, framing and photographic style.", + "steps": 40, + "seed": 774, + "width": 1024, + "height": 1024, + "images": [ + "/poc/samples/varied-v2/nf4/editorial-40.png" + ] + }, + { + "label": "poster-edit-40", + "prompt": "Edit only the text at the top: replace \"THE SECRET GARDEN\" with \"THE LEMON HOUSE\". Keep the same dark green serif type style and centered placement. Preserve all other text exactly, the lemon branch illustration, paper background, border, spacing and overall design.", + "steps": 40, + "seed": 885, + "width": 1024, + "height": 1024, + "images": [ + "/poc/samples/varied-v2/nf4/botanical-poster-40.png" + ] + }, + { + "label": "city-edit-40", + "prompt": "Transform this rainy Tokyo street photograph into a richly textured hand-painted watercolor illustration. Preserve the layout of the buildings, the ramen restaurant on the left, all three bicycles on the right, the central person with the transparent umbrella and yellow raincoat, the overhead wires, blue-hour lighting, and puddle reflections. Keep the RAMEN sign legible.", + "steps": 40, + "seed": 996, + "width": 1024, + "height": 1024, + "images": [ + "/poc/samples/varied-v2/nf4/rainy-city-40.png" + ] + } +] diff --git a/reproduction/experiments/fidelity-v3/selection.json b/reproduction/experiments/fidelity-v3/selection.json new file mode 100644 index 0000000000000000000000000000000000000000..0ad5a2ae4b12e5f3c5b2688f7c604af4b67351fb --- /dev/null +++ b/reproduction/experiments/fidelity-v3/selection.json @@ -0,0 +1,15 @@ +{ + "selected": "nunchaku-v3-r128", + "checkpoint": "/cache/qwen-nunchaku-v3-r128", + "manifest_sha256": "cf69e83a646981d22ee6df8d5239b46a50df25d8eb73c9f0478feae87323e6cb", + "compose_override": "compose.nunchaku-v3.yaml", + "decision": "Retain calibrated rank128; selective MLP rank512 is experimental, not promoted.", + "reason": "Selective512 improves numerical prediction error but has new free-running image count and geometry regressions. Rank128 is faster and smaller and already has a complete42-image visual review.", + "counterexamples": [ + "food-spatial-40: seven blueberries and two knives", + "fantasy-cutaway-25: two cats", + "fantasy-cutaway-40: distorted telescope", + "expanded-community-kitchen-25: six bread rolls" + ], + "limitations": "Neither checkpoint is established near-lossless or uniformly superior; serving verification recorded separately." +} diff --git a/reproduction/experiments/fidelity-v3/teacher-priority-bf16.json b/reproduction/experiments/fidelity-v3/teacher-priority-bf16.json new file mode 100644 index 0000000000000000000000000000000000000000..45b304ecb39fd765c9f23c5b860915e2fc2d1a0f --- /dev/null +++ b/reproduction/experiments/fidelity-v3/teacher-priority-bf16.json @@ -0,0 +1,26 @@ +[ + { + "label": "food-spatial-40", + "prompt": "Overhead high-end food photograph of a neatly arranged breakfast on a pale stone table. A large round white plate is centered. On the plate, exactly three pancakes form a vertical stack, topped by exactly five blueberries and a small square of butter. To the left of the plate is a silver fork, to the right is a silver knife with its blade facing the plate. A clear glass of orange juice sits at the upper right, a small white bowl containing sliced strawberries sits at the upper left, and a folded rust-colored linen napkin lies beneath the fork. A little maple syrup runs down only the right edge of the pancake stack. Soft natural light from the upper left, realistic food texture, all objects fully inside frame.", + "steps": 40, + "seed": 445, + "width": 1024, + "height": 1024 + }, + { + "label": "fantasy-cutaway-40", + "prompt": "Intricate isometric cutaway illustration of a cozy three-level library built inside an enormous ancient tree, in a hand-painted storybook style. Ground floor: a round green entrance door, a circular rug, and a sleeping orange cat beside a small wood stove. Middle floor: curved bookshelves filled with colorful books, a writing desk by an oval window, and a brass telescope pointing outside. Top floor: a glass-domed reading nook with two red armchairs and a hanging lantern. A single wooden spiral staircase visibly connects all three floors. Exposed roots wrap around mossy rocks, tiny mushrooms cluster near the door, and the leafy crown surrounds the dome without hiding it. Consistent isometric perspective, warm interior lighting, cool twilight forest background, detailed but readable, no text.", + "steps": 40, + "seed": 554, + "width": 1024, + "height": 1024 + }, + { + "label": "editorial-40", + "prompt": "Candid editorial photograph in a sunlit pottery studio. Exactly two adults: an older woman with short silver hair on the left demonstrates shaping a clay bowl on a pottery wheel; a younger man with curly dark hair on the right watches and holds a folded blue towel with both hands. Their hands are anatomically natural and clearly visible, with clay on the woman's fingertips. Wooden shelves hold irregular ceramic vases, one trailing green plant hangs high in the back right, and soft morning light enters from a large window on the left. Waist-up environmental composition, documentary photography, natural skin texture, restrained warm colors, realistic depth of field. No lettering.", + "steps": 40, + "seed": 118, + "width": 1024, + "height": 1024 + } +] diff --git a/reproduction/experiments/fidelity-v3/teacher-priority.json b/reproduction/experiments/fidelity-v3/teacher-priority.json new file mode 100644 index 0000000000000000000000000000000000000000..fa789804b7ef4365c89d4e03309da001ab9884cd --- /dev/null +++ b/reproduction/experiments/fidelity-v3/teacher-priority.json @@ -0,0 +1,58 @@ +[ + { + "label": "food-spatial-40", + "prompt": "Overhead high-end food photograph of a neatly arranged breakfast on a pale stone table. A large round white plate is centered. On the plate, exactly three pancakes form a vertical stack, topped by exactly five blueberries and a small square of butter. To the left of the plate is a silver fork, to the right is a silver knife with its blade facing the plate. A clear glass of orange juice sits at the upper right, a small white bowl containing sliced strawberries sits at the upper left, and a folded rust-colored linen napkin lies beneath the fork. A little maple syrup runs down only the right edge of the pancake stack. Soft natural light from the upper left, realistic food texture, all objects fully inside frame.", + "steps": 40, + "seed": 445, + "width": 1024, + "height": 1024 + }, + { + "label": "fantasy-cutaway-40", + "prompt": "Intricate isometric cutaway illustration of a cozy three-level library built inside an enormous ancient tree, in a hand-painted storybook style. Ground floor: a round green entrance door, a circular rug, and a sleeping orange cat beside a small wood stove. Middle floor: curved bookshelves filled with colorful books, a writing desk by an oval window, and a brass telescope pointing outside. Top floor: a glass-domed reading nook with two red armchairs and a hanging lantern. A single wooden spiral staircase visibly connects all three floors. Exposed roots wrap around mossy rocks, tiny mushrooms cluster near the door, and the leafy crown surrounds the dome without hiding it. Consistent isometric perspective, warm interior lighting, cool twilight forest background, detailed but readable, no text.", + "steps": 40, + "seed": 554, + "width": 1024, + "height": 1024 + }, + { + "label": "editorial-40", + "prompt": "Candid editorial photograph in a sunlit pottery studio. Exactly two adults: an older woman with short silver hair on the left demonstrates shaping a clay bowl on a pottery wheel; a younger man with curly dark hair on the right watches and holds a folded blue towel with both hands. Their hands are anatomically natural and clearly visible, with clay on the woman's fingertips. Wooden shelves hold irregular ceramic vases, one trailing green plant hangs high in the back right, and soft morning light enters from a large window on the left. Waist-up environmental composition, documentary photography, natural skin texture, restrained warm colors, realistic depth of field. No lettering.", + "steps": 40, + "seed": 118, + "width": 1024, + "height": 1024 + }, + { + "label": "expanded-community-kitchen-40", + "prompt": "Documentary photograph for a community cooking program, wide eye-level composition in a bright, practical teaching kitchen. Exactly six adults, all visible from at least the waist up. At the left counter, a woman with short black hair in a green apron chops carrots on a wooden board, with one hand holding the knife and the other safely curled on the carrot. At center, a silver-haired man in a white apron holds a ladle over a large soup pot. Immediately to his right, a woman in a mustard cardigan passes a rectangular tray of four bread rolls to a younger man in a navy shirt, who receives it with both hands; their hands meet opposite edges of the same tray. At the far right, a man in a red apron rinses a colander of leafy greens at the sink. Behind the central counter, a woman wearing glasses and a purple sweater reads a paper recipe card. Natural varied expressions, clear individual faces, plausible arms and hands attached to their owners, no extra people, realistic morning window light and stainless-steel reflections. No readable signage. The image should feel like a real coordinated cooking session, not a posed group portrait.", + "steps": 40, + "seed": 21011, + "width": 1024, + "height": 1024 + }, + { + "label": "expanded-chess-awards-40", + "prompt": "Editorial event photograph of a small chess championship award presentation on a low indoor stage. Exactly five adults. At center, a smiling woman with a short dark bob sits in a clearly recognizable manual wheelchair; she and the presenter standing immediately to her left jointly hold one gold trophy between them. The presenter is an older man wearing a charcoal suit. Immediately to the winner's right, the runner-up, a young man in a cream sweater, wears one silver medal and applauds. At the far left, a female photographer kneels with a camera held to her eye and aims toward the winner. At the far right, a host in a burgundy suit holds a microphone and a small cue card. All five people fit fully in frame, with believable hands, wheelchair wheels and footrests. Behind them, a clean white stage banner reads exactly \"CITY CHESS FINAL\" with \"2026\" centered on the line below. A small chess table with a board stands at the rear left. Warm indoor event lighting, realistic fabric and skin, no audience or additional figures.", + "steps": 40, + "seed": 21022, + "width": 1024, + "height": 1024 + }, + { + "label": "expanded-bilingual-festival-40", + "prompt": "A polished bilingual neighborhood film festival poster, square format, flat front-on print artwork with generous margins, warm cream paper and deep navy typography accented by vermilion. The large centered English title at the top reads exactly \"RIVERLIGHT FILM WEEK\". Directly beneath it, an equally clear Chinese title reads exactly \"河畔电影周\". The next centered line reads exactly \"18–20 SEPTEMBER 2026\". Below a thin vermilion rule, a tidy two-column program occupies the middle: the left column contains times, and the right column contains one English film title with its Chinese title directly underneath. Row one: \"18:00\", \"A Quiet River\", \"静静的河\". Row two: \"19:30\", \"Night Market\", \"夜市\". Row three: \"21:00\", \"Home Again\", \"再次回家\". Maintain aligned columns and clearly separated rows. At the bottom, centered on separate lines, print exactly \"RIVERSIDE CINEMA / 河畔影院\" and \"FREE ENTRY / 免费入场\". A small restrained illustration of a red bridge over two blue river lines sits between the program and footer without touching any letters. Sharp readable Latin and Chinese lettering, consistent type hierarchy, no additional text, no mockup perspective.", + "steps": 40, + "seed": 21044, + "width": 1024, + "height": 1024 + }, + { + "label": "expanded-library-comic-40", + "prompt": "A polished, friendly three-panel comic strip in a square canvas, with exactly three equal vertical panels arranged left to right, clear white gutters and a thin dark border around each panel. Keep the same two adult characters and library reading-room setting in all panels: Maya has a chin-length black bob, round glasses and a red sweater; Leo has short curly brown hair and a blue button-up shirt. Each panel contains exactly two white speech bubbles, with clear tails pointing to the correct speaker. Panel 1: Maya stands beside a wooden reading chair, looks worried and searches her bag. Maya says exactly \"Where are my keys?\" Leo, standing to her right, says exactly \"Check your pocket.\" Panel 2: Maya turns out an empty coat pocket and says exactly \"Not here!\" Leo points down at a small key ring visible under that same chair and says exactly \"Under the chair.\" Panel 3: Maya stands upright holding that key ring up, smiles and says exactly \"Found them. Thanks!\" Leo smiles and says exactly \"You're welcome.\" Crisp ink lines, restrained warm colors, expressive natural poses, consistent faces and clothing, believable hands, a coherent sequence with the chair in the same room. Readable dialogue with no narration captions and no additional panels or characters.", + "steps": 40, + "seed": 21055, + "width": 1024, + "height": 1024 + } +] diff --git a/reproduction/experiments/fidelity-v3/teacher-remaining.json b/reproduction/experiments/fidelity-v3/teacher-remaining.json new file mode 100644 index 0000000000000000000000000000000000000000..5ed46fc456d803aa38ca04f49cadb4d0eb548511 --- /dev/null +++ b/reproduction/experiments/fidelity-v3/teacher-remaining.json @@ -0,0 +1,358 @@ +[ + { + "label": "expanded-community-kitchen-40", + "prompt": "Documentary photograph for a community cooking program, wide eye-level composition in a bright, practical teaching kitchen. Exactly six adults, all visible from at least the waist up. At the left counter, a woman with short black hair in a green apron chops carrots on a wooden board, with one hand holding the knife and the other safely curled on the carrot. At center, a silver-haired man in a white apron holds a ladle over a large soup pot. Immediately to his right, a woman in a mustard cardigan passes a rectangular tray of four bread rolls to a younger man in a navy shirt, who receives it with both hands; their hands meet opposite edges of the same tray. At the far right, a man in a red apron rinses a colander of leafy greens at the sink. Behind the central counter, a woman wearing glasses and a purple sweater reads a paper recipe card. Natural varied expressions, clear individual faces, plausible arms and hands attached to their owners, no extra people, realistic morning window light and stainless-steel reflections. No readable signage. The image should feel like a real coordinated cooking session, not a posed group portrait.", + "steps": 40, + "seed": 21011, + "width": 1024, + "height": 1024 + }, + { + "label": "expanded-chess-awards-40", + "prompt": "Editorial event photograph of a small chess championship award presentation on a low indoor stage. Exactly five adults. At center, a smiling woman with a short dark bob sits in a clearly recognizable manual wheelchair; she and the presenter standing immediately to her left jointly hold one gold trophy between them. The presenter is an older man wearing a charcoal suit. Immediately to the winner's right, the runner-up, a young man in a cream sweater, wears one silver medal and applauds. At the far left, a female photographer kneels with a camera held to her eye and aims toward the winner. At the far right, a host in a burgundy suit holds a microphone and a small cue card. All five people fit fully in frame, with believable hands, wheelchair wheels and footrests. Behind them, a clean white stage banner reads exactly \"CITY CHESS FINAL\" with \"2026\" centered on the line below. A small chess table with a board stands at the rear left. Warm indoor event lighting, realistic fabric and skin, no audience or additional figures.", + "steps": 40, + "seed": 21022, + "width": 1024, + "height": 1024 + }, + { + "label": "expanded-station-reunion-40", + "prompt": "Natural editorial photograph of a family reunion beside a stationary passenger train on a quiet outdoor railway platform. Exactly four people, full bodies in frame, no other passengers. On the left is a woman in her thirties with an oval face, dark curly hair tied in a low bun, small round gold earrings, a navy knee-length coat and a mustard yellow scarf. She holds the left hand of a young girl standing at center; the girl wears a red raincoat and carries a small stuffed rabbit in her free hand. On the right, a man with close-cropped brown hair, a light gray jacket and a red backpack bends slightly to embrace an older gray-haired woman wearing a teal coat; the older woman faces him with one arm around his shoulder. The mother and girl watch them with happy, understated expressions. One upright green suitcase stands beside the mother's left leg. Anatomically plausible hands and embraces with clear ownership of each arm, natural eye contact, soft overcast daylight, authentic train windows and platform paving. Do not crop feet or add duplicate people. No readable logos or signs.", + "steps": 40, + "seed": 21033, + "width": 1024, + "height": 1024 + }, + { + "label": "expanded-bilingual-festival-40", + "prompt": "A polished bilingual neighborhood film festival poster, square format, flat front-on print artwork with generous margins, warm cream paper and deep navy typography accented by vermilion. The large centered English title at the top reads exactly \"RIVERLIGHT FILM WEEK\". Directly beneath it, an equally clear Chinese title reads exactly \"河畔电影周\". The next centered line reads exactly \"18–20 SEPTEMBER 2026\". Below a thin vermilion rule, a tidy two-column program occupies the middle: the left column contains times, and the right column contains one English film title with its Chinese title directly underneath. Row one: \"18:00\", \"A Quiet River\", \"静静的河\". Row two: \"19:30\", \"Night Market\", \"夜市\". Row three: \"21:00\", \"Home Again\", \"再次回家\". Maintain aligned columns and clearly separated rows. At the bottom, centered on separate lines, print exactly \"RIVERSIDE CINEMA / 河畔影院\" and \"FREE ENTRY / 免费入场\". A small restrained illustration of a red bridge over two blue river lines sits between the program and footer without touching any letters. Sharp readable Latin and Chinese lettering, consistent type hierarchy, no additional text, no mockup perspective.", + "steps": 40, + "seed": 21044, + "width": 1024, + "height": 1024 + }, + { + "label": "expanded-library-comic-40", + "prompt": "A polished, friendly three-panel comic strip in a square canvas, with exactly three equal vertical panels arranged left to right, clear white gutters and a thin dark border around each panel. Keep the same two adult characters and library reading-room setting in all panels: Maya has a chin-length black bob, round glasses and a red sweater; Leo has short curly brown hair and a blue button-up shirt. Each panel contains exactly two white speech bubbles, with clear tails pointing to the correct speaker. Panel 1: Maya stands beside a wooden reading chair, looks worried and searches her bag. Maya says exactly \"Where are my keys?\" Leo, standing to her right, says exactly \"Check your pocket.\" Panel 2: Maya turns out an empty coat pocket and says exactly \"Not here!\" Leo points down at a small key ring visible under that same chair and says exactly \"Under the chair.\" Panel 3: Maya stands upright holding that key ring up, smiles and says exactly \"Found them. Thanks!\" Leo smiles and says exactly \"You're welcome.\" Crisp ink lines, restrained warm colors, expressive natural poses, consistent faces and clothing, believable hands, a coherent sequence with the chair in the same room. Readable dialogue with no narration captions and no additional panels or characters.", + "steps": 40, + "seed": 21055, + "width": 1024, + "height": 1024 + }, + { + "label": "expanded-museum-plan-40", + "prompt": "A clear visitor floor-plan diagram for a small museum, square format, flat orthographic top-down vector design on white, with thick charcoal exterior walls, thinner interior walls and clearly visible doorway gaps. Title centered above the plan: \"MUSEUM VISITOR MAP\". Inside one rectangular building, arrange five labeled spaces: \"GALLERY A\" in a large upper-left room, \"GALLERY B\" in a large upper-right room, \"MAIN HALL\" in a wide horizontal central space spanning the building, \"CAFE\" in the lower-left room, and \"SHOP\" in the lower-right room. A narrow unlabeled central entrance aisle runs from a doorway at the bottom edge into the main hall, separating the cafe and shop. Label that bottom doorway \"ENTRANCE\" just outside the building. Each of the four corner rooms has one doorway directly into the main hall. Give Gallery A a pale blue fill, Gallery B pale green, the cafe pale orange and the shop pale violet; keep the hall and entrance aisle white. Put exactly two simple rectangular display plinth symbols in each gallery, exactly three round table symbols in the cafe, and exactly two parallel shelf symbols in the shop. A single continuous blue visitor-route line with arrowheads begins outside the entrance, goes through the entrance aisle and main hall, and ends inside Gallery A, passing through actual doorway openings without crossing any walls. Small bottom legend: a blue arrow icon followed by exactly \"SUGGESTED ROUTE\". Spacious legible labeling, precise geometry, no perspective, no people and no extra rooms.", + "steps": 40, + "seed": 21066, + "width": 1024, + "height": 1024 + }, + { + "label": "expanded-ridgeline-campaign-40", + "prompt": "Premium studio product campaign for an insulated water bottle, square print advertisement. Center one tall matte cobalt-blue metal bottle on a low pale stone rectangular pedestal. The bottle has a black screw cap, one narrow orange vertical stripe on its front, and a small white mountain-triangle emblem near its base. On the tabletop to the left of the pedestal is one short yellow enamel cup; to the right is one short turquoise enamel cup. There are exactly three products in total: the bottle and the two cups, all fully visible and separate. Clean warm gray seamless backdrop, soft directional daylight from the upper left, realistic brushed metal and enamel, restrained shadows. At the top, large elegant dark navy sans-serif lettering reads exactly \"TAKE THE LONG WAY\". Directly underneath it, smaller lettering reads exactly \"RIDGELINE / EVERYDAY ADVENTURE\". At the bottom center, small text reads exactly \"750 mL · BUILT TO REUSE\". Plenty of negative space around the type, crisp product edges, no additional props, hands or extra words. Front three-quarter product view with the orange stripe and mountain emblem clearly visible.", + "steps": 40, + "seed": 21077, + "width": 1024, + "height": 1024 + }, + { + "label": "expanded-newsroom-interview-40", + "prompt": "Behind-the-scenes editorial photograph of a radio newsroom recording an interview, wide eye-level square composition. Exactly four adults with clearly separated faces and bodies. On the left, a woman journalist with a dark ponytail, white shirt and black headphones sits at a small table and holds an open notebook in her left hand while gesturing toward the guest with her right. On the right, a middle-aged man with a trimmed beard, brown jacket and silver headphones sits opposite her, speaking toward one black microphone on an articulated desk boom; a second matching microphone points toward the journalist. Behind a glass partition at center, a sound engineer in a blue sweater sits at a mixing console with one hand on a fader. Next to the engineer, a producer wearing a green shirt stands holding up a white card that reads exactly \"2 MIN\". Above the studio window, a red illuminated sign reads exactly \"ON AIR\". The two seated people look at each other, the engineer watches the controls, and the producer faces the studio. Plausible microphone arms and cables, convincing glass with subtle reflections that do not duplicate faces, believable fingers, soft warm studio lighting, documentary realism and no other people.", + "steps": 40, + "seed": 21088, + "width": 1024, + "height": 1024 + }, + { + "label": "expanded-community-kitchen-25", + "prompt": "Documentary photograph for a community cooking program, wide eye-level composition in a bright, practical teaching kitchen. Exactly six adults, all visible from at least the waist up. At the left counter, a woman with short black hair in a green apron chops carrots on a wooden board, with one hand holding the knife and the other safely curled on the carrot. At center, a silver-haired man in a white apron holds a ladle over a large soup pot. Immediately to his right, a woman in a mustard cardigan passes a rectangular tray of four bread rolls to a younger man in a navy shirt, who receives it with both hands; their hands meet opposite edges of the same tray. At the far right, a man in a red apron rinses a colander of leafy greens at the sink. Behind the central counter, a woman wearing glasses and a purple sweater reads a paper recipe card. Natural varied expressions, clear individual faces, plausible arms and hands attached to their owners, no extra people, realistic morning window light and stainless-steel reflections. No readable signage. The image should feel like a real coordinated cooking session, not a posed group portrait.", + "steps": 25, + "seed": 21011, + "width": 1024, + "height": 1024 + }, + { + "label": "expanded-chess-awards-25", + "prompt": "Editorial event photograph of a small chess championship award presentation on a low indoor stage. Exactly five adults. At center, a smiling woman with a short dark bob sits in a clearly recognizable manual wheelchair; she and the presenter standing immediately to her left jointly hold one gold trophy between them. The presenter is an older man wearing a charcoal suit. Immediately to the winner's right, the runner-up, a young man in a cream sweater, wears one silver medal and applauds. At the far left, a female photographer kneels with a camera held to her eye and aims toward the winner. At the far right, a host in a burgundy suit holds a microphone and a small cue card. All five people fit fully in frame, with believable hands, wheelchair wheels and footrests. Behind them, a clean white stage banner reads exactly \"CITY CHESS FINAL\" with \"2026\" centered on the line below. A small chess table with a board stands at the rear left. Warm indoor event lighting, realistic fabric and skin, no audience or additional figures.", + "steps": 25, + "seed": 21022, + "width": 1024, + "height": 1024 + }, + { + "label": "expanded-station-reunion-25", + "prompt": "Natural editorial photograph of a family reunion beside a stationary passenger train on a quiet outdoor railway platform. Exactly four people, full bodies in frame, no other passengers. On the left is a woman in her thirties with an oval face, dark curly hair tied in a low bun, small round gold earrings, a navy knee-length coat and a mustard yellow scarf. She holds the left hand of a young girl standing at center; the girl wears a red raincoat and carries a small stuffed rabbit in her free hand. On the right, a man with close-cropped brown hair, a light gray jacket and a red backpack bends slightly to embrace an older gray-haired woman wearing a teal coat; the older woman faces him with one arm around his shoulder. The mother and girl watch them with happy, understated expressions. One upright green suitcase stands beside the mother's left leg. Anatomically plausible hands and embraces with clear ownership of each arm, natural eye contact, soft overcast daylight, authentic train windows and platform paving. Do not crop feet or add duplicate people. No readable logos or signs.", + "steps": 25, + "seed": 21033, + "width": 1024, + "height": 1024 + }, + { + "label": "expanded-bilingual-festival-25", + "prompt": "A polished bilingual neighborhood film festival poster, square format, flat front-on print artwork with generous margins, warm cream paper and deep navy typography accented by vermilion. The large centered English title at the top reads exactly \"RIVERLIGHT FILM WEEK\". Directly beneath it, an equally clear Chinese title reads exactly \"河畔电影周\". The next centered line reads exactly \"18–20 SEPTEMBER 2026\". Below a thin vermilion rule, a tidy two-column program occupies the middle: the left column contains times, and the right column contains one English film title with its Chinese title directly underneath. Row one: \"18:00\", \"A Quiet River\", \"静静的河\". Row two: \"19:30\", \"Night Market\", \"夜市\". Row three: \"21:00\", \"Home Again\", \"再次回家\". Maintain aligned columns and clearly separated rows. At the bottom, centered on separate lines, print exactly \"RIVERSIDE CINEMA / 河畔影院\" and \"FREE ENTRY / 免费入场\". A small restrained illustration of a red bridge over two blue river lines sits between the program and footer without touching any letters. Sharp readable Latin and Chinese lettering, consistent type hierarchy, no additional text, no mockup perspective.", + "steps": 25, + "seed": 21044, + "width": 1024, + "height": 1024 + }, + { + "label": "expanded-library-comic-25", + "prompt": "A polished, friendly three-panel comic strip in a square canvas, with exactly three equal vertical panels arranged left to right, clear white gutters and a thin dark border around each panel. Keep the same two adult characters and library reading-room setting in all panels: Maya has a chin-length black bob, round glasses and a red sweater; Leo has short curly brown hair and a blue button-up shirt. Each panel contains exactly two white speech bubbles, with clear tails pointing to the correct speaker. Panel 1: Maya stands beside a wooden reading chair, looks worried and searches her bag. Maya says exactly \"Where are my keys?\" Leo, standing to her right, says exactly \"Check your pocket.\" Panel 2: Maya turns out an empty coat pocket and says exactly \"Not here!\" Leo points down at a small key ring visible under that same chair and says exactly \"Under the chair.\" Panel 3: Maya stands upright holding that key ring up, smiles and says exactly \"Found them. Thanks!\" Leo smiles and says exactly \"You're welcome.\" Crisp ink lines, restrained warm colors, expressive natural poses, consistent faces and clothing, believable hands, a coherent sequence with the chair in the same room. Readable dialogue with no narration captions and no additional panels or characters.", + "steps": 25, + "seed": 21055, + "width": 1024, + "height": 1024 + }, + { + "label": "expanded-museum-plan-25", + "prompt": "A clear visitor floor-plan diagram for a small museum, square format, flat orthographic top-down vector design on white, with thick charcoal exterior walls, thinner interior walls and clearly visible doorway gaps. Title centered above the plan: \"MUSEUM VISITOR MAP\". Inside one rectangular building, arrange five labeled spaces: \"GALLERY A\" in a large upper-left room, \"GALLERY B\" in a large upper-right room, \"MAIN HALL\" in a wide horizontal central space spanning the building, \"CAFE\" in the lower-left room, and \"SHOP\" in the lower-right room. A narrow unlabeled central entrance aisle runs from a doorway at the bottom edge into the main hall, separating the cafe and shop. Label that bottom doorway \"ENTRANCE\" just outside the building. Each of the four corner rooms has one doorway directly into the main hall. Give Gallery A a pale blue fill, Gallery B pale green, the cafe pale orange and the shop pale violet; keep the hall and entrance aisle white. Put exactly two simple rectangular display plinth symbols in each gallery, exactly three round table symbols in the cafe, and exactly two parallel shelf symbols in the shop. A single continuous blue visitor-route line with arrowheads begins outside the entrance, goes through the entrance aisle and main hall, and ends inside Gallery A, passing through actual doorway openings without crossing any walls. Small bottom legend: a blue arrow icon followed by exactly \"SUGGESTED ROUTE\". Spacious legible labeling, precise geometry, no perspective, no people and no extra rooms.", + "steps": 25, + "seed": 21066, + "width": 1024, + "height": 1024 + }, + { + "label": "expanded-ridgeline-campaign-25", + "prompt": "Premium studio product campaign for an insulated water bottle, square print advertisement. Center one tall matte cobalt-blue metal bottle on a low pale stone rectangular pedestal. The bottle has a black screw cap, one narrow orange vertical stripe on its front, and a small white mountain-triangle emblem near its base. On the tabletop to the left of the pedestal is one short yellow enamel cup; to the right is one short turquoise enamel cup. There are exactly three products in total: the bottle and the two cups, all fully visible and separate. Clean warm gray seamless backdrop, soft directional daylight from the upper left, realistic brushed metal and enamel, restrained shadows. At the top, large elegant dark navy sans-serif lettering reads exactly \"TAKE THE LONG WAY\". Directly underneath it, smaller lettering reads exactly \"RIDGELINE / EVERYDAY ADVENTURE\". At the bottom center, small text reads exactly \"750 mL · BUILT TO REUSE\". Plenty of negative space around the type, crisp product edges, no additional props, hands or extra words. Front three-quarter product view with the orange stripe and mountain emblem clearly visible.", + "steps": 25, + "seed": 21077, + "width": 1024, + "height": 1024 + }, + { + "label": "expanded-newsroom-interview-25", + "prompt": "Behind-the-scenes editorial photograph of a radio newsroom recording an interview, wide eye-level square composition. Exactly four adults with clearly separated faces and bodies. On the left, a woman journalist with a dark ponytail, white shirt and black headphones sits at a small table and holds an open notebook in her left hand while gesturing toward the guest with her right. On the right, a middle-aged man with a trimmed beard, brown jacket and silver headphones sits opposite her, speaking toward one black microphone on an articulated desk boom; a second matching microphone points toward the journalist. Behind a glass partition at center, a sound engineer in a blue sweater sits at a mixing console with one hand on a fader. Next to the engineer, a producer wearing a green shirt stands holding up a white card that reads exactly \"2 MIN\". Above the studio window, a red illuminated sign reads exactly \"ON AIR\". The two seated people look at each other, the engineer watches the controls, and the producer faces the studio. Plausible microphone arms and cables, convincing glass with subtle reflections that do not duplicate faces, believable fingers, soft warm studio lighting, documentary realism and no other people.", + "steps": 25, + "seed": 21088, + "width": 1024, + "height": 1024 + }, + { + "label": "expanded-station-coat-edit-25", + "prompt": "Change only the mother's navy knee-length coat to a rich burgundy red fabric of the same cut, length, folds and texture. She is the woman on the left with dark curly hair in a low bun, gold earrings and a mustard yellow scarf. Preserve her exact face and identity, expression, gaze, body pose and handholding with the girl. Preserve the yellow scarf, all other clothing, the girl and stuffed rabbit, the father and grandmother's embrace, every person's hands and face, the green suitcase, train, platform, lighting, crop and photographic style. This is a localized garment recoloring, with the rest of the photograph retained.", + "steps": 25, + "seed": 21101, + "width": 1024, + "height": 1024, + "images": [ + "/poc/samples/fidelity-v3/bf16-dit/expanded-station-reunion-40.png" + ] + }, + { + "label": "expanded-station-coat-edit-40", + "prompt": "Change only the mother's navy knee-length coat to a rich burgundy red fabric of the same cut, length, folds and texture. She is the woman on the left with dark curly hair in a low bun, gold earrings and a mustard yellow scarf. Preserve her exact face and identity, expression, gaze, body pose and handholding with the girl. Preserve the yellow scarf, all other clothing, the girl and stuffed rabbit, the father and grandmother's embrace, every person's hands and face, the green suitcase, train, platform, lighting, crop and photographic style. This is a localized garment recoloring, with the rest of the photograph retained.", + "steps": 40, + "seed": 21101, + "width": 1024, + "height": 1024, + "images": [ + "/poc/samples/fidelity-v3/bf16-dit/expanded-station-reunion-40.png" + ] + }, + { + "label": "expanded-festival-type-edit-25", + "prompt": "Update only two text lines in this festival poster. Replace the top English headline \"RIVERLIGHT FILM WEEK\" with exactly \"RIVERLIGHT FILM NIGHTS\", keeping it centered in the same headline area with the same navy font style and a sensible size adjustment if needed. Replace the date line \"18–20 SEPTEMBER 2026\" with exactly \"25–27 SEPTEMBER 2026\". Preserve the Chinese headline \"河畔电影周\" exactly, every program time and English/Chinese film title, both footer lines, the two-column alignment, row spacing, bridge illustration, rules, colors, margins and paper background. Do not translate or reword any other text, and do not redesign the poster.", + "steps": 25, + "seed": 21112, + "width": 1024, + "height": 1024, + "images": [ + "/poc/samples/fidelity-v3/bf16-dit/expanded-bilingual-festival-40.png" + ] + }, + { + "label": "expanded-festival-type-edit-40", + "prompt": "Update only two text lines in this festival poster. Replace the top English headline \"RIVERLIGHT FILM WEEK\" with exactly \"RIVERLIGHT FILM NIGHTS\", keeping it centered in the same headline area with the same navy font style and a sensible size adjustment if needed. Replace the date line \"18–20 SEPTEMBER 2026\" with exactly \"25–27 SEPTEMBER 2026\". Preserve the Chinese headline \"河畔电影周\" exactly, every program time and English/Chinese film title, both footer lines, the two-column alignment, row spacing, bridge illustration, rules, colors, margins and paper background. Do not translate or reword any other text, and do not redesign the poster.", + "steps": 40, + "seed": 21112, + "width": 1024, + "height": 1024, + "images": [ + "/poc/samples/fidelity-v3/bf16-dit/expanded-bilingual-festival-40.png" + ] + }, + { + "label": "expanded-multiref-portrait-25", + "prompt": "Create a natural lifestyle campaign photograph using both references. From the first image, use only the mother on the left: preserve her recognizable oval face, dark curly hair tied in a low bun, small round gold earrings, navy coat and mustard yellow scarf. From the second image, use the central cobalt-blue insulated bottle: preserve its tall shape, black screw cap, single narrow orange front stripe and small white mountain-triangle emblem near the base. Show this same woman seated alone on a wooden railway-platform bench, waist-up, smiling gently toward the camera while holding that bottle upright with both hands around its lower half; the cap, orange stripe and emblem should remain visible. Put one green suitcase on the ground beside the bench, partly visible in frame. Use soft overcast daylight and a stationary passenger train softly out of focus behind her. Realistic hands and fingers, recognizable identity, coherent scale and contact shadows. Do not include the other people from the first reference, the cups or pedestal from the second reference, or any poster lettering. The intended result is one coherent photograph with exactly one woman and one bottle.", + "steps": 25, + "seed": 21123, + "width": 1024, + "height": 1024, + "images": [ + "/poc/samples/fidelity-v3/bf16-dit/expanded-station-reunion-40.png", + "/poc/samples/fidelity-v3/bf16-dit/expanded-ridgeline-campaign-40.png" + ] + }, + { + "label": "expanded-multiref-portrait-40", + "prompt": "Create a natural lifestyle campaign photograph using both references. From the first image, use only the mother on the left: preserve her recognizable oval face, dark curly hair tied in a low bun, small round gold earrings, navy coat and mustard yellow scarf. From the second image, use the central cobalt-blue insulated bottle: preserve its tall shape, black screw cap, single narrow orange front stripe and small white mountain-triangle emblem near the base. Show this same woman seated alone on a wooden railway-platform bench, waist-up, smiling gently toward the camera while holding that bottle upright with both hands around its lower half; the cap, orange stripe and emblem should remain visible. Put one green suitcase on the ground beside the bench, partly visible in frame. Use soft overcast daylight and a stationary passenger train softly out of focus behind her. Realistic hands and fingers, recognizable identity, coherent scale and contact shadows. Do not include the other people from the first reference, the cups or pedestal from the second reference, or any poster lettering. The intended result is one coherent photograph with exactly one woman and one bottle.", + "steps": 40, + "seed": 21123, + "width": 1024, + "height": 1024, + "images": [ + "/poc/samples/fidelity-v3/bf16-dit/expanded-station-reunion-40.png", + "/poc/samples/fidelity-v3/bf16-dit/expanded-ridgeline-campaign-40.png" + ] + }, + { + "label": "expanded-product-relocation-edit-25", + "prompt": "Edit the arrangement of the small cups in this product campaign while preserving the central bottle and all typography. Remove the yellow cup that is currently on the left, restoring the tabletop and its lighting naturally. Move the turquoise cup from the right side to the far left of the pedestal, fully visible with a small clear gap from the pedestal; leave the former right-hand position empty. The final image must contain exactly two products: the unchanged cobalt-blue bottle on its original pedestal and the one turquoise cup on the left. Preserve the bottle's shape, cap, orange stripe, white mountain emblem, location and scale; preserve the pedestal, backdrop, framing, light direction, and the exact text \"TAKE THE LONG WAY\", \"RIDGELINE / EVERYDAY ADVENTURE\", and \"750 mL · BUILT TO REUSE\". Give the moved cup a physically consistent contact shadow. Do not introduce any extra objects or redesign the campaign.", + "steps": 25, + "seed": 21134, + "width": 1024, + "height": 1024, + "images": [ + "/poc/samples/fidelity-v3/bf16-dit/expanded-ridgeline-campaign-40.png" + ] + }, + { + "label": "expanded-product-relocation-edit-40", + "prompt": "Edit the arrangement of the small cups in this product campaign while preserving the central bottle and all typography. Remove the yellow cup that is currently on the left, restoring the tabletop and its lighting naturally. Move the turquoise cup from the right side to the far left of the pedestal, fully visible with a small clear gap from the pedestal; leave the former right-hand position empty. The final image must contain exactly two products: the unchanged cobalt-blue bottle on its original pedestal and the one turquoise cup on the left. Preserve the bottle's shape, cap, orange stripe, white mountain emblem, location and scale; preserve the pedestal, backdrop, framing, light direction, and the exact text \"TAKE THE LONG WAY\", \"RIDGELINE / EVERYDAY ADVENTURE\", and \"750 mL · BUILT TO REUSE\". Give the moved cup a physically consistent contact shadow. Do not introduce any extra objects or redesign the campaign.", + "steps": 40, + "seed": 21134, + "width": 1024, + "height": 1024, + "images": [ + "/poc/samples/fidelity-v3/bf16-dit/expanded-ridgeline-campaign-40.png" + ] + }, + { + "label": "editorial-25", + "prompt": "Candid editorial photograph in a sunlit pottery studio. Exactly two adults: an older woman with short silver hair on the left demonstrates shaping a clay bowl on a pottery wheel; a younger man with curly dark hair on the right watches and holds a folded blue towel with both hands. Their hands are anatomically natural and clearly visible, with clay on the woman's fingertips. Wooden shelves hold irregular ceramic vases, one trailing green plant hangs high in the back right, and soft morning light enters from a large window on the left. Waist-up environmental composition, documentary photography, natural skin texture, restrained warm colors, realistic depth of field. No lettering.", + "steps": 25, + "seed": 118, + "width": 1024, + "height": 1024 + }, + { + "label": "rainy-city-40", + "prompt": "Cinematic street-level architectural photograph of a narrow Tokyo side street at blue hour after rain. A tiny warmly lit ramen restaurant occupies the left foreground, with a striped navy awning and a vertical sign that reads \"RAMEN\". Exactly three bicycles are parked along the right wall. A person holding a transparent umbrella walks away at center, wearing a mustard yellow raincoat. Overhead wires cross between weathered three-story buildings, red lanterns glow farther down the street, and puddles reflect the signs. Strong one-point perspective, realistic glass and wet asphalt, fine distant details, balanced blue and amber lighting.", + "steps": 40, + "seed": 227, + "width": 1024, + "height": 1024 + }, + { + "label": "rainy-city-25", + "prompt": "Cinematic street-level architectural photograph of a narrow Tokyo side street at blue hour after rain. A tiny warmly lit ramen restaurant occupies the left foreground, with a striped navy awning and a vertical sign that reads \"RAMEN\". Exactly three bicycles are parked along the right wall. A person holding a transparent umbrella walks away at center, wearing a mustard yellow raincoat. Overhead wires cross between weathered three-story buildings, red lanterns glow farther down the street, and puddles reflect the signs. Strong one-point perspective, realistic glass and wet asphalt, fine distant details, balanced blue and amber lighting.", + "steps": 25, + "seed": 227, + "width": 1024, + "height": 1024 + }, + { + "label": "botanical-poster-25", + "prompt": "Sophisticated botanical exhibition poster on warm ivory paper, flat front-on view, crisp print design. At the top the large dark green serif headline reads exactly \"THE SECRET GARDEN\" on two centered lines. Under it a smaller line reads exactly \"BOTANICAL STUDIES / 2026\". The center contains a delicate detailed watercolor illustration of a lemon branch with exactly two yellow lemons, white blossoms and dark green leaves, surrounded by generous negative space. Along the bottom three evenly spaced text blocks read \"12 OCTOBER\", \"10 AM - 6 PM\", and \"GLASSHOUSE No. 4\". A thin rectangular green border surrounds the entire design. Elegant typography, no other text, no mockup, no shadows outside the paper.", + "steps": 25, + "seed": 336, + "width": 1024, + "height": 1024 + }, + { + "label": "botanical-poster-40", + "prompt": "Sophisticated botanical exhibition poster on warm ivory paper, flat front-on view, crisp print design. At the top the large dark green serif headline reads exactly \"THE SECRET GARDEN\" on two centered lines. Under it a smaller line reads exactly \"BOTANICAL STUDIES / 2026\". The center contains a delicate detailed watercolor illustration of a lemon branch with exactly two yellow lemons, white blossoms and dark green leaves, surrounded by generous negative space. Along the bottom three evenly spaced text blocks read \"12 OCTOBER\", \"10 AM - 6 PM\", and \"GLASSHOUSE No. 4\". A thin rectangular green border surrounds the entire design. Elegant typography, no other text, no mockup, no shadows outside the paper.", + "steps": 40, + "seed": 336, + "width": 1024, + "height": 1024 + }, + { + "label": "food-spatial-25", + "prompt": "Overhead high-end food photograph of a neatly arranged breakfast on a pale stone table. A large round white plate is centered. On the plate, exactly three pancakes form a vertical stack, topped by exactly five blueberries and a small square of butter. To the left of the plate is a silver fork, to the right is a silver knife with its blade facing the plate. A clear glass of orange juice sits at the upper right, a small white bowl containing sliced strawberries sits at the upper left, and a folded rust-colored linen napkin lies beneath the fork. A little maple syrup runs down only the right edge of the pancake stack. Soft natural light from the upper left, realistic food texture, all objects fully inside frame.", + "steps": 25, + "seed": 445, + "width": 1024, + "height": 1024 + }, + { + "label": "fantasy-cutaway-25", + "prompt": "Intricate isometric cutaway illustration of a cozy three-level library built inside an enormous ancient tree, in a hand-painted storybook style. Ground floor: a round green entrance door, a circular rug, and a sleeping orange cat beside a small wood stove. Middle floor: curved bookshelves filled with colorful books, a writing desk by an oval window, and a brass telescope pointing outside. Top floor: a glass-domed reading nook with two red armchairs and a hanging lantern. A single wooden spiral staircase visibly connects all three floors. Exposed roots wrap around mossy rocks, tiny mushrooms cluster near the door, and the leafy crown surrounds the dome without hiding it. Consistent isometric perspective, warm interior lighting, cool twilight forest background, detailed but readable, no text.", + "steps": 25, + "seed": 554, + "width": 1024, + "height": 1024 + }, + { + "label": "glass-product-40", + "prompt": "Luxury studio still life photograph on a polished dark green marble plinth. A rectangular clear glass perfume bottle half filled with pale amber liquid stands in the center, with a brushed gold cylindrical cap and a small cream label reading exactly \"LUMEN\". A thin curved strip of orange peel lies in front of the bottle. Behind and to the left is a translucent ribbed glass sphere; behind and to the right are two glossy dark green leaves. Hard afternoon sunlight from the upper left creates realistic caustics, transparent overlapping shadows, precise refraction through the bottle and a soft reflection on the marble. Deep emerald background, controlled commercial composition, realistic materials, no additional text.", + "steps": 40, + "seed": 663, + "width": 1024, + "height": 1024 + }, + { + "label": "glass-product-25", + "prompt": "Luxury studio still life photograph on a polished dark green marble plinth. A rectangular clear glass perfume bottle half filled with pale amber liquid stands in the center, with a brushed gold cylindrical cap and a small cream label reading exactly \"LUMEN\". A thin curved strip of orange peel lies in front of the bottle. Behind and to the left is a translucent ribbed glass sphere; behind and to the right are two glossy dark green leaves. Hard afternoon sunlight from the upper left creates realistic caustics, transparent overlapping shadows, precise refraction through the bottle and a soft reflection on the marble. Deep emerald background, controlled commercial composition, realistic materials, no additional text.", + "steps": 25, + "seed": 663, + "width": 1024, + "height": 1024 + }, + { + "label": "editorial-edit-25", + "prompt": "Replace only the blue towel held by the younger man with a bright red towel of the same size and folded shape. Preserve both people's identities, facial expressions, hands, clothing, clay bowl, studio shelves, window light, framing and photographic style.", + "steps": 25, + "seed": 774, + "width": 1024, + "height": 1024, + "images": [ + "/poc/samples/varied-v2/nf4/editorial-40.png" + ] + }, + { + "label": "editorial-edit-40", + "prompt": "Replace only the blue towel held by the younger man with a bright red towel of the same size and folded shape. Preserve both people's identities, facial expressions, hands, clothing, clay bowl, studio shelves, window light, framing and photographic style.", + "steps": 40, + "seed": 774, + "width": 1024, + "height": 1024, + "images": [ + "/poc/samples/varied-v2/nf4/editorial-40.png" + ] + }, + { + "label": "poster-edit-25", + "prompt": "Edit only the text at the top: replace \"THE SECRET GARDEN\" with \"THE LEMON HOUSE\". Keep the same dark green serif type style and centered placement. Preserve all other text exactly, the lemon branch illustration, paper background, border, spacing and overall design.", + "steps": 25, + "seed": 885, + "width": 1024, + "height": 1024, + "images": [ + "/poc/samples/varied-v2/nf4/botanical-poster-40.png" + ] + }, + { + "label": "poster-edit-40", + "prompt": "Edit only the text at the top: replace \"THE SECRET GARDEN\" with \"THE LEMON HOUSE\". Keep the same dark green serif type style and centered placement. Preserve all other text exactly, the lemon branch illustration, paper background, border, spacing and overall design.", + "steps": 40, + "seed": 885, + "width": 1024, + "height": 1024, + "images": [ + "/poc/samples/varied-v2/nf4/botanical-poster-40.png" + ] + }, + { + "label": "city-edit-25", + "prompt": "Transform this rainy Tokyo street photograph into a richly textured hand-painted watercolor illustration. Preserve the layout of the buildings, the ramen restaurant on the left, all three bicycles on the right, the central person with the transparent umbrella and yellow raincoat, the overhead wires, blue-hour lighting, and puddle reflections. Keep the RAMEN sign legible.", + "steps": 25, + "seed": 996, + "width": 1024, + "height": 1024, + "images": [ + "/poc/samples/varied-v2/nf4/rainy-city-40.png" + ] + }, + { + "label": "city-edit-40", + "prompt": "Transform this rainy Tokyo street photograph into a richly textured hand-painted watercolor illustration. Preserve the layout of the buildings, the ramen restaurant on the left, all three bicycles on the right, the central person with the transparent umbrella and yellow raincoat, the overhead wires, blue-hour lighting, and puddle reflections. Keep the RAMEN sign legible.", + "steps": 40, + "seed": 996, + "width": 1024, + "height": 1024, + "images": [ + "/poc/samples/varied-v2/nf4/rainy-city-40.png" + ] + } +] diff --git a/reproduction/lean_encoder.py b/reproduction/lean_encoder.py new file mode 100644 index 0000000000000000000000000000000000000000..e2c8e9208235f7c35959ee777a0517b4566a43dd --- /dev/null +++ b/reproduction/lean_encoder.py @@ -0,0 +1,163 @@ +"""Optional Qwen3-VL encoder-only adapter for the pinned Qwen Image 2.1 POC. + +Save a normal (including quantized) checkpoint BEFORE calling enable_lean_encoder. +The adapted instance no longer has the architecture's lm_head; do not pass it to +save_pretrained or pipeline.save_pretrained. Reload the normal checkpoint and +reapply this adapter at runtime. No decoder layers or normalization are changed. + +Typical use, after loading and before serving concurrent requests:: + + report = verify_encoder_parity(te, processor, prompt="A red teapot") + assert report["passed"], report + enable_lean_encoder(te) + +The parity helper is optional, runs two encoder forwards, and temporarily changes +forward dispatch. Run it only while the engine is idle. The application must +serialize enablement and parity checks with generation requests. +""" +from types import MethodType + + +def _base_model_forward(self, *args, **kwargs): + # Image generation reads features, never autoregressive vocabulary logits. + kwargs["use_cache"] = False + return self.model(*args, **kwargs) + + +def _dispatch_attribute(te): + # Accelerate wraps forward and calls _old_forward. Replace its inner function + # to preserve onload/offload hooks; future remove/reinstall cycles retain it. + return "_old_forward" if hasattr(te, "_hf_hook") and hasattr(te, "_old_forward") else "forward" + + +def enable_lean_encoder(te): + """Remove only lm_head and forward through the existing multimodal base model. + + Returns small metadata; mutates te in place and is idempotent. This preserves + te.model.language_model so Diffusers' final-RMSNorm hook keeps working. + Existing Accelerate CPU-offload hooks are retained. Save checkpoints first. + """ + if getattr(te, "_qwen21_lean_encoder", False): + return {"enabled": True, "already_enabled": True} + if not hasattr(te, "model") or not hasattr(te.model, "language_model"): + raise TypeError("Expected Qwen3VLForConditionalGeneration with model.language_model") + if not hasattr(te, "lm_head"): + raise ValueError("lm_head is missing; load an unmodified checkpoint before enabling this adapter") + removed_parameters = sum(p.numel() for p in te.lm_head.parameters()) + dispatch = _dispatch_attribute(te) + setattr(te, dispatch, MethodType(_base_model_forward, te)) + te.config.use_cache = False + te.config.text_config.use_cache = False + del te.lm_head + te._qwen21_lean_encoder = True + return { + "enabled": True, + "already_enabled": False, + "removed_parameters": removed_parameters, + "dispatch_attribute": dispatch, + "checkpoint_save_supported": False, + } + + +def _prepare_inputs(processor, prompt, images, device): + from PIL import Image + system = "Comprehend and analyze the provided prompt." + vision_images = [] + for img in images or []: + if not isinstance(img, Image.Image): + raise TypeError("Parity references must be PIL images, already resized as for inference") + if img.mode == "RGBA": + white = Image.new("RGB", img.size, (255, 255, 255)) + white.paste(img, mask=img.getchannel("A")) + img = white + vision_images.append(img) + prefix = " ".join( + f"<|vision_start|><|image_pad|><|vision_end|>" + for i in range(1, len(vision_images) + 1) + ) + text = ( + f"<|im_start|>system\n{system}<|im_end|>\n" + f"<|im_start|>user\n{prefix}{prompt or ' '}<|im_end|>\n" + "<|im_start|>assistant\n" + ) + kwargs = {"text": [text], "padding": True, "padding_side": "left", "return_tensors": "pt"} + if vision_images: + kwargs["images"] = vision_images + encoded = processor(**kwargs).to(device) + forward = {"input_ids": encoded.input_ids, "attention_mask": encoded.attention_mask} + for key in ("pixel_values", "image_grid_thw", "mm_token_type_ids"): + if key in encoded: + forward[key] = encoded[key] + return forward + + +def verify_encoder_parity(te, processor=None, prompt="A red teapot", images=None, + model_inputs=None, device=None, atol=0.0, rtol=0.0): + """Compare full final pre-norm features with/without vocabulary projection. + + Does not remove lm_head or permanently change dispatch/configuration. Supply + either model_inputs (prepared forward kwargs, including reference pixels) or + processor/prompt/optional PIL images. Images should already have the same + resize used by the pipeline. CPU-offloaded encoders use their hook's execution + device by default. Returns a parity report; enable only when passed is true. + + Exact equality is the default because both paths execute the same base model. + Tolerances can be supplied explicitly if the runtime is nondeterministic. + The comparison mimics the pinned pipeline's norm hook and output_hidden_states + request. It compares full features rather than only the prompt-trimmed suffix. + """ + import torch + if getattr(te, "_qwen21_lean_encoder", False) or not hasattr(te, "lm_head"): + raise ValueError("Verify parity before enabling lean encoding") + if not hasattr(te, "model") or not hasattr(te.model, "language_model"): + raise TypeError("Expected Qwen3VLForConditionalGeneration with model.language_model") + if device is None: + device = getattr(getattr(te, "_hf_hook", None), "execution_device", None) + if device is None: + device = next(te.parameters()).device + if model_inputs is None: + if processor is None: + raise ValueError("Provide processor or prepared model_inputs") + model_inputs = _prepare_inputs(processor, prompt, images, device) + else: + model_inputs = { + k: v.to(device) if isinstance(v, torch.Tensor) else v + for k, v in dict(model_inputs).items() + } + model_inputs = {**model_inputs, "use_cache": False, "output_hidden_states": True, "return_dict": True} + # These conditional-generation-only options do not belong to base-model input. + for key in ("labels", "logits_to_keep"): + model_inputs.pop(key, None) + dispatch = _dispatch_attribute(te) + original_forward = getattr(te, dispatch) + was_training = te.training + norm = te.model.language_model.norm + handle = norm.register_forward_hook(lambda module, args, output: args[0]) + try: + te.eval() + with torch.inference_mode(): + baseline = te(**model_inputs).hidden_states[-1].detach().float().cpu().clone() + setattr(te, dispatch, MethodType(_base_model_forward, te)) + candidate = te(**model_inputs).hidden_states[-1].detach().float().cpu() + finally: + setattr(te, dispatch, original_forward) + handle.remove() + if was_training: + te.train() + if baseline.shape != candidate.shape: + return {"passed": False, "reason": "shape_mismatch", "baseline_shape": list(baseline.shape), + "candidate_shape": list(candidate.shape)} + delta = candidate - baseline + finite = bool(torch.isfinite(baseline).all() and torch.isfinite(candidate).all()) + return { + "passed": finite and bool(torch.allclose(baseline, candidate, atol=atol, rtol=rtol)), + "finite": finite, + "shape": list(baseline.shape), + "max_abs_difference": float(delta.abs().max()), + "rms_difference": float(delta.square().mean().sqrt()), + "baseline_rms": float(baseline.square().mean().sqrt()), + "atol": atol, + "rtol": rtol, + "reference_count": len(images or []) if images is not None else None, + "dispatch_attribute": dispatch, + } diff --git a/reproduction/nunchaku_backend/README.md b/reproduction/nunchaku_backend/README.md new file mode 100644 index 0000000000000000000000000000000000000000..f4309273027f8e12bcb55b6b9104e7e82c4c6582 --- /dev/null +++ b/reproduction/nunchaku_backend/README.md @@ -0,0 +1,54 @@ +# Calibrated Qwen Image 2.1 Nunchaku backend + +The selected backend is **calibrated v3 rank 128**, using Nunchaku's actual signed INT4 residual weights/activations and BF16 low-rank correction. The pinned Diffusers Qwen2.1 model keeps attention, prefix KV caching, positional embeddings and boundary projections. Only its 224 block linears are replaced. This is a custom backend, not an official Nunchaku Qwen2.1 release or a near-lossless quality claim. + +The superseded NF4-trajectory absmax collector and one-pass whole-model exporter have been removed from this active package. Their unchanged source and the complete pre-cleanup backend snapshot are in [the dated archive](../archive/2026-09-21-superseded-implementations/README.md), with file hashes. Existing samples, reports and checkpoint manifests are preserved. + +## Runtime dependency boundary + +Loading a completed calibrated checkpoint requires only these project modules: + +- `__init__.py`: backend format and public loader. +- `runtime.py`: strict packed-checkpoint validation and Nunchaku layer replacement. +- `layout.py`: block-linear name pattern; no Torch import. + +Runtime also requires the pinned Torch, Diffusers, Transformers, Nunchaku and safetensors dependencies. It does **not** need DeepCompressor, the archived collector/exporter, or optimization modules. Optional Engine BF16 overrides continue to use `hybrid_v3.py` and its shared checkpoint helpers; that optional path was retained for compatibility and is not the selected rank-128 configuration. + +## Calibration and selected conversion recipe + +The final calibration collected original BF16 DiT activations on 40-step trajectories at steps **0, 10, 20, 30, 39**, with the NF4 text encoder held fixed. Disjoint train/validation/held-out jobs are in `experiments/fidelity-v3/calibration-jobs.json`. Collection saves per-layer activation files and provenance. GPU commands below must run in the isolated POC environment on the reserved non-Klein GPU. + +```bash +python -m nunchaku_backend.collect_v3 \ + --jobs /poc/experiments/fidelity-v3/calibration-jobs.json \ + --out /cache/qwen21-activation-v3-40 --steps 40 + +python -m nunchaku_backend.export_v3 \ + --model-path /cache/huggingface/hub/models--Qwen--Qwen-Image-2.1/snapshots/b3179ad355be050328e483a9dfdd9e60cd62adfa \ + --activations /cache/qwen21-activation-v3-40 \ + --baseline-calibration /cache/qwen21-calibration \ + --out /cache/qwen-nunchaku-v3-r128-reproduction \ + --ranks 128 --alphas 0.25 0.5 0.75 \ + --families activation_only smoothquant --weighting none \ + --iterations 16 --final-gptq --gptq-damp 0.01 \ + --factorization balanced --ridge 0.01 --baseline-rank 32 \ + --seed 1947 --niter 4 --oversample 16 --threads 8 \ + --device cuda:0 --objective-backend auto +``` + +Use the preserved baseline calibration artifact for exact historical reproduction: its SHA-256 is `7f032d0930881f8352309e10084b3e5e4418ac81fb1c69a9bc7457f5fa731018`. It influences the baseline candidate, even though new search candidates use BF16 activation rows. Without that artifact, omitting the flag is a new conversion recipe and must not be described as reproducing the saved checkpoint. The old collector can be run from a separate restored archive workspace when historical re-collection is needed; do not restore it into a deployment package. + +The final activation fingerprint is `be755525bef86b1be059e0b5667fb8f38305472a6775506b3363d9477f630b25`. It includes paths as well as archive contents, so preserve paths when comparing identity. The saved manifest in `results/nunchaku-v3-manifest.json.gz` is authoritative for settings. Collection/re-export is costly and was not rerun during cleanup; mathematical extraction equivalence and CPU tests verify this refactor, not bitwise GPU reproduction. + +## Active implementation + +- `collect_v3.py` / `teacher_v3.py`: bounded teacher activation capture. +- `export_v3.py` / `optimize_v3.py` / `gptq_v3.py`: activation-output scoring, smoothing search, alternating low-rank/residual refinement, output correction and optional final GPTQ. +- `packing.py`: pinned DeepCompressor packing adapter and group size. +- `baseline_candidate.py`: unchanged one-pass **internal fallback**. The final optimizer evaluates this candidate; deleting it would change the recipe. There is no active one-pass whole-checkpoint CLI. +- `checkpoint_io.py`: lazy safetensors reads and atomic JSON writes. +- `convert_activation_reference.py` / `convert_reference.py`: numerical kernel reference, retained for validation. + +DeepCompressor packing commit: `69f3473f5e1c1504bae35cc50c7858ef900a9b17`; Diffusers commit: `80c7ed262aeffbeb43ef13ae04baeb9b84515a69`. Conversion expects the pinned DeepCompressor checkout on `PYTHONPATH` (`/opt/deepcompressor` in the POC image). Loading does not. + +Rank-512 upgrade, mixed-BF16 and denoiser/parent-MLP probes remain optional research tools with their existing module paths so historical diagnostics and the Engine's optional overrides remain usable. They are not selected release defaults. See [the fidelity report](../experiments/fidelity-v3/REPORT.md) and [rank-512 visual review](../research/quality-v3-rank512-original-generation.md) for measured limitations. Cleanup does not change old source-hash manifests; new refinement manifests name the extracted modules. diff --git a/reproduction/nunchaku_backend/__init__.py b/reproduction/nunchaku_backend/__init__.py new file mode 100644 index 0000000000000000000000000000000000000000..50cdb99fc9455a960155ace1cacf35c7274e4537 --- /dev/null +++ b/reproduction/nunchaku_backend/__init__.py @@ -0,0 +1,11 @@ +"""Experimental Qwen Image2.1 SVDQuant/Nunchaku integration. + +This package does not alias NF4: its runtime uses packed signed INT4 weights, +INT4 activations, and a BF16 low-rank branch through Nunchaku CUDA kernels. +""" +FORMAT = "qwen21-nunchaku-svdq-int4-v1" + + +def load_transformer(*args, **kwargs): + from .runtime import load_transformer as implementation + return implementation(*args, **kwargs) diff --git a/reproduction/nunchaku_backend/baseline_candidate.py b/reproduction/nunchaku_backend/baseline_candidate.py new file mode 100644 index 0000000000000000000000000000000000000000..40a6fbdad2026286d326d459fa2a38797ffa69fe --- /dev/null +++ b/reproduction/nunchaku_backend/baseline_candidate.py @@ -0,0 +1,265 @@ +"""One-pass baseline candidate retained for faithful calibrated-v3 reproduction. + +This is an internal fallback evaluated by optimize_v3, not a supported standalone +whole-model exporter. The superseded converter and CLI are in the dated archive. +""" +from __future__ import annotations + +import time +from typing import Any +import torch +from .packing import DEEPCOMPRESSOR_COMMIT, GROUP_SIZE, _upstream_converter + +def _working_float(tensor: torch.Tensor, name: str, device: torch.device) -> torch.Tensor: + if not tensor.is_floating_point(): + raise ValueError(f"{name} must be floating point") + result = tensor.detach().to(device=device, dtype=torch.float32) + if not torch.isfinite(result).all(): + raise ValueError(f"{name} contains nonfinite values") + return result + + +def _smoothing( + weight: torch.Tensor, + smooth: torch.Tensor | None, + input_absmax: torch.Tensor | None, + alpha: float, + dtype: torch.dtype, +) -> tuple[torch.Tensor, str]: + if smooth is not None and input_absmax is not None: + raise ValueError("Pass smooth or input_absmax, not both") + if not 0 <= alpha <= 1: + raise ValueError("smooth_alpha must be in [0, 1]") + ic = weight.shape[1] + if smooth is not None: + factor = _working_float(smooth, "smooth", weight.device) + if factor.shape != (ic,) or not (factor > 0).all(): + raise ValueError("smooth must be a positive vector of length in_features") + source = "caller_provided" + elif input_absmax is not None: + amax = _working_float(input_absmax, "input_absmax", weight.device) + if amax.shape != (ic,) or not (amax >= 0).all(): + raise ValueError("input_absmax must be a nonnegative vector of length in_features") + # SmoothQuant-style starting point, without per-layer alpha search. + wmax = weight.abs().amax(dim=0).clamp_min(1e-5) + factor = (amax.clamp_min(1e-5).pow(alpha) / wmax.pow(1 - alpha)).clamp(1e-4, 1e4) + source = "activation_absmax" + else: + factor = torch.ones(ic, dtype=torch.float32, device=weight.device) + source = "identity_uncalibrated" + # Decompose using the exact factors that the kernel will load. + factor = factor.to(dtype).float() + if not torch.isfinite(factor).all() or not (factor > 0).all(): + raise ValueError("smooth cannot be represented as positive finite values in parameter dtype") + return factor, source + + +@torch.no_grad() +def convert_linear_weight( + weight: torch.Tensor, + bias: torch.Tensor | None = None, + rank: int = 32, + *, + smooth: torch.Tensor | None = None, + input_absmax: torch.Tensor | None = None, + smooth_alpha: float = 0.5, + svd_method: str = "randomized", + seed: int = 0, + niter: int = 4, + oversample: int = 16, + torch_dtype: torch.dtype = torch.bfloat16, + return_reference: bool = False, + conversion_device: str | torch.device = "cpu", +) -> tuple: + """Return packed ``SVDQW4A4Linear`` state and JSON-compatible statistics. + + Input/output widths must be multiples of 128; rank must be a positive + multiple of 16. Weight shape is [out_features, in_features]. No source + tensors are mutated. Smoothing has the convention X/s and W*s; the + low-rank down matrix is divided by s because that kernel branch sees X. + + ``input_absmax`` is the per-input-channel maximum absolute activation + collected from representative denoising calls. No activations means an + explicitly uncalibrated conversion. Provided statistics do not establish + representative coverage or image quality. ``return_reference=True`` adds + a third dictionary of unpacked tensors for numerical kernel probing. + Conversion runs on CPU unless ``conversion_device='cuda:0'`` is explicitly + selected. All returned tensors remain on the conversion device. + """ + start = time.perf_counter() + pack = _upstream_converter() + if torch_dtype not in (torch.bfloat16, torch.float16): + raise ValueError("torch_dtype must be bfloat16 or float16") + device = torch.device(conversion_device) + if device.type not in ("cpu", "cuda"): + raise ValueError("conversion_device must be cpu or an explicit CUDA device") + if device.type == "cuda" and device.index is None: + raise ValueError("Specify a CUDA device index, e.g. conversion_device='cuda:0'") + original = _working_float(weight, "weight", device) + if original.ndim != 2: + raise ValueError("weight must be two-dimensional") + oc, ic = original.shape + if not oc or not ic or oc % 128 or ic % 128: + raise ValueError("in_features and out_features must be positive multiples of 128") + if rank <= 0 or rank % 16 or rank > min(oc, ic): + raise ValueError("rank must be a positive multiple of 16, at most min(in_features, out_features)") + if niter < 0 or oversample < 0: + raise ValueError("niter and oversample must be nonnegative") + if bias is not None: + bias = _working_float(bias, "bias", device) + if bias.shape != (oc,): + raise ValueError("bias must have shape [out_features]") + bias = bias.to(torch_dtype) + if not torch.isfinite(bias).all(): + raise ValueError("bias overflows parameter dtype") + + factor, source = _smoothing(original, smooth, input_absmax, smooth_alpha, torch_dtype) + smoothed = original * factor.unsqueeze(0) + if not torch.isfinite(smoothed).all(): + raise ValueError("smoothed weight overflowed") + if svd_method == "full": + u, singular, vh = torch.linalg.svd(smoothed, full_matrices=False) + u, singular, v = u[:, :rank], singular[:rank], vh[:rank, :].T + elif svd_method == "randomized": + # CPU mode never touches CUDA generators. Explicit CUDA conversion + # saves/restores only that device's RNG, plus the CPU stream. + with torch.random.fork_rng(devices=[device.index] if device.type == "cuda" else []): + torch.random.default_generator.manual_seed(seed) + if device.type == "cuda": + torch.cuda.default_generators[device.index].manual_seed(seed) + u, singular, v = torch.svd_lowrank(smoothed, q=min(rank + oversample, oc, ic), niter=niter) + u, singular, v = u[:, :rank], singular[:rank], v[:, :rank] + else: + raise ValueError("svd_method must be 'randomized' or 'full'") + + root = singular.clamp_min(0).sqrt() + up = (u * root.unsqueeze(0)).to(torch_dtype) + down_smoothed = (root.unsqueeze(1) * v.T).to(torch_dtype) + down = (down_smoothed.float() / factor.unsqueeze(0)).to(torch_dtype) + if not torch.isfinite(up).all() or not torch.isfinite(down).all(): + raise ValueError("low-rank factors overflow parameter dtype") + # Include the final inverse-smoothing BF16 rounding in the residual too. + # The exported low-rank branch represents up@down in original coordinates. + residual = smoothed - up.float() @ (down.float() * factor.unsqueeze(0)) + + grouped = residual.reshape(oc, ic // GROUP_SIZE, GROUP_SIZE) + scales = (grouped.abs().amax(dim=-1) / 7).clamp_min(torch.finfo(torch_dtype).tiny).to(torch_dtype) + if not torch.isfinite(scales).all(): + raise ValueError("residual scales overflow parameter dtype") + quantized = (grouped / scales.float().unsqueeze(-1)).round().clamp(-7, 7) + dequantized = (quantized * scales.float().unsqueeze(-1)).reshape(oc, ic).to(torch_dtype) + # Upstream handles all tensor-core tile permutations, nibble packing, + # packed scales/bias/smoothing, and separate low-rank matrix layouts. + qweight, wscales, packed_bias, packed_smooth, lora, subscale = pack( + dequantized, + scales.reshape(oc, 1, ic // GROUP_SIZE, 1), + bias=bias, + smooth=factor.to(torch_dtype), + lora=(down, up), + float_point=False, + ) + assert lora is not None and subscale is None + state = { + "qweight": qweight.contiguous(), + "wscales": wscales.contiguous(), + "smooth_factor": packed_smooth.contiguous(), + "smooth_factor_orig": packed_smooth.clone().contiguous(), + "proj_down": lora[0].contiguous(), + "proj_up": lora[1].contiguous(), + } + if bias is not None: + state["bias"] = packed_bias.contiguous() + + # Representation error uses integer*stored_scale, matching the kernel's + # weight values (the dequantized BF16 intermediary is only for export). + error_sq, weight_sq = 0.0, 0.0 + for offset in range(0, oc, 256): + sl = slice(offset, offset + 256) + recovered = (quantized[sl] * scales[sl].float().unsqueeze(-1)).reshape(-1, ic) + recovered = recovered / factor + up[sl].float() @ down.float() + error_sq += (recovered - original[sl]).double().square().sum().item() + weight_sq += original[sl].double().square().sum().item() + stats = { + "algorithm": "one_pass_smoothed_svd_signed_int4", + "deepcompressor_commit": DEEPCOMPRESSOR_COMMIT, + "in_features": ic, + "out_features": oc, + "rank": rank, + "group_size": GROUP_SIZE, + "precision": "int4", + "torch_dtype": str(torch_dtype), + "conversion_device": str(device), + "smoothing_source": source, + "calibrated": input_absmax is not None, + "caller_smoothing_provided": smooth is not None, + "smooth_alpha": smooth_alpha if input_absmax is not None else None, + "smooth_min": factor.min().item(), + "smooth_max": factor.max().item(), + "svd_method": svd_method, + "seed": seed, + "niter": niter, + "oversample": oversample, + "weight_relative_l2_error": (error_sq / weight_sq) ** 0.5 if weight_sq else 0.0, + "packed_bytes": sum(t.numel() * t.element_size() for t in state.values()), + "conversion_seconds": time.perf_counter() - start, + "limitations": "No activation rounding simulation, alpha search, GPTQ, iterative residual fit, or image quality guarantee.", + } + if return_reference: + reference = { + "residual_dequant": (quantized * scales.float().unsqueeze(-1)).reshape(oc, ic).contiguous(), + "down_unpacked": down.contiguous(), + "up_unpacked": up.contiguous(), + "smooth": factor.to(torch_dtype).contiguous(), + } + if bias is not None: + reference["bias"] = bias.contiguous() + return state, stats, reference + return state, stats + + +def convert_linear(linear: torch.nn.Linear, **kwargs): + """Convenience wrapper; the original module is left unchanged.""" + return convert_linear_weight(linear.weight, linear.bias, **kwargs) + + +@torch.no_grad() +def reference_forward( + x: torch.Tensor, + reference: dict[str, torch.Tensor], + *, + quantize_activations: bool = True, +) -> torch.Tensor: + """FP32 numerical reference using exported weights and kernel A4 rules. + + Smoothing rounds to input BF16/FP16; per-token group64 activation scales + are calculated in FP32. Rounding to integer uses the *unrounded* FP32 + scale, while dequantization uses the BF16/FP16 stored scale. The low-rank + branch consumes original input. Residual, bias and low-rank activation + rounding follow kernel epilogue boundaries. Returns compute-dtype-rounded + values in FP32. CUDA approximate reciprocals and accumulation order are + not bit-exact here. All-zero activation groups dequantize to zero. + """ + if x.dtype not in (torch.bfloat16, torch.float16): + raise ValueError("reference input must match the kernel BF16/FP16 input dtype") + dtype = x.dtype + smooth = reference["smooth"].to(device=x.device, dtype=torch.float32) + original = x.float() + smoothed = (original / smooth).to(dtype).float() + if smoothed.shape[-1] % GROUP_SIZE: + raise ValueError("input width must be divisible by 64") + if quantize_activations: + grouped = smoothed.reshape(*smoothed.shape[:-1], smoothed.shape[-1] // GROUP_SIZE, GROUP_SIZE) + # Mirrors gemm_w4a4.cuh: max*(1/7), FP32 reciprocal for rounding, + # cvt.rni (ties-to-even), signed saturation; store scales in BF16. + scale = grouped.abs().amax(-1, keepdim=True) * (1.0 / 7.0) + reciprocal = torch.where(scale > 0, scale.reciprocal(), torch.zeros_like(scale)) + q = (grouped * reciprocal).round().clamp(-8, 7) + smoothed = (q * scale.to(dtype).float()).reshape_as(smoothed) + residual = reference["residual_dequant"].to(device=x.device, dtype=torch.float32) + down = reference["down_unpacked"].to(device=x.device, dtype=torch.float32) + up = reference["up_unpacked"].to(device=x.device, dtype=torch.float32) + output = (smoothed @ residual.T).to(dtype).float() + if "bias" in reference: + output = (output + reference["bias"].to(device=x.device, dtype=torch.float32)).to(dtype).float() + hidden = (original @ down.T).to(dtype).float() + return (output + hidden @ up.T).to(dtype).float() diff --git a/reproduction/nunchaku_backend/baseline_candidate_test.py b/reproduction/nunchaku_backend/baseline_candidate_test.py new file mode 100644 index 0000000000000000000000000000000000000000..7d8d389ba2a06e7b15ea090a3537fda9979d30d9 --- /dev/null +++ b/reproduction/nunchaku_backend/baseline_candidate_test.py @@ -0,0 +1,141 @@ +"""CPU layout, numerical, smoothing and validation tests (no Nunchaku import).""" + +import json +import unittest + +import torch +from deepcompressor.backend.nunchaku.utils import NunchakuWeightPacker + +if __package__: + from .baseline_candidate import convert_linear_weight, reference_forward +else: + from baseline_candidate import convert_linear_weight, reference_forward + + +def unpack_qweight(packed): + """Invert upstream's documented INT4 warp128 memory order independently.""" + oc, half_ic = packed.shape + ic = half_ic * 2 + words = packed.contiguous().view(torch.int32) + q = ((words.unsqueeze(-1) >> torch.arange(0, 32, 4, dtype=torch.int32)) & 15) + q = q.reshape(oc // 128, ic // 64, 1, 8, 8, 4, 2, 2, 1, 8) + q = q.permute(0, 3, 6, 4, 8, 1, 2, 7, 5, 9).contiguous().reshape(oc, ic) + return torch.where(q >= 8, q - 16, q) + + +def unpack_scale(packed, oc): + groups = packed.numel() // oc + scale = packed.reshape(oc // 128, groups, 1, 8, 4, 2, 2) + return scale.permute(0, 2, 3, 5, 4, 6, 1).contiguous().reshape(oc, groups) + + +def unpack_state(state): + oc, half_ic = state["qweight"].shape + ic = half_ic * 2 + q = unpack_qweight(state["qweight"]) + scales = unpack_scale(state["wscales"], oc).float() + residual = (q.reshape(oc, ic // 64, 64).float() * scales.unsqueeze(-1)).reshape(oc, ic) + smooth = unpack_scale(state["smooth_factor"], ic).reshape(ic).float() + packer = NunchakuWeightPacker(4) + down = packer.unpack_lowrank_weight(state["proj_down"], down=True).float() + up = packer.unpack_lowrank_weight(state["proj_up"], down=False).float() + return residual, smooth, down, up + + +class ConversionTests(unittest.TestCase): + def setUp(self): + torch.manual_seed(21) + torch.set_num_threads(2) + + def test_layout_roundtrip_signed_nibbles_and_scales(self): + packer = NunchakuWeightPacker(4) + original = torch.arange(256 * 384, dtype=torch.int32).reshape(256, 384) % 16 - 8 + packed = packer.pack_weight(original.clone()) + self.assertTrue(torch.equal(unpack_qweight(packed), original)) + scales = torch.arange(256 * 6).reshape(256, 1, 6, 1).to(torch.bfloat16) + self.assertTrue(torch.equal(unpack_scale(packer.pack_scale(scales, 64), 256), scales.reshape(256, 6))) + + def test_rank32_reconstructs_dominant_lowrank_better_than_rtn(self): + weight = (torch.randn(256, 16) @ torch.randn(16, 256) + torch.randn(256, 256) * 0.02).bfloat16() + original = weight.clone() + bias = torch.randn(256).bfloat16() + state, stats = convert_linear_weight(weight, bias, rank=32) + residual, smooth, down, up = unpack_state(state) + recovered = residual / smooth + up @ down + error = (recovered - weight.float()).norm() / weight.float().norm() + grouped = weight.float().reshape(256, 4, 64) + scale = grouped.abs().amax(-1, keepdim=True) / 7 + rtn = ((grouped / scale).round() * scale).reshape_as(weight) + rtn_error = (rtn - weight.float()).norm() / weight.float().norm() + self.assertLess(error, rtn_error * 0.15) + self.assertAlmostEqual(error.item(), stats["weight_relative_l2_error"], places=6) + self.assertTrue(torch.equal(weight, original)) + self.assertTrue(torch.equal(unpack_scale(state["bias"], 256).flatten(), bias)) + self.assertFalse(stats["calibrated"]) + self.assertEqual(state["wscales"].shape, (4, 256)) + self.assertEqual(state["proj_down"].shape, (256, 32)) + json.dumps(stats, allow_nan=False) + + def test_nonuniform_smoothing_uses_original_input_for_lowrank(self): + weight = (torch.randn(128, 8) @ torch.randn(8, 256)).bfloat16() + smooth_input = torch.logspace(-1, 1, 256) + state, stats = convert_linear_weight(weight, smooth=smooth_input, rank=32, svd_method="full") + residual, smooth, down, up = unpack_state(state) + x = torch.randn(17, 256) + expected = x @ weight.float().T + actual = (x / smooth) @ residual.T + (x @ down.T) @ up.T + error = (actual - expected).norm() / expected.norm() + self.assertLess(error, 0.015) + # A regression to smoothing the lowrank branch must be detectable. + wrong = (x / smooth) @ residual.T + ((x / smooth) @ down.T) @ up.T + self.assertGreater(((wrong - expected).norm() / expected.norm()).item(), 0.5) + self.assertFalse(stats["calibrated"]) + self.assertTrue(stats["caller_smoothing_provided"]) + + def test_activation_statistics_and_rng_are_preserved(self): + weight = torch.randn(128, 256).bfloat16() + absmax = torch.logspace(-2, 2, 256) + rng = torch.random.get_rng_state().clone() + state, stats = convert_linear_weight(weight, input_absmax=absmax, seed=187) + self.assertTrue(torch.equal(torch.random.get_rng_state(), rng)) + self.assertEqual(stats["smoothing_source"], "activation_absmax") + expected = (absmax.sqrt() / weight.float().abs().amax(0).sqrt()).clamp(1e-4, 1e4).bfloat16().float() + self.assertTrue(torch.equal(unpack_scale(state["smooth_factor"], 256).flatten().float(), expected)) + self.assertNotIn("bias", state) + + def test_zero_weight_and_fail_closed_shapes(self): + weight = torch.zeros(128, 128, dtype=torch.bfloat16) + state, stats = convert_linear_weight(weight) + self.assertEqual(stats["weight_relative_l2_error"], 0) + self.assertTrue((state["qweight"] == 0).all()) + for kwargs in ({"rank": 17}, {"smooth": torch.zeros(128)}, {"input_absmax": -torch.ones(128)}, {"smooth_alpha": 1.1}): + with self.assertRaises(ValueError): + convert_linear_weight(weight, **kwargs) + with self.assertRaises(ValueError): + convert_linear_weight(torch.zeros(127, 128)) + with self.assertRaises(ValueError): + convert_linear_weight(torch.full((128, 128), float("nan"))) + + def test_reference_tensors_match_packed_export(self): + weight = torch.randn(128, 256).bfloat16() + bias = torch.randn(128).bfloat16() + state, stats, reference = convert_linear_weight(weight, bias, smooth=torch.logspace(-1, 1, 256), return_reference=True) + residual, smooth, down, up = unpack_state(state) + self.assertTrue(torch.equal(residual, reference["residual_dequant"])) + self.assertTrue(torch.equal(down, reference["down_unpacked"].float())) + self.assertTrue(torch.equal(up, reference["up_unpacked"].float())) + x = torch.randn(1, 17, 256).bfloat16() + base = ((x.float() / smooth).bfloat16().float() @ residual.T).bfloat16().float() + base = (base + bias.float()).bfloat16().float() + hidden = (x.float() @ down.T).bfloat16().float() + expected = (base + hidden @ up.T).bfloat16().float() + self.assertTrue(torch.equal(reference_forward(x, reference, quantize_activations=False), expected)) + actual = reference_forward(x, reference) + self.assertTrue(torch.isfinite(actual).all()) + self.assertEqual(actual.shape, (1, 17, 128)) + zeros = reference_forward(torch.zeros_like(x), reference) + self.assertTrue(torch.equal(zeros, bias.float().expand_as(zeros))) + + +if __name__ == "__main__": + unittest.main(verbosity=2) diff --git a/reproduction/nunchaku_backend/cache_repeat_diagnostic.py b/reproduction/nunchaku_backend/cache_repeat_diagnostic.py new file mode 100644 index 0000000000000000000000000000000000000000..c832df211a8ca93ff50f9075064735de1fdc374f --- /dev/null +++ b/reproduction/nunchaku_backend/cache_repeat_diagnostic.py @@ -0,0 +1,146 @@ +"""Four controlled generation runs; instrumentation only, no runtime edits.""" +import argparse +import hashlib +import json +from pathlib import Path +import time + + +def tensor_fingerprint(tensor): + if tensor is None: + return None + import torch + return {"shape": list(tensor.shape), "stride": list(tensor.stride()), + "dtype": str(tensor.dtype), "device": str(tensor.device), + "sha256": hashlib.sha256(tensor.detach().contiguous().cpu().view(torch.uint8).numpy().tobytes()).hexdigest()} + + +def prompt_cache_fingerprint(engine): + return [{"key_sha256": hashlib.sha256(repr(key).encode()).hexdigest(), + "tensors": [tensor_fingerprint(tensor) for tensor in value]} + for key, value in engine.prompt_cache.items()] + + +def pixel_comparison(first, second): + import numpy as np + from PIL import Image + with Image.open(first) as image: + a = np.array(image.convert("RGBA")) + with Image.open(second) as image: + b = np.array(image.convert("RGBA")) + if a.shape != b.shape: + raise ValueError("Repeated image shape changed") + error = a.astype(np.float64) - b.astype(np.float64) + rgb = error[:, :, :3] + return {"first": str(first), "second": str(second), + "encoded_bytes_equal": Path(first).read_bytes() == Path(second).read_bytes(), + "pixels_equal": bool(np.array_equal(a, b)), + "rgb_mae_255": float(np.abs(rgb).mean()), "rgb_rmse_255": float(np.mean(rgb ** 2) ** .5), + "max_rgb_difference_255": int(np.abs(rgb).max()), + "changed_rgb_pixels_fraction": float(np.any(rgb != 0, axis=-1).mean()), + "changed_alpha_pixels": int(np.count_nonzero(error[:, :, 3]))} + + +def compare_fingerprints(a, b): + if a is None or b is None: + return a is None and b is None + return a["shape"] == b["shape"] and a["dtype"] == b["dtype"] and a["sha256"] == b["sha256"] + + +def main(): + parser = argparse.ArgumentParser(description=__doc__) + parser.add_argument("--checkpoint", default="/cache/qwen-nunchaku-v3-r128") + parser.add_argument("--prequant", default="/cache/qwen-nf4") + parser.add_argument("--jobs", type=Path, default=Path("/poc/experiments/fidelity-v3/quality-jobs-all.json")) + parser.add_argument("--job", default="expanded-bilingual-festival-25") + parser.add_argument("--out", type=Path, default=Path("/poc/results/repeatability-diagnostic")) + parser.add_argument("--dry-run", action="store_true") + args = parser.parse_args() + from runner import Engine, _parse_cli_args + import torch + jobs = {job["label"]: job for job in json.loads(args.jobs.read_text())} + job = {**jobs[args.job], "cfg": 1.0, "kv_cache": True} + if job.get("images"): + raise ValueError("This controlled diagnosis intentionally isolates text-to-image prompt caching") + if args.out.exists(): + raise ValueError("Use a new diagnostic output directory; prior evidence must be preserved") + engine_args = _parse_cli_args(["--backend", "nunchaku", "--nunchaku-checkpoint", args.checkpoint, + "--prequant", args.prequant, "--sample-dir", str(args.out / "samples"), + "--cache", "--no-tiling", "--lean-encoder", "--stage-offload", "--release-kv"]) + if engine_args.restore_roles or engine_args.bf16_source: + raise ValueError("Remove inherited BF16 restoration environment for this exact rank128 diagnosis") + phases = [("no-cache-1", False), ("no-cache-2", False), ("cache-miss", True), ("cache-hit", True)] + if args.dry_run: + print(json.dumps({"dry_run": True, "job": job, "engine_args": vars(engine_args), "phases": phases, + "cuda_initialized": torch.cuda.is_initialized(), "gpu_work_executed": False})) + return + args.out.mkdir(parents=True) + report_path = args.out / "diagnostic.json" + report = {"complete": False, "job": job, "engine_args": vars(engine_args), "runs": [], + "timing_limitation": "Tensor hashing adds CPU/GPU transfers and synchronization; diagnostic times are not benchmark latencies", + "selection": "two no-cache runs then cache-miss and cache-hit; one Engine, identical explicit generation seed"} + def save(): + temporary = report_path.with_suffix(".tmp") + temporary.write_text(json.dumps(report, indent=2) + "\n") + temporary.replace(report_path) + save() + engine = Engine(engine_args) + current = {} + original_prompt = engine.pipe._get_qwen_prompt_embeds + def prompt_wrapper(*positional, **kwargs): + outputs = original_prompt(*positional, **kwargs) + current.setdefault("prompt_calls", []).append([tensor_fingerprint(tensor) for tensor in outputs]) + return outputs + engine.pipe._get_qwen_prompt_embeds = prompt_wrapper + def before_transformer(module, positional, kwargs): + current["transformer_calls"] = current.get("transformer_calls", 0) + 1 + if current["transformer_calls"] == 1: + current["first_transformer_inputs"] = {name: tensor_fingerprint(kwargs.get(name)) for name in + ("hidden_states", "encoder_hidden_states", "encoder_hidden_states_mask", "img_mask", "timestep")} + current["first_transformer_metadata"] = {"img_shapes": kwargs.get("img_shapes"), "kv_cache_mode": kwargs.get("kv_cache_mode")} + def after_transformer(module, positional, output): + if current["transformer_calls"] == 1: + value = output[0] if isinstance(output, (tuple, list)) else output.sample + current["first_transformer_output"] = tensor_fingerprint(value) + pre = engine.pipe.transformer.register_forward_pre_hook(before_transformer, with_kwargs=True) + post = engine.pipe.transformer.register_forward_hook(after_transformer) + try: + for phase, enabled in phases: + engine.args.cache = enabled + if phase != "cache-hit": + engine.prompt_cache.clear() + engine.vae_cache.clear() + current = {"phase": phase, "cache_enabled": enabled, "cache_before": prompt_cache_fingerprint(engine)} + started = time.perf_counter() + metrics = engine.generate({**job, "label": "repeatability-" + phase}) + current["cache_after"] = prompt_cache_fingerprint(engine) + path = Path(metrics["path"]) + current.update(metrics=json.loads(json.dumps(metrics)), path=str(path), sha256=hashlib.sha256(path.read_bytes()).hexdigest(), + wall_seconds=time.perf_counter() - started) + report["runs"].append(current) + save() + print(json.dumps({"phase": phase, "sha256": current["sha256"], "prompt_cache_hits": metrics.get("prompt_cache_hits", 0)}), flush=True) + comparisons = [] + for first, second in ((0, 1), (1, 2), (2, 3), (0, 3)): + a, b = report["runs"][first], report["runs"][second] + row = pixel_comparison(a["path"], b["path"]) + row.update(phases=[a["phase"], b["phase"]], + prompt_tensor_equal=[compare_fingerprints(x, y) for x, y in zip(a["prompt_calls"][0], b["prompt_calls"][0])], + first_transformer_input_equal={name: compare_fingerprints(a["first_transformer_inputs"][name], b["first_transformer_inputs"][name]) for name in a["first_transformer_inputs"]}, + first_transformer_output_equal=compare_fingerprints(a["first_transformer_output"], b["first_transformer_output"])) + comparisons.append(row) + report["comparisons"] = comparisons + report["cache_hit_stored_tensors_unchanged"] = report["runs"][3]["cache_before"] == report["runs"][3]["cache_after"] + report["cache_miss_after_matches_hit_before"] = report["runs"][2]["cache_after"] == report["runs"][3]["cache_before"] + report["complete"] = True + save() + print(json.dumps({"complete": True, "comparisons": comparisons, + "cache_hit_stored_tensors_unchanged": report["cache_hit_stored_tensors_unchanged"]}), flush=True) + finally: + pre.remove() + post.remove() + engine.pipe._get_qwen_prompt_embeds = original_prompt + + +if __name__ == "__main__": + main() diff --git a/reproduction/nunchaku_backend/checkpoint_io.py b/reproduction/nunchaku_backend/checkpoint_io.py new file mode 100644 index 0000000000000000000000000000000000000000..36a2fcbe9f3d1d19dc50392a823c31369d2b84d5 --- /dev/null +++ b/reproduction/nunchaku_backend/checkpoint_io.py @@ -0,0 +1,28 @@ +"""Shared streaming checkpoint I/O; no conversion CLI or eager Torch import.""" +import json + +def _json_write(path, content): + temporary = path.with_name(path.name + ".tmp") + temporary.write_text(json.dumps(content, indent=2, sort_keys=True) + "\n") + temporary.replace(path) + + +def _source_index(directory): + from safetensors import safe_open + result = {} + for path in sorted(directory.glob("*.safetensors")): + with safe_open(str(path), framework="pt", device="cpu") as handle: + for key in handle.keys(): + if key in result: + raise ValueError(f"Duplicate source tensor {key}; provide one model variant only") + result[key] = path + if not result: + raise ValueError(f"No source safetensors in {directory}") + return result + + +def _read_tensor(index, key): + from safetensors import safe_open + with safe_open(str(index[key]), framework="pt", device="cpu") as handle: + return handle.get_tensor(key) + diff --git a/reproduction/nunchaku_backend/cleanup_test.py b/reproduction/nunchaku_backend/cleanup_test.py new file mode 100644 index 0000000000000000000000000000000000000000..9d66f21fbcc23de2a27679cfeb928802e8a018bf --- /dev/null +++ b/reproduction/nunchaku_backend/cleanup_test.py @@ -0,0 +1,74 @@ +"""CPU-only safety checks for archived CLI removal and runtime packaging.""" +import ast +import hashlib +import json +from pathlib import Path +import shutil +import subprocess +import sys +import tempfile +import unittest + + +PACKAGE = Path(__file__).resolve().parent +ARCHIVE = PACKAGE.parent / "archive/2026-09-21-superseded-implementations" + + +def functions(path): + return {n.name: ast.dump(n, include_attributes=False) for n in ast.parse(path.read_text()).body + if isinstance(n, (ast.FunctionDef, ast.AsyncFunctionDef))} + + +class CleanupTests(unittest.TestCase): + def test_historical_sources_are_byte_for_byte_preserved(self): + hashes = json.loads((ARCHIVE / "SHA256SUMS.json").read_text()) + self.assertGreaterEqual(len(hashes), 50) + for name, expected in hashes.items(): + with self.subTest(name=name): + self.assertEqual(hashlib.sha256((ARCHIVE / name).read_bytes()).hexdigest(), expected) + + def test_candidate_packing_and_io_math_unchanged(self): + old = ARCHIVE / "nunchaku_backend" + original = functions(old / "convert.py") + extracted = functions(PACKAGE / "baseline_candidate.py") | functions(PACKAGE / "packing.py") + self.assertEqual(original, extracted) + old_io = functions(old / "export_checkpoint.py") + new_io = functions(PACKAGE / "checkpoint_io.py") + self.assertEqual(set(new_io), {"_json_write", "_source_index", "_read_tensor"}) + for name in new_io: + self.assertEqual(old_io[name], new_io[name]) + + def test_active_relative_imports_resolve_without_archived_modules(self): + forbidden = {"convert", "export_checkpoint", "calibration", "calibrate"} + for path in PACKAGE.glob("*.py"): + for node in ast.walk(ast.parse(path.read_text())): + if isinstance(node, ast.ImportFrom) and node.level and node.module: + first = node.module.split(".")[0] + with self.subTest(file=path.name, dependency=first): + self.assertNotIn(first, forbidden) + self.assertTrue((PACKAGE / (first + ".py")).is_file() or (PACKAGE / first).is_dir()) + for name in forbidden: + self.assertFalse((PACKAGE / (name + ".py")).exists()) + + def test_runtime_imports_with_only_three_release_files(self): + with tempfile.TemporaryDirectory() as directory: + target = Path(directory) / "nunchaku_backend" + target.mkdir() + for name in ("__init__.py", "runtime.py", "layout.py"): + shutil.copyfile(PACKAGE / name, target / name) + code = ("import sys; sys.path.insert(0, " + repr(directory) + "); " + "import nunchaku_backend.runtime; from nunchaku_backend.layout import BLOCK_LINEAR_PATTERN; " + "assert BLOCK_LINEAR_PATTERN.fullmatch('transformer_blocks.31.img_mlp.proj'); " + "assert not BLOCK_LINEAR_PATTERN.fullmatch('proj_out'); " + "assert 'torch' not in sys.modules; assert 'deepcompressor' not in sys.modules") + subprocess.run([sys.executable, "-I", "-c", code], check=True, capture_output=True, text=True) + + def test_final_collection_and_export_cli_entry_points_still_parse(self): + for module in ("collect_v3", "export_v3"): + result = subprocess.run([sys.executable, "-m", "nunchaku_backend." + module, "--help"], + cwd=PACKAGE.parent, check=True, capture_output=True, text=True) + self.assertIn("--out", result.stdout) + + +if __name__ == "__main__": + unittest.main() diff --git a/reproduction/nunchaku_backend/collect_v3.py b/reproduction/nunchaku_backend/collect_v3.py new file mode 100644 index 0000000000000000000000000000000000000000..0a169ea3031e0161265dc6e6ba6e5b849e6b250a --- /dev/null +++ b/reproduction/nunchaku_backend/collect_v3.py @@ -0,0 +1,81 @@ +"""Collect deterministic BF16 teacher activations with disjoint prompt splits.""" +import argparse,hashlib,json,math,time +from pathlib import Path +from .layout import BLOCK_LINEAR_PATTERN + + +def main(): + p=argparse.ArgumentParser();p.add_argument('--jobs',required=True);p.add_argument('--out',required=True);p.add_argument('--steps',type=int,default=16) + args=p.parse_args() + import torch + from PIL import Image + from safetensors.torch import save_file + from .teacher_v3 import make_teacher + engine=make_teacher('/poc/samples/fidelity-v3/calibration-unused') + jobs=json.loads(Path(args.jobs).read_text());splits={'train':512,'validation':256,'heldout':256} + selected_steps=sorted({0,args.steps//4,args.steps//2,3*args.steps//4,args.steps-1}) + split_counts={s:sum(j['split']==s for j in jobs) for s in splits} + samples={};amax={};provenance=[];handles=[] + state={'job':None,'step':-1,'job_index':0,'token_groups':None} + def transformer_pre(module,inputs,kwargs): + state['step']+=1;state['token_groups']=None + if state['step']==0: + mask=kwargs['img_mask'][0].detach().cpu().bool() + expanded=mask.repeat_interleave(torch.where(mask,4,1)) + target=math.prod(kwargs['img_shapes'][0][-1]);prefix=len(expanded)-target + positions=torch.arange(len(expanded)) + state['token_groups']=[positions[(~expanded)&(positions=prefix]] + handles.append(engine.pipe.transformer.register_forward_pre_hook(transformer_pre,with_kwargs=True)) + def hook(name): + def capture(module,inputs): + if state['step'] not in selected_steps:return + job=state['job'];split=job['split'];x=inputs[0].detach().reshape(-1,inputs[0].shape[-1]) + if split=='train': + current=x.abs().amax(0).float().cpu() + amax[name]=torch.maximum(amax[name],current) if name in amax else current + quota=math.ceil(splits[split]/(split_counts[split]*len(selected_steps))) + # Same row selection for projections sharing a token sequence. + seed=123456+state['job_index']*1000+state['step'] + g=torch.Generator(device='cpu').manual_seed(seed) + groups=state['token_groups'] + if groups is not None: + groups=[group for group in groups if len(group)] + selected=[];remaining=quota + for i,group in enumerate(groups): + take=min(len(group),math.ceil(remaining/(len(groups)-i))) + selected.append(group[torch.randperm(len(group),generator=g)[:take]]) + remaining-=take + idx=torch.cat(selected).to(x.device) + else:idx=torch.randperm(x.shape[0],generator=g)[:quota].to(x.device) + samples.setdefault(name,{}).setdefault(split,[]).append(x.index_select(0,idx).cpu().contiguous()) + return capture + for name,module in engine.pipe.transformer.named_modules(): + if BLOCK_LINEAR_PATTERN.fullmatch(name):handles.append(module.register_forward_pre_hook(hook(name))) + try: + for i,job in enumerate(jobs): + state.update(job=job,step=-1,job_index=i) + engine.metrics={};engine._kv_caches={} + refs=[Image.open(path).copy() for path in job.get('images',[])] + t=time.perf_counter() + output=engine.pipe(prompt=job['prompt'],image=refs or None,width=job.get('width',1024),height=job.get('height',1024),output_resolution=1024, + num_inference_steps=args.steps,true_cfg_scale=1.,use_kv_cache=True,output_type='latent',generator=torch.Generator(device='cuda').manual_seed(job['seed'])) + torch.cuda.synchronize();del output;engine._kv_caches.clear();torch.cuda.empty_cache() + record={**job,'seconds':time.perf_counter()-t,'captured_steps':selected_steps} + provenance.append(record);print(json.dumps({'event':'teacher_calibration_job',**record}),flush=True) + out=Path(args.out);out.mkdir(parents=True,exist_ok=True);files={} + for name,ss in samples.items(): + values={split:torch.cat(ss[split])[:count].contiguous() for split,count in splits.items()} + values['input_absmax']=amax[name] + for split,count in splits.items(): + if values[split].shape[0]!=count:raise RuntimeError(f'Insufficient rows {name} {split}') + if not torch.isfinite(values[split]).all():raise RuntimeError('Nonfinite teacher activation') + filename=name+'.safetensors';save_file(values,str(out/filename)) + files[name]={'file':filename,'rows':{s:len(values[s]) for s in splits},'in_features':values['train'].shape[-1]} + if len(files)!=224:raise RuntimeError('Incomplete teacher calibration') + report={'format':'qwen21-activation-v3','teacher':'BF16 DiT, same NF4 encoder; streamed-group1 offload','model_revision':'b3179ad355be050328e483a9dfdd9e60cd62adfa','steps':args.steps,'sampling':'equal job/timestep quotas; first step stratified across text, reference-image and target-image tokens; later steps uniform target tokens','split_limitation':'Calibration editing prompts differ across splits but share the same source photograph. Final image-evaluation references differ from calibration.','jobs':provenance,'layers':files,'created_unix':time.time(),'evaluation_prompts_and_references_held_out':True} + (out/'manifest.json').write_text(json.dumps(report,indent=2)+'\n');print(json.dumps({'event':'teacher_calibration_complete','layers':len(files),'out':str(out)}),flush=True) + finally: + for h in handles:h.remove() + +if __name__=='__main__':main() diff --git a/reproduction/nunchaku_backend/compare_iterations_v3.py b/reproduction/nunchaku_backend/compare_iterations_v3.py new file mode 100644 index 0000000000000000000000000000000000000000..c4652ba6ac5a2de6ae2e3f896ed14619f2d92826 --- /dev/null +++ b/reproduction/nunchaku_backend/compare_iterations_v3.py @@ -0,0 +1,79 @@ +"""Compare isolated longer-fit probes against the ACTUAL saved V3 metrics.""" +import argparse +import gzip +import json +from pathlib import Path + + +def read_manifest(path): + opener = gzip.open if str(path).endswith(".gz") else open + with opener(path, "rt") as handle: + return json.load(handle) + + +def layer_stats(manifest): + return {name: row for shard in manifest["files"].values() for name, row in shard.get("layer_stats", {}).items()} + + +def compare(v3, probes): + original = layer_stats(v3) + results = [] + for probe in probes: + identity = probe["conversion_identity"] + for key in ("activation_fingerprint", "baseline_calibration_sha256", "source_config_sha256"): + if identity[key] != v3["conversion_identity"][key]: + raise ValueError(f"Probe does not match V3 {key}") + for name, current in layer_stats(probe).items(): + old = original[name] + settings = current["search"] + if settings["seed"] != old["search"]["seed"]: + raise ValueError(f"Probe seed changed: {name}") + selected = old["selected"] + if settings.get("fixed_smoothing") != selected["family"] or (selected["family"] != "identity" and settings["alphas"] != [selected["alpha"]]): + raise ValueError(f"Probe smoothing differs from V3 selection: {name}") + if settings["ranks"] != [selected["rank"]] or settings["weighting"] != [selected["weighting"]]: + raise ValueError(f"Probe rank/weighting changed: {name}") + for key_name in ("niter", "oversample", "ridge", "factorization", "final_gptq", "gptq_damp"): + if settings[key_name] != old["search"][key_name]: + raise ValueError(f"Probe {key_name} changed: {name}") + def key(row): + return tuple(row.get(k) for k in ("family", "alpha", "rank", "weighting", "iteration", "output_correction", "factorization")) + old_history = {key(r): r for r in old["history"] if not r.get("gptq") and r["family"] != "one_pass_baseline"} + current_history = [r for r in current["history"] if not r.get("gptq") and r["family"] != "one_pass_baseline"] + prefix_differences = [abs(r["validation"]["mse"] - old_history[key(r)]["validation"]["mse"]) / + max(old_history[key(r)]["validation"]["mse"], 1e-30) + for r in current_history if key(r) in old_history] + if not prefix_differences: + raise ValueError(f"No matching deterministic recurrence prefix: {name}") + validation_ratio = current["validation"]["mse"] / max(old["validation"]["mse"], 1e-30) + heldout_ratio = current["heldout"]["mse"] / max(old["heldout"]["mse"], 1e-30) + results.append({"layer": name, "seed": settings["seed"], "v3_selected": old["selected"], + "probe_selected": current["selected"], "v3_hit_cap": selected["iteration"] == 15, + "max_iteration_evaluated": max(r["iteration"] for r in current_history), + "prefix_compared_candidates": len(prefix_differences), + "prefix_max_relative_validation_mse_difference": max(prefix_differences), + "v3_validation": old["validation"], "probe_validation": current["validation"], + "validation_mse_ratio_to_actual_v3": validation_ratio, + "v3_heldout": old["heldout"], "probe_heldout": current["heldout"], + "heldout_mse_ratio_to_actual_v3": heldout_ratio, + "validation_prefers_probe": validation_ratio < 1, + "heldout_improvement_exceeds_5pct": heldout_ratio < 0.95, + "probe_seconds": current["seconds"]}) + return {"comparison": "100-iteration-limit fixed-recipe replay versus actual saved V3 metrics", "layers": results, + "limitations": "Validation chooses candidates; heldout is reporting only. Early stopping is unchanged and can end well before100. These are deterministic replays fromiteration0, not restored optimizer states. Final GPTQ is tried only on the selected raw candidate, so old V3 may still beat the new probe; retain V3 in that case. No full export or deployment is authorized by this report."} + + +def main(): + parser = argparse.ArgumentParser() + parser.add_argument("--v3", type=Path, required=True) + parser.add_argument("--probes", nargs="+", type=Path, required=True) + parser.add_argument("--out", type=Path, required=True) + args = parser.parse_args() + result = compare(read_manifest(args.v3), [read_manifest(path) for path in args.probes]) + args.out.parent.mkdir(parents=True, exist_ok=True) + args.out.write_text(json.dumps(result, indent=2) + "\n") + print(json.dumps(result, indent=2)) + + +if __name__ == "__main__": + main() diff --git a/reproduction/nunchaku_backend/convert_activation_reference.py b/reproduction/nunchaku_backend/convert_activation_reference.py new file mode 100644 index 0000000000000000000000000000000000000000..e8f99c9e67be3ab524aff2e2c435836149390d5d --- /dev/null +++ b/reproduction/nunchaku_backend/convert_activation_reference.py @@ -0,0 +1,102 @@ +"""Independent activation-quantization reference using Nunchaku's PTX math. + +No Nunchaku kernel or packed layout is reused. This launches only when called +explicitly on CUDA tensors. Source: nunchaku-ai/nunchaku v1.2.1, +src/kernels/zgemm/gemm_utils.cuh h2div/cuda_frcp/quantize_float2 and +src/kernels/zgemm/gemm_w4a4.cuh quantize_w4a4_from_fpsum_warp. +""" + +from __future__ import annotations + +import torch + +try: + import triton + import triton.language as tl +except ImportError: + triton = None + tl = None + + +if triton is not None: + + @triton.jit + def _quantize_activation_ptx(X, Smooth, QA, SA, K: tl.constexpr, GROUPS: tl.constexpr): + block = tl.program_id(0) + row = block // GROUPS + group = block % GROUPS + channel = group * 64 + tl.arange(0, 64) + values = tl.load(X + row * K + channel).to(tl.float32) + smooth = tl.load(Smooth + channel).to(tl.float32) + + # CUDA h2div uses __fdividef, then float22half2. Explicit PTX avoids + # Triton changing division to a differently-rounded reciprocal/mul. + divided = tl.inline_asm_elementwise( + "div.approx.ftz.f32 $0, $1, $2;", + constraints="=f,f,f", args=[values, smooth], dtype=tl.float32, + is_pure=True, pack=1, + ) + divided = divided.to(X.dtype.element_ty).to(tl.float32) + amax = tl.max(tl.abs(divided), axis=0) + + # Match C++ float RECPI_QVALUE_MAX_SIGNED = 1 / 7.0f. Keeping mul + # and reciprocal as distinct PTX operations prevents reassociation. + scale = tl.inline_asm_elementwise( + "mul.rn.f32 $0, $1, $2;", + constraints="=f,f,f", + args=[amax, tl.full((), 0.14285714285714285, tl.float32)], + dtype=tl.float32, is_pure=True, pack=1, + ) + reciprocal = tl.inline_asm_elementwise( + "rcp.approx.ftz.f32 $0, $1;", + constraints="=f,f", args=[scale], dtype=tl.float32, + is_pure=True, pack=1, + ) + normalized = tl.inline_asm_elementwise( + "mul.rn.f32 $0, $1, $2;", + constraints="=f,f,f", args=[divided, reciprocal], dtype=tl.float32, + is_pure=True, pack=1, + ) + integer = tl.inline_asm_elementwise( + "cvt.rni.s32.f32 $0, $1;", + constraints="=r,f", args=[normalized], dtype=tl.int32, + is_pure=True, pack=1, + ) + integer = tl.maximum(tl.minimum(integer, 7), -8) + # An all-zero group intentionally follows PTX's NaN->INT_MIN->-8 + # conversion; its stored scale is zero, so it dequantizes to zero. + tl.store(QA + block * 64 + tl.arange(0, 64), integer.to(tl.float32)) + tl.store(SA + block, scale) + + +@torch.no_grad() +def quantize_activations_ptx(x: torch.Tensor, smooth: torch.Tensor) -> tuple[torch.Tensor, torch.Tensor]: + """Return unpacked QA FP32[N,K/64,64], scales BF16/FP16[N,K/64]. + + Inputs must already be on CUDA. The function independently reproduces + Nunchaku's signed A4 mathematics; it does not call its quantizer or read + its packed tensors. This source was syntax-checked on CPU; actual GPU + validation belongs to the caller's numerical probe. + """ + if triton is None: + raise ImportError("Triton is required for the explicit PTX activation reference") + if x.device.type != "cuda" or smooth.device.type != "cuda": + raise ValueError("Activation reference requires explicit CUDA input and smoothing tensors") + if x.dtype not in (torch.bfloat16, torch.float16): + raise ValueError("Input must be BF16 or FP16") + if x.ndim < 2 or x.shape[-1] % 64: + raise ValueError("Input must have at least two dimensions and width divisible by 64") + if smooth.shape != (x.shape[-1],) or smooth.device != x.device or smooth.dtype != x.dtype: + raise ValueError("Smoothing must have matching device/dtype and shape [input_width]") + k = x.shape[-1] + flattened = x.reshape(-1, k).contiguous() + groups = k // 64 + qa = torch.empty((flattened.shape[0], groups, 64), device=x.device, dtype=torch.float32) + sa = torch.empty((flattened.shape[0], groups), device=x.device, dtype=x.dtype) + if flattened.shape[0] == 0: + return qa, sa + _quantize_activation_ptx[(flattened.shape[0] * groups,)]( + flattened, smooth.contiguous(), qa, sa, K=k, GROUPS=groups, + num_warps=4, num_stages=1, + ) + return qa, sa diff --git a/reproduction/nunchaku_backend/convert_recipe_v3.md b/reproduction/nunchaku_backend/convert_recipe_v3.md new file mode 100644 index 0000000000000000000000000000000000000000..f36c44cc681d32ecf468d6c8e80be0ef3f8d7633 --- /dev/null +++ b/reproduction/nunchaku_backend/convert_recipe_v3.md @@ -0,0 +1,160 @@ +# Pinned DeepCompressor SVDQuant recipe and Qwen v3 implications + +Source review of DeepCompressor commit +`69f3473f5e1c1504bae35cc50c7858ef900a9b17`. This is a description of source +behavior, not an image-quality or latency measurement. No stable converter, +checkpoint, or GPU workload was modified for this review. + +## What the official example configuration actually requests + +The base SVDQuant example selects rank 32, output-error calibration, up to +100 low-rank iterations, and early stopping. Smoothing uses absolute channel +maxima for activations and weights, a grid search, low-rank-aware evaluation, +and `fuse_when_possible=false`. + +Its `alpha: 0.5`, `beta: -2`, `num_grids: 20` do **not** specify one alpha=0.5 +conversion. Those flags generate 39 smoothing candidates: + +- identity `(alpha,beta)=(0,0)`; +- 19 activation-only candidates `(alpha,0)` with alpha=0.05 through 0.95; +- 19 balanced candidates `(alpha,1-alpha)` on the same grid. + +The scale is `amax(X)^alpha / amax(W)^beta`. Invalid/zero scale handling in +the source differs from our bounded positive-clamp implementation. The fast +configuration uses 10 grids and 64 calibration samples; the general default +uses 128 calibration samples. + +Sources: [base SVDQuant YAML](https://github.com/nunchaku-ai/deepcompressor/blob/69f3473f5e1c1504bae35cc50c7858ef900a9b17/examples/diffusion/configs/svdquant/__default__.yaml), +[alpha/beta candidate generation](https://github.com/nunchaku-ai/deepcompressor/blob/69f3473f5e1c1504bae35cc50c7858ef900a9b17/deepcompressor/calib/config/smooth.py#L134), +[fast override](https://github.com/nunchaku-ai/deepcompressor/blob/69f3473f5e1c1504bae35cc50c7858ef900a9b17/examples/diffusion/configs/svdquant/fast.yaml). + +## Exact iterative low-rank algorithm + +In the following, W already includes any preceding smoothing. Q denotes a +dequantized quantized residual; L denotes the effective low-rank weight. + +```text +Q = 0 # default compensate=false +# alternate optional initialization: Q = Quantize(W) when compensate=true +best = none +repeat up to num_iters: + L = rank_r_SVD(W - Q) + Q = Quantize(W - L, kernel=None) + evaluate candidate module with weight Q, low-rank branch L, + and configured activation quantization + retain L if output-error sum <= best_error + stop on the first worse candidate when early_stop=true +return best L +``` + +This is alternating quantization and SVD, not gradient training or an +activation-weighted SVD. The next SVD acts on `W - previous_Q`, so rerunning +SVD on W is not equivalent. Candidate generation continues from the current Q, +not from the best Q, unless early stopping ends the loop. + +The actual `LowRankBranch` implementation computes full FP64 SVD and stores +`down=Vh[:r]`, `up=U[:,:r]*singular_values[:r]` in the original parameter +dtype. It does not split sqrt(singular_values) between the factors. Both +factorizations have the same ideal matrix product, but their BF16 hidden +activation rounding differs. The current custom converter's randomized SVD +and balanced factors are implementation choices, not exact copies of this +reference recipe. + +For the output-error objective, the candidate module's residual weight is Q. +A branch hook captures input before activation-quantizer hooks and adds the +low-rank output afterward. The branch therefore sees unquantized input. +For weight/product objectives the candidate representation can instead be +assembled as Q+L; the default SVDQuant recipe uses output error. + +The winning branch is later subtracted from the real module weight, and the +remaining weight is quantized during the final quantization phase. There is +no guarantee that selecting by weight reconstruction error picks the same +branch as selecting by output error. + +Sources: [low-rank calibrator](https://github.com/nunchaku-ai/deepcompressor/blob/69f3473f5e1c1504bae35cc50c7858ef900a9b17/deepcompressor/calib/lowrank.py#L80), +[SVD branch factorization](https://github.com/nunchaku-ai/deepcompressor/blob/69f3473f5e1c1504bae35cc50c7858ef900a9b17/deepcompressor/nn/patch/lowrank.py#L31), +[input-capturing branch hook](https://github.com/nunchaku-ai/deepcompressor/blob/69f3473f5e1c1504bae35cc50c7858ef900a9b17/deepcompressor/utils/hooks/branch.py#L28), +[applying the selected branch](https://github.com/nunchaku-ai/deepcompressor/blob/69f3473f5e1c1504bae35cc50c7858ef900a9b17/deepcompressor/app/diffusion/quant/weight.py#L105). + +## What smoothing and output-error evaluation include + +For each smoothing candidate, weights are scaled by s and inputs are divided +by s. With `allow_low_rank=true`, a low-rank branch and quantized residual are +constructed for the candidate. This inner smoothing evaluation uses +`quantize_with_low_rank`, which performs one SVD/residual quantization, not the +whole up-to-100-iteration low-rank search. The activation quantizer is installed +after the low-rank input capture. + +The generic output-error search evaluates the selected module on cached +inputs and sums squared differences from its original outputs. For Q/K +projections, the diffusion wrapper can evaluate the enclosing attention or +parallel transformer module instead of only the linear projection. Shared-input +Q/K/V weights can be concatenated for a shared down projection when +`exclusive=false`. Other linears are usually evaluated individually. + +The pinned recipe's PyTorch quantization simulation does not reproduce our +observed Nunchaku groupwise BF16 accumulation and PTX reciprocal bit for bit. +Using the verified custom numerical reference to select Qwen candidates is a +justified adaptation; it must be described as such. + +Sources: [low-rank-aware smoothing module evaluation](https://github.com/nunchaku-ai/deepcompressor/blob/69f3473f5e1c1504bae35cc50c7858ef900a9b17/deepcompressor/calib/smooth.py#L578), +[one-pass quantize_with_low_rank](https://github.com/nunchaku-ai/deepcompressor/blob/69f3473f5e1c1504bae35cc50c7858ef900a9b17/deepcompressor/quantizer/processor.py#L215), +[output-error sum](https://github.com/nunchaku-ai/deepcompressor/blob/69f3473f5e1c1504bae35cc50c7858ef900a9b17/deepcompressor/calib/search.py#L817), +[attention evaluation scope](https://github.com/nunchaku-ai/deepcompressor/blob/69f3473f5e1c1504bae35cc50c7858ef900a9b17/deepcompressor/app/diffusion/quant/weight.py#L50). + +## GPTQ is optional and comes later + +The optional GPTQ configuration uses damping 0.01, block size 128, up to +250 inversion attempts, and a Hessian **sample accumulation chunk** of 512. +`hessian_block_size=512` does not make the Hessian block diagonal: the source +still allocates the full K×K Hessian. + +The implementation builds `H = sum(2/N * X.T @ X)` from cached input samples, +handles zero diagonal channels, sorts channels by descending Hessian diagonal, +damps the diagonal, and takes an upper Cholesky factor of the inverse. +It quantizes columns sequentially, propagates each column's quantization error +through the inverse-Hessian factor, propagates completed block error to the +remaining columns, then restores the original channel order. Quantization +scales are addressed using original column-group indices even while columns +are permuted. + +Both smoothing and iterative low-rank candidate quantization explicitly pass +`kernel=None`; GPTQ therefore is not active inside those searches. The final +weight quantization call can use the configured GPTQ kernel after the selected +branch has been subtracted. It minimizes weight-product error using input +covariance; it does not directly optimize the error from rounding inputs to A4. + +For Qwen's K=12288 MLP output, one full FP32 Hessian is 576 MiB, before copies, +Cholesky workspace or weights. SVD, covariance storage and inversion deserve +their own timing and peak-memory measurements. + +Sources: [optional GPTQ YAML](https://github.com/nunchaku-ai/deepcompressor/blob/69f3473f5e1c1504bae35cc50c7858ef900a9b17/examples/diffusion/configs/svdquant/gptq.yaml), +[GPTQ kernel](https://github.com/nunchaku-ai/deepcompressor/blob/69f3473f5e1c1504bae35cc50c7858ef900a9b17/deepcompressor/quantizer/kernel/gptq.py#L129), +[final quantization call](https://github.com/nunchaku-ai/deepcompressor/blob/69f3473f5e1c1504bae35cc50c7858ef900a9b17/deepcompressor/app/diffusion/quant/weight.py#L239). + +## Practical Qwen v3 priorities + +1. Search identity, activation-only smoothing and balanced smoothing with + representative generation/editing activations. Keep separate held-out + prompts/timesteps for selection verification; a large token count from a + single request is not equivalent to diverse requests. +2. Score the complete A4 residual plus BF16 low-rank output, including actual + rounding boundaries. Preserve the best candidate and iteration, not the + last. Add alternating `L=SVD(W-Q)` iterations after one-pass smoothing + selection; start with a bounded iteration budget and measured early stopping. +3. Compare upstream-style `(U*S,Vh)` and balanced factorization under that + output metric. Retain the final BF16 inverse-smoothing correction. Consider + higher ranks only with measured held-out improvement and VRAM/latency cost. +4. Diagnose the worst linears separately. If activation error dominates, GPTQ + alone will not solve it. A held-out-scored fit of the low-rank output to + residual output error can be tested, but that is an additional method, not + the pinned DeepCompressor algorithm. +5. Add optional final residual GPTQ only after the simpler calibrated baseline + passes. Preserve original channel order and group64 scales for Nunchaku; + build the Hessian in the correct smoothed coordinate system. + +The official INT4 example also enables activation shifting and unsigned +quantization. Do not copy that by toggling generic +`SVDQW4A4Linear.act_unsigned`: its standalone quantizer instantiates signed +activation quantization and does not accept that flag. Unsigned inference needs +a compatible quantizer and correct shift/bias handling, not just a GEMM flag. diff --git a/reproduction/nunchaku_backend/convert_reference.py b/reproduction/nunchaku_backend/convert_reference.py new file mode 100644 index 0000000000000000000000000000000000000000..169eaaf5921c1a19fbd5e6872676ead7968dde6d --- /dev/null +++ b/reproduction/nunchaku_backend/convert_reference.py @@ -0,0 +1,72 @@ +"""Groupwise BF16 CUDA arithmetic reference; does not alter exported weights. + +Nunchaku gemm_w4a4_kernel selects USE_FP32_ACCUM=false. The kernel rounds the +INT32 dot and the product of scales to BF16/FP16, then accumulates with a +BF16/FP16 FMA after *each* group64. This differs from one FP32 GEMM rounded +only at the end. See gemm_base.cuh apply_scales(fpsum_warp&) and +gemm_w4a4.cuh gemm_w4a4_kernel at Nunchaku commit +302e0e97024ebd68688fe890e5df83731edf7b54. +""" + +from __future__ import annotations + +import torch + + +@torch.no_grad() +def reference_forward_groupwise(x: torch.Tensor, reference: dict[str, torch.Tensor]) -> torch.Tensor: + """Emulate signed INT4 group64 arithmetic and LoRA epilogue boundaries. + + Returns compute-dtype-rounded FP32 output. CUDA reciprocal approximation + and low-rank atomic reduction order still need not be bit-identical. + This accepts the existing converter's unpacked reference without changing + the conversion/checkpoint. Weight scales recover exactly from nonzero + groups because our symmetric converter always includes a +/-7 extremum. + A guard rejects reference tensors where that assumption does not hold. + """ + if x.dtype not in (torch.bfloat16, torch.float16): + raise ValueError("BF16/FP16 input required") + dtype, device = x.dtype, x.device + ic = x.shape[-1] + if ic % 64: + raise ValueError("Input width must be a multiple of 64") + smooth = reference["smooth"].to(device=device, dtype=torch.float32) + original = x.float() + smoothed = (original / smooth).to(dtype).float() + grouped = smoothed.reshape(-1, ic // 64, 64) + scales_fp32 = grouped.abs().amax(-1) * (1.0 / 7.0) + reciprocal = torch.where(scales_fp32 > 0, scales_fp32.reciprocal(), torch.zeros_like(scales_fp32)) + qa = (grouped * reciprocal.unsqueeze(-1)).round().clamp(-8, 7) + sa = scales_fp32.to(dtype).float() + if device.type == "cuda": + from .convert_activation_reference import quantize_activations_ptx + qa, sa = quantize_activations_ptx(x, smooth.to(dtype)) + sa = sa.float() + + residual = reference["residual_dequant"].to(device=device, dtype=torch.float32) + oc = residual.shape[0] + grouped_weight = residual.reshape(oc, ic // 64, 64) + sw = (grouped_weight.abs().amax(-1) / 7.0).to(dtype).float() + sw = torch.where(sw > 0, sw, torch.ones_like(sw)) + qw = (grouped_weight / sw.unsqueeze(-1)).round() + if (qw.abs() > 7).any() or not torch.equal(qw * sw.unsqueeze(-1), grouped_weight): + raise ValueError("Reference residual does not match symmetric group64 +/-7 quantization") + + output = torch.zeros((grouped.shape[0], oc), dtype=torch.float32, device=device) + for group in range(ic // 64): + # Small integer values are exact in FP32 dot accumulation here. + int_dot = qa[:, group] @ qw[:, group].T + rounded_dot = int_dot.to(dtype).float() + scale_product = (sa[:, group, None] * sw[None, :, group]).to(dtype).float() + # CUDA __hfma2 has a single output rounding, not BF16 mul then add. + # FP64 represents these BF16/FP16 operands closely enough to avoid + # an extra FP32 rounding before the final compute-dtype conversion. + output = (rounded_dot.double() * scale_product.double() + output.double()).to(dtype).float() + + output = output.reshape(*x.shape[:-1], oc) + if "bias" in reference: + output = (output + reference["bias"].to(device=device, dtype=torch.float32)).to(dtype).float() + down = reference["down_unpacked"].to(device=device, dtype=torch.float32) + up = reference["up_unpacked"].to(device=device, dtype=torch.float32) + hidden = (original @ down.T).to(dtype).float() + return (output + hidden @ up.T).to(dtype).float() diff --git a/reproduction/nunchaku_backend/convert_reference_notes.md b/reproduction/nunchaku_backend/convert_reference_notes.md new file mode 100644 index 0000000000000000000000000000000000000000..9ad4fa306f8f1e7dc8ba99e50a40a673b400eb8e --- /dev/null +++ b/reproduction/nunchaku_backend/convert_reference_notes.md @@ -0,0 +1,54 @@ +# Groupwise Nunchaku arithmetic reference + +The first GPU probe reported finite output but actual/reference relative L2 +differences of roughly 0.009–0.016, increasing with input width. The original +FP32 reference omitted a concrete kernel behavior: the INT4 residual branch +uses BF16 accumulation after every 64-channel group. + +`convert_reference.py::reference_forward_groupwise` is a standalone replacement +reference. It does not modify conversion or existing checkpoints. The GPU +probe still needs rerunning against this reference before attributing the +observed discrepancy to accumulation alone. + +For each group, the source computes: + +```text +integer_dot = INT32 dot(qA_group, qW_group) +dot = BF16(integer_dot) +scale = BF16(BF16_activation_scale * BF16_weight_scale) +running = BF16_FMA(dot, scale, running) +``` + +Groups accumulate sequentially from 0 through K/64 - 1. The two-stage loop +prefetches into alternating slots but does not reorder the arithmetic. +After that, the existing BF16 bias and low-rank epilogue boundaries apply. + +The emulator uses FP64 multiplication plus addition followed by BF16 rounding +to represent a single BF16 fused multiply-add without introducing a separate +rounded multiply. It recovers INT4 weight groups and scales from the existing +unpacked residual, using the converter's nonzero-group +/-7 extremum. An exact +reconstruction guard rejects tensors for which that assumption fails. + +CPU tests passed with Torch 2.14.0. One regression constructs sequential group +contributions near 256, 1 and 1: BF16 running accumulation gives 256, while +the old full FP32 sum rounded at the end gives 258. CUDA reciprocal +approximation and low-rank atomic reduction order can still create residual +differences; this reference is not claimed bit-exact before a GPU comparison. + +Source anchors at Nunchaku commit 302e0e97024ebd68688fe890e5df83731edf7b54: + +- [Kernel explicitly selects USE_FP32_ACCUM=false, near line 1080](https://github.com/nunchaku-ai/nunchaku/blob/302e0e97024ebd68688fe890e5df83731edf7b54/src/kernels/zgemm/gemm_w4a4.cuh#L1080) +- [Per-group BF16 dot conversion, scale multiplication and fused addition](https://github.com/nunchaku-ai/nunchaku/blob/302e0e97024ebd68688fe890e5df83731edf7b54/src/kernels/zgemm/gemm_base.cuh#L369) +- [Sequential two-stage group accumulation](https://github.com/nunchaku-ai/nunchaku/blob/302e0e97024ebd68688fe890e5df83731edf7b54/src/kernels/zgemm/gemm_w4a4.cuh#L884) + +## Subsequent PTX validation + +The groupwise correction reduced the mismatch but retained activation-boundary +rounding differences. The CUDA reference now delegates activation math to +`convert_activation_reference.py`, an independent Triton implementation using +PTX `div.approx.ftz.f32`, `rcp.approx.ftz.f32`, and integer ties-to-even conversion. +This resolved the discrepancy without checkpoint changes or tolerance increases: +all nine GPU cases passed, relative L2 from 1.19e-6 to 5.57e-5. The worst absolute +error point in every case differed by one BF16 ULP. The full result is preserved +in `../results/nunchaku-kernel-probe-ptx.json`; earlier failed references remain +alongside it for comparison. diff --git a/reproduction/nunchaku_backend/convert_reference_test.py b/reproduction/nunchaku_backend/convert_reference_test.py new file mode 100644 index 0000000000000000000000000000000000000000..7ccffc44e556267a9baa579f224b05a27cca96de --- /dev/null +++ b/reproduction/nunchaku_backend/convert_reference_test.py @@ -0,0 +1,59 @@ +"""CPU sanity checks for the groupwise CUDA arithmetic emulator.""" + +import unittest + +import torch + +if __package__: + from .baseline_candidate import convert_linear_weight, reference_forward + from .convert_reference import reference_forward_groupwise +else: + from baseline_candidate import convert_linear_weight, reference_forward + from convert_reference import reference_forward_groupwise + + +class GroupwiseReferenceTests(unittest.TestCase): + def test_emulator_handles_export_zero_groups_and_nonuniform_smoothing(self): + torch.set_num_threads(2) + torch.manual_seed(88) + weight = torch.randn(128, 256).bfloat16() + bias = torch.randn(128).bfloat16() + _, _, ref = convert_linear_weight(weight, bias, smooth=torch.logspace(-1, 1, 256), return_reference=True) + x = torch.randn(1, 17, 256).bfloat16() + out = reference_forward_groupwise(x, ref) + self.assertEqual(out.shape, (1, 17, 128)) + self.assertTrue(torch.isfinite(out).all()) + self.assertTrue(torch.equal(reference_forward_groupwise(torch.zeros_like(x), ref), bias.float().expand_as(out))) + # Group rounding should produce a small, nonzero change from the + # previous approximation, not a different scale or packing order. + old = reference_forward(x, ref) + difference = (out - old).norm() / old.norm() + self.assertGreater(difference.item(), 0) + self.assertLess(difference.item(), 0.02) + + def test_per_group_accumulator_rounding_is_observable(self): + # Group0 contributes 256; groups1/2 each add1. BF16 running FMA + # ties-to-even rounds 257 back to256 twice. Full FP32 sum gives258. + q = torch.zeros(128, 192) + q[:, 0] = 7 + q[:, 64] = 7 + q[:, 128] = 7 + sw = torch.tensor([256.0 / 49, 1.0 / 49, 1.0 / 49]).bfloat16().float() + ref = { + "residual_dequant": (q.reshape(128, 3, 64) * sw[None, :, None]).reshape(128, 192), + "smooth": torch.ones(192, dtype=torch.bfloat16), + "down_unpacked": torch.zeros(32, 192, dtype=torch.bfloat16), + "up_unpacked": torch.zeros(128, 32, dtype=torch.bfloat16), + } + x = torch.zeros(1, 1, 192, dtype=torch.bfloat16) + x[..., 0] = 7 + x[..., 64] = 7 + x[..., 128] = 7 + out = reference_forward_groupwise(x, ref) + old = reference_forward(x, ref) + self.assertTrue((out == 256).all()) + self.assertTrue((old == 258).all()) + + +if __name__ == "__main__": + unittest.main(verbosity=2) diff --git a/reproduction/nunchaku_backend/denoiser_probe_v3.py b/reproduction/nunchaku_backend/denoiser_probe_v3.py new file mode 100644 index 0000000000000000000000000000000000000000..cc4aebcf7d2bf769e7cf8b07f9a5865a298055ea --- /dev/null +++ b/reproduction/nunchaku_backend/denoiser_probe_v3.py @@ -0,0 +1,391 @@ +"""Bounded BF16-versus-Nunchaku denoiser prediction diagnostic. + +Run only in the isolated POC container after its GPU has been reserved:: + + python -m nunchaku_backend.denoiser_probe_v3 capture \ + --jobs experiments/fidelity-v3/jobs.json --label heldout-example \ + --out /cache/denoiser-probe-v3 + python -m nunchaku_backend.denoiser_probe_v3 compare \ + --capture /cache/denoiser-probe-v3 \ + --checkpoint v2=/cache/qwen21-nunchaku-r32 \ + --checkpoint v3=/cache/qwen21-nunchaku-v3 \ + --out results/denoiser-probe-v3.json + +Capture retains exactly first/middle/last transformer inputs and target-token +predictions. It NEVER stores teacher KV. Each cached-step comparison rebuilds +a fresh prefix using the selected quantized checkpoint's first-step forward. +Later inputs remain teacher-trajectory inputs: this measures conditional +prediction error, not accumulated free-running trajectory/image error. + +Importing this file does not import Torch or start GPU work. The `capture` and +`compare` subcommands explicitly execute on the container's cuda:0 only. +""" +from __future__ import annotations + +import argparse +import gc +import hashlib +import json +import math +from pathlib import Path +import time + + +FORMAT = "qwen21-denoiser-teacher-capture-v3" + + +def _write_json(path, value): + path = Path(path) + path.parent.mkdir(parents=True, exist_ok=True) + temporary = path.with_suffix(path.suffix + ".tmp") + temporary.write_text(json.dumps(value, indent=2, sort_keys=True) + "\n") + temporary.replace(path) + + +def _hash_file(path): + digest = hashlib.sha256() + with Path(path).open("rb") as handle: + for block in iter(lambda: handle.read(1024 * 1024), b""): + digest.update(block) + return digest.hexdigest() + + +def _tree(value, torch, device): + """Copy supported forward data; reject accidentally captured cache objects.""" + if isinstance(value, torch.Tensor): + return value.detach().to(device=device, copy=True).contiguous() + if isinstance(value, dict): + if "kv_cache" in value: + raise ValueError("A teacher KV cache must never enter a saved input tree") + return {key: _tree(item, torch, device) for key, item in value.items()} + if isinstance(value, tuple): + return tuple(_tree(item, torch, device) for item in value) + if isinstance(value, list): + return [_tree(item, torch, device) for item in value] + if value is None or isinstance(value, (str, int, float, bool)): + return value + raise TypeError(f"Unsupported forward data: {type(value).__name__}") + + +def _tensor_bytes(value, torch): + if isinstance(value, torch.Tensor): + return value.numel() * value.element_size() + if isinstance(value, dict): + return sum(_tensor_bytes(item, torch) for item in value.values()) + if isinstance(value, (list, tuple)): + return sum(_tensor_bytes(item, torch) for item in value) + return 0 + + +def _target_prediction(output, kwargs): + """Match pipeline noise_pred[:, -latents.size(1):], including edit prefill.""" + prediction = output[0] if isinstance(output, tuple) else output.sample + target_tokens = math.prod(kwargs["img_shapes"][0][-1]) + if prediction.ndim != 3 or prediction.shape[1] < target_tokens: + raise ValueError(f"Invalid denoiser prediction shape {prediction.shape}; target {target_tokens}") + return prediction[:, -target_tokens:] + + +def _clear_cache(cache): + if cache is not None: + for layer in cache.layer_caches: + layer.k = None + layer.v = None + + +def _cache_summary(cache): + layers = cache.layer_caches + if not layers or any(layer.k is None or layer.v is None for layer in layers): + raise RuntimeError("Quantized prefill did not populate every prefix KV layer") + return { + "layers": len(layers), + "tensor_bytes": sum(t.numel() * t.element_size() for layer in layers for t in (layer.k, layer.v)), + "first_key_shape": list(layers[0].k.shape), + "dtype": str(layers[0].k.dtype), + } + + +def _metrics(actual, teacher, torch): + if actual.shape != teacher.shape: + raise ValueError(f"Prediction shapes differ: {actual.shape} versus {teacher.shape}") + actual = actual.detach().to(device="cpu", dtype=torch.float64).reshape(-1) + teacher = teacher.detach().to(device="cpu", dtype=torch.float64).reshape(-1) + if not bool(torch.isfinite(actual).all() and torch.isfinite(teacher).all()): + raise RuntimeError("Non-finite denoiser prediction") + delta = actual - teacher + pred_norm, teacher_norm = actual.norm().item(), teacher.norm().item() + return { + "relative_l2": delta.norm().item() / max(teacher_norm, 1e-30), + "cosine_similarity": float(torch.dot(actual, teacher)) / max(pred_norm * teacher_norm, 1e-30), + "max_abs_error": delta.abs().max().item(), + "mse": delta.square().mean().item(), + "teacher_l2_norm": teacher_norm, + "prediction_l2_norm": pred_norm, + "teacher_rms": teacher.square().mean().sqrt().item(), + "prediction_rms": actual.square().mean().sqrt().item(), + "teacher_max_abs": teacher.abs().max().item(), + "prediction_max_abs": actual.abs().max().item(), + "elements": actual.numel(), + "finite": True, + } + + +def capture(args): + import torch + from PIL import Image + from .teacher_v3 import make_teacher + + jobs = json.loads(Path(args.jobs).read_text()) + matches = [job for job in jobs if job.get("label") == args.label] + if len(matches) != 1: + raise ValueError("Choose exactly one job by its unique --label") + job = matches[0] + steps = int(job.get("steps", 40)) + if steps < 3 or float(job.get("cfg", 1.0)) != 1.0 or job.get("kv_cache", True) is not True: + raise ValueError("This bounded probe requires at least 3 steps, CFG 1, and KV caching") + if (job.get("width", 1024), job.get("height", 1024)) != (1024, 1024): + raise ValueError("This held-out diagnostic deliberately requires a 1024×1024 job") + directory = Path(args.out) + directory.mkdir(parents=True, exist_ok=True) + if any(directory.iterdir()): + raise ValueError("Capture output must be a new empty directory") + if args.max_capture_mib <= 0: + raise ValueError("Capture memory bound must be positive") + selected = sorted({0, steps // 2, steps - 1}) + metadata = { + "format": FORMAT, "complete": False, "job": job, + "jobs_file": str(Path(args.jobs).resolve()), "jobs_sha256": _hash_file(args.jobs), + "selected_steps_zero_based": selected, "teacher": "original BF16 DiT; fixed NF4 encoder", + "teacher_offload": "streamed group1" if not args.no_stream else "synchronous group4", + "cache_policy": "Teacher KV never saved; independently rebuilt by every compared backend", + "maximum_capture_mib": args.max_capture_mib, "captures": [], + "reference_sha256": {str(path): _hash_file(path) for path in job.get("images", [])}, + "claim": "Conditional prediction diagnostic; not free-running image quality or clean inference timing", + } + _write_json(directory / "capture.json", metadata) + engine = make_teacher(stream=not args.no_stream) + engine.metrics = {} + engine._kv_caches = {} + transformer = engine.pipe.transformer + if not transformer.config.causal_condition: + raise ValueError("Prefix replay requires causal_condition=True") + pending = None + call_index = -1 + bytes_captured = 0 + + def before(module, positional, kwargs): + nonlocal call_index, pending + call_index += 1 + if positional: + raise ValueError("Pinned pipeline must pass transformer inputs by keyword") + expected = "extract" if call_index == 0 else "cached" + if kwargs.get("kv_cache_mode") != expected or kwargs.get("kv_cache") is None: + raise ValueError("Unexpected teacher call/cache sequence; CFG or pipeline contract changed") + if call_index not in selected: + return + allowed = {key: value for key, value in kwargs.items() if key != "kv_cache"} + # Check size before copying; three bounded snapshots are the only + # trajectory data retained, and each is written immediately to disk. + input_bytes = _tensor_bytes(allowed, torch) + if bytes_captured + input_bytes > args.max_capture_mib * 2**20: + raise MemoryError("Teacher capture input exceeds the configured bound") + pending = _tree(allowed, torch, "cpu") + + def after(module, positional, kwargs, output): + nonlocal pending, bytes_captured + if call_index not in selected: + return + if pending is None: + raise RuntimeError("Missing bounded input snapshot") + prediction = _target_prediction(output, kwargs) + new_bytes = _tensor_bytes(pending, torch) + prediction.numel() * prediction.element_size() + if bytes_captured + new_bytes > args.max_capture_mib * 2**20: + raise MemoryError("Teacher capture output exceeds the configured bound") + payload = {"kwargs": pending, "teacher_target_prediction": prediction.detach().cpu().clone(), + "step_zero_based": call_index} + filename = f"step-{call_index:03d}.pt" + torch.save(payload, directory / filename) + bytes_captured += new_bytes + metadata["captures"].append({ + "step_zero_based": call_index, "file": filename, "sha256": _hash_file(directory / filename), + "timestep": pending["timestep"].tolist(), "cache_mode": pending["kv_cache_mode"], + "target_shape": list(prediction.shape), "tensor_bytes": new_bytes, + }) + _write_json(directory / "capture.json", metadata) + print(json.dumps({"event": "denoiser_teacher_capture", **metadata["captures"][-1]}), flush=True) + pending = None + + handles = [transformer.register_forward_pre_hook(before, with_kwargs=True), + transformer.register_forward_hook(after, with_kwargs=True)] + references = [] + for path in job.get("images", []): + with Image.open(path) as image: + references.append(image.copy()) + started = time.perf_counter() + try: + with torch.inference_mode(): + result = engine.pipe( + prompt=job["prompt"], image=references or None, width=1024, height=1024, + output_resolution=1024, num_inference_steps=steps, true_cfg_scale=1.0, + generator=torch.Generator(device="cuda:0").manual_seed(job.get("seed", 42)), + use_kv_cache=True, output_type="latent", + ) + del result + torch.cuda.synchronize() + if call_index + 1 != steps or [row["step_zero_based"] for row in metadata["captures"]] != selected: + raise RuntimeError("Teacher trajectory/capture count does not match requested steps") + metadata.update(complete=True, tensor_bytes=bytes_captured, + instrumented_seconds=time.perf_counter() - started, + transformer_calls=call_index + 1) + _write_json(directory / "capture.json", metadata) + finally: + for handle in handles: + handle.remove() + for cache in engine._kv_caches.values(): + _clear_cache(cache) + engine._kv_caches.clear() + engine.pipe.maybe_free_model_hooks() + pending = None + gc.collect() + torch.cuda.empty_cache() + + +def _load_capture(directory, record, torch): + path = directory / record["file"] + if path.parent.resolve() != directory.resolve() or _hash_file(path) != record["sha256"]: + raise ValueError("Capture path or checksum mismatch") + payload = torch.load(path, map_location="cpu", weights_only=True) + if "kv_cache" in payload["kwargs"]: + raise ValueError("Refuse to reuse captured teacher KV") + if payload["step_zero_based"] != record["step_zero_based"]: + raise ValueError("Capture step metadata mismatch") + return payload + + +def _checkpoint_spec(spec): + name, separator, path = spec.partition("=") + if not separator or not name or not path: + raise ValueError("Checkpoint must be NAME=/absolute/checkpoint/path") + return name, Path(path) + + +def compare(args): + import torch + from diffusers.models.transformers.transformer_qwenimage21 import QwenImage21KVCache + from .runtime import load_transformer + + directory = Path(args.capture) + metadata = json.loads((directory / "capture.json").read_text()) + if metadata.get("format") != FORMAT or metadata.get("complete") is not True: + raise ValueError("Teacher capture is incomplete or incompatible") + records = sorted(metadata["captures"], key=lambda row: row["step_zero_based"]) + if len(records) != 3 or records[0]["step_zero_based"] != 0: + raise ValueError("Expected exactly three captures, beginning at step0") + checkpoints = [_checkpoint_spec(spec) for spec in args.checkpoint] + if len({name for name, _ in checkpoints}) != len(checkpoints): + raise ValueError("Checkpoint names must be unique") + restore_roles = getattr(args, "restore_role", []) + restore_names = getattr(args, "restore_name", []) + bf16_source = getattr(args, "bf16_source", None) + if bool(restore_roles or restore_names) != bool(bf16_source): + raise ValueError("Hybrid diagnostics require both --bf16-source and at least one restoration selector") + if bf16_source: + from .hybrid_v3 import _selected_names + _selected_names(restore_roles, restore_names) + report = { + "format": "qwen21-denoiser-comparison-v3", "complete": False, + "capture_manifest_sha256": _hash_file(directory / "capture.json"), + "teacher_capture": metadata, "backends": [], + "cache_policy": "Fresh backend-owned prefix rebuilt from step0 separately before every cached-step probe", + "claim": "Conditional denoiser prediction error on teacher trajectory; final image evaluation still required", + } + _write_json(args.out, report) + for name, checkpoint in checkpoints: + hybrid_report = None + model = load_transformer(checkpoint, device="cpu" if bf16_source else "cuda:0") + if bf16_source: + from .hybrid_v3 import restore_bf16_projections + hybrid_report = restore_bf16_projections(model, bf16_source, roles=restore_roles, names=restore_names) + model.to("cuda:0") + if not model.config.causal_condition: + raise ValueError("Compared checkpoint is not causal") + backend = {"name": name, "checkpoint": str(checkpoint), + "manifest_sha256": _hash_file(checkpoint / "manifest.json"), "probes": []} + if hybrid_report is not None: + backend["hybrid_override"] = hybrid_report + report["backends"].append(backend) + try: + for record in records: + cache = QwenImage21KVCache(len(model.transformer_blocks)) + inputs = None + prefill = None + prediction = None + started = time.perf_counter() + torch.cuda.reset_peak_memory_stats() + try: + with torch.inference_mode(): + if record["step_zero_based"] != 0: + prefill = _load_capture(directory, records[0], torch) + prefill_inputs = _tree(prefill["kwargs"], torch, "cuda:0") + prefill_inputs.update(kv_cache=cache, kv_cache_mode="extract", return_dict=False) + with model.cache_context("cond"): + prefill_output = model(**prefill_inputs) + _cache_summary(cache) # Must be fully backend-populated before decode. + del prefill_output, prefill_inputs, prefill + prefill = None + payload = _load_capture(directory, record, torch) + inputs = _tree(payload["kwargs"], torch, "cuda:0") + mode = "extract" if record["step_zero_based"] == 0 else "cached" + if inputs["kv_cache_mode"] != mode: + raise ValueError("Captured cache mode does not match selected step") + inputs.update(kv_cache=cache, kv_cache_mode=mode, return_dict=False) + with model.cache_context("cond"): + output = model(**inputs) + prediction = _target_prediction(output, inputs) + metrics = _metrics(prediction, payload["teacher_target_prediction"], torch) + cache_info = _cache_summary(cache) + torch.cuda.synchronize() + row = {"step_zero_based": record["step_zero_based"], "timestep": record["timestep"], + "cache_mode": mode, "fresh_prefill_calls": 1, "cache": cache_info, + "metrics": metrics, "instrumented_seconds_including_prefill": time.perf_counter() - started, + "peak_allocated_mib": torch.cuda.max_memory_allocated() / 2**20, + "peak_reserved_mib": torch.cuda.max_memory_reserved() / 2**20} + backend["probes"].append(row) + print(json.dumps({"event": "denoiser_probe", "backend": name, **row}), flush=True) + _write_json(args.out, report) + del output, payload + finally: + _clear_cache(cache) + cache = inputs = prefill = prediction = None + gc.collect() + torch.cuda.empty_cache() + finally: + del model + gc.collect() + torch.cuda.empty_cache() + report["complete"] = True + _write_json(args.out, report) + + +def main(): + parser = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter) + commands = parser.add_subparsers(dest="command", required=True) + cap = commands.add_parser("capture", help="Run one BF16 teacher trajectory and store only three probes") + cap.add_argument("--jobs", required=True) + cap.add_argument("--label", required=True) + cap.add_argument("--out", required=True) + cap.add_argument("--no-stream", action="store_true") + cap.add_argument("--max-capture-mib", type=int, default=512) + comp = commands.add_parser("compare", help="Compare checkpoints sequentially with independent prefix caches") + comp.add_argument("--capture", required=True) + comp.add_argument("--checkpoint", action="append", required=True) + comp.add_argument("--out", required=True) + comp.add_argument("--bf16-source", help="Optional original BF16 snapshot for hybrid projection diagnostics") + comp.add_argument("--restore-role", action="append", default=[], help="Restore a role in all 32 blocks, e.g. attn.to_q") + comp.add_argument("--restore-name", action="append", default=[], help="Restore one exact transformer projection name") + args = parser.parse_args() + (capture if args.command == "capture" else compare)(args) + + +if __name__ == "__main__": + main() diff --git a/reproduction/nunchaku_backend/diagnose_kernel.py b/reproduction/nunchaku_backend/diagnose_kernel.py new file mode 100644 index 0000000000000000000000000000000000000000..a3f51e549d1025a854e43b430bdc268cd319f406 --- /dev/null +++ b/reproduction/nunchaku_backend/diagnose_kernel.py @@ -0,0 +1,45 @@ +"""GPU-only isolated residual/low-rank diagnosis after a kernel probe failure.""" +import argparse +import json +from pathlib import Path + + +def main(): + p=argparse.ArgumentParser() + p.add_argument('--model-path',type=Path,required=True) + p.add_argument('--calibration',type=Path,required=True) + p.add_argument('--out',type=Path,required=True) + a=p.parse_args() + import torch + from safetensors.torch import load_file + from nunchaku.models.linear import SVDQW4A4Linear + from .baseline_candidate import convert_linear_weight,reference_forward + from .checkpoint_io import _source_index,_read_tensor + from .kernel_probe import _errors + torch.backends.cuda.matmul.allow_tf32=False + torch.set_num_threads(4) + name='transformer_blocks.0.img_mlp.out' + index=_source_index(a.model_path/'transformer') + weight=_read_tensor(index,name+'.weight') + amax=load_file(str(a.calibration/'activation_stats.safetensors'))[name+'.input_absmax'] + packed,stats,reference=convert_linear_weight(weight,rank=32,input_absmax=amax,seed=102,return_reference=True,conversion_device='cuda:0') + layer=SVDQW4A4Linear(weight.shape[1],weight.shape[0],rank=32,bias=False,torch_dtype=torch.bfloat16,device='cuda:0') + layer.load_state_dict(packed) + layer.eval().requires_grad_(False) + reference={k:v.cuda() for k,v in reference.items()} + gen=torch.Generator(device='cpu').manual_seed(173) + x=torch.randn(1,257,weight.shape[1],generator=gen).to(torch.bfloat16).cuda() + rows={} + with torch.inference_mode(): + rows['full']=_errors(layer(x),reference_forward(x,reference)) + layer.proj_up.zero_() + ref0={**reference,'up_unpacked':torch.zeros_like(reference['up_unpacked'])} + rows['residual_only']=_errors(layer(x),reference_forward(x,ref0)) + layer.load_state_dict(packed) + layer.qweight.zero_() + ref1={**reference,'residual_dequant':torch.zeros_like(reference['residual_dequant'])} + rows['lowrank_only']=_errors(layer(x),reference_forward(x,ref1)) + a.out.write_text(json.dumps(rows,indent=2)+'\n') + print(json.dumps(rows),flush=True) + +if __name__=='__main__':main() diff --git a/reproduction/nunchaku_backend/export_v3.py b/reproduction/nunchaku_backend/export_v3.py new file mode 100644 index 0000000000000000000000000000000000000000..c042289bb2d54dfc903e5423e3576f5ff7fa34f5 --- /dev/null +++ b/reproduction/nunchaku_backend/export_v3.py @@ -0,0 +1,249 @@ +"""Export isolated activation-output calibrated checkpoints, or probe layers. + +Activation layout: one root containing keys ``.train``, +``.validation``, ``.heldout`` in any *.safetensors shards; +alternatively pass three split directories with ``.inputs`` keys. +Each tensor is BF16 [sample_rows,in_features]. Different jobs must supply +the three splits; this exporter does not invent independent test data. +""" +from __future__ import annotations + +import argparse +import hashlib +import json +from pathlib import Path +import time + +from . import FORMAT +from .layout import BLOCK_LINEAR_PATTERN +from .checkpoint_io import _json_write, _source_index, _read_tensor + + +def _hash_file(path): + digest = hashlib.sha256() + with path.open("rb") as handle: + while block := handle.read(8 * 1024 * 1024): + digest.update(block) + return digest.hexdigest() + + +class ActivationReader: + """Lazy single-layer reads keep large activation archives out of RAM.""" + def __init__(self, root=None, train=None, validation=None, heldout=None): + from safetensors import safe_open + self.entries = {} + self.provenance = {"files": {}, "metadata": {}} + if root is not None: + if any(v is not None for v in (train, validation, heldout)): + raise ValueError("Use --activations or three split directories, not both") + roots = [(None, Path(root))] + elif all(v is not None for v in (train, validation, heldout)): + roots = [("train", Path(train)), ("validation", Path(validation)), ("heldout", Path(heldout))] + else: + raise ValueError("Supply activation root or all three split directories") + for split, directory in roots: + if not directory.is_dir(): + raise ValueError(f"Missing activation directory: {directory}") + for metadata_path in (directory / "activations.json", directory / "calibration.json", directory / "manifest.json"): + if metadata_path.exists(): + self.provenance["metadata"][str(metadata_path)] = json.loads(metadata_path.read_text()) + for path in sorted(directory.rglob("*.safetensors")): + self.provenance["files"][str(path)] = _hash_file(path) + with safe_open(str(path), framework="pt", device="cpu") as handle: + for key in handle.keys(): + layer, _, suffix = key.rpartition(".") + # Parent's collector uses one .safetensors per + # layer, with short train/validation/heldout keys. + if key in ("train", "validation", "heldout", "input_absmax") and BLOCK_LINEAR_PATTERN.fullmatch(path.stem): + layer, suffix = path.stem, key + if split is not None and suffix == "inputs": + key_split = split + elif suffix in ("train", "validation", "heldout") and (split is None or suffix == split): + key_split = suffix + elif suffix == "input_absmax" and split in (None, "train"): + key_split = suffix + else: + continue + if not BLOCK_LINEAR_PATTERN.fullmatch(layer): + continue + identity = (layer, key_split) + if identity in self.entries: + raise ValueError(f"Duplicate activation tensor {identity}") + self.entries[identity] = (path, key) + self.fingerprint = hashlib.sha256(json.dumps(self.provenance, sort_keys=True).encode()).hexdigest() + + def require(self, names): + missing = [(name, split) for name in names for split in ("train", "validation", "heldout") + if (name, split) not in self.entries] + if missing: + raise ValueError(f"Missing activation splits: {missing[:12]} (total {len(missing)})") + + def read(self, name, split): + from safetensors import safe_open + path, key = self.entries[name, split] + with safe_open(str(path), framework="pt", device="cpu") as handle: + return handle.get_tensor(key) + + +def main(): + parser = argparse.ArgumentParser(description=__doc__) + parser.add_argument("--model-path", type=Path, required=True) + parser.add_argument("--out", type=Path, required=True) + parser.add_argument("--activations", type=Path) + parser.add_argument("--activation-train", type=Path) + parser.add_argument("--activation-validation", type=Path) + parser.add_argument("--activation-heldout", type=Path) + parser.add_argument("--baseline-calibration", type=Path, help="Original absmax archive for apples-to-apples v1 baseline") + parser.add_argument("--ranks", type=int, nargs="+", default=[32, 64, 128]) + parser.add_argument("--alphas", type=float, nargs="+", default=[0.25, 0.5, 0.75]) + parser.add_argument("--families", nargs="+", choices=["smoothquant", "activation_only"], default=["smoothquant", "activation_only"]) + parser.add_argument("--weighting", nargs="+", choices=["none", "rms"], default=["none", "rms"]) + parser.add_argument("--iterations", type=int, default=8) + parser.add_argument("--ridge", type=float, default=0.01) + parser.add_argument("--no-output-correction", action="store_true") + parser.add_argument("--final-gptq", action="store_true") + parser.add_argument("--gptq-damp", type=float, default=0.01) + parser.add_argument("--factorization", choices=["balanced", "up_singular"], default="balanced") + parser.add_argument("--fixed-smoothing", choices=["identity", "activation_only", "smoothquant"], + help="Probe one exact smoothing family/alpha; excludes identity unless selected") + parser.add_argument("--baseline-rank", type=int, default=32) + parser.add_argument("--seed", type=int, default=1947) + parser.add_argument("--niter", type=int, default=4) + parser.add_argument("--oversample", type=int, default=16) + parser.add_argument("--threads", type=int, default=8) + parser.add_argument("--device", default="cpu") + parser.add_argument("--objective-backend", choices=["auto", "nunchaku", "reference"], default="auto") + parser.add_argument("--resume", action="store_true") + parser.add_argument("--max-blocks", type=int) + parser.add_argument("--layers", nargs="+", help="Bounded per-layer probe; checkpoint remains incomplete") + parser.add_argument("--verbose-candidates", action="store_true") + args = parser.parse_args() + import torch + from importlib.metadata import version + from safetensors.torch import save_file, load_file + from .optimize_v3 import optimize_linear_weight + from .packing import DEEPCOMPRESSOR_COMMIT + torch.set_num_threads(args.threads) + directory = args.model_path / "transformer" if (args.model_path / "transformer").is_dir() else args.model_path + config = json.loads((directory / "config.json").read_text()) + if config.get("quantization_config") or config.get("_class_name") != "QwenImage21Transformer2DModel" or config.get("num_layers", 32) != 32: + raise ValueError("Source must be original BF16 Qwen Image2.1 transformer") + index = _source_index(directory) + all_names = sorted({name[:-7] for name in index if name.endswith(".weight") and BLOCK_LINEAR_PATTERN.fullmatch(name[:-7])}, + key=lambda name: (int(name.split(".")[1]), name)) + if len(all_names) != 224: + raise ValueError(f"Expected224 BF16 block linears, found{len(all_names)}") + names = args.layers or all_names + if len(set(names)) != len(names) or set(names) - set(all_names): + raise ValueError("--layers contains duplicates or unknown layer names") + reader = ActivationReader(args.activations, args.activation_train, args.activation_validation, args.activation_heldout) + reader.require(names) + baseline = {} + baseline_hash = None + if args.baseline_calibration: + path = args.baseline_calibration + if path.is_dir(): + path = path / "activation_stats.safetensors" + baseline = load_file(str(path), device="cpu") + baseline_hash = _hash_file(path) + if any(name + ".input_absmax" not in baseline for name in names): + raise ValueError("Baseline calibration archive is incomplete") + settings = {"ranks": args.ranks, "alphas": args.alphas, "families": args.families, "weighting": args.weighting, + "iterations": args.iterations, "ridge": args.ridge, "output_correction": not args.no_output_correction, + "baseline_rank": args.baseline_rank, "seed": args.seed, "niter": args.niter, "oversample": args.oversample, + "device": args.device, "objective_backend": args.objective_backend, "layers": args.layers} + settings.update({"final_gptq": args.final_gptq, "gptq_damp": args.gptq_damp, "factorization": args.factorization, + "fixed_smoothing": args.fixed_smoothing}) + identity = {"algorithm": "activation-output-v3", "source_config_sha256": _hash_file(directory / "config.json"), + "activation_fingerprint": reader.fingerprint, "baseline_calibration_sha256": baseline_hash, + "deepcompressor_commit": DEEPCOMPRESSOR_COMMIT, "settings": settings} + args.out.mkdir(parents=True, exist_ok=True) + manifest_path = args.out / "manifest.json" + if manifest_path.exists(): + manifest = json.loads(manifest_path.read_text()) + if not args.resume or manifest.get("conversion_identity") != identity: + raise ValueError("Existing export requires --resume and identical inputs/settings") + for filename in manifest.get("files", {}): + if not (args.out / filename).is_file(): + raise ValueError(f"Resume checkpoint shard missing: {filename}") + else: + if any(args.out.iterdir()): + raise ValueError("Use a new empty output directory") + manifest = {"backend_format": FORMAT, "complete": False, "model": "Qwen/Qwen-Image-2.1", + "model_revision": "b3179ad355be050328e483a9dfdd9e60cd62adfa", + "diffusers_commit": "80c7ed262aeffbeb43ef13ae04baeb9b84515a69", + "conversion_identity": identity, "calibration": reader.provenance, "created_unix": time.time(), + "algorithm": "v3 output-selected iterative SVDQuant with optional RMS weighting and ridge output correction", + "runtime": {"torch": torch.__version__, "nunchaku": version("nunchaku") if args.device.startswith("cuda") else "not_loaded_cpu_reference"}, + "limitations": "Experimental per-linear output calibration, not block-output or end-to-end equivalence; optional final GPTQ selected by validation.", + "layers": {}, "files": {}} + _json_write(args.out / "config.json", config) + _json_write(manifest_path, manifest) + # Per-layer shards support interrupted long searches without repeating an + # entire block; the existing runtime supports arbitrary shard counts. + if not args.layers and "model-boundary.safetensors" not in manifest["files"]: + converted = {name + suffix for name in all_names for suffix in (".weight", ".bias")} + boundaries = {name: _read_tensor(index, name).contiguous() for name in index if name not in converted} + filename = "model-boundary.safetensors" + save_file(boundaries, str(args.out / filename)) + manifest["files"][filename] = {"keys": sorted(boundaries), "bytes": sum(t.numel() * t.element_size() for t in boundaries.values())} + del boundaries + _json_write(manifest_path, manifest) + first_block = None + for name in names: + if name in manifest["layers"]: + continue + block = int(name.split(".")[1]) + if first_block is None: + first_block = block + if args.max_blocks and block >= first_block + args.max_blocks: + break + ordinal = all_names.index(name) + weight = _read_tensor(index, name + ".weight") + bias = _read_tensor(index, name + ".bias") if name + ".bias" in index else None + if weight.dtype != torch.bfloat16: + raise ValueError(f"Source weight is {weight.dtype}, expected original BF16: {name}") + tensors = {split: reader.read(name, split) for split in ("train", "validation", "heldout")} + train_absmax = reader.read(name, "input_absmax") if (name, "input_absmax") in reader.entries else None + def progress(row): + print(json.dumps({"event": "candidate", "layer": name, **row}), flush=True) + packed, stats, reference = optimize_linear_weight( + weight, tensors["train"], tensors["validation"], tensors["heldout"], bias, + ranks=tuple(args.ranks), alphas=tuple(args.alphas), smoothing_families=tuple(args.families), + weighting=tuple(args.weighting), iterations=args.iterations, ridge=args.ridge, + output_correction=not args.no_output_correction, baseline_rank=args.baseline_rank, + baseline_absmax=baseline.get(name + ".input_absmax"), baseline_seed=args.seed + ordinal, + train_absmax=train_absmax, + final_gptq=args.final_gptq, gptq_damp=args.gptq_damp, + factorization=args.factorization, + fixed_smoothing=args.fixed_smoothing, + seed=args.seed + ordinal, niter=args.niter, oversample=args.oversample, + conversion_device=args.device, objective_backend=args.objective_backend, + progress=progress if args.verbose_candidates else None, + ) + state = {name + "." + key: value.detach().cpu().contiguous() for key, value in packed.items()} + filename = f"model-layer-{ordinal:03d}.safetensors" + temporary = args.out / (filename + ".tmp") + save_file(state, str(temporary)) + temporary.replace(args.out / filename) + manifest["files"][filename] = {"keys": sorted(state), "bytes": sum(t.numel() * t.element_size() for t in state.values()), + "layer_stats": {name: stats}} + manifest["layers"][name] = {"in_features": weight.shape[1], "out_features": weight.shape[0], + "rank": stats["selected"]["rank"], "bias": bias is not None, "precision": "int4"} + _json_write(manifest_path, manifest) + print(json.dumps({"event": "optimized_linear", "name": name, **{k: v for k, v in stats.items() if k != "history"}}), flush=True) + del packed, state, reference, weight, bias, tensors + manifest["complete"] = len(manifest["layers"]) == 224 and not args.layers + if manifest["complete"]: + _json_write(args.out / "model.safetensors.index.json", { + "metadata": {"total_size": sum(info["bytes"] for info in manifest["files"].values())}, + "weight_map": {key: filename for filename, info in manifest["files"].items() for key in info["keys"]}, + }) + manifest["completed_unix"] = time.time() + _json_write(manifest_path, manifest) + print(json.dumps({"event": "export_v3_status", "out": str(args.out), "complete": manifest["complete"], + "layers": len(manifest["layers"])}), flush=True) + + +if __name__ == "__main__": + main() diff --git a/reproduction/nunchaku_backend/export_v3_test.py b/reproduction/nunchaku_backend/export_v3_test.py new file mode 100644 index 0000000000000000000000000000000000000000..e59c22c54ee7e36e66e8facf1281dcee0cac939e --- /dev/null +++ b/reproduction/nunchaku_backend/export_v3_test.py @@ -0,0 +1,44 @@ +"""Small contract checks for the collector/exporter activation interface.""" +import json +from pathlib import Path +import tempfile +import unittest + +import torch +from safetensors.torch import save_file + +from .export_v3 import ActivationReader + + +class ActivationReaderTests(unittest.TestCase): + def test_parent_collector_layout_and_training_absmax(self): + name = "transformer_blocks.15.img_mlp.gate_layer" + with tempfile.TemporaryDirectory() as temporary: + path = Path(temporary) + tensors = {split: torch.full((8, 128), index, dtype=torch.bfloat16) + for index, split in enumerate(("train", "validation", "heldout"))} + tensors["input_absmax"] = torch.ones(128) + save_file(tensors, str(path / (name + ".safetensors"))) + (path / "manifest.json").write_text(json.dumps({"source": "bf16_teacher", "jobs_disjoint": True})) + reader = ActivationReader(path) + reader.require([name]) + self.assertEqual(len(reader.entries), 4) + self.assertEqual(len(reader.fingerprint), 64) + for key, value in tensors.items(): + self.assertTrue(torch.equal(reader.read(name, key), value)) + self.assertTrue(reader.provenance["metadata"]) + with self.assertRaises(ValueError): + reader.require(["transformer_blocks.0.attn.to_q"]) + + def test_duplicate_split_fails(self): + name = "transformer_blocks.0.attn.to_q" + with tempfile.TemporaryDirectory() as temporary: + path = Path(temporary) + save_file({"train": torch.ones(8, 128)}, str(path / (name + ".safetensors"))) + save_file({name + ".train": torch.ones(8, 128)}, str(path / "duplicate.safetensors")) + with self.assertRaises(ValueError): + ActivationReader(path) + + +if __name__ == "__main__": + unittest.main() diff --git a/reproduction/nunchaku_backend/gptq_v3.py b/reproduction/nunchaku_backend/gptq_v3.py new file mode 100644 index 0000000000000000000000000000000000000000..142403e4bca47f3bfdcb23671acb913c5e39aa38 --- /dev/null +++ b/reproduction/nunchaku_backend/gptq_v3.py @@ -0,0 +1,164 @@ +"""Optional Hessian-aware residual quantization; isolated from stable export. + +Follows the inverse-Hessian Cholesky error-propagation recurrence in pinned +DeepCompressor quantizer/kernel/gptq.py, without activation-order permutation. +Scales remain attached to original input groups of 64. GPTQ is a candidate: +validate its full A4+low-rank output before selecting it for a checkpoint. +""" + +from __future__ import annotations + +import time + +import torch + + +def _quantize_rtn(weight, scales, group_size): + grouped = weight.reshape(weight.shape[0], -1, group_size) + integers = (grouped / scales.float().unsqueeze(-1)).round().clamp(-7, 7) + return (integers * scales.float().unsqueeze(-1)).reshape_as(weight) + + +@torch.no_grad() +def quantize_residual_gptq( + weight_smoothed: torch.Tensor, + smoothed_train_inputs: torch.Tensor, + group_size: int = 64, + damp: float = 0.01, + block_size: int = 128, + *, + conversion_device: str | torch.device = "cpu", + max_damp_tries: int = 8, + hessian_chunk_size: int = 512, + return_diagnostics: bool = False, +) -> tuple: + """Return (dequantized_residual_FP32, scales_BF16), optionally diagnostics. + + ``weight_smoothed`` is the dense residual AFTER subtracting the selected + low-rank branch, in smoothed coordinates. Inputs must use those same + smoothed coordinates. There is no low-rank refit or activation A4 fitting. + + CPU is the default; CUDA requires ``conversion_device='cuda:'``. + The full FP32 KxK Hessian and factorization workspace are allocated. + ``hessian_chunk_size`` limits only sample accumulation, not Hessian width. + Input tensors are not mutated, and original channel/group order is kept. + + Zero-observation columns preserve their ordinary RTN values instead of + zeroing unknown weights. Entirely zero calibration data returns RTN. + Cholesky retries increase damping tenfold; exhaustion fails explicitly. + """ + started = time.perf_counter() + device = torch.device(conversion_device) + if device.type not in ("cpu", "cuda") or (device.type == "cuda" and device.index is None): + raise ValueError("Use cpu or an explicitly indexed CUDA conversion device") + if group_size != 64: + raise ValueError("This Nunchaku residual exporter supports only group_size=64") + if not isinstance(block_size, int) or block_size <= 0: + raise ValueError("block_size must be a positive integer") + if not isinstance(max_damp_tries, int) or max_damp_tries < 1 or hessian_chunk_size < 1: + raise ValueError("Damping retries and Hessian sample chunk size must be positive") + if not 0 < damp < float("inf"): + raise ValueError("damp must be finite and positive") + if not weight_smoothed.is_floating_point() or weight_smoothed.ndim != 2: + raise ValueError("weight_smoothed must be a floating point [output,input] matrix") + oc, ic = weight_smoothed.shape + if not oc or not ic or ic % group_size: + raise ValueError("Weight dimensions must be positive and input width divisible by64") + if (not smoothed_train_inputs.is_floating_point() or smoothed_train_inputs.ndim != 2 + or smoothed_train_inputs.shape[1] != ic or not smoothed_train_inputs.shape[0]): + raise ValueError("smoothed_train_inputs must be a nonempty floating point [samples,input] matrix") + weight = weight_smoothed.detach().to(device=device, dtype=torch.float32) + inputs = smoothed_train_inputs.detach().to(device=device, dtype=torch.float32) + if not torch.isfinite(weight).all() or not torch.isfinite(inputs).all(): + raise ValueError("Weight and calibration inputs must be finite") + scales = (weight.reshape(oc, -1, group_size).abs().amax(-1) / 7).clamp_min( + torch.finfo(torch.bfloat16).tiny + ).to(torch.bfloat16) + if not torch.isfinite(scales).all(): + raise ValueError("Residual scales overflow BF16") + rtn = _quantize_rtn(weight, scales, group_size) + n = inputs.shape[0] + hessian = torch.zeros((ic, ic), dtype=torch.float32, device=device) + for start in range(0, n, hessian_chunk_size): + chunk = inputs[start:start + hessian_chunk_size] + hessian.addmm_(chunk.T, chunk, beta=1.0, alpha=2.0 / n) + if not torch.isfinite(hessian).all(): + raise ValueError("Input covariance overflowed FP32") + diagonal = hessian.diagonal() + dead = diagonal == 0 + mean_diagonal = diagonal.mean().item() + diagnostics = { + "algorithm": "gptq_residual_original_order", + "group_size": group_size, + "block_size": block_size, + "calibration_rows": n, + "input_features": ic, + "conversion_device": str(device), + "hessian_dtype": "torch.float32", + "hessian_bytes": ic * ic * 4, + "requested_damp": damp, + "zero_observation_columns": int(dead.sum().item()), + "act_order": False, + "limitations": "Fixed BF16 scales; full Hessian; no A4 activation optimization; require held-out full-layer validation.", + } + if mean_diagonal == 0: + result = rtn + diagnostics.update({"status": "zero_inputs_rtn", "effective_damp": 0.0, "damping_tries": 0}) + else: + # A zero covariance column has no calibrated influence. Giving its + # diagonal a positive value decouples it while preserving RTN weights. + diagonal[dead] = mean_diagonal + upper = None + effective_damp = damp + for attempt in range(max_damp_tries): + regularized = hessian.clone() + regularized.diagonal().add_(effective_damp * mean_diagonal) + lower, info = torch.linalg.cholesky_ex(regularized, check_errors=False) + del regularized + if int(info.item()) == 0 and torch.isfinite(lower).all(): + inverse = torch.cholesky_inverse(lower) + candidate, upper_info = torch.linalg.cholesky_ex(inverse, upper=True, check_errors=False) + del inverse + if int(upper_info.item()) == 0 and torch.isfinite(candidate).all(): + upper = candidate + del lower + break + del candidate + del lower + effective_damp *= 10 + if upper is None: + raise RuntimeError("GPTQ inverse-Hessian factorization failed after damping retries") + del hessian + working = weight.clone() + integers = torch.empty_like(working) + for start in range(0, ic, block_size): + end = min(start + block_size, ic) + current = working[:, start:end].clone() + errors = torch.zeros_like(current) + factor_block = upper[start:end, start:end] + for local in range(end - start): + channel = start + local + column = current[:, local].clone() + column_scale = scales[:, channel // group_size].float() + q = (column / column_scale).round().clamp(-7, 7) + integers[:, channel] = q + error = (column - q * column_scale) / factor_block[local, local] + errors[:, local] = error + current[:, local:] -= error[:, None] * factor_block[local, local:][None, :] + if end < ic: + working[:, end:] -= errors @ upper[start:end, end:] + result = (integers.reshape(oc, -1, group_size) * scales.float().unsqueeze(-1)).reshape(oc, ic) + if not torch.isfinite(result).all(): + raise RuntimeError("GPTQ produced a nonfinite residual") + diagnostics.update({"status": "quantized", "effective_damp": effective_damp, "damping_tries": attempt + 1}) + if return_diagnostics: + teacher = inputs @ weight.T + denom = teacher.double().square().mean().item() + for label, candidate in (("gptq", result), ("rtn", rtn)): + error = (inputs @ candidate.T).double() - teacher.double() + mse = error.square().mean().item() + diagnostics[label + "_train_mse"] = mse + diagnostics[label + "_train_relative_l2"] = (mse / max(denom, 1e-30)) ** 0.5 + diagnostics["seconds"] = time.perf_counter() - started + return result.contiguous(), scales.contiguous(), diagnostics + return result.contiguous(), scales.contiguous() diff --git a/reproduction/nunchaku_backend/gptq_v3_test.py b/reproduction/nunchaku_backend/gptq_v3_test.py new file mode 100644 index 0000000000000000000000000000000000000000..78f50f99ceb2759a5f40f69c0b693547d51310ea --- /dev/null +++ b/reproduction/nunchaku_backend/gptq_v3_test.py @@ -0,0 +1,98 @@ +"""CPU tests for fixed-scale, original-order residual GPTQ.""" + +import json +import unittest + +import torch + +from .gptq_v3 import quantize_residual_gptq + + +def rtn(weight, scales): + return ( + (weight.reshape(weight.shape[0], -1, 64) / scales.float().unsqueeze(-1)).round().clamp(-7, 7) + * scales.float().unsqueeze(-1) + ).reshape_as(weight) + + +class GptqResidualTests(unittest.TestCase): + def setUp(self): + torch.set_num_threads(2) + + def test_correlated_rank_deficient_calibration_improves_unseen_outputs(self): + generator = torch.Generator().manual_seed(216) + weight = torch.randn(128, 128, generator=generator) * 0.05 + mixing = torch.randn(12, 128, generator=generator) + train = torch.randn(48, 12, generator=generator) @ mixing + torch.randn(48, 128, generator=generator) * 0.1 + heldout = torch.randn(96, 12, generator=generator) @ mixing + torch.randn(96, 128, generator=generator) * 0.1 + original_weight, original_inputs = weight.clone(), train.clone() + quantized, scales, diagnostic = quantize_residual_gptq(weight, train, return_diagnostics=True) + baseline = rtn(weight, scales) + teacher = heldout @ weight.T + gptq_mse = ((heldout @ quantized.T).double() - teacher.double()).square().mean() + rtn_mse = ((heldout @ baseline.T).double() - teacher.double()).square().mean() + self.assertLess(gptq_mse.item(), rtn_mse.item() * 0.3) + self.assertLess(diagnostic["gptq_train_mse"], diagnostic["rtn_train_mse"]) + self.assertTrue(torch.equal(weight, original_weight)) + self.assertTrue(torch.equal(train, original_inputs)) + self.assertFalse(diagnostic["act_order"]) + self.assertEqual(scales.dtype, torch.bfloat16) + # The output stays exactly on each original group's integer grid. + integer = (quantized.reshape(128, 2, 64) / scales.float().unsqueeze(-1)).round() + self.assertTrue((integer.abs() <= 7).all()) + self.assertTrue(torch.equal((integer * scales.float().unsqueeze(-1)).reshape_as(weight), quantized)) + expected_scales = (weight.reshape(128, 2, 64).abs().amax(-1) / 7).to(torch.bfloat16) + self.assertTrue(torch.equal(scales, expected_scales)) + json.dumps(diagnostic, allow_nan=False) + + def test_diagonal_hessian_reduces_to_rtn_and_keeps_group_order(self): + generator = torch.Generator().manual_seed(81) + weight = torch.randn(128, 128, generator=generator) + # Greatly different group ranges expose accidental scale reordering. + weight[:, 64:] *= 32 + train = torch.eye(128) * torch.linspace(0.5, 2, 128) + quantized, scales = quantize_residual_gptq(weight, train, block_size=37) + self.assertTrue(torch.equal(quantized, rtn(weight, scales))) + + def test_zero_inputs_and_unobserved_columns_preserve_rtn(self): + generator = torch.Generator().manual_seed(82) + weight = torch.randn(128, 128, generator=generator) + zeros = torch.zeros(17, 128) + quantized, scales, diagnostic = quantize_residual_gptq(weight, zeros, return_diagnostics=True) + self.assertTrue(torch.equal(quantized, rtn(weight, scales))) + self.assertEqual(diagnostic["status"], "zero_inputs_rtn") + self.assertEqual(diagnostic["zero_observation_columns"], 128) + partial = torch.eye(128)[:24] + quantized, scales, diagnostic = quantize_residual_gptq(weight, partial, return_diagnostics=True) + self.assertEqual(diagnostic["zero_observation_columns"], 104) + self.assertTrue(torch.equal(quantized, rtn(weight, scales))) + self.assertTrue((quantized[:, 24:] != 0).any()) + + def test_duplicate_rows_and_zero_weights_remain_finite(self): + generator = torch.Generator().manual_seed(83) + train = torch.randn(1, 128, generator=generator).expand(8, 128) + weight = torch.randn(128, 128, generator=generator) + quantized, scales, diagnostic = quantize_residual_gptq(weight, train, damp=0.001, return_diagnostics=True) + self.assertTrue(torch.isfinite(quantized).all()) + self.assertTrue(torch.isfinite(scales).all()) + self.assertGreaterEqual(diagnostic["damping_tries"], 1) + quantized, scales = quantize_residual_gptq(torch.zeros_like(weight), train) + self.assertTrue((quantized == 0).all()) + self.assertTrue((scales > 0).all()) + + def test_invalid_configuration_and_nonfinite_data_fail_closed(self): + weight = torch.ones(128, 128) + train = torch.ones(8, 128) + for kwargs in ({"damp": 0}, {"damp": float("nan")}, {"block_size": 0}, + {"group_size": 32}, {"max_damp_tries": 0}, {"conversion_device": "cuda"}): + with self.subTest(kwargs=kwargs): + with self.assertRaises(ValueError): + quantize_residual_gptq(weight, train, **kwargs) + with self.assertRaises(ValueError): + quantize_residual_gptq(weight, torch.full_like(train, float("nan"))) + with self.assertRaises(ValueError): + quantize_residual_gptq(weight, train[:0]) + + +if __name__ == "__main__": + unittest.main(verbosity=2) diff --git a/reproduction/nunchaku_backend/hybrid_v3.py b/reproduction/nunchaku_backend/hybrid_v3.py new file mode 100644 index 0000000000000000000000000000000000000000..8eabf84f1ab4a5892402626009d1a2a9d789be81 --- /dev/null +++ b/reproduction/nunchaku_backend/hybrid_v3.py @@ -0,0 +1,263 @@ +"""Optional BF16 projection restoration for isolated Nunchaku diagnostics. + +This does not alter a checkpoint on disk or the stable runner/runtime. The +utility takes a fully CPU-resident loaded model and replaces selected packed +projection modules with original BF16 nn.Linear modules. Reported byte costs +are model-state bytes, not measured peak GPU memory. + +Example, after reserving the POC GPU:: + + python -m nunchaku_backend.hybrid_v3 generate \ + --checkpoint /cache/qwen21-nunchaku-v3 --source /cache/original-snapshot \ + --restore-role attn.to_q --restore-role attn.to_k \ + --jobs experiments/fidelity-v3/jobs-heldout.json \ + --sample-dir samples/fidelity-v3/hybrid-qk + +`self-test` runs only small synthetic CPU modules and local temporary files. +Importing this module does not import Torch or perform GPU work. +""" +from __future__ import annotations + +import argparse +import hashlib +import json +from pathlib import Path +import time + + +ROLES = ("attn.to_q", "attn.to_k", "attn.to_v", "attn.to_out.0", + "img_mlp.proj", "img_mlp.gate_layer", "img_mlp.out") + + +def _selected_names(roles, names, num_blocks=32): + from .layout import BLOCK_LINEAR_PATTERN + roles, names = tuple(roles or ()), tuple(names or ()) + if not roles and not names: + raise ValueError("Specify at least one projection role or exact projection name") + if any(role not in ROLES for role in roles): + raise ValueError(f"Unsupported role; choose from {ROLES}") + selected = {f"transformer_blocks.{block}.{role}" for role in roles for block in range(num_blocks)} + for name in names: + if not BLOCK_LINEAR_PATTERN.fullmatch(name) or not 0 <= int(name.split(".")[1]) < num_blocks: + raise ValueError(f"Not an exact Qwen Image 2.1 block projection: {name}") + selected.add(name) + return sorted(selected, key=lambda name: (int(name.split(".")[1]), name)) + + +def restore_bf16_projections(model, source, *, roles=(), names=()): + """Replace selected SVDQW4A4Linear modules atomically after CPU validation. + + All source tensors must be original BF16 weights. A quantized source, + incompatible architecture, non-CPU model, unexpected bias, or dimensional + mismatch is rejected before any projection is swapped. The source can be + either the original snapshot root or its transformer subdirectory. + """ + import torch + from .checkpoint_io import _source_index, _read_tensor + + started = time.perf_counter() + if type(model).__name__ != "QwenImage21Transformer2DModel": + raise ValueError("Loaded model must be QwenImage21Transformer2DModel") + if len(model.transformer_blocks) != 32: + raise ValueError("Expected exactly 32 transformer blocks") + non_cpu = [name for name, tensor in (*model.named_parameters(), *model.named_buffers()) + if tensor.device.type != "cpu"] + if non_cpu: + raise ValueError(f"Restore projections before moving the model to CUDA: {non_cpu[:5]}") + if getattr(model, "_qwen21_hybrid", None) is not None: + raise ValueError("Model already has hybrid overrides; load a fresh checkpoint for each diagnostic") + directory = Path(source).resolve() + if (directory / "transformer").is_dir(): + directory /= "transformer" + config_file = directory / "config.json" + config = json.loads(config_file.read_text()) + if config.get("_class_name") != "QwenImage21Transformer2DModel" or config.get("num_layers", 32) != 32: + raise ValueError("Source is not the original 32-block Qwen Image 2.1 transformer") + if config.get("quantization_config"): + raise ValueError("Restore only from original BF16 weights, never an NF4/quantized checkpoint") + selected = _selected_names(roles, names) + index = _source_index(directory) + prepared, records = [], [] + for name in selected: + old = model.get_submodule(name) + if type(old).__name__ != "SVDQW4A4Linear" or not hasattr(old, "qweight"): + raise ValueError(f"Expected an unmodified Nunchaku projection at {name}") + # The supported Engine uses a whole-transformer CPU-offload hook, so + # its next forward recursively moves the newly installed children. + # Reject a per-linear placement hook that swapping would silently lose. + if getattr(old, "_hf_hook", None) is not None: + raise ValueError(f"Unsupported per-projection Accelerate hook at {name}") + weight_key, bias_key = name + ".weight", name + ".bias" + if weight_key not in index: + raise ValueError(f"Missing original weight: {weight_key}") + weight = _read_tensor(index, weight_key) + has_bias = getattr(old, "bias", None) is not None + if has_bias != (bias_key in index): + raise ValueError(f"Source/model bias presence mismatch at {name}") + bias = _read_tensor(index, bias_key) if has_bias else None + if weight.dtype != torch.bfloat16 or weight.device.type != "cpu": + raise ValueError(f"Original weight must be CPU BF16: {name}") + if tuple(weight.shape) != (old.out_features, old.in_features): + raise ValueError(f"Source/model dimensions differ at {name}: {tuple(weight.shape)}") + if bias is not None and (bias.dtype != torch.bfloat16 or bias.shape != (old.out_features,)): + raise ValueError(f"Source bias dtype/shape mismatch at {name}") + if not bool(torch.isfinite(weight).all()) or bias is not None and not bool(torch.isfinite(bias).all()): + raise ValueError(f"Non-finite source weight/bias at {name}") + # Build on meta to avoid a redundant randomly initialized allocation; + # Parameters below own the CPU tensors read directly from the source. + with torch.device("meta"): + replacement = torch.nn.Linear(old.in_features, old.out_features, bias=has_bias, + dtype=torch.bfloat16) + replacement.weight = torch.nn.Parameter(weight.contiguous(), requires_grad=False) + if bias is not None: + replacement.bias = torch.nn.Parameter(bias.contiguous(), requires_grad=False) + replacement.eval() + packed_bytes = sum(value.numel() * value.element_size() for value in old.state_dict().values()) + restored_bytes = sum(value.numel() * value.element_size() for value in replacement.state_dict().values()) + records.append({"name": name, "in_features": old.in_features, "out_features": old.out_features, + "bias": has_bias, "original_rank": getattr(old, "rank", None), + "replaced_packed_state_bytes": packed_bytes, "bf16_state_bytes": restored_bytes, + "net_state_bytes_delta": restored_bytes - packed_bytes, + "source_weight_file": str(index[weight_key].name)}) + prepared.append((name, replacement)) + # All selected tensors/modules are verified before the first mutation. + for name, replacement in prepared: + parent, child = name.rsplit(".", 1) + model.get_submodule(parent)._modules[child] = replacement + report = { + "kind": "nunchaku-with-original-bf16-projections", "source_transformer": str(directory), + "source_config_sha256": hashlib.sha256(config_file.read_bytes()).hexdigest(), + "roles": sorted(set(roles or ())), "requested_exact_names": sorted(set(names or ())), + "restored_names": selected, "restored_count": len(selected), "layers": records, + "bf16_state_bytes": sum(row["bf16_state_bytes"] for row in records), + "replaced_packed_state_bytes": sum(row["replaced_packed_state_bytes"] for row in records), + "net_state_bytes_delta": sum(row["net_state_bytes_delta"] for row in records), + "preparation_seconds": time.perf_counter() - started, + "byte_cost_scope": "Model state tensor bytes only; not measured GPU peak or speed", + "status": "Prepared diagnostic override; no quality improvement established", + } + model._qwen21_hybrid = report + if hasattr(model, "_qwen21_backend"): + model._qwen21_backend = {**model._qwen21_backend, "hybrid_override": report} + return report + + +def generate(args): + from runner import Engine, default_args + from .denoiser_probe_v3 import _write_json + + directory = Path(args.sample_dir) + if directory.exists() and any(directory.iterdir()): + raise ValueError("Use a new separate sample directory for each hybrid configuration") + jobs = json.loads(Path(args.jobs).read_text()) + if not isinstance(jobs, list) or not jobs: + raise ValueError("Jobs must be a nonempty JSON list") + # Fail before loading any GPU component if selectors are misspelled. + _selected_names(args.restore_role, args.restore_name) + settings = default_args() + settings.backend = "nunchaku" + settings.nunchaku_checkpoint = args.checkpoint + settings.prequant = args.prequant + settings.compile = False + settings.flex = False + settings.sample_dir = str(directory) + settings.cache = False + engine = Engine(settings) + engine.pipe.transformer.to("cpu") + report = restore_bf16_projections(engine.pipe.transformer, args.source, + roles=args.restore_role, names=args.restore_name) + engine.args.hybrid_roles = report["roles"] + engine.args.hybrid_names = report["restored_names"] + engine.args.hybrid_source = report["source_transformer"] + engine.args.hybrid_state_bytes_delta = report["net_state_bytes_delta"] + engine.args.hybrid_precision = "BF16 restored projections; Nunchaku W4A4 remaining projections" + report["checkpoint"] = args.checkpoint + report["jobs_file"] = str(Path(args.jobs).resolve()) + report["jobs_sha256"] = hashlib.sha256(Path(args.jobs).read_bytes()).hexdigest() + _write_json(directory / "hybrid.json", report) + print(json.dumps({"event": "hybrid_prepared", **report}), flush=True) + # Existing whole-model offload hooks remain installed. Engine.generate's + # measured peak includes the additional restored BF16 weights. + for job in jobs: + engine.generate(job) + + +def self_test(): + """Small real-Torch CPU test: exact restoration, validation and byte costs.""" + import tempfile + import torch + from safetensors.torch import save_file + + class SVDQW4A4Linear(torch.nn.Module): + def __init__(self): + super().__init__() + self.in_features, self.out_features, self.rank = 128, 128, 32 + self.register_buffer("qweight", torch.zeros((128, 64), dtype=torch.int8)) + self.register_buffer("wscales", torch.ones((128, 2), dtype=torch.bfloat16)) + self.bias = torch.nn.Parameter(torch.zeros(128, dtype=torch.bfloat16), requires_grad=False) + + class QwenImage21Transformer2DModel(torch.nn.Module): + def __init__(self): + super().__init__() + self.transformer_blocks = torch.nn.ModuleList() + for _ in range(32): + block = torch.nn.Module() + block.attn = torch.nn.Module() + block.attn.to_q = SVDQW4A4Linear() + block.attn.to_k = SVDQW4A4Linear() + self.transformer_blocks.append(block) + + assert len(_selected_names(["attn.to_q", "attn.to_k"], [])) == 64 + with tempfile.TemporaryDirectory() as temporary: + directory = Path(temporary) + (directory / "config.json").write_text(json.dumps({"_class_name": "QwenImage21Transformer2DModel", "num_layers": 32})) + weight = torch.arange(128 * 128, dtype=torch.float32).reshape(128, 128).div(16384).to(torch.bfloat16) + bias = torch.arange(128, dtype=torch.float32).div(128).to(torch.bfloat16) + name = "transformer_blocks.0.attn.to_q" + save_file({name + ".weight": weight, name + ".bias": bias}, str(directory / "model.safetensors")) + model = QwenImage21Transformer2DModel() + untouched = model.transformer_blocks[0].attn.to_k + report = restore_bf16_projections(model, directory, names=[name]) + restored = model.get_submodule(name) + assert isinstance(restored, torch.nn.Linear) + assert torch.equal(restored.weight, weight) and torch.equal(restored.bias, bias) + assert restored.weight.device.type == "cpu" and not restored.weight.requires_grad + assert model.transformer_blocks[0].attn.to_k is untouched + assert report["restored_count"] == 1 + assert report["bf16_state_bytes"] == 2 * (128 * 128 + 128) + assert report["net_state_bytes_delta"] == report["bf16_state_bytes"] - report["replaced_packed_state_bytes"] + x = torch.ones((2, 128), dtype=torch.bfloat16) + assert torch.equal(restored(x), torch.nn.functional.linear(x, weight, bias)) + # A late validation error must not partially replace earlier modules. + model = QwenImage21Transformer2DModel() + original = model.get_submodule(name) + try: + restore_bf16_projections(model, directory, names=[name, "transformer_blocks.1.attn.to_q"]) + except ValueError: + pass + else: + raise AssertionError("Missing source tensor was not rejected") + assert model.get_submodule(name) is original + print(json.dumps({"self_test": "passed", "device": "cpu", "checks": [ + "role expansion", "exact BF16 weight/bias", "unchanged other projection", "frozen CPU module", + "exact linear output", "byte accounting", "atomic rejection of missing source"]})) + + +def main(): + parser = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter) + commands = parser.add_subparsers(dest="command", required=True) + commands.add_parser("self-test", help="Run a small CPU-only synthetic restoration test") + gen = commands.add_parser("generate", help="Run explicitly selected hybrid diagnostics on the POC GPU") + gen.add_argument("--checkpoint", required=True) + gen.add_argument("--source", required=True) + gen.add_argument("--prequant", default="/cache/qwen-nf4") + gen.add_argument("--restore-role", choices=ROLES, action="append", default=[]) + gen.add_argument("--restore-name", action="append", default=[]) + gen.add_argument("--jobs", required=True) + gen.add_argument("--sample-dir", required=True) + args = parser.parse_args() + self_test() if args.command == "self-test" else generate(args) + + +if __name__ == "__main__": + main() diff --git a/reproduction/nunchaku_backend/iterations100_probe.md b/reproduction/nunchaku_backend/iterations100_probe.md new file mode 100644 index 0000000000000000000000000000000000000000..fa4bde563db135da83bb27d2e7f7224a51f85662 --- /dev/null +++ b/reproduction/nunchaku_backend/iterations100_probe.md @@ -0,0 +1,38 @@ +# Fixed-recipe 100-iteration probe — completed; full refinement deferred + +The three-layer GPU probe completed. It did not show a meaningful held-out improvement, so the main agent decided **not to run full-model refinement**. The original V3 checkpoint remains the comparison baseline. Measured results are recorded in [`nunchaku-v3-iterations100-comparison.json`](../results/nunchaku-v3-iterations100-comparison.json). + +| Layer | Selected iteration, V3 → probe (zero-based) | Validation MSE change | Held-out MSE change | Probe time | +|---|---:|---:|---:|---:| +| block22 MLP output, non-cap control | 13 → 13 | +0.00025% | −0.00015% | 7.13 s | +| block11 attention output | 15 → 22 | −0.548% | +3.871% | 4.98 s | +| block1 Q | 15 → 18 | +0.286% | +0.120% | 4.79 s | + +Negative MSE changes indicate improvement. The control reproduced its original result within numerical variation. Attention output improved validation slightly but regressed on held-out inputs; Q regressed on both. Early stopping ended evaluation at iterations 14, 24, and 19 respectively, before the new limit of 100. These three layers do not prove that every layer would fail to improve, but they provide no evidence supporting the cost of a full replay. Overlapping candidate-history validation MSE differed by at most 0.00746% relative; the held-out regressions exceed that replay variation. + +Run `bash /poc/nunchaku_backend/probe_iterations_v3.sh` inside the existing POC image on GPU0 **after the main agent releases that GPU**. The script performs three isolated per-layer exports and a CPU comparison. It does not launch a full export, modify the full V3 checkpoint, or change serving configuration. + +| Layer | Fixed V3 smoothing | Exact per-layer seed | V3 selected iteration | Reason | +|---|---|---:|---:|---| +| block22 MLP output | activation-only alpha.25 | 2106 | 13 | Highest heldout linear error; useful non-cap control | +| block11 attention output | activation-only alpha.25 | 2025 | 15 | High error and hit16-iteration cap | +| block1 Q | activation-only alpha.75 | 1956 | 15 | Earliest cap-hitting Q layer; block0 selectediteration10 | + +All use rank128, no RMS weighting, balanced factors, randomized SVD niter4/oversample16, ridge.01, final GPTQ damp.01, and `/cache/qwen21-activation-v3-40`. Global exporter seed1947 is intentionally used: the exporter adds the original full-model ordinal to produce the exact seeds above. Passing2106 directly as the global seed would be incorrect. + +The new optional `--fixed-smoothing` switch excludes other smoothing families and the identity candidate. The existing one-pass baseline guard remains. All default behavior is unchanged when the switch is omitted; stable `convert.py` and `runtime.py` are untouched. + +This is a deterministic replay fromiteration0, extending the cap to100, because V3 stores the selected checkpoint rather than the final optimizer recurrence state. Early stopping remains unchanged. Block22 may therefore stop at the same early iteration and show no gain; this is a meaningful negative control. The report checks matching calibration/model fingerprints, recipe settings, seeds, and the numerical agreement of overlapping candidate history. + +Completed outputs: + +- `/cache/qwen-v3-iterations100/block22-mlp-out/manifest.json` +- `/cache/qwen-v3-iterations100/block11-attn-out/manifest.json` +- `/cache/qwen-v3-iterations100/block01-query/manifest.json` +- `/poc/results/nunchaku-v3-iterations100-comparison.json` + +Override the output root with `QWEN_ITER_PROBE_ROOT` to preserve an earlier run; an existing output directory is rejected. The report compares validation and heldout MSE to the **actual saved V3 metrics**, not the regenerated rank32 baseline, and marks whether heldout improvement exceeds5%. That5% flag is descriptive, not a deployment criterion. Validation alone selects candidates; heldout remains untouched by selection. Final GPTQ is attempted only on the selected raw candidate, so a longer fit is not guaranteed to beat the old GPTQ-selected V3 checkpoint. Retain V3 if the new validation score is worse. + +Measured per-layer processing totaled 16.90 seconds, excluding additional startup/file-hashing time. No automatic full-model rerun is included. + +Preparation validation: shell syntax passed; 15 CPU tests passed, including fixed-family exclusion. A saved-V3 identity comparison returned MSE ratio 1 and zero prefix difference, and mismatched fingerprints were rejected. GPU execution was subsequently coordinated and launched by the main agent. diff --git a/reproduction/nunchaku_backend/kernel_probe.py b/reproduction/nunchaku_backend/kernel_probe.py new file mode 100644 index 0000000000000000000000000000000000000000..580101314a959b461bd770716e89f35650cae973 --- /dev/null +++ b/reproduction/nunchaku_backend/kernel_probe.py @@ -0,0 +1,143 @@ +"""Numerically validate generic Nunchaku kernels before image experiments. + +This CLI DOES execute GPU work. Run only in the reserved GPU0 POC container. +It compares actual W4A4 CUDA output with an independently reconstructed A4/W4 +plus low-rank reference, and separately measures error versus original BF16. +Passing this probe checks packing/scaling correctness, not image quality. +""" +import argparse +import json +import math +from pathlib import Path +import statistics +import time + + +def _errors(actual, expected): + import torch + a, e = actual.float(), expected.float() + delta = a - e + rms = e.square().mean().sqrt() + worst_index = int(delta.abs().reshape(-1).argmax()) + worst_expected = float(e.reshape(-1)[worst_index]) + worst_actual = float(a.reshape(-1)[worst_index]) + bf16_spacing = 2.0 ** (math.frexp(abs(worst_expected))[1] - 8) if worst_expected else 2.0 ** -133 + return {"finite": bool(torch.isfinite(a).all() and torch.isfinite(e).all()), + "relative_l2": float(delta.norm() / e.norm().clamp_min(1e-12)), + "rmse": float(delta.square().mean().sqrt()), "reference_rms": float(rms), + "max_abs": float(delta.abs().max()), + "max_abs_over_reference_rms": float(delta.abs().max() / rms.clamp_min(1e-12)), + "worst_element": {"flat_index": worst_index, "expected": worst_expected, "actual": worst_actual, + "relative_abs_error": abs(worst_actual-worst_expected)/max(abs(worst_expected),1e-12), + "bf16_spacing": bf16_spacing, "approx_bf16_ulps": abs(worst_actual-worst_expected)/bf16_spacing}} + + +def main(): + parser = argparse.ArgumentParser() + parser.add_argument("--out", type=Path, required=True) + parser.add_argument("--model-path", type=Path) + parser.add_argument("--calibration", type=Path) + parser.add_argument("--rank", type=int, default=32) + parser.add_argument("--sequences", type=int, nargs="+", default=[17, 257, 4096]) + parser.add_argument("--device", default="cuda:0") + parser.add_argument("--conversion-device", default="cpu") + parser.add_argument("--threads", type=int, default=8) + parser.add_argument("--max-kernel-relative-l2", type=float, default=0.02) + parser.add_argument("--max-kernel-normalized-error", type=float, default=0.15) + args = parser.parse_args() + import torch + import torch.nn.functional as F + from safetensors.torch import load_file + from nunchaku.models.linear import SVDQW4A4Linear + from .baseline_candidate import convert_linear_weight + from .convert_reference import reference_forward_groupwise as reference_forward + from .runtime import probe_backend + from .checkpoint_io import _source_index, _read_tensor + torch.set_num_threads(args.threads) + torch.backends.cuda.matmul.allow_tf32 = False + device = torch.device(args.device) + if device.type != "cuda" or device.index is None: + parser.error("Specify the explicitly reserved CUDA device, e.g. cuda:0") + generator = torch.Generator(device="cpu").manual_seed(9127) + calibration = {} + if args.calibration: + path = args.calibration / "activation_stats.safetensors" if args.calibration.is_dir() else args.calibration + calibration = load_file(str(path), device="cpu") + if args.model_path: + directory = args.model_path / "transformer" if (args.model_path / "transformer").is_dir() else args.model_path + index = _source_index(directory) + names = ["transformer_blocks.0.attn.to_q", "transformer_blocks.0.img_mlp.gate_layer", "transformer_blocks.0.img_mlp.out"] + cases = [(name, _read_tensor(index, name + ".weight")) for name in names] + else: + cases = [] + for oc, ic in [(256, 128), (4096, 4096), (12288, 4096), (4096, 12288)]: + w = torch.randn(oc, ic, generator=generator) / ic ** 0.5 + w[:, ::257] *= 4 # exercise nonuniform input-channel smoothing + cases.append((f"synthetic_{oc}x{ic}", w.to(torch.bfloat16))) + report = {"runtime": probe_backend(), "device": str(device), "gpu": torch.cuda.get_device_name(device), + "rank": args.rank, "weight_source": str(args.model_path) if args.model_path else "synthetic", + "pass_thresholds": {"relative_l2": args.max_kernel_relative_l2, + "max_abs_over_reference_rms": args.max_kernel_normalized_error}, + "reference": "nunchaku-bf16-groupwise-ptx-a4-v2", + "reference_limitations": "Group64 BF16 accumulation plus independent PTX activation reference; low-rank reduction order need not be bit-exact.", + "cases": []} + all_passed = True + for ordinal, (name, weight) in enumerate(cases): + amax = calibration.get(name + ".input_absmax") + smooth = None + if amax is None: + smooth = torch.exp(torch.linspace(-1, 1, weight.shape[1])).to(torch.bfloat16) + packed, conversion, reference = convert_linear_weight( + weight, rank=args.rank, input_absmax=amax, smooth=smooth, + seed=100 + ordinal, return_reference=True, conversion_device=args.conversion_device, + ) + module = SVDQW4A4Linear(weight.shape[1], weight.shape[0], rank=args.rank, bias=False, + precision="int4", torch_dtype=torch.bfloat16, device="cpu") + module.load_state_dict({k: v.cpu() for k, v in packed.items()}, strict=True) + module.eval().requires_grad_(False).to(device) + gpu_weight = weight.to(device) + gpu_reference = {k: v.to(device) for k, v in reference.items()} + del reference, packed + for seq in args.sequences: + x = torch.randn(1, seq, weight.shape[1], generator=generator).to(torch.bfloat16).to(device) + with torch.inference_mode(): + expected = reference_forward(x, gpu_reference).to(torch.bfloat16) + baseline = F.linear(x, gpu_weight) + actual = module(x) + torch.cuda.synchronize(device) + kernel_error = _errors(actual, expected) + bf16_error = _errors(actual, baseline) + passed = (kernel_error["finite"] and kernel_error["relative_l2"] <= args.max_kernel_relative_l2 + and kernel_error["max_abs_over_reference_rms"] <= args.max_kernel_normalized_error) + samples = [] + for _ in range(3): + started = time.perf_counter() + module(x) + torch.cuda.synchronize(device) + samples.append(time.perf_counter() - started) + row = {"name": name, "shape": [weight.shape[0], weight.shape[1]], "sequence": seq, + "passed": passed, "actual_vs_w4a4_reference": kernel_error, + "actual_vs_original_bf16": bf16_error, "kernel_median_seconds": statistics.median(samples), + "conversion": conversion} + report["cases"].append(row) + all_passed &= passed + print(json.dumps({"event": "kernel_probe_case", **row}), flush=True) + del x, expected, baseline, actual + # Zero groups exercise quantizer edge handling independently of tolerance. + with torch.inference_mode(): + zero = module(torch.zeros(1, 17, weight.shape[1], dtype=torch.bfloat16, device=device)) + zero_passed = bool(torch.isfinite(zero).all() and (zero == 0).all()) + all_passed &= zero_passed + report["cases"].append({"name": name, "zero_input_passed": zero_passed}) + del zero, module, gpu_weight, gpu_reference + torch.cuda.empty_cache() + report["passed"] = all_passed + args.out.parent.mkdir(parents=True, exist_ok=True) + args.out.write_text(json.dumps(report, indent=2) + "\n") + print(json.dumps({"event": "kernel_probe_complete", "passed": all_passed, "out": str(args.out)}), flush=True) + if not all_passed: + raise SystemExit("Nunchaku numerical probe failed; do not proceed to image quality claims") + + +if __name__ == "__main__": + main() diff --git a/reproduction/nunchaku_backend/layout.py b/reproduction/nunchaku_backend/layout.py new file mode 100644 index 0000000000000000000000000000000000000000..cbab8ceb7d98057779d983b6f7deaea3c7f00d40 --- /dev/null +++ b/reproduction/nunchaku_backend/layout.py @@ -0,0 +1,7 @@ +"""Shared Qwen2.1 block-linear selection; safe to import without Torch.""" +import re + +BLOCK_LINEAR_PATTERN = re.compile( + r"^transformer_blocks\.\d+\.(?:attn\.(?:to_q|to_k|to_v|to_out\.0)|img_mlp\.(?:proj|out|gate_layer))$" +) + diff --git a/reproduction/nunchaku_backend/mlp_parent_probe_v3.py b/reproduction/nunchaku_backend/mlp_parent_probe_v3.py new file mode 100644 index 0000000000000000000000000000000000000000..95ec6921609a2740b4f859b4aff87630d135dcc3 --- /dev/null +++ b/reproduction/nunchaku_backend/mlp_parent_probe_v3.py @@ -0,0 +1,257 @@ +"""Evaluate saved rank candidates through Qwen2.1's complete SwiGLU MLP. + +The exact pinned expression is out(SiLU(gate_layer(x)) * proj(x)). Source +Nunchaku gate/output projections stay fixed while only the saved projection +candidate changes. The reference uses all three original BF16 projections. +No fitting, checkpoint export, Engine mutation or serving change is performed. + +Explicit GPU command, only after root releases the POC GPU:: + + python -m nunchaku_backend.mlp_parent_probe_v3 \ + --source-checkpoint /cache/qwen-nunchaku-v3 \ + --model-path /cache/original-snapshot \ + --activations /cache/qwen21-activation-v3-40 \ + --rank-probe /cache/qwen-v3-mlpproj-rank-probe \ + --out results/mlp-parent-rank-v3.json --device cuda:0 + +All candidates are scored on validation first. Diagnostic rankings are frozen +before heldout rows are read. Heldout is reported afterward and never used to +choose a rank. These are tokenwise MLP results on teacher inputs, not residual +block, denoiser, accumulated trajectory, or image-quality equivalence. +""" +from __future__ import annotations + +import argparse +import gc +import json +import math +from pathlib import Path +import time + + +def validation_rankings(rows): + """Source/earlier candidates win ties; heldout fields are never read.""" + if not rows or rows[0]["name"] != "source": + raise ValueError("First candidate must be the actual packed source") + eligible = [row for row in rows if row.get("requested_rank_matches", True)] + result = {} + for objective in ("parent", "projection"): + for row in eligible: + metric = row["validation"][objective] + if not metric.get("finite", True) or not math.isfinite(float(metric["mse"])) or metric["mse"] < 0: + raise ValueError("Invalid validation MSE") + ordered = sorted(eligible, key=lambda row: row["validation"][objective]["mse"]) + result[objective + "_ranking"] = [row["name"] for row in ordered] + result[objective + "_preferred"] = ordered[0]["name"] + result["rankings_disagree"] = result["parent_preferred"] != result["projection_preferred"] + result["decision_uses_heldout"] = False + return result + + +def _candidate_rows(report, names): + selected = {} + for name in names: + rows = sorted([row for row in report["results"] if row["layer"] == name], + key=lambda row: row["requested_rank"]) + if not rows or len({row["requested_rank"] for row in rows}) != len(rows): + raise ValueError(f"Missing or duplicate saved candidate ranks for {name}") + selected[name] = rows + return selected + + +def _saved_candidate(directory, row, torch): + from safetensors.torch import load_file + from .export_v3 import _hash_file + path = directory / row["candidate_file"] + if path.parent.resolve() != directory.resolve() or not path.is_file(): + raise ValueError("Saved candidate path must be a file inside the rank-probe directory") + if _hash_file(path) != row["candidate_sha256"]: + raise ValueError("Saved candidate checksum changed") + state = load_file(str(path), device="cpu") + if any(t.device.type != "cpu" for t in state.values()): + raise ValueError("Candidate must be read on CPU") + return state + + +def _quantized_module(state, info, device): + import torch + from nunchaku.models.linear import SVDQW4A4Linear + rank = info["rank"] + if not isinstance(rank, int) or isinstance(rank, bool) or rank <= 0 or rank % 16: + raise ValueError("Invalid saved low-rank dimension") + module = SVDQW4A4Linear(info["in_features"], info["out_features"], rank=rank, + bias=info["bias"], precision="int4", act_unsigned=False, + torch_dtype=torch.bfloat16, device="cpu") + expected = module.state_dict() + if set(state) != set(expected): + raise ValueError("Packed candidate keys differ from the exact Nunchaku layer schema") + for key, value in state.items(): + if value.shape != expected[key].shape or value.dtype != expected[key].dtype: + raise ValueError(f"Packed candidate shape/dtype mismatch: {key}") + if value.is_floating_point() and not bool(torch.isfinite(value).all()): + raise ValueError(f"Non-finite packed candidate: {key}") + module.load_state_dict(state, strict=True) + return module.eval().requires_grad_(False).to(device) + + +def evaluate_outputs(parent, teacher, inputs, torch): + """Use the actual parent module, preserving BF16 SiLU/product boundaries.""" + from .denoiser_probe_v3 import _metrics + with torch.inference_mode(): + actual_projection = parent.proj(inputs) + reference_projection = teacher.proj(inputs) + actual_parent = parent(inputs) + reference_parent = teacher(inputs) + return {"projection": _metrics(actual_projection, reference_projection, torch), + "parent": _metrics(actual_parent, reference_parent, torch)} + + +def main(): + parser = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter) + parser.add_argument("--source-checkpoint", type=Path, required=True) + parser.add_argument("--model-path", type=Path, required=True) + parser.add_argument("--activations", type=Path, required=True) + parser.add_argument("--rank-probe", type=Path, required=True) + parser.add_argument("--out", type=Path, required=True) + parser.add_argument("--device", choices=["cuda:0"], required=True) + parser.add_argument("--blocks", type=int, nargs="+", default=[4, 11, 28]) + parser.add_argument("--max-rows", type=int, default=1024, + help="Reject larger activation splits; never silently subsample stored rows") + args = parser.parse_args() + if not args.blocks or len(set(args.blocks)) != len(args.blocks) or any(not 0 <= block < 32 for block in args.blocks): + parser.error("Blocks must be unique indices in [0,31]") + if args.max_rows < 1: + parser.error("max-rows must be positive") + if args.out.exists(): + raise ValueError("Use a new report path; this diagnostic never overwrites existing results") + import torch + from diffusers.models.transformers.transformer_qwenimage21 import QwenImage21SwiGLUFeedForward + from .runtime import _read_manifest + from .checkpoint_io import _source_index, _read_tensor, _json_write + from .export_v3 import ActivationReader, _hash_file + from .refine_checkpoint_v3 import _source_layer_state + + torch.set_num_threads(8) + source, probe = args.source_checkpoint.resolve(), args.rank_probe.resolve() + original = args.model_path / "transformer" if (args.model_path / "transformer").is_dir() else args.model_path + source_manifest = _read_manifest(source) + source_hash = _hash_file(source / "manifest.json") + probe_report = json.loads((probe / "comparison.json").read_text()) + if probe_report.get("complete") is not True or probe_report.get("source_manifest_sha256") != source_hash: + raise ValueError("Rank probe must be complete and derived from this exact packed checkpoint") + config_hash = _hash_file(original / "config.json") + config = json.loads((original / "config.json").read_text()) + if config.get("_class_name") != "QwenImage21Transformer2DModel" or config.get("quantization_config"): + raise ValueError("Teacher must use original unquantized Qwen Image 2.1 weights") + if source_manifest["conversion_identity"]["source_config_sha256"] != config_hash: + raise ValueError("Packed checkpoint and original teacher config fingerprints differ") + names = [f"transformer_blocks.{block}.img_mlp.proj" for block in args.blocks] + candidates = _candidate_rows(probe_report, names) + reader = ActivationReader(args.activations) + reader.require(names) + if reader.fingerprint != probe_report["activation_fingerprint"]: + raise ValueError("Activation archive differs from the rank-probe calibration archive") + source_index, original_index = _source_index(source), _source_index(original) + report = { + "format": "qwen21-parent-mlp-rank-probe-v3", "complete": False, + "source_checkpoint": str(source), "source_manifest_sha256": source_hash, + "rank_probe": str(probe), "rank_probe_report_sha256": _hash_file(probe / "comparison.json"), + "activation_fingerprint": reader.fingerprint, "teacher_config_sha256": config_hash, + "device": args.device, "blocks": [], "fitting_performed": False, + "composition": "out(SiLU(gate_layer(x)) * proj(x)); pinned Diffusers QwenImage21SwiGLUFeedForward", + "fixed_modules": "Original source-checkpoint Nunchaku gate_layer and out; only saved proj candidate varies", + "teacher": "All three original BF16 MLP projections on identical BF16 proj-input rows", + "selection_policy": "Diagnostic parent/raw-projection validation rankings frozen before heldout inputs are read", + "limitations": "MLP-only teacher-input evaluation; no block residual/tanh gate, attention, denoiser, or image-quality equivalence", + } + args.out.parent.mkdir(parents=True, exist_ok=True) + _json_write(args.out, report) + for name in names: + prefix = name.rsplit(".", 1)[0] + info = source_manifest["layers"][name] + source_roles = {role: prefix + "." + role for role in ("proj", "gate_layer", "out")} + teacher_state = {} + for role, full_name in source_roles.items(): + if source_manifest["layers"][full_name]["bias"] or full_name + ".bias" in original_index: + raise ValueError("Pinned Qwen Image 2.1 SwiGLU projections must be bias-free") + weight = _read_tensor(original_index, full_name + ".weight") + shape = (info["in_features"], info["out_features"]) if role == "out" else (info["out_features"], info["in_features"]) + if weight.dtype != torch.bfloat16 or tuple(weight.shape) != shape: + raise ValueError(f"Teacher shape/dtype mismatch for {full_name}") + teacher_state[role + ".weight"] = weight + with torch.device("meta"): + teacher = QwenImage21SwiGLUFeedForward(info["in_features"], info["out_features"]) + parent = QwenImage21SwiGLUFeedForward(info["in_features"], info["out_features"]) + teacher.load_state_dict(teacher_state, strict=True, assign=True) + teacher = teacher.eval().requires_grad_(False).to(args.device) + modules = {} + fixed_hashes = {} + for role in ("gate_layer", "out"): + full_name = source_roles[role] + state = _source_layer_state(source_index, full_name) + setattr(parent, role, _quantized_module(state, source_manifest["layers"][full_name], args.device)) + for key in state: + file = source_index[full_name + "." + key] + fixed_hashes[str(file)] = _hash_file(file) + source_state = _source_layer_state(source_index, name) + modules["source"] = _quantized_module(source_state, info, args.device) + rows = [{"name": "source", "actual_rank": info["rank"], "requested_rank_matches": True}] + for candidate in candidates[name]: + key = f"requested-rank-{candidate['requested_rank']}" + actual_rank = candidate["actual_selected_rank"] + state = _saved_candidate(probe, candidate, torch) + modules[key] = _quantized_module(state, {**info, "rank": actual_rank}, args.device) + rows.append({"name": key, "requested_rank": candidate["requested_rank"], "actual_rank": actual_rank, + "requested_rank_matches": actual_rank == candidate["requested_rank"], + "candidate_file": candidate["candidate_file"], "candidate_sha256": candidate["candidate_sha256"], + "raw_projection_probe_prefers_candidate": candidate["validation_prefers_candidate"]}) + parent.eval().requires_grad_(False) + block_report = {"layer": name, "fixed_module_shards_sha256": fixed_hashes, "candidates": rows} + report["blocks"].append(block_report) + for split in ("validation", "heldout"): + if split == "heldout": + # Do this before reading heldout; no heldout result participates. + block_report["validation_decision"] = validation_rankings(rows) + _json_write(args.out, report) + inputs = reader.read(name, split) + if inputs.dtype != torch.bfloat16 or inputs.ndim != 2 or inputs.shape[1] != info["in_features"]: + raise ValueError("Need BF16 [rows,proj-in-features] teacher inputs") + if not 0 < inputs.shape[0] <= args.max_rows or not bool(torch.isfinite(inputs).all()): + raise ValueError("Invalid/nonfinite activation rows or bounded-row limit exceeded") + # The collector observes the same pre-MLP input for both branches. + # Confirm this provenance where the gate activation archive exists. + if (source_roles["gate_layer"], split) in reader.entries: + if not torch.equal(inputs, reader.read(source_roles["gate_layer"], split)): + raise ValueError("Gate/projection calibration rows are not identical shared MLP inputs") + x = inputs.to(args.device).unsqueeze(0) + block_report[split + "_rows"] = inputs.shape[0] + for row in rows: + parent.proj = modules[row["name"]] + started = time.perf_counter() + row[split] = evaluate_outputs(parent, teacher, x, torch) + torch.cuda.synchronize() + row[split]["instrumented_seconds"] = time.perf_counter() - started + print(json.dumps({"event": "parent_mlp_probe", "layer": name, "candidate": row["name"], "split": split, + "parent_relative_l2": row[split]["parent"]["relative_l2"], + "projection_relative_l2": row[split]["projection"]["relative_l2"]}), flush=True) + _json_write(args.out, report) + del x, inputs + for row in rows: + for split in ("validation", "heldout"): + for objective in ("parent", "projection"): + row[split][objective + "_mse_ratio_to_source"] = ( + row[split][objective]["mse"] / max(rows[0][split][objective]["mse"], 1e-30)) + if block_report["validation_decision"] != validation_rankings(rows): + raise RuntimeError("Validation decision changed after heldout reporting") + _json_write(args.out, report) + del modules, parent, teacher, teacher_state, source_state, state + gc.collect() + torch.cuda.empty_cache() + if _hash_file(source / "manifest.json") != source_hash or _hash_file(probe / "comparison.json") != report["rank_probe_report_sha256"]: + raise ValueError("Source metadata changed during evaluation") + report["complete"] = True + _json_write(args.out, report) + + +if __name__ == "__main__": + main() diff --git a/reproduction/nunchaku_backend/mlp_parent_probe_v3_test.py b/reproduction/nunchaku_backend/mlp_parent_probe_v3_test.py new file mode 100644 index 0000000000000000000000000000000000000000..72e36cef1a0ba9957fd4c2832e1bd3bc597fdab0 --- /dev/null +++ b/reproduction/nunchaku_backend/mlp_parent_probe_v3_test.py @@ -0,0 +1,82 @@ +"""Bounded CPU checks for saved-candidate parent-MLP evaluation.""" +import copy +import importlib.util +import unittest + +from .mlp_parent_probe_v3 import _candidate_rows, evaluate_outputs, validation_rankings + + +def row(name, parent, projection, **extra): + return {"name": name, "validation": {"parent": {"mse": parent}, "projection": {"mse": projection}}, **extra} + + +class SelectionTests(unittest.TestCase): + def test_parent_and_raw_projection_can_choose_different_ranks(self): + result = validation_rankings([row("source", 4, 4), row("256", 1, 3), row("512", 2, 1)]) + self.assertEqual(result["parent_preferred"], "256") + self.assertEqual(result["projection_preferred"], "512") + self.assertTrue(result["rankings_disagree"]) + + def test_heldout_cannot_change_validation_selection(self): + rows = [row("source", 4, 4), row("256", 1, 1)] + original = validation_rankings(rows) + rows[0]["heldout"] = {"parent": {"mse": 0}, "projection": {"mse": 0}} + rows[1]["heldout"] = {"parent": {"mse": 1e10}, "projection": {"mse": 1e10}} + self.assertEqual(original, validation_rankings(rows)) + self.assertFalse(original["decision_uses_heldout"]) + + def test_source_wins_ties_and_guard_fallback_is_not_rank_upgrade(self): + result = validation_rankings([row("source", 1, 1), row("256", 1, 1), + row("512-actual32", 0, 0, requested_rank_matches=False)]) + self.assertEqual(result["parent_ranking"], ["source", "256"]) + self.assertEqual(result["projection_preferred"], "source") + + def test_nonfinite_negative_and_missing_source_rejected(self): + for rows in ([row("source", float("nan"), 1)], [row("source", -1, 1)], + [row("256", 1, 1)], []): + with self.subTest(rows=rows), self.assertRaises(ValueError): + validation_rankings(rows) + + def test_candidate_filter_order_and_missing_duplicate_guards(self): + report = {"results": [{"layer": "a", "requested_rank": 512}, + {"layer": "b", "requested_rank": 128}, + {"layer": "a", "requested_rank": 256}]} + self.assertEqual([r["requested_rank"] for r in _candidate_rows(report, ["a"])["a"]], [256, 512]) + with self.assertRaises(ValueError): + _candidate_rows(report, ["missing"]) + report["results"].append({"layer": "a", "requested_rank": 256}) + with self.assertRaises(ValueError): + _candidate_rows(report, ["a"]) + + +@unittest.skipUnless(importlib.util.find_spec("torch") and importlib.util.find_spec("diffusers"), + "Real parent CPU checks need the existing POC Torch/Diffusers image") +class ExactParentTests(unittest.TestCase): + def test_actual_pinned_parent_composition_and_different_sensitivity(self): + import torch + from diffusers.models.transformers.transformer_qwenimage21 import QwenImage21SwiGLUFeedForward + teacher = QwenImage21SwiGLUFeedForward(2, 2).to(dtype=torch.bfloat16).eval() + with torch.no_grad(): + teacher.proj.weight.copy_(torch.eye(2)) + teacher.gate_layer.weight.copy_(torch.eye(2)) + teacher.out.weight.copy_(torch.tensor([[1., 0.], [0., 0.]])) + inputs = torch.tensor([[[1., 2.], [-1., 1.]]], dtype=torch.bfloat16) + parent = copy.deepcopy(teacher) + with torch.no_grad(): + parent.proj.weight[1, 1] += 2 + first = evaluate_outputs(parent, teacher, inputs, torch) + self.assertGreater(first["projection"]["mse"], 0) + self.assertEqual(first["parent"]["mse"], 0) + with torch.no_grad(): + manual = teacher.out(torch.nn.functional.silu(teacher.gate_layer(inputs)) * teacher.proj(inputs)) + self.assertTrue(torch.equal(manual, teacher(inputs))) + parent = copy.deepcopy(teacher) + with torch.no_grad(): + parent.proj.weight[0, 0] += .25 + second = evaluate_outputs(parent, teacher, inputs, torch) + self.assertLess(second["projection"]["mse"], first["projection"]["mse"]) + self.assertGreater(second["parent"]["mse"], first["parent"]["mse"]) + + +if __name__ == "__main__": + unittest.main() diff --git a/reproduction/nunchaku_backend/optimize_v3.py b/reproduction/nunchaku_backend/optimize_v3.py new file mode 100644 index 0000000000000000000000000000000000000000..b7c309b2383a8a9375a8c18dbd388c72ea85bfd8 --- /dev/null +++ b/reproduction/nunchaku_backend/optimize_v3.py @@ -0,0 +1,348 @@ +"""Activation-output calibrated SVDQuant experiment, isolated from v1/v2. + +The upstream recurrence is L=SVD(W-Q), Q=INT4(W-L), selecting by +quantized module output error. This implementation adds optional diagonal +activation weighting and a ridge-stabilized low-rank output correction. +Neither extension is represented as the exact upstream recipe. + +CUDA scoring uses the real Nunchaku kernel; CPU scoring uses the independent +groupwise reference. Held-out tensors NEVER affect candidate selection. +""" +from __future__ import annotations + +import time +from dataclasses import dataclass +from typing import Callable + +import torch +import torch.nn.functional as F + +from .packing import GROUP_SIZE, _upstream_converter +from .baseline_candidate import convert_linear_weight + + +@dataclass +class Candidate: + residual: torch.Tensor # FP32 exact integer * stored BF16 scale + scales: torch.Tensor # [O,I/64] BF16 + down: torch.Tensor # [rank,I], original-input coordinates + up: torch.Tensor # [O,rank] + smooth: torch.Tensor # [I] BF16 + metadata: dict + + def reference(self, bias=None): + result = {"residual_dequant": self.residual, "down_unpacked": self.down, + "up_unpacked": self.up, "smooth": self.smooth, "weight_scales": self.scales} + if bias is not None: + result["bias"] = bias + return result + + +def pack_candidate(candidate: Candidate, bias=None): + """Use the unchanged upstream tile/layout packer, never an alternate codec.""" + oc, ic = candidate.residual.shape + qweight, scales, packed_bias, smooth, lora, subscale = _upstream_converter()( + candidate.residual.to(candidate.up.dtype), + candidate.scales.reshape(oc, 1, ic // GROUP_SIZE, 1), + bias=bias, smooth=candidate.smooth, lora=(candidate.down, candidate.up), + float_point=False, + ) + assert lora is not None and subscale is None + state = {"qweight": qweight, "wscales": scales, "smooth_factor": smooth, + "smooth_factor_orig": smooth.clone(), "proj_down": lora[0], "proj_up": lora[1]} + if bias is not None: + state["bias"] = packed_bias + return {key: value.contiguous() for key, value in state.items()} + + +def output_metrics(actual, teacher): + error = actual.double() - teacher.double() + mse = error.square().mean().item() + energy = teacher.double().square().mean().item() + return {"mse": mse, "relative_l2": (mse / max(energy, 1e-30)) ** 0.5, + "max_abs": error.abs().max().item(), "teacher_rms": energy ** 0.5, + "finite": bool(torch.isfinite(actual).all())} + + +def _forward(candidate, x, bias, backend): + if backend == "nunchaku": + from nunchaku.models.linear import SVDQW4A4Linear + oc, ic = candidate.residual.shape + module = SVDQW4A4Linear(ic, oc, rank=candidate.down.shape[0], bias=bias is not None, + precision="int4", torch_dtype=x.dtype, device=x.device) + module.load_state_dict(pack_candidate(candidate, bias), strict=True) + module.eval().requires_grad_(False) + return module(x.unsqueeze(0)).squeeze(0).float() + if backend == "reference": + return _reference_explicit_scales(candidate, x, bias) + raise ValueError("objective_backend must be auto, nunchaku, or reference") + + +def _reference_explicit_scales(candidate, x, bias): + """Same validated groupwise arithmetic, using actual stored weight scales. + + GPTQ groups need not contain a +/-7 extremum, so inferring scales from + residual maxima (as the original RTN-only reference does) is invalid. + Kept here to avoid modifying the stable converter/reference modules. + """ + dtype = x.dtype + grouped = (x.float() / candidate.smooth.float()).to(dtype).float().reshape(x.shape[0], -1, 64) + scale_fp32 = grouped.abs().amax(-1) * (1.0 / 7.0) + inv = torch.where(scale_fp32 > 0, scale_fp32.reciprocal(), torch.zeros_like(scale_fp32)) + qa = (grouped * inv.unsqueeze(-1)).round().clamp(-8, 7) + sa = scale_fp32.to(dtype).float() + if x.device.type == "cuda": + from .convert_activation_reference import quantize_activations_ptx + qa, sa = quantize_activations_ptx(x, candidate.smooth) + sa = sa.float() + oc, ic = candidate.residual.shape + sw = candidate.scales.float() + qw = (candidate.residual.reshape(oc, ic // 64, 64) / sw.unsqueeze(-1)).round() + if (qw.abs() > 7).any() or not torch.equal(qw * sw.unsqueeze(-1), candidate.residual.reshape_as(qw)): + raise ValueError("Candidate does not match stored group64 INT4 scales") + output = torch.zeros(x.shape[0], oc, dtype=torch.float32, device=x.device) + for group in range(ic // 64): + dot = (qa[:, group] @ qw[:, group].T).to(dtype).float() + scale = (sa[:, group, None] * sw[None, :, group]).to(dtype).float() + output = (dot.double() * scale.double() + output.double()).to(dtype).float() + if bias is not None: + output = (output + bias.float()).to(dtype).float() + hidden = (x.float() @ candidate.down.float().T).to(dtype).float() + return (output + hidden @ candidate.up.float().T).to(dtype).float() + + +def _weight_only_proxy(candidate, x, bias): + """Diagnostic A16 residual branch, NOT a deployed W4A16 kernel. + + This omits activation INT4 rounding and uses BF16 GEMM rather than the + Nunchaku per-group BF16 accumulator. Error differences therefore include + both effects and must not be called an exact activation-error partition. + """ + dtype = x.dtype + smoothed = (x.float() / candidate.smooth.float()).to(dtype) + residual = F.linear(smoothed, candidate.residual.to(dtype), bias) + lowrank = F.linear(F.linear(x, candidate.down), candidate.up) + return (residual + lowrank).float() + + +def _factors(matrix, column_weight, rank, smooth, seed, niter, oversample, factorization="balanced"): + """Randomized SVD; inverse weighting returns original-input coordinates.""" + device = matrix.device + with torch.random.fork_rng(devices=[device.index] if device.type == "cuda" else []): + torch.random.default_generator.manual_seed(seed) + if device.type == "cuda": + torch.cuda.default_generators[device.index].manual_seed(seed) + u, singular, v = torch.svd_lowrank(matrix * column_weight, + q=min(rank + oversample, *matrix.shape), niter=niter) + if factorization == "up_singular": + return ((v[:, :rank].T / column_weight / smooth.float()).to(torch.bfloat16), + (u[:, :rank] * singular[:rank]).to(torch.bfloat16)) + if factorization != "balanced": + raise ValueError("factorization must be balanced or up_singular") + root = singular[:rank].clamp_min(0).sqrt() + up = (u[:, :rank] * root).to(torch.bfloat16) + down = ((root[:, None] * v[:, :rank].T) / column_weight / smooth.float()).to(torch.bfloat16) + return down, up + + +def _quantize_residual(weight_smoothed, down, up, smooth): + oc, ic = weight_smoothed.shape + residual = weight_smoothed - up.float() @ (down.float() * smooth.float()) + grouped = residual.reshape(oc, ic // GROUP_SIZE, GROUP_SIZE) + scales = (grouped.abs().amax(-1) / 7).clamp_min(torch.finfo(torch.bfloat16).tiny).to(torch.bfloat16) + integers = (grouped / scales.float().unsqueeze(-1)).round().clamp(-7, 7) + return (integers * scales.float().unsqueeze(-1)).reshape(oc, ic), scales + + +def _smooth_candidates(weight, train, alphas, families, train_absmax=None): + xmax = (train.float().abs().amax(0) if train_absmax is None else train_absmax).clamp_min(1e-5) + wmax = weight.abs().amax(0).clamp_min(1e-5) + yield torch.ones_like(xmax, dtype=torch.bfloat16), {"family": "identity", "alpha": None} + for family in families: + if family not in ("smoothquant", "activation_only"): + raise ValueError(f"Unknown smoothing family {family}") + for alpha in alphas: + if not 0 <= alpha <= 1: + raise ValueError("Smoothing alpha must be in [0,1]") + smooth = xmax.pow(alpha) + if family == "smoothquant": + smooth = smooth / wmax.pow(1 - alpha) + # A global scalar has no exact arithmetic effect, but normalization + # avoids overflow/underflow before the runtime's BF16 divisions. + smooth = (smooth / smooth.log().mean().exp()).clamp(1e-4, 1e4).to(torch.bfloat16) + yield smooth, {"family": family, "alpha": alpha} + + +@torch.no_grad() +def optimize_linear_weight( + weight, train_inputs, validation_inputs, heldout_inputs=None, bias=None, *, + ranks=(32, 64, 128), alphas=(0.25, 0.5, 0.75), smoothing_families=("smoothquant", "activation_only"), + iterations=8, weighting=("none", "rms"), ridge=0.01, output_correction=True, + seed=1947, niter=4, oversample=16, conversion_device="cpu", objective_backend="auto", + baseline_rank=32, baseline_absmax=None, baseline_seed=None, train_absmax=None, + final_gptq=False, gptq_damp=0.01, + factorization="balanced", + fixed_smoothing=None, + progress: Callable[[dict], None] | None = None, +): + """Return (packed_state, JSON_stats, unpacked_reference). + + Train, validation and heldout are independent [sample,input_channel] BF16 + matrices. Their provenance must be enforced by the caller. The heldout + matrix is used only AFTER a winner has been selected by validation MSE. + Baseline one-pass rank32 is always a candidate (so selection cannot worsen + validation error). This guarantee does not extend to unseen activations. + Actual bias is included in all teacher and quantized outputs. + """ + started = time.perf_counter() + device = torch.device(conversion_device) + if device.type not in ("cpu", "cuda") or device.type == "cuda" and device.index is None: + raise ValueError("Use cpu or explicitly indexed CUDA device") + if iterations < 1 or ridge <= 0: + raise ValueError("iterations must be positive and ridge must be positive") + if fixed_smoothing not in (None, "identity", "activation_only", "smoothquant"): + raise ValueError("fixed_smoothing must be identity, activation_only, smoothquant, or None") + if fixed_smoothing not in (None, "identity") and (len(alphas) != 1 or tuple(smoothing_families) != (fixed_smoothing,)): + raise ValueError("Fixed smoothing requires exactly one alpha and the matching single family") + backend = ("nunchaku" if device.type == "cuda" else "reference") if objective_backend == "auto" else objective_backend + if backend == "nunchaku" and device.type != "cuda": + raise ValueError("Actual Nunchaku scoring requires an explicitly reserved CUDA device") + weight = weight.detach().to(device=device, dtype=torch.float32) + if weight.ndim != 2 or any(d % 128 or d <= 0 for d in weight.shape) or not torch.isfinite(weight).all(): + raise ValueError("Weight must be finite [O,I] with dimensions divisible by128") + oc, ic = weight.shape + if any(r <= 0 or r % 16 or r > min(oc, ic) for r in (*ranks, baseline_rank)): + raise ValueError("Ranks must be positive multiples of16 within matrix dimensions") + def inputs(value, label): + if value is None: + return None + if value.ndim != 2 or value.shape[1] != ic or value.shape[0] < 2 or not value.is_floating_point(): + raise ValueError(f"{label} must be floating point [rows>=2,{ic}]") + result = value.detach().to(device=device, dtype=torch.bfloat16).contiguous() + if not torch.isfinite(result).all(): + raise ValueError(f"Nonfinite {label}") + return result + train = inputs(train_inputs, "train") + validation = inputs(validation_inputs, "validation") + if train_absmax is not None: + train_absmax = train_absmax.detach().to(device=device, dtype=torch.float32) + if train_absmax.shape != (ic,) or not torch.isfinite(train_absmax).all() or (train_absmax < 0).any(): + raise ValueError("train_absmax must be a finite nonnegative per-input-channel vector") + # Heldout is not even transferred/validated until selection is complete. + if bias is not None: + bias = bias.detach().to(device=device, dtype=torch.bfloat16) + if bias.shape != (oc,) or not torch.isfinite(bias).all(): + raise ValueError("Invalid bias") + teacher_weight = weight.to(torch.bfloat16) + teacher_train = F.linear(train, teacher_weight, bias).float() + teacher_validation = F.linear(validation, teacher_weight, bias).float() + history = [] + def score(candidate, label=None): + result = output_metrics(_forward(candidate, validation, bias, backend), teacher_validation) + if not result["finite"]: + raise ValueError("Nonfinite quantized candidate output") + row = {**candidate.metadata, "validation": result} + if label: + row["event"] = label + history.append(row) + if progress: + progress(row) + return result + original_absmax = baseline_absmax if baseline_absmax is not None else (train_absmax if train_absmax is not None else train.float().abs().amax(0)) + _, baseline_conversion, baseline_reference = convert_linear_weight( + weight, bias, rank=baseline_rank, input_absmax=original_absmax, seed=seed if baseline_seed is None else baseline_seed, + niter=niter, oversample=oversample, return_reference=True, conversion_device=device) + bres = baseline_reference["residual_dequant"] + bscales = (bres.reshape(oc, ic // 64, 64).abs().amax(-1) / 7).to(torch.bfloat16) + # Zero groups preserve the converter's positive tiny stored scale. + bscales = bscales.clamp_min(torch.finfo(torch.bfloat16).tiny) + baseline = Candidate(bres, bscales, baseline_reference["down_unpacked"], baseline_reference["up_unpacked"], + baseline_reference["smooth"], {"family": "one_pass_baseline", "rank": baseline_rank, + "alpha": 0.5, "iteration": 0, "weighting": "none", "output_correction": 0.0}) + best = baseline + best_metric = baseline_metric = score(baseline, "baseline") + rms = train.float().square().mean(0).sqrt() + rms = rms.clamp_min(rms.mean().clamp_min(1e-6) * 0.01) + for smooth, smoothing_meta in _smooth_candidates(weight, train, alphas, smoothing_families, train_absmax): + if fixed_smoothing is not None and smoothing_meta["family"] != fixed_smoothing: + continue + smoothed = weight * smooth.float() + for rank in ranks: + for method in weighting: + if method not in ("none", "rms"): + raise ValueError("weighting must contain none or rms") + column_weight = torch.ones_like(rms) if method == "none" else rms / smooth.float() + column_weight = column_weight / column_weight.mean() + residual_quant = torch.zeros_like(smoothed) + previous_error = float("inf") + for iteration in range(iterations): + down, up = _factors(smoothed - residual_quant, column_weight, rank, smooth, + seed, niter, oversample, factorization) + residual_quant, scales = _quantize_residual(smoothed, down, up, smooth) + candidate = Candidate(residual_quant, scales, down, up, smooth, + {**smoothing_meta, "rank": rank, "iteration": iteration, + "weighting": method, "output_correction": 0.0, "factorization": factorization}) + metric = score(candidate) + if metric["mse"] < best_metric["mse"]: + best, best_metric = candidate, metric + # Official recurrence early stops on the first worsening. + if metric["mse"] > previous_error: + break + previous_error = metric["mse"] + # Fit only the BF16 up projection on train inputs. Ridge + # penalizes changes from the current SVD factors; validation + # rejects corrections that overfit or upset kernel rounding. + if output_correction: + actual_train = _forward(candidate, train, bias, backend) + hidden = (train.float() @ down.float().T).to(torch.bfloat16).float() + gram = hidden.T @ hidden + damping = ridge * gram.diag().mean().clamp_min(1e-12) + gram.diagonal().add_(damping) + correction = torch.linalg.solve(gram, hidden.T @ (teacher_train - actual_train)).T + for mix in (0.25, 0.5, 1.0): + corrected = Candidate(residual_quant, scales, down, + (up.float() + mix * correction).to(torch.bfloat16), smooth, + {**candidate.metadata, "output_correction": mix, "ridge": ridge}) + corrected_metric = score(corrected) + if corrected_metric["mse"] < best_metric["mse"]: + best, best_metric = corrected, corrected_metric + gptq_diagnostics = None + if final_gptq: + from .gptq_v3 import quantize_residual_gptq + target = weight * best.smooth.float() - best.up.float() @ (best.down.float() * best.smooth.float()) + smoothed_train = (train.float() / best.smooth.float()).to(torch.bfloat16) + quantized, scales, gptq_diagnostics = quantize_residual_gptq( + target, smoothed_train, damp=gptq_damp, conversion_device=device, return_diagnostics=True) + candidate = Candidate(quantized, scales, best.down, best.up, best.smooth, + {**best.metadata, "gptq": True, "gptq_damp": gptq_damp}) + metric = score(candidate, "final_gptq_candidate") + if metric["mse"] < best_metric["mse"]: + best, best_metric = candidate, metric + # Selection frozen. The following is reporting only; no test-set gate. + stats = {"algorithm": "v3_activation_output_selected_iterative_svdquant", "selected": best.metadata, + "objective_backend": backend, "teacher": "original_weight_bf16_linear", + "baseline": {"validation": baseline_metric, "conversion": baseline_conversion}, + "validation": best_metric, "validation_mse_ratio_to_baseline": best_metric["mse"] / max(baseline_metric["mse"], 1e-30), + "train_rows": train.shape[0], "validation_rows": validation.shape[0], "candidate_count": len(history), + "smoothing_amax_source": "full_training_observations" if train_absmax is not None else "sampled_training_rows", + "search": {"ranks": list(ranks), "alphas": list(alphas), "smoothing_families": list(smoothing_families), + "iterations": iterations, "weighting": list(weighting), "ridge": ridge, + "niter": niter, "oversample": oversample, "seed": seed, "final_gptq": final_gptq, + "gptq_damp": gptq_damp}, "history": history, "gptq_diagnostics": gptq_diagnostics, + "limitations": "Independent train/validation/test provenance supplied by caller. Per-linear output error is not end-to-end image equivalence. Randomized SVD, diagonal weighting and ridge correction differ from official full-SVD recipe. Optional final GPTQ retains original channel order."} + test = inputs(heldout_inputs, "heldout") + stats["search"]["factorization"] = factorization + stats["search"]["fixed_smoothing"] = fixed_smoothing + if test is not None: + teacher_test = F.linear(test, teacher_weight, bias).float() + stats["heldout_rows"] = test.shape[0] + stats["heldout"] = output_metrics(_forward(best, test, bias, backend), teacher_test) + stats["baseline"]["heldout"] = output_metrics(_forward(baseline, test, bias, backend), teacher_test) + stats["heldout_mse_ratio_to_baseline"] = stats["heldout"]["mse"] / max(stats["baseline"]["heldout"]["mse"], 1e-30) + stats["heldout_weight_only_proxy"] = output_metrics(_weight_only_proxy(best, test, bias), teacher_test) + stats["baseline"]["heldout_weight_only_proxy"] = output_metrics(_weight_only_proxy(baseline, test, bias), teacher_test) + stats["weight_only_proxy_limitation"] = "A16 residual branch with BF16 GEMM, not a real W4A16 kernel; differs in activation quantization AND group accumulator arithmetic." + state = pack_candidate(best, bias) + stats["packed_bytes"] = sum(t.numel() * t.element_size() for t in state.values()) + stats["seconds"] = time.perf_counter() - started + return state, stats, best.reference(bias) diff --git a/reproduction/nunchaku_backend/packing.py b/reproduction/nunchaku_backend/packing.py new file mode 100644 index 0000000000000000000000000000000000000000..b2f6e81735cfc906e5288ab4fa631d0bacd73904 --- /dev/null +++ b/reproduction/nunchaku_backend/packing.py @@ -0,0 +1,17 @@ +"""Shared pinned Nunchaku packing adapter; no calibration or CLI.""" + +DEEPCOMPRESSOR_COMMIT = "69f3473f5e1c1504bae35cc50c7858ef900a9b17" +GROUP_SIZE = 64 + + +def _upstream_converter(): + try: + from deepcompressor.backend.nunchaku.utils import convert_to_nunchaku_w4x4y16_linear_weight + except ImportError as exc: + raise ImportError( + "Nunchaku packing requires the DeepCompressor checkout at " + f"{DEEPCOMPRESSOR_COMMIT} on PYTHONPATH, plus torch and safetensors." + ) from exc + return convert_to_nunchaku_w4x4y16_linear_weight + + diff --git a/reproduction/nunchaku_backend/probe_iterations_v3.sh b/reproduction/nunchaku_backend/probe_iterations_v3.sh new file mode 100644 index 0000000000000000000000000000000000000000..75ee106e8585bc7466b34d1f01be229ef69350f7 --- /dev/null +++ b/reproduction/nunchaku_backend/probe_iterations_v3.sh @@ -0,0 +1,35 @@ +#!/usr/bin/env bash +# Run INSIDE the POC GPU0 container, only after the main agent releases GPU0. +# Isolated probes only. This never replaces or extends the full V3 checkpoint. +set -euo pipefail + +QWEN_ITER_PROBE_ROOT="${QWEN_ITER_PROBE_ROOT:-/cache/qwen-v3-iterations100}" +QWEN_ITER_SOURCE="${QWEN_ITER_SOURCE:-/cache/qwen-nunchaku-v3-r128}" +QWEN_ITER_DATA="${QWEN_ITER_DATA:-/cache/qwen21-activation-v3-40}" +QWEN_ITER_REPORT="${QWEN_ITER_REPORT:-/poc/results/nunchaku-v3-iterations100-comparison.json}" + +common=( + --model-path /cache/huggingface/hub/models--Qwen--Qwen-Image-2.1/snapshots/b3179ad355be050328e483a9dfdd9e60cd62adfa + --activations "$QWEN_ITER_DATA" + --baseline-calibration /cache/qwen21-calibration + --device cuda:0 --ranks 128 --weighting none --iterations 100 + --families activation_only --fixed-smoothing activation_only + --final-gptq --gptq-damp .01 --factorization balanced + --seed 1947 --niter 4 --oversample 16 --ridge .01 +) + +# Global seed1947 + stable full-model ordinal gives exactly2106/2025/1956. +# Block22 is a high-error NON-cap control: V3 selectediteration13. +python -m nunchaku_backend.export_v3 "${common[@]}" --alphas .25 \ + --layers transformer_blocks.22.img_mlp.out --out "$QWEN_ITER_PROBE_ROOT/block22-mlp-out" +python -m nunchaku_backend.export_v3 "${common[@]}" --alphas .25 \ + --layers transformer_blocks.11.attn.to_out.0 --out "$QWEN_ITER_PROBE_ROOT/block11-attn-out" +python -m nunchaku_backend.export_v3 "${common[@]}" --alphas .75 \ + --layers transformer_blocks.1.attn.to_q --out "$QWEN_ITER_PROBE_ROOT/block01-query" + +python -m nunchaku_backend.compare_iterations_v3 \ + --v3 "$QWEN_ITER_SOURCE/manifest.json" \ + --probes "$QWEN_ITER_PROBE_ROOT/block22-mlp-out/manifest.json" \ + "$QWEN_ITER_PROBE_ROOT/block11-attn-out/manifest.json" \ + "$QWEN_ITER_PROBE_ROOT/block01-query/manifest.json" \ + --out "$QWEN_ITER_REPORT" diff --git a/reproduction/nunchaku_backend/rank_probe_v3.md b/reproduction/nunchaku_backend/rank_probe_v3.md new file mode 100644 index 0000000000000000000000000000000000000000..d8eb91e257b7b3597d047a2fa8aa053844214cd6 --- /dev/null +++ b/reproduction/nunchaku_backend/rank_probe_v3.md @@ -0,0 +1,58 @@ +# MLP projection rank 256/512 probe — completed + +All six actual CUDA-kernel probes completed successfully. Both ranks improved validation and held-out MSE against the actual packed V3 rank-128 layers. Rank 512 improved more in all three cases. These linear measurements support a staged selective upgrade; parent-MLP and whole-denoiser evaluation, followed by image comparison, remain necessary before accepting quality. + +| Layer | Candidate rank | Validation MSE / actual V3 | Held-out MSE / actual V3 | Held-out relative L2 | Fit/evaluation time | +|---|---:|---:|---:|---:|---:| +| block 4 MLP projection | 256 | 0.839 | 0.837 | 8.990% | 2.52 s | +| block 4 MLP projection | 512 | 0.614 | 0.614 | 7.699% | 3.05 s | +| block 11 MLP projection | 256 | 0.866 | 0.850 | 8.240% | 2.29 s | +| block 11 MLP projection | 512 | 0.659 | 0.645 | 7.177% | 3.10 s | +| block 28 MLP projection | 256 | 0.861 | 0.852 | 10.108% | 2.37 s | +| block 28 MLP projection | 512 | 0.649 | 0.645 | 8.796% | 3.12 s | + +Ratios below 1 indicate improvement; they are MSE ratios, not relative-L2 ratios. Total measured per-case processing was 16.46 seconds, excluding shared startup/provenance checks. All requested ranks won their internal rank-32 guards and improved actual-source validation. The [full measured report](../results/rank-probe-v3-comparison.json) preserves candidate histories, exact scores, timings and provenance. + +The development denoiser role sweep identified MLP projection as a promising sensitivity target. This bounded probe keeps its INT4 residual and increases the BF16 low-rank branch. It does not change the runner, export a full model, or launch itself. + +The installed Nunchaku version was checked with CPU constructors before GPU calls. Official v1.2.1 source defines a low-rank maximum of 1024 and 16-wide tiles in [lora.cuh](https://github.com/nunchaku-tech/nunchaku/blob/v1.2.1/src/kernels/zgemm/lora.cuh#L20), with dynamic rank and multiple-of-16 assertions in both GEMM and input quantization [launch code](https://github.com/nunchaku-tech/nunchaku/blob/v1.2.1/src/kernels/zgemm/gemm_w4a4_launch_impl.cuh#L193). The completed probes now verify actual Nunchaku CUDA execution for ranks 256 and 512 on the RTX 4070 Ti SUPER, beyond constructor compatibility. End-to-end inference latency remains a separate measurement. + +Representatives are selected reproducibly as the highest saved V3 **validation** relative-L2 MLP projection in each block-depth third. Held-out errors and evaluation images do not select these layers: + +| Layer | Depth range | V3 validation relative L2 | Fixed smoothing | Exact seed | +|---|---|---:|---|---:| +| block 4 MLP projection | 0–10 | 9.733% | activation-only alpha 0.5 | 1981 | +| block 11 MLP projection | 11–21 | 8.738% | activation-only alpha 0.5 | 2030 | +| block 28 MLP projection | 22–31 | 10.892% | activation-only alpha 0.5 | 2149 | + +Each candidate uses the actual V3 selected smoothing family, alpha, weighting, factorization, seed and solver settings, with rank changed to 256 or 512. The limit stays at 16 iterations with existing early stopping, training output correction, and optional final GPTQ. Each rank is fitted separately; held-out results do not feed the optimizer or rank selection. All use the original BF16 teacher and `/cache/qwen21-activation-v3-40` train/validation/heldout splits. At rank 512, only 512 training activation rows are available, so extra capacity may overfit despite ridge regularization; held-out results are essential. + +The comparison loads the **actual packed V3 tensors**, runs their Nunchaku kernel afresh on validation and held-out inputs, and compares proposed packed tensors to those outputs against the same BF16 teacher. It also records saved V3 metrics for replay consistency. Validation freezes a per-candidate preference before held-out inputs are read. The old one-pass rank-32 guard may win inside the optimizer; such a result is clearly labeled and is never accepted as the requested higher-rank recipe. A16 proxy diagnostics change both activation rounding and accumulation and are not a deployed W4A16 path. + +| Rank | BF16 low-rank weights per 4096→12288 layer | Extra over rank 128 per layer | Extra across all 32 MLP projections | +|---:|---:|---:|---:| +| 128, existing | 4 MiB | 0 | 0 | +| 256 | 8 MiB | 4 MiB | 128 MiB | +| 512 | 16 MiB | 12 MiB | 384 MiB | + +These are exact tensor-storage deltas, not measured peak VRAM. INT4 residual storage is unchanged. Temporary low-rank activations also grow with rank and token rows; inference latency must be measured before any selective upgrade. + +The original V3 transformer state measured 4,655,177,728 bytes (4.335 GiB). If all 32 projections select rank 512, adding 384 MiB gives a planned transformer state of 5,057,830,912 bytes (4.710 GiB). This is model-state storage planning only: it excludes the text encoder, VAE, KV cache, other activations, CUDA workspaces, and allocator overhead, and is not a VRAM measurement. + +Run inside the POC image **only after the main agent reserves GPU0**: + +```sh +python -m nunchaku_backend.rank_probe_v3 \ + --source-checkpoint /cache/qwen-nunchaku-v3-r128 \ + --model-path /cache/huggingface/hub/models--Qwen--Qwen-Image-2.1/snapshots/b3179ad355be050328e483a9dfdd9e60cd62adfa \ + --activations /cache/qwen21-activation-v3-40 \ + --baseline-calibration /cache/qwen21-calibration \ + --out /cache/qwen-v3-mlpproj-rank-probe \ + --device cuda:0 --ranks 256 512 --iterations 16 +``` + +Append `--dry-run` for CPU-only source/provenance validation and installed-kernel constructor checks without CUDA initialization or output-directory creation. To stage a repeat, use `--ranks 256` with a distinct output directory, then run `--ranks 512` separately. Existing output directories are rejected; there is no automatic full-model conversion or deployment. + +CPU preparation passed in the actual image with GPUs hidden: installed `1.2.1+cu12.8torch2.8` constructors produced the expected rank-256/512 BF16 state shapes and exact 8/16 MiB low-rank storage, with `cuda_initialized:false`. Actual-source dry-run selected blocks 4/11/28 and all six exact recipe variants. The nine new selection/cost/recipe tests plus 30 existing tests passed together in 0.378 seconds. The [saved dry-run report](../results/nunchaku-v3-mlpproj-rank-dryrun.json) records those settings. GPU execution was subsequently coordinated and launched by the main agent. + +Outputs are `/cache/qwen-v3-mlpproj-rank-probe/comparison.json` and six isolated `.rank.safetensors` candidate states with short parameter names. These are probe artifacts, not loadable full-model checkpoints. Reports include source hashes, exact recipes, candidate histories, fresh actual-V3 comparisons, validation decisions, held-out errors, state bytes, and timings. Original V3 shards are checked for immutability after execution. diff --git a/reproduction/nunchaku_backend/rank_probe_v3.py b/reproduction/nunchaku_backend/rank_probe_v3.py new file mode 100644 index 0000000000000000000000000000000000000000..5df85bbe96b8eb294a4925eeddc9cc5358b4ff68 --- /dev/null +++ b/reproduction/nunchaku_backend/rank_probe_v3.py @@ -0,0 +1,225 @@ +"""Bounded higher-rank MLP projection probes; never exports a full model. + +Select depth representatives using V3 validation statistics only. Compare each +candidate to the actual packed V3 layer on identical BF16 teacher activations. +Heldout is evaluated only after the validation decision. No serving mutation. +""" +from __future__ import annotations + +import argparse +import json +import math +from pathlib import Path +import time + +from .compare_iterations_v3 import layer_stats +from .checkpoint_io import _json_write, _source_index, _read_tensor +from .export_v3 import ActivationReader, _hash_file +from .refine_checkpoint_v3 import ( + assert_separate_output, choose_refinement, fixed_recipe, + require_matching_provenance, _actual_forward, _source_layer_state, +) + + +def select_probe_layers(manifest): + """Highest validation relative-L2 MLP projection in each depth third.""" + stats = layer_stats(manifest) + selected = [] + for low, high in ((0, 10), (11, 21), (22, 31)): + rows = [(name, row) for name, row in stats.items() + if name.endswith(".img_mlp.proj") and low <= int(name.split(".")[1]) <= high] + if not rows: + raise ValueError(f"No MLP projection validation data for blocks {low}..{high}") + if any(not row["validation"].get("finite", True) or + not math.isfinite(float(row["validation"]["relative_l2"])) or + float(row["validation"]["relative_l2"]) < 0 for _, row in rows): + raise ValueError("Nonfinite source validation statistics") + # Stable tie-break, independent of source dictionary insertion order. + selected.append(sorted(rows, key=lambda pair: (-pair[1]["validation"]["relative_l2"], pair[0]))[0][0]) + return selected + + +def _validate_rank(rank): + if not isinstance(rank, int) or isinstance(rank, bool) or rank <= 0 or rank % 16 or rank > 1024: + raise ValueError("Nunchaku rank must be a positive multiple of 16, at most 1024") + + +def lowrank_cost(in_features, out_features, rank, source_rank=128, role_layers=32): + """BF16 up/down storage only; excludes activation scratch and allocator.""" + _validate_rank(rank) + _validate_rank(source_rank) + if rank > min(in_features, out_features) or min(in_features, out_features, role_layers) <= 0: + raise ValueError("Rank exceeds dimensions or dimensions/count are invalid") + unit = 2 * (in_features + out_features) + return {"lowrank_bytes": unit * rank, + "source_lowrank_bytes": unit * source_rank, + "delta_bytes_per_layer": unit * (rank - source_rank), + "delta_bytes_all_role_layers": unit * (rank - source_rank) * role_layers, + "role_layers": role_layers} + + +def probe_recipe(stats, rank, iterations=16): + _validate_rank(rank) + recipe = fixed_recipe(stats, iterations) + recipe["ranks"] = (rank,) + return recipe + + +def cpu_constructor_probe(info, ranks): + """Validate installed Python state layouts without initializing CUDA.""" + import importlib.metadata + import torch + from nunchaku.models.linear import SVDQW4A4Linear + results = [] + for rank in ranks: + cost = lowrank_cost(info["in_features"], info["out_features"], rank, info["rank"]) + module = SVDQW4A4Linear(info["in_features"], info["out_features"], rank=rank, + bias=info["bias"], precision="int4", torch_dtype=torch.bfloat16, device="cpu") + state = module.state_dict() + actual = sum(state[key].numel() * state[key].element_size() for key in ("proj_down", "proj_up")) + if actual != cost["lowrank_bytes"] or any(t.device.type != "cpu" for t in state.values()): + raise ValueError("Installed Nunchaku constructor state differs from expected rank layout") + results.append({"rank": rank, "proj_down": list(state["proj_down"].shape), + "proj_up": list(state["proj_up"].shape), "lowrank_bytes": actual}) + del module, state + return {"nunchaku_version": importlib.metadata.version("nunchaku"), "ranks": results, + "cuda_initialized": torch.cuda.is_initialized(), "forward_executed": False} + + +def main(): + parser = argparse.ArgumentParser(description=__doc__) + parser.add_argument("--source-checkpoint", type=Path, required=True) + parser.add_argument("--model-path", type=Path, required=True) + parser.add_argument("--activations", type=Path, required=True) + parser.add_argument("--baseline-calibration", type=Path, required=True) + parser.add_argument("--out", type=Path, required=True) + parser.add_argument("--device", required=True) + parser.add_argument("--ranks", type=int, nargs="+", default=[256, 512]) + parser.add_argument("--iterations", type=int, default=16) + parser.add_argument("--threads", type=int, default=8) + parser.add_argument("--dry-run", action="store_true") + args = parser.parse_args() + assert_separate_output(args.source_checkpoint, args.out) + if not args.ranks or len(set(args.ranks)) != len(args.ranks): + parser.error("Supply unique candidate ranks") + import torch + from safetensors.torch import load_file, save_file + from .runtime import _read_manifest + from .optimize_v3 import Candidate, optimize_linear_weight, output_metrics, _weight_only_proxy + torch.set_num_threads(args.threads) + device = torch.device(args.device) + if device.type != "cuda" or device.index is None: + parser.error("Use an explicitly reserved indexed CUDA device") + source = args.source_checkpoint.resolve() + manifest_path = source / "manifest.json" + manifest = _read_manifest(source) + manifest_hash = _hash_file(manifest_path) + stats = layer_stats(manifest) + names = select_probe_layers(manifest) + for name in names: + if stats[name]["selected"]["rank"] != manifest["layers"][name]["rank"]: + raise ValueError(f"Source selected-recipe and packed-layer ranks disagree: {name}") + directory = args.model_path / "transformer" if (args.model_path / "transformer").is_dir() else args.model_path + baseline_path = args.baseline_calibration / "activation_stats.safetensors" if args.baseline_calibration.is_dir() else args.baseline_calibration + activations = ActivationReader(args.activations) + activations.require(names) + require_matching_provenance(manifest, activations.fingerprint, _hash_file(directory / "config.json"), _hash_file(baseline_path)) + if json.loads((source / "config.json").read_text()) != json.loads((directory / "config.json").read_text()): + raise ValueError("Source packed model and BF16 teacher configs differ") + recipes = {name: {rank: probe_recipe(stats[name], rank, args.iterations) for rank in args.ranks} for name in names} + costs = {name: {rank: lowrank_cost(manifest["layers"][name]["in_features"], manifest["layers"][name]["out_features"], + rank, manifest["layers"][name]["rank"]) for rank in args.ranks} for name in names} + constructor = cpu_constructor_probe(manifest["layers"][names[0]], args.ranks) + report = {"complete": False, "source_manifest_sha256": manifest_hash, + "activation_fingerprint": activations.fingerprint, "source_checkpoint": str(source), + "selection": "highest saved V3 validation relative_l2 in blocks0..10,11..21,22..31; no heldout/image evaluation used", + "selected_layers": names, "recipes": recipes, "lowrank_costs": costs, + "cpu_constructor": constructor, "iterations": args.iterations, "candidate_ranks": args.ranks, + "limitations": "Per-linear BF16 teacher objective, not whole-denoiser or image equivalence. Low-rank storage is not measured VRAM. At rank512 the512 training rows may limit generalization; ridge remains. No full export or runtime mutation.", + "results": []} + if args.dry_run: + report.update(dry_run=True, gpu_work_executed=False, cuda_initialized=torch.cuda.is_initialized()) + print(json.dumps(report, indent=2)) + return + if args.out.exists(): + raise ValueError("Use a new isolated output directory; this probe never overwrites/resumes") + source_index = _source_index(source) + model_index = _source_index(directory) + source_paths = sorted({source_index[key] for key in source_index if any(key.startswith(name + ".") for name in names)}) + source_hashes = {str(path): _hash_file(path) for path in source_paths} + source_metadata_hashes = {str(source / filename): _hash_file(source / filename) + for filename in ("config.json", "model.safetensors.index.json")} + report["source_layer_shards_sha256"] = source_hashes + report["source_metadata_sha256"] = source_metadata_hashes + report["code_sha256"] = {filename: _hash_file(Path(__file__).parent / filename) for filename in + ("rank_probe_v3.py", "optimize_v3.py", "gptq_v3.py", "baseline_candidate.py", "packing.py", "checkpoint_io.py", "layout.py", "runtime.py")} + args.out.mkdir(parents=True) + _json_write(args.out / "comparison.json", report) + baseline = load_file(str(baseline_path), device="cpu") + for name in names: + info = manifest["layers"][name] + state = _source_layer_state(source_index, name) + weight = _read_tensor(model_index, name + ".weight") + bias = _read_tensor(model_index, name + ".bias") if name + ".bias" in model_index else None + if weight.dtype != torch.bfloat16 or tuple(weight.shape) != (info["out_features"], info["in_features"]): + raise ValueError(f"BF16 teacher weight mismatch: {name}") + train = activations.read(name, "train") + validation = activations.read(name, "validation").to(device) + train_absmax = activations.read(name, "input_absmax") + teacher_weight = weight.to(device) + teacher_bias = bias.to(device) if bias is not None else None + with torch.inference_mode(): + teacher_validation = torch.nn.functional.linear(validation, teacher_weight, teacher_bias).float() + source_validation = output_metrics(_actual_forward(state, info, validation, device), teacher_validation) + for rank in args.ranks: + started = time.perf_counter() + candidate_state, candidate_stats, reference = optimize_linear_weight( + weight, train, validation, None, bias, **recipes[name][rank], + train_absmax=train_absmax, baseline_absmax=baseline[name + ".input_absmax"], + conversion_device=device, objective_backend="nunchaku", + output_correction=manifest["conversion_identity"]["settings"].get("output_correction", True)) + selected = candidate_stats["selected"] + candidate_info = {**info, "rank": selected["rank"]} + same_requested_recipe = (selected["rank"] == rank and selected["family"] == stats[name]["selected"]["family"] + and selected.get("alpha") == stats[name]["selected"].get("alpha")) + actual_validation = output_metrics(_actual_forward(candidate_state, candidate_info, validation, device), teacher_validation) + accepted, reason = choose_refinement(source_validation, actual_validation) + if not same_requested_recipe: + accepted, reason = False, "one_pass_guard_won_instead_of_requested_rank_recipe" + # Decision frozen; heldout is reporting only, including across ranks. + heldout = activations.read(name, "heldout").to(device) + teacher_heldout = torch.nn.functional.linear(heldout, teacher_weight, teacher_bias).float() + source_heldout = output_metrics(_actual_forward(state, info, heldout, device), teacher_heldout) + actual_heldout = output_metrics(_actual_forward(candidate_state, candidate_info, heldout, device), teacher_heldout) + candidate = Candidate(reference["residual_dequant"], reference["weight_scales"], reference["down_unpacked"], + reference["up_unpacked"], reference["smooth"], selected) + proxy = output_metrics(_weight_only_proxy(candidate, heldout, teacher_bias), teacher_heldout) + filename = f"{name}.rank{rank}.safetensors" + save_file({key: value.detach().cpu().contiguous() for key, value in candidate_state.items()}, str(args.out / filename)) + torch.cuda.synchronize(device) + row = {"layer": name, "requested_rank": rank, "actual_selected_rank": selected["rank"], + "source_selected": stats[name]["selected"], "candidate_selected": selected, + "validation_prefers_candidate": accepted, "reason": reason, "decision_uses_heldout": False, + "source_actual_validation": source_validation, "candidate_actual_validation": actual_validation, + "source_actual_heldout": source_heldout, "candidate_actual_heldout": actual_heldout, + "source_saved_validation": stats[name]["validation"], "source_saved_heldout": stats[name]["heldout"], + "validation_mse_ratio_to_actual_v3": actual_validation["mse"] / max(source_validation["mse"], 1e-30), + "heldout_mse_ratio_to_actual_v3": actual_heldout["mse"] / max(source_heldout["mse"], 1e-30), + "heldout_weight_only_proxy": proxy, "weight_only_proxy_limitation": "A16 residual BF16 GEMM changes activation rounding and accumulator arithmetic; not a deployed W4A16 path", + "candidate_optimizer_stats": candidate_stats, "requested_lowrank_cost": costs[name][rank], + "packed_bytes": sum(t.numel() * t.element_size() for t in candidate_state.values()), + "candidate_file": filename, "candidate_sha256": _hash_file(args.out / filename), + "seconds": time.perf_counter() - started} + report["results"].append(row) + _json_write(args.out / "comparison.json", report) + print(json.dumps({key: row[key] for key in ("layer", "requested_rank", "actual_selected_rank", "validation_mse_ratio_to_actual_v3", "heldout_mse_ratio_to_actual_v3", "seconds")}), flush=True) + del candidate_state, candidate_stats, reference, candidate, heldout, teacher_heldout + del state, weight, bias, train, validation, teacher_weight, teacher_bias, teacher_validation + if _hash_file(manifest_path) != manifest_hash or any(_hash_file(Path(path)) != digest for path, digest in {**source_hashes, **source_metadata_hashes}.items()): + raise ValueError("Source changed during probe") + report["complete"] = True + _json_write(args.out / "comparison.json", report) + + +if __name__ == "__main__": + main() diff --git a/reproduction/nunchaku_backend/rank_probe_v3_test.py b/reproduction/nunchaku_backend/rank_probe_v3_test.py new file mode 100644 index 0000000000000000000000000000000000000000..a45c90f7991f843373bd1da25e64b40121cb2ed6 --- /dev/null +++ b/reproduction/nunchaku_backend/rank_probe_v3_test.py @@ -0,0 +1,143 @@ +"""CPU-only selection, provenance and weight-cost checks for the rank probe.""" + +import copy +import unittest + +from .rank_probe_v3 import lowrank_cost, probe_recipe, select_probe_layers +from .refine_checkpoint_v3 import fixed_recipe + + +def _stats(): + return { + "selected": {"family": "activation_only", "alpha": 0.75, "rank": 128, + "weighting": "rms", "factorization": "up_singular"}, + "search": {"seed": 1977, "niter": 4, "oversample": 16, "iterations": 16, + "ridge": 0.01, "final_gptq": True, "gptq_damp": 0.01, + "factorization": "balanced"}, + "baseline": {"conversion": {"rank": 32}}, + "validation": {"relative_l2": 0.1, "mse": 0.01, "finite": True}, + "heldout": {"relative_l2": 0.3, "mse": 0.09, "finite": True}, + } + + +def _manifest(): + suffixes = ("attn.to_q", "attn.to_k", "attn.to_v", "attn.to_out.0", + "img_mlp.proj", "img_mlp.out", "img_mlp.gate_layer") + rows, layers = {}, {} + for block in reversed(range(32)): + for suffix in suffixes: + name = f"transformer_blocks.{block}.{suffix}" + row = _stats() + row["validation"]["relative_l2"] = 1000 if suffix != "img_mlp.proj" else 0.01 + block / 1000 + row["heldout"]["relative_l2"] = 100 - block + rows[name] = row + layers[name] = {"in_features": 4096, "out_features": 12288 if suffix == "img_mlp.proj" else 4096, + "rank": 128, "bias": False, "precision": "int4"} + for block, score in ((2, 0.9), (17, 0.8), (25, 0.7)): + rows[f"transformer_blocks.{block}.img_mlp.proj"]["validation"]["relative_l2"] = score + return {"complete": True, "layers": layers, "files": {"all.safetensors": {"layer_stats": rows}}} + + +class RankProbeTests(unittest.TestCase): + def test_selects_worst_validation_projection_in_each_block_range(self): + manifest = _manifest() + original = copy.deepcopy(manifest) + selected = select_probe_layers(manifest) + self.assertEqual(selected, [f"transformer_blocks.{b}.img_mlp.proj" for b in (2, 17, 25)]) + self.assertEqual(manifest, original) + + def test_selection_never_uses_heldout_or_other_roles(self): + manifest = _manifest() + selected = select_probe_layers(manifest) + for i, row in enumerate(manifest["files"]["all.safetensors"]["layer_stats"].values()): + row["heldout"] = {"relative_l2": float("nan") if i % 2 else 1e12, + "mse": float("inf"), "finite": False} + self.assertEqual(select_probe_layers(manifest), selected) + + def test_ranges_include_exact_boundary_blocks(self): + manifest = _manifest() + rows = manifest["files"]["all.safetensors"]["layer_stats"] + for block in range(32): + rows[f"transformer_blocks.{block}.img_mlp.proj"]["validation"]["relative_l2"] = 0.01 + for block in (10, 11, 22): + rows[f"transformer_blocks.{block}.img_mlp.proj"]["validation"]["relative_l2"] = 1 + self.assertEqual(select_probe_layers(manifest), [f"transformer_blocks.{b}.img_mlp.proj" for b in (10, 11, 22)]) + + def test_selection_rejects_invalid_validation_and_missing_depth_range(self): + for metric in ({"relative_l2": float("nan"), "finite": True}, + {"relative_l2": float("inf"), "finite": True}, + {"relative_l2": -0.1, "finite": True}, + {"relative_l2": 0.1, "finite": False}): + with self.subTest(metric=metric): + manifest = _manifest() + rows = manifest["files"]["all.safetensors"]["layer_stats"] + rows["transformer_blocks.2.img_mlp.proj"]["validation"] = metric + with self.assertRaises(ValueError): + select_probe_layers(manifest) + manifest = _manifest() + rows = manifest["files"]["all.safetensors"]["layer_stats"] + for block in range(11, 22): + del rows[f"transformer_blocks.{block}.img_mlp.proj"] + with self.assertRaisesRegex(ValueError, "11..21"): + select_probe_layers(manifest) + + def test_validation_ties_are_independent_of_manifest_order(self): + manifest = _manifest() + rows = manifest["files"]["all.safetensors"]["layer_stats"] + for name, row in rows.items(): + if name.endswith(".img_mlp.proj"): + row["validation"]["relative_l2"] = 0.1 + before = select_probe_layers(manifest) + manifest["files"]["all.safetensors"]["layer_stats"] = dict(reversed(list(rows.items()))) + self.assertEqual(select_probe_layers(manifest), before) + self.assertEqual(before, [f"transformer_blocks.{b}.img_mlp.proj" for b in (0, 11, 22)]) + + def test_weight_cost_rejects_impossible_rank_and_empty_dimensions(self): + for args in ((128, 128, 256), (0, 128, 128), (128, -1, 128)): + with self.subTest(args=args), self.assertRaises(ValueError): + lowrank_cost(*args) + with self.assertRaises(ValueError): + lowrank_cost(4096, 12288, 256, role_layers=0) + + def test_bf16_lowrank_weight_cost_matches_exact_parameter_counts(self): + cost = lowrank_cost(4096, 12288, 256) + self.assertEqual(cost["lowrank_bytes"], 2 * (4096 * 256 + 12288 * 256)) + self.assertEqual(cost["delta_bytes_per_layer"], 4 * 1024 * 1024) + self.assertEqual(cost["delta_bytes_all_role_layers"], 128 * 1024 * 1024) + self.assertEqual(lowrank_cost(4096, 12288, 128)["delta_bytes_per_layer"], 0) + largest = lowrank_cost(4096, 12288, 1024) + self.assertEqual(largest["delta_bytes_all_role_layers"], 896 * 1024 * 1024) + custom = lowrank_cost(4096, 12288, 256, source_rank=64, role_layers=3) + self.assertEqual(custom["delta_bytes_all_role_layers"], 3 * 2 * (4096 + 12288) * (256 - 64)) + + def test_rank_change_preserves_every_other_source_recipe_setting(self): + stats = _stats() + saved = copy.deepcopy(stats) + expected = fixed_recipe(stats, iterations=16) + expected["ranks"] = (512,) + self.assertEqual(probe_recipe(stats, 512), expected) + self.assertEqual(stats, saved) + actual = probe_recipe(stats, 256, iterations=31) + self.assertEqual(actual["iterations"], 31) + self.assertEqual(actual["seed"], 1977) + self.assertEqual(actual["baseline_seed"], 1977) + self.assertEqual(actual["fixed_smoothing"], "activation_only") + self.assertEqual(tuple(actual["alphas"]), (0.75,)) + self.assertEqual(actual["factorization"], "up_singular") + + def test_probe_recipe_ignores_heldout_and_accepts_kernel_rank_limits(self): + stats = _stats() + before = probe_recipe(stats, 1024) + stats["heldout"] = {"rank": 16, "seed": 0, "mse": float("nan")} + self.assertEqual(probe_recipe(stats, 1024), before) + self.assertEqual(tuple(probe_recipe(stats, 16)["ranks"]), (16,)) + for invalid in (0, -16, 17, 1040, True, 128.0): + with self.subTest(rank=invalid): + with self.assertRaises(ValueError): + probe_recipe(stats, invalid) + with self.assertRaises(ValueError): + probe_recipe(stats, 128, iterations=0) + + +if __name__ == "__main__": + unittest.main(verbosity=2) diff --git a/reproduction/nunchaku_backend/rank_upgrade_helpers.py b/reproduction/nunchaku_backend/rank_upgrade_helpers.py new file mode 100644 index 0000000000000000000000000000000000000000..f522222ea10b348c4d810eaeb45c3a20c53abd14 --- /dev/null +++ b/reproduction/nunchaku_backend/rank_upgrade_helpers.py @@ -0,0 +1,105 @@ +"""Pure selection and tensor-schema guards for isolated MLP rank upgrades. + +These helpers neither run inference nor mutate checkpoint tensors. Validation +MSE alone selects a candidate; heldout fields are deliberately never read. +""" +from __future__ import annotations + +import copy +import math + + +def _rank(value): + if not isinstance(value, int) or isinstance(value, bool) or not 0 < value <= 1024 or value % 16: + raise ValueError("Rank must be an integer multiple of 16 in [16,1024]") + + +def require_unique_ranks(ranks, source_rank=128): + """Validate strictly higher candidate ranks, retaining the caller's order.""" + _rank(source_rank) + result = tuple(ranks) + if not result: + raise ValueError("At least one candidate rank is required") + for value in result: + _rank(value) + if value <= source_rank: + raise ValueError("Candidate ranks must exceed the source rank") + if len(set(result)) != len(result): + raise ValueError("Candidate ranks must be unique") + return result + + +def _valid_mse(metrics): + try: + value = float(metrics["mse"]) + except (KeyError, TypeError, ValueError, OverflowError): + return None + if not metrics.get("finite", True) or not math.isfinite(value) or value < 0: + return None + return value + + +def select_validation_candidate(source_metrics, candidates): + """Return the strictly best candidate index; source and earlier ties win. + +The caller supplies actual validation output metrics. Invalid source metrics +raise because they cannot provide a comparison baseline; invalid candidate +metrics are rejected. Rank determines no preference beyond list order. +""" + best = _valid_mse(source_metrics) + if best is None: + raise ValueError("Source validation MSE must be finite and nonnegative") + winner = None + for index, candidate in enumerate(candidates): + _rank(candidate["rank"]) + score = _valid_mse(candidate["validation"]) + if score is not None and score < best: + best, winner = score, index + return winner + + +def validate_rank_state(source_state, candidate_state, source_info, candidate_rank): + """Allow changed low-rank dimensions only; preserve all other tensor schema. + +Values in any tensor may change after recalibration. Only proj_down/proj_up +may change shape. Both states must match the canonical generic Nunchaku INT4 +layout, preventing a malformed source from legitimizing a malformed candidate. +""" + import torch + + require_unique_ranks((candidate_rank,), source_info["rank"]) + inputs, outputs = source_info["in_features"], source_info["out_features"] + if any(not isinstance(dim, int) or isinstance(dim, bool) or dim <= 0 or dim % 128 + for dim in (inputs, outputs)): + raise ValueError("Nunchaku input/output dimensions must be positive multiples of 128") + if max(source_info["rank"], candidate_rank) > min(inputs, outputs): + raise ValueError("Rank exceeds layer dimensions") + if not isinstance(source_info["bias"], bool): + raise ValueError("Source bias metadata must be boolean") + if source_info.get("precision", "int4") != "int4": + raise ValueError("Only signed INT4 checkpoint schema is supported") + shapes = { + "qweight": (outputs, inputs // 2), + "wscales": (inputs // 64, outputs), + "smooth_factor": (inputs,), + "smooth_factor_orig": (inputs,), + } + if source_info["bias"]: + shapes["bias"] = (outputs,) + for label, state, rank in (("source", source_state, source_info["rank"]), + ("candidate", candidate_state, candidate_rank)): + expected = {**shapes, "proj_down": (inputs, rank), "proj_up": (outputs, rank)} + if set(state) != set(expected): + raise ValueError(f"{label} state keys do not match layer schema") + for key, shape in expected.items(): + tensor = state[key] + dtype = torch.int8 if key == "qweight" else torch.bfloat16 + if not isinstance(tensor, torch.Tensor) or tuple(tensor.shape) != shape or tensor.dtype != dtype: + raise ValueError(f"{label} tensor schema is invalid: {key}") + if tensor.is_meta: + raise ValueError(f"{label} tensor has no stored data: {key}") + if label == "candidate" and tensor.is_floating_point() and not torch.isfinite(tensor).all(): + raise ValueError(f"Candidate parameter is nonfinite: {key}") + result = copy.deepcopy(source_info) + result["rank"] = candidate_rank + return result diff --git a/reproduction/nunchaku_backend/rank_upgrade_helpers_test.py b/reproduction/nunchaku_backend/rank_upgrade_helpers_test.py new file mode 100644 index 0000000000000000000000000000000000000000..41e51f85647ff1977ad0c8764401881d289149b5 --- /dev/null +++ b/reproduction/nunchaku_backend/rank_upgrade_helpers_test.py @@ -0,0 +1,141 @@ +"""CPU-only schema and heldout-independent rank-upgrade selection checks.""" +import copy +import unittest + +import torch + +from .rank_upgrade_helpers import require_unique_ranks, select_validation_candidate, validate_rank_state + + +def _info(bias=True): + return {"in_features": 128, "out_features": 256, "rank": 16, + "bias": bias, "precision": "int4", "provenance": {"source": "immutable"}} + + +def _state(rank, bias=True): + state = { + "qweight": torch.zeros((256, 64), dtype=torch.int8), + "wscales": torch.ones((2, 256), dtype=torch.bfloat16), + "smooth_factor": torch.ones(128, dtype=torch.bfloat16), + "smooth_factor_orig": torch.ones(128, dtype=torch.bfloat16), + "proj_down": torch.zeros((128, rank), dtype=torch.bfloat16), + "proj_up": torch.zeros((256, rank), dtype=torch.bfloat16), + } + if bias: + state["bias"] = torch.zeros(256, dtype=torch.bfloat16) + return state + + +def _metrics(mse, finite=True): + return {"mse": mse, "finite": finite} + + +class RankUpgradeHelperTests(unittest.TestCase): + def test_valid_rank_upgrade_changes_info_copy_only(self): + for bias in (True, False): + with self.subTest(bias=bias): + source, candidate, info = _state(16, bias), _state(32, bias), _info(bias) + original_info = copy.deepcopy(info) + originals = {k: (v, v.clone()) for k, v in source.items()} + candidate["wscales"].mul_(2) # Schema fixed, recalibrated values permitted. + result = validate_rank_state(source, candidate, info, 32) + self.assertEqual(result, {**info, "rank": 32}) + result["provenance"]["source"] = "modified_copy" + self.assertEqual(info, original_info) + for key, (tensor, snapshot) in originals.items(): + self.assertIs(source[key], tensor) + self.assertTrue(torch.equal(source[key], snapshot)) + + def test_every_non_lora_shape_is_fixed_and_lora_shape_must_match_rank(self): + for key in _state(32): + with self.subTest(key=key): + candidate = _state(32) + candidate[key] = candidate[key].reshape(-1)[:1] + with self.assertRaises(ValueError): + validate_rank_state(_state(16), candidate, _info(), 32) + + def test_missing_extra_bias_and_unrecognized_keys_fail(self): + for key in _state(32): + candidate = _state(32) + del candidate[key] + with self.subTest(missing=key), self.assertRaises(ValueError): + validate_rank_state(_state(16), candidate, _info(), 32) + candidate = _state(32) + candidate["extra"] = torch.ones(1, dtype=torch.bfloat16) + with self.assertRaises(ValueError): + validate_rank_state(_state(16), candidate, _info(), 32) + with self.assertRaises(ValueError): + validate_rank_state(_state(16, False), _state(32, True), _info(False), 32) + + def test_dtype_and_finiteness_checks_cover_all_parameters(self): + for key in _state(32): + candidate = _state(32) + candidate[key] = candidate[key].float() + with self.subTest(dtype=key), self.assertRaises(ValueError): + validate_rank_state(_state(16), candidate, _info(), 32) + if key == "qweight": + continue + for value in (float("nan"), float("inf"), -float("inf")): + candidate = _state(32) + candidate[key].flatten()[0] = value + with self.subTest(finite=key, value=value), self.assertRaises(ValueError): + validate_rank_state(_state(16), candidate, _info(), 32) + + def test_invalid_source_schema_and_impossible_rank_fail(self): + source = _state(16) + source["proj_up"] = torch.zeros((256, 32), dtype=torch.bfloat16) + with self.assertRaises(ValueError): + validate_rank_state(source, _state(32), _info(), 32) + for rank in (16, 256): + with self.subTest(rank=rank), self.assertRaises(ValueError): + validate_rank_state(_state(16), _state(rank), _info(), rank) + + def test_source_wins_exact_ties_and_regressions(self): + candidates = [{"rank": 256, "validation": _metrics(0.25)}, + {"rank": 512, "validation": _metrics(0.3)}] + self.assertIsNone(select_validation_candidate(_metrics(0.25), candidates)) + self.assertIsNone(select_validation_candidate(_metrics(0), candidates)) + self.assertIsNone(select_validation_candidate(_metrics(0.25), [])) + + def test_strict_best_validation_and_earlier_candidate_tie(self): + candidates = [{"rank": rank, "validation": _metrics(mse)} + for rank, mse in ((256, 0.2), (512, 0.1), (1024, 0.1))] + saved = copy.deepcopy(candidates) + self.assertEqual(select_validation_candidate(_metrics(0.3), candidates), 1) + self.assertEqual(candidates, saved) + candidates[2]["validation"]["mse"] = 0.099 + self.assertEqual(select_validation_candidate(_metrics(0.3), candidates), 2) + + def test_heldout_never_participates_in_selection(self): + candidates = [{"rank": 256, "validation": _metrics(0.2), "heldout": _metrics(1e12)}, + {"rank": 512, "validation": _metrics(0.3), "heldout": _metrics(0)}] + source = {**_metrics(0.4), "heldout": _metrics(float("nan"))} + self.assertEqual(select_validation_candidate(source, candidates), 0) + candidates[0]["heldout"] = {"finite": False, "mse": float("nan")} + candidates[1]["heldout"] = {"rank": 1024, "validation": _metrics(-100)} + self.assertEqual(select_validation_candidate(source, candidates), 0) + + def test_invalid_source_raises_invalid_candidates_are_rejected(self): + invalid = [_metrics(float("nan")), _metrics(float("inf")), _metrics(-1), + _metrics(0.1, False), {}, _metrics("invalid")] + for metrics in invalid: + with self.subTest(metrics=metrics): + with self.assertRaises(ValueError): + select_validation_candidate(metrics, []) + self.assertEqual(select_validation_candidate(_metrics(1), [ + {"rank": 256, "validation": metrics}, + {"rank": 512, "validation": _metrics(0.5)}]), 1) + + def test_unique_upgrade_rank_bounds_and_order(self): + self.assertEqual(require_unique_ranks([512, 256, 1024]), (512, 256, 1024)) + self.assertEqual(require_unique_ranks([32], source_rank=16), (32,)) + for ranks in ([], [128], [64], [256, 256], [1040], [257], [True], [256.0], [-16]): + with self.subTest(ranks=ranks), self.assertRaises(ValueError): + require_unique_ranks(ranks) + for rank in (0, 17, True, 128.0, 1040): + with self.subTest(source_rank=rank), self.assertRaises(ValueError): + require_unique_ranks([256], source_rank=rank) + + +if __name__ == "__main__": + unittest.main(verbosity=2) diff --git a/reproduction/nunchaku_backend/refine_checkpoint_v3.py b/reproduction/nunchaku_backend/refine_checkpoint_v3.py new file mode 100644 index 0000000000000000000000000000000000000000..49b527cb8fd3337656d3d11943f1557c159ad19f --- /dev/null +++ b/reproduction/nunchaku_backend/refine_checkpoint_v3.py @@ -0,0 +1,401 @@ +"""Optional fixed-recipe refinement, guarded by the ACTUAL packed checkpoint. + +This is prepared tooling, not an instruction to launch another full export. +GPU execution requires an explicit --device after the main agent reserves it. +The immutable source wins ties and all validation regressions. Heldout scores +never participate in selection. All output goes to a distinct directory. +""" +from __future__ import annotations + +import argparse +import copy +import json +import math +from pathlib import Path +import shutil +import time + +from . import FORMAT +from .layout import BLOCK_LINEAR_PATTERN +from .compare_iterations_v3 import layer_stats +from .checkpoint_io import _json_write, _source_index, _read_tensor +from .export_v3 import ActivationReader, _hash_file + + +def choose_refinement(source_metrics, candidate_metrics, min_relative_improvement=0.0): + """Use validation metrics only; never inspect any heldout values.""" + if not 0 <= min_relative_improvement < 1: + raise ValueError("Minimum relative improvement must be in [0,1)") + source_mse = float(source_metrics["mse"]) + candidate_mse = float(candidate_metrics["mse"]) + if not source_metrics.get("finite", True) or not math.isfinite(source_mse) or source_mse < 0: + raise ValueError("The actual source checkpoint has invalid validation output") + if not candidate_metrics.get("finite", True) or not math.isfinite(candidate_mse) or candidate_mse < 0: + return False, "candidate_nonfinite_or_invalid" + if candidate_mse < source_mse * (1 - min_relative_improvement): + return True, "validation_improved" + return False, "source_retained_tie_regression_or_below_threshold" + + +def select_state(source_state, candidate_state, accepted): + """Fallback returns the exact original mapping/tensors without alteration.""" + if not accepted: + return source_state + if set(source_state) != set(candidate_state): + raise ValueError("Candidate state keys differ from the source checkpoint") + import torch + for key, original in source_state.items(): + current = candidate_state[key] + if original.shape != current.shape or original.dtype != current.dtype: + raise ValueError(f"Candidate tensor schema changed: {key}") + if current.is_floating_point() and not torch.isfinite(current).all(): + raise ValueError(f"Candidate parameter is nonfinite: {key}") + return candidate_state + + +def assert_separate_output(source_dir, out_dir): + source = Path(source_dir).resolve() + output = Path(out_dir).resolve() + if source == output or source in output.parents or output in source.parents: + raise ValueError("Output must be separate from, and not an ancestor/descendant of, the immutable source") + + +def fixed_recipe(stats, iterations=100): + """Replay the selected recipe, not a fresh smoothing/rank grid search.""" + if iterations <= 0: + raise ValueError("Iteration limit must be positive") + selected, search = stats["selected"], stats["search"] + family = selected["family"] + if family == "one_pass_baseline": + # Baseline smoothing is unnormalized. Mapping it into the search + # normalization would not be an exact replay of its original recipe. + raise NotImplementedError("Retain one-pass-baseline selected layers unchanged") + if family not in ("identity", "activation_only", "smoothquant"): + raise ValueError(f"Unsupported source smoothing family {family}") + factorization = selected["factorization"] if "factorization" in selected else search["factorization"] + if selected["weighting"] not in ("none", "rms") or factorization not in ("balanced", "up_singular"): + raise ValueError("Unsupported source weighting or factorization") + for label, value in (("rank", selected["rank"]), ("baseline rank", stats["baseline"]["conversion"]["rank"])): + if not isinstance(value, int) or isinstance(value, bool) or value <= 0 or value % 16: + raise ValueError(f"Invalid source {label}") + for label in ("seed", "niter", "oversample"): + if not isinstance(search[label], int) or isinstance(search[label], bool) or search[label] < 0: + raise ValueError(f"Invalid source {label}") + if family != "identity" and (not math.isfinite(float(selected["alpha"])) or not 0 <= float(selected["alpha"]) <= 1): + raise ValueError("Invalid source smoothing alpha") + for label in ("ridge", "gptq_damp"): + if not math.isfinite(float(search[label])) or float(search[label]) <= 0: + raise ValueError(f"Invalid source {label}") + if not isinstance(search["final_gptq"], bool): + raise ValueError("Source final_gptq flag must be boolean") + return { + "ranks": (int(selected["rank"]),), "iterations": iterations, + "alphas": () if family == "identity" else (float(selected["alpha"]),), + "smoothing_families": () if family == "identity" else (family,), + "fixed_smoothing": family, "weighting": (selected["weighting"],), + "factorization": factorization, "seed": int(search["seed"]), + "baseline_seed": int(search["seed"]), + "baseline_rank": int(stats["baseline"]["conversion"]["rank"]), + "niter": int(search["niter"]), "oversample": int(search["oversample"]), + "ridge": float(search["ridge"]), "final_gptq": bool(search["final_gptq"]), + "gptq_damp": float(search["gptq_damp"]), + } + + +def require_matching_provenance(source_manifest, activation_fingerprint, model_config_hash, baseline_hash): + identity = source_manifest["conversion_identity"] + for key, actual in (("activation_fingerprint", activation_fingerprint), + ("source_config_sha256", model_config_hash), + ("baseline_calibration_sha256", baseline_hash)): + if identity.get(key) != actual: + raise ValueError(f"Refinement must use the source checkpoint's exact {key}") + + +def _actual_forward(state, info, inputs, device): + """No unpacking approximation: evaluate the actual Nunchaku stored tensors.""" + import torch + from nunchaku.models.linear import SVDQW4A4Linear + module = SVDQW4A4Linear(info["in_features"], info["out_features"], rank=info["rank"], + bias=info["bias"], precision="int4", torch_dtype=torch.bfloat16, device=device) + module.load_state_dict(state, strict=True) + module.eval().requires_grad_(False) + with torch.inference_mode(): + return module(inputs.unsqueeze(0)).squeeze(0).float() + + +def _source_layer_state(index, name): + prefix = name + "." + return {key[len(prefix):]: _read_tensor(index, key) for key in index if key.startswith(prefix)} + + +def write_layer_shard(output, filename, name, state, source_index=None, retain_source=False): + """Copy a complete source shard byte-for-byte on the normal V3 fallback.""" + from safetensors.torch import save_file + keys = {name + "." + key for key in state} + temporary = output / (filename + ".tmp") + copied = False + if retain_source and source_index is not None: + paths = {source_index[key] for key in keys} + if len(paths) == 1: + path = next(iter(paths)) + all_file_keys = {key for key, value in source_index.items() if value == path} + if all_file_keys == keys: + shutil.copyfile(path, temporary) + copied = True + if not copied: + save_file({name + "." + key: value.detach().cpu().contiguous() for key, value in state.items()}, str(temporary)) + temporary.replace(output / filename) + return {"keys": sorted(keys), "bytes": sum(t.numel() * t.element_size() for t in state.values()), + "sha256": _hash_file(output / filename), "source_file_copied_verbatim": copied} + + +def main(): + parser = argparse.ArgumentParser(description=__doc__) + parser.add_argument("--source-checkpoint", type=Path, required=True) + parser.add_argument("--model-path", type=Path, required=True) + parser.add_argument("--activations", type=Path, required=True) + parser.add_argument("--baseline-calibration", type=Path, required=True) + parser.add_argument("--out", type=Path, required=True) + parser.add_argument("--device", required=True, help="Explicit reserved device, e.g. cuda:0") + parser.add_argument("--iterations", type=int, default=100) + parser.add_argument("--min-validation-improvement", type=float, default=0.0) + parser.add_argument("--threads", type=int, default=8) + parser.add_argument("--resume", action="store_true") + parser.add_argument("--max-layers", type=int, help="Bounded verification export; remains incomplete") + parser.add_argument("--dry-run", action="store_true", help="CPU-only provenance/recipe validation; writes no checkpoint") + args = parser.parse_args() + assert_separate_output(args.source_checkpoint, args.out) + if args.iterations < 1 or not 0 <= args.min_validation_improvement < 1: + parser.error("Invalid iteration limit or validation improvement threshold") + import torch + from safetensors.torch import load_file, save_file + from .optimize_v3 import Candidate, optimize_linear_weight, output_metrics, _weight_only_proxy + from .runtime import _read_manifest, probe_backend + torch.set_num_threads(args.threads) + device = torch.device(args.device) + if device.type != "cuda" or device.index is None: + parser.error("Actual packed-checkpoint validation requires an explicitly reserved indexed CUDA device") + source = args.source_checkpoint.resolve() + source_manifest_path = source / "manifest.json" + source_manifest = _read_manifest(source) + source_manifest_hash = _hash_file(source_manifest_path) + source_stats = layer_stats(source_manifest) + names = sorted(source_manifest["layers"], key=lambda n: (int(n.split(".")[1]), n)) + if len(source_stats) != 224 or set(source_stats) != set(names): + raise ValueError("Source must have complete selected-recipe V3 statistics for all224 layers") + directory = args.model_path / "transformer" if (args.model_path / "transformer").is_dir() else args.model_path + model_config_hash = _hash_file(directory / "config.json") + source_config_hash = _hash_file(source / "config.json") + source_index_hash = _hash_file(source / "model.safetensors.index.json") + if json.loads((source / "config.json").read_text()) != json.loads((directory / "config.json").read_text()): + raise ValueError("Actual source checkpoint config differs from the original BF16 model config") + baseline_file = args.baseline_calibration / "activation_stats.safetensors" if args.baseline_calibration.is_dir() else args.baseline_calibration + baseline_hash = _hash_file(baseline_file) + activations = ActivationReader(args.activations) + activations.require(names) + require_matching_provenance(source_manifest, activations.fingerprint, model_config_hash, baseline_hash) + recipes = {} + retained_without_replay = [] + for name in names: + if source_stats[name]["selected"]["rank"] != source_manifest["layers"][name]["rank"]: + raise ValueError(f"Source selected-recipe rank and packed-layer rank disagree: {name}") + if source_manifest["layers"][name]["rank"] > min(source_manifest["layers"][name]["in_features"], source_manifest["layers"][name]["out_features"]): + raise ValueError(f"Source rank exceeds layer dimensions: {name}") + try: + recipes[name] = fixed_recipe(source_stats[name], args.iterations) + except NotImplementedError: + retained_without_replay.append(name) + if args.dry_run: + print(json.dumps({"dry_run": True, "gpu_work_executed": False, "source_manifest_sha256": source_manifest_hash, + "replayed_layers": len(recipes), "retained_without_replay": retained_without_replay, + "iterations": args.iterations, "out": str(args.out), "cuda_initialized": torch.cuda.is_initialized()})) + return + + source_index = _source_index(source) + source_weight_map = json.loads((source / "model.safetensors.index.json").read_text())["weight_map"] + if set(source_weight_map) != set(source_index) or any(source_weight_map[key] != path.name for key, path in source_index.items()): + raise ValueError("Actual source weight index differs from its checkpoint shards") + model_index = _source_index(directory) + expected_model_names = {key[:-7] for key in model_index if key.endswith(".weight") and BLOCK_LINEAR_PATTERN.fullmatch(key[:-7])} + if expected_model_names != set(names): + raise ValueError("Source checkpoint and original BF16 transformer layer names differ") + source_files = {str(path.relative_to(source)): _hash_file(path) for path in sorted(set(source_index.values()))} + identity = {"algorithm": "v3_fixed_recipe_refinement_actual_source_guard", "source_checkpoint": str(source), + "source_manifest_sha256": source_manifest_hash, "source_file_sha256": source_files, + "source_checkpoint_config_sha256": source_config_hash, "source_checkpoint_index_sha256": source_index_hash, + "activation_fingerprint": activations.fingerprint, "source_config_sha256": model_config_hash, + "baseline_calibration_sha256": baseline_hash, "iterations": args.iterations, + "min_validation_improvement": args.min_validation_improvement, "device": str(device)} + identity["settings"] = {"iterations": args.iterations, "selected_recipe_refinement": True, + "output_correction": source_manifest["conversion_identity"]["settings"].get("output_correction", True)} + identity["code_sha256"] = {filename: _hash_file(Path(__file__).parent / filename) for filename in + ("refine_checkpoint_v3.py", "optimize_v3.py", "gptq_v3.py", "baseline_candidate.py", "packing.py", "checkpoint_io.py", "layout.py", "runtime.py")} + args.out.mkdir(parents=True, exist_ok=True) + manifest_path = args.out / "manifest.json" + if manifest_path.exists(): + manifest = json.loads(manifest_path.read_text()) + if not args.resume or manifest.get("conversion_identity") != identity: + raise ValueError("Resume requires the identical immutable source, data, and settings") + for filename, info in manifest["files"].items(): + if not (args.out / filename).is_file() or _hash_file(args.out / filename) != info["sha256"]: + raise ValueError(f"Missing or modified resume shard {filename}") + for item in manifest.get("refinement_reports", {}).values(): + path = args.out / item["file"] + if not path.is_file() or _hash_file(path) != item["sha256"]: + raise ValueError(f"Missing or modified resume report {item['file']}") + else: + if any(args.out.iterdir()): + raise ValueError("Use a new empty output directory") + manifest = {"backend_format": FORMAT, "complete": False, "model": source_manifest["model"], + "model_revision": source_manifest["model_revision"], "diffusers_commit": source_manifest["diffusers_commit"], + "created_unix": time.time(), "conversion_identity": identity, "runtime": probe_backend(), + "algorithm": "fixed selected V3 recipes replayed with a higher iteration limit; actual packed-source validation guard", + "calibration": activations.provenance, "layers": {}, "files": {}, "refinement_reports": {}, + "limitations": "Independent heldout reporting only. Per-linear improvements do not establish image equivalence. Original V3 tensors are retained on validation ties/regressions."} + shutil.copyfile(source / "config.json", args.out / "config.json") + _json_write(manifest_path, manifest) + baseline = load_file(str(baseline_file), device="cpu") + if any(name + ".input_absmax" not in baseline for name in names): + raise ValueError("Original one-pass baseline calibration is incomplete") + boundary_name = "model-boundary.safetensors" + if boundary_name not in manifest["files"]: + prefixes = tuple(name + "." for name in names) + boundary_keys = {key for key in source_index if not key.startswith(prefixes)} + boundary_paths = {source_index[key] for key in boundary_keys} + copied = False + if len(boundary_paths) == 1: + path = next(iter(boundary_paths)) + if {key for key, value in source_index.items() if value == path} == boundary_keys: + shutil.copyfile(path, args.out / boundary_name) + copied = True + boundary = {key: _read_tensor(source_index, key) for key in boundary_keys} + if not copied: + save_file(boundary, str(args.out / boundary_name)) + manifest["files"][boundary_name] = {"keys": sorted(boundary_keys), "bytes": sum(t.numel() * t.element_size() for t in boundary.values()), + "sha256": _hash_file(args.out / boundary_name), "source_file_copied_verbatim": copied} + del boundary + _json_write(manifest_path, manifest) + + completed_now = 0 + for ordinal, name in enumerate(names): + if name in manifest["layers"]: + continue + started = time.perf_counter() + original_state = _source_layer_state(source_index, name) + info = source_manifest["layers"][name] + old_stats = source_stats[name] + if name not in recipes: + accepted, reason = False, "one_pass_baseline_recipe_retained_without_approximate_mapping" + chosen_state = original_state + chosen_stats = copy.deepcopy(old_stats) + report = {"accepted": False, "reason": reason, "source_selected": old_stats["selected"]} + else: + weight = _read_tensor(model_index, name + ".weight") + bias = _read_tensor(model_index, name + ".bias") if name + ".bias" in model_index else None + if weight.dtype != torch.bfloat16 or tuple(weight.shape) != (info["out_features"], info["in_features"]): + raise ValueError(f"Original BF16 weight mismatch: {name}") + train = activations.read(name, "train") + validation = activations.read(name, "validation").to(device) + train_absmax = activations.read(name, "input_absmax") if (name, "input_absmax") in activations.entries else None + teacher_weight = weight.to(device) + teacher_bias = bias.to(device) if bias is not None else None + with torch.inference_mode(): + teacher_validation = torch.nn.functional.linear(validation, teacher_weight, teacher_bias).float() + source_validation = output_metrics(_actual_forward(original_state, info, validation, device), teacher_validation) + candidate_state, candidate_stats, reference = optimize_linear_weight( + weight, train, validation, None, bias, **recipes[name], train_absmax=train_absmax, + baseline_absmax=baseline[name + ".input_absmax"], conversion_device=device, + objective_backend="nunchaku", output_correction=source_manifest["conversion_identity"]["settings"].get("output_correction", True), + ) + # A regenerated rank32 guard is not the requested refinement + # recipe. Never deploy it as a substitute for actual V3. + same_recipe = (candidate_stats["selected"]["rank"] == info["rank"] and + candidate_stats["selected"]["family"] == old_stats["selected"]["family"] and + candidate_stats["selected"]["alpha"] == old_stats["selected"]["alpha"]) + if same_recipe: + select_state(original_state, candidate_state, True) # schema/finite validation + candidate_validation = output_metrics(_actual_forward(candidate_state, info, validation, device), teacher_validation) + accepted, reason = choose_refinement(source_validation, candidate_validation, args.min_validation_improvement) + if all(torch.equal(value, candidate_state[key].detach().cpu()) for key, value in original_state.items()): + accepted, reason = False, "candidate_tensors_identical_to_source" + else: + candidate_validation = candidate_stats["validation"] + accepted, reason = False, "candidate_did_not_preserve_selected_recipe" + chosen_state = select_state(original_state, candidate_state, accepted) + # Decision frozen before reading or evaluating heldout tensors. + heldout = activations.read(name, "heldout").to(device) + teacher_heldout = torch.nn.functional.linear(heldout, teacher_weight, teacher_bias).float() + source_heldout = output_metrics(_actual_forward(original_state, info, heldout, device), teacher_heldout) + candidate_heldout = output_metrics(_actual_forward(candidate_state, info, heldout, device), teacher_heldout) if same_recipe else None + chosen_stats = copy.deepcopy(candidate_stats if accepted else old_stats) + chosen_stats["validation"] = candidate_validation if accepted else source_validation + chosen_stats["heldout"] = candidate_heldout if accepted else source_heldout + # Preserve the original documented rank32 comparison for + # existing summary tooling; the new actual-source comparison + # is explicit in the separate refinement report. + chosen_stats["baseline"] = copy.deepcopy(old_stats["baseline"]) + chosen_stats["baseline_metrics_provenance"] = "Copied original one-pass rank32 metrics from immutable V3 manifest on the identical calibration archive" + chosen_stats["heldout_mse_ratio_to_baseline"] = chosen_stats["heldout"]["mse"] / max(old_stats["baseline"]["heldout"]["mse"], 1e-30) + chosen_stats["validation_mse_ratio_to_baseline"] = chosen_stats["validation"]["mse"] / max(old_stats["baseline"]["validation"]["mse"], 1e-30) + if accepted: + candidate = Candidate(reference["residual_dequant"], reference["weight_scales"], reference["down_unpacked"], + reference["up_unpacked"], reference["smooth"], candidate_stats["selected"]) + chosen_stats["heldout_weight_only_proxy"] = output_metrics(_weight_only_proxy(candidate, heldout, teacher_bias), teacher_heldout) + del candidate + report = {"accepted": accepted, "reason": reason, "source_selected": old_stats["selected"], + "candidate_selected": candidate_stats["selected"], "source_actual_validation": source_validation, + "candidate_actual_validation": candidate_validation, "source_actual_heldout": source_heldout, + "candidate_actual_heldout": candidate_heldout, "candidate_optimizer_stats": candidate_stats, + "validation_mse_ratio_to_actual_source": candidate_validation["mse"] / max(source_validation["mse"], 1e-30), + "heldout_mse_ratio_to_actual_source": candidate_heldout["mse"] / max(source_heldout["mse"], 1e-30) if candidate_heldout else None, + "decision_uses_heldout": False, "source_reported_validation": old_stats["validation"]} + del teacher_weight, teacher_bias, train, validation, heldout, teacher_validation, teacher_heldout, reference + filename = f"model-layer-{ordinal:03d}.safetensors" + file_info = write_layer_shard(args.out, filename, name, chosen_state, source_index, retain_source=not accepted) + report.update({"layer": name, "seconds": time.perf_counter() - started, + "source_tensors_preserved": not accepted, "source_manifest_sha256": source_manifest_hash}) + report_filename = f"refinement-layer-{ordinal:03d}.json" + _json_write(args.out / report_filename, report) + chosen_stats["refinement"] = {key: value for key, value in report.items() if key != "candidate_optimizer_stats"} + file_info["layer_stats"] = {name: chosen_stats} + manifest["files"][filename] = file_info + manifest["layers"][name] = copy.deepcopy(info) + manifest["refinement_reports"][name] = {"file": report_filename, "sha256": _hash_file(args.out / report_filename), + "accepted": accepted, "reason": reason} + _json_write(manifest_path, manifest) + print(json.dumps({"event": "refined_layer", "layer": name, "accepted": accepted, "reason": reason, + "validation_ratio_to_actual_source": report.get("validation_mse_ratio_to_actual_source"), + "heldout_ratio_to_actual_source": report.get("heldout_mse_ratio_to_actual_source"), + "seconds": report["seconds"]}), flush=True) + del original_state, chosen_state, chosen_stats + if name in recipes: + del candidate_state, candidate_stats, weight, bias + completed_now += 1 + if args.max_layers and completed_now >= args.max_layers: + break + # Detect external changes to the source rather than silently attributing + # the resulting checkpoint to an immutable source that no longer matches. + if _hash_file(source_manifest_path) != source_manifest_hash: + raise ValueError("Source manifest changed during refinement") + if _hash_file(source / "config.json") != source_config_hash or _hash_file(source / "model.safetensors.index.json") != source_index_hash: + raise ValueError("Source checkpoint config or weight index changed during refinement") + for filename, digest in source_files.items(): + if _hash_file(source / filename) != digest: + raise ValueError(f"Source checkpoint shard changed during refinement: {filename}") + if len(manifest["layers"]) == 224: + _json_write(args.out / "model.safetensors.index.json", { + "metadata": {"total_size": sum(item["bytes"] for item in manifest["files"].values())}, + "weight_map": {key: filename for filename, item in manifest["files"].items() for key in item["keys"]}, + }) + if not manifest["complete"]: + manifest["completed_unix"] = time.time() + manifest["complete"] = True + manifest["refinement_summary"] = {"layers": len(manifest["layers"]), + "accepted": sum(row["accepted"] for row in manifest["refinement_reports"].values()), + "retained_source": sum(not row["accepted"] for row in manifest["refinement_reports"].values())} + _json_write(manifest_path, manifest) + print(json.dumps({"event": "refinement_status", "complete": manifest["complete"], **manifest["refinement_summary"]}), flush=True) + + +if __name__ == "__main__": + main() diff --git a/reproduction/nunchaku_backend/refine_v3_test.py b/reproduction/nunchaku_backend/refine_v3_test.py new file mode 100644 index 0000000000000000000000000000000000000000..7e81fe6ba58e6ecc7c0a7e458422f15b034cded9 --- /dev/null +++ b/reproduction/nunchaku_backend/refine_v3_test.py @@ -0,0 +1,289 @@ +"""CPU safeguard tests for fixed-recipe checkpoint refinement.""" + +import copy +import contextlib +import hashlib +import io +import json +from pathlib import Path +import tempfile +import unittest +from unittest.mock import patch + +import torch + +from .refine_checkpoint_v3 import ( + assert_separate_output, + choose_refinement, + fixed_recipe, + require_matching_provenance, + select_state, + write_layer_shard, +) + + +def metric(mse, finite=True): + return {"mse": mse, "finite": finite, "relative_l2": float(mse) ** 0.5 if mse >= 0 else 0.0} + + +def source_stats(): + return { + "selected": {"family": "activation_only", "alpha": 0.75, "rank": 128, + "weighting": "rms", "factorization": "up_singular", + "iteration": 2, "output_correction": 0.0}, + "search": {"seed": 1953, "niter": 4, "oversample": 16, "iterations": 3, + "ridge": 0.01, "final_gptq": True, "gptq_damp": 0.01, + "factorization": "balanced"}, + "validation": metric(1.0), "heldout": metric(0.9), + "baseline": {"conversion": {"rank": 32}}, + } + + +class RefinementSafeguardTests(unittest.TestCase): + def test_selection_requires_strict_validation_improvement(self): + for candidate in (metric(1.0), metric(1.1)): + accepted, reason = choose_refinement(metric(1.0), candidate) + self.assertFalse(accepted) + self.assertTrue(reason) + accepted, reason = choose_refinement(metric(1.0), metric(0.8)) + self.assertTrue(accepted) + self.assertTrue(reason) + self.assertFalse(choose_refinement(metric(0.0), metric(0.0))[0]) + + def test_selection_ignores_attached_heldout_scores(self): + source = {**metric(1.0), "heldout": metric(0.0001)} + candidate = {**metric(0.8), "heldout": metric(1e9)} + first = choose_refinement(source, candidate) + self.assertTrue(first[0]) + source["heldout"] = metric(float("nan")) + candidate["heldout"] = metric(0.0) + self.assertEqual(first, choose_refinement(source, candidate)) + + def test_selection_rejects_nonfinite_and_honors_minimum_gain(self): + for candidate in (metric(float("nan")), metric(float("inf")), metric(0.1, finite=False), metric(-0.1)): + accepted, reason = choose_refinement(metric(1.0), candidate) + self.assertFalse(accepted) + self.assertTrue(reason) + self.assertFalse(choose_refinement(metric(10.0), metric(9.5), min_relative_improvement=0.1)[0]) + self.assertTrue(choose_refinement(metric(10.0), metric(8.0), min_relative_improvement=0.1)[0]) + for source in (metric(float("nan")), metric(float("inf")), metric(1, finite=False), metric(-1)): + with self.assertRaises(ValueError): + choose_refinement(source, metric(0.1)) + for threshold in (-0.1, 1, float("nan")): + with self.assertRaises(ValueError): + choose_refinement(metric(1), metric(0.1), threshold) + + def test_state_fallback_preserves_original_tensor_identity_and_bytes(self): + source = {"qweight": torch.arange(16, dtype=torch.int8).reshape(4, 4), + "proj_up": torch.arange(8).reshape(4, 2).bfloat16()} + candidate = {"qweight": -source["qweight"], "proj_up": source["proj_up"] + 3} + saved = {key: tensor.clone() for key, tensor in source.items()} + result = select_state(source, candidate, accepted=False) + self.assertEqual(result.keys(), source.keys()) + for key in source: + self.assertIs(result[key], source[key]) + self.assertTrue(torch.equal(source[key], saved[key])) + self.assertIs(select_state(source, {}, accepted=False), source) + accepted = select_state(source, candidate, accepted=True) + for key in candidate: + self.assertTrue(torch.equal(accepted[key], candidate[key])) + self.assertTrue(torch.equal(source[key], saved[key])) + + def test_candidate_schema_cannot_change_during_acceptance(self): + source = {"qweight": torch.zeros(4, 4, dtype=torch.int8), "proj_up": torch.zeros(4, 2, dtype=torch.bfloat16)} + cases = [ + {"qweight": source["qweight"]}, + {**source, "extra": torch.ones(1)}, + {**source, "qweight": torch.zeros(4, 3, dtype=torch.int8)}, + {**source, "proj_up": torch.zeros(4, 2, dtype=torch.float32)}, + {**source, "proj_up": torch.full((4, 2), float("nan"), dtype=torch.bfloat16)}, + ] + for candidate in cases: + with self.subTest(keys=list(candidate)): + with self.assertRaises(ValueError): + select_state(source, candidate, accepted=True) + + def test_output_directory_cannot_overlap_source_even_through_symlinks(self): + with tempfile.TemporaryDirectory() as directory: + root = Path(directory) + source = root / "model" + source.mkdir() + child = source / "new-output" + for output in (source, root, child, source / ".." / "model"): + with self.subTest(output=output): + with self.assertRaises(ValueError): + assert_separate_output(source, output) + alias = root / "alias" + alias.symlink_to(source, target_is_directory=True) + with self.assertRaises(ValueError): + assert_separate_output(source, alias) + with self.assertRaises(ValueError): + assert_separate_output(source, alias / "child") + assert_separate_output(source, root / "model-refined") + + def test_recipe_is_fixed_and_preserves_source_provenance(self): + source = source_stats() + saved = copy.deepcopy(source) + recipe = fixed_recipe(source) + self.assertEqual(tuple(recipe["ranks"]), (128,)) + self.assertEqual(tuple(recipe["alphas"]), (0.75,)) + self.assertEqual(tuple(recipe["smoothing_families"]), ("activation_only",)) + self.assertEqual(tuple(recipe["weighting"]), ("rms",)) + self.assertEqual(recipe["fixed_smoothing"], "activation_only") + self.assertEqual(recipe["factorization"], "up_singular") + self.assertEqual(recipe["seed"], 1953) + self.assertEqual(recipe["niter"], 4) + self.assertEqual(recipe["oversample"], 16) + self.assertEqual(recipe["iterations"], 100) + self.assertEqual(source, saved) + + def test_recipe_does_not_use_heldout_metrics(self): + source = source_stats() + recipe = fixed_recipe(source) + source["heldout"] = {"mse": float("nan"), "rank": 16, "seed": 0, + "selected": {"family": "identity"}} + self.assertEqual(recipe, fixed_recipe(source)) + + def test_selected_factorization_is_sufficient_without_search_fallback(self): + source = source_stats() + source["search"].pop("factorization") + self.assertEqual(fixed_recipe(source)["factorization"], "up_singular") + + def test_invalid_selected_recipe_fails_before_gpu_execution(self): + cases = [("selected", "rank", 127), ("selected", "alpha", float("nan")), + ("selected", "alpha", 1.1), ("selected", "weighting", "unknown"), + ("selected", "factorization", "unknown"), ("search", "niter", -1), + ("search", "seed", 1.2), ("search", "ridge", 0), + ("search", "gptq_damp", float("inf")), ("search", "final_gptq", "false")] + for section, key, value in cases: + source = source_stats() + source[section][key] = value + with self.subTest(section=section, key=key), self.assertRaises(ValueError): + fixed_recipe(source) + + def test_provenance_requires_exact_source_inputs(self): + identity = {"activation_fingerprint": "activation-sha", "source_config_sha256": "config-sha", + "baseline_calibration_sha256": "baseline-sha"} + manifest = {"conversion_identity": identity, "heldout": {"mse": float("nan")}} + require_matching_provenance(manifest, "activation-sha", "config-sha", "baseline-sha") + for values in (("other", "config-sha", "baseline-sha"), + ("activation-sha", "other", "baseline-sha"), + ("activation-sha", "config-sha", "other")): + with self.assertRaises(ValueError): + require_matching_provenance(manifest, *values) + + def test_retained_safetensors_shard_is_byte_identical(self): + from safetensors import safe_open + from safetensors.torch import load_file, save_file + name = "transformer_blocks.0.attn.to_q" + state = {"qweight": torch.arange(16, dtype=torch.int8).reshape(4, 4), + "proj_up": torch.arange(8).reshape(4, 2).bfloat16()} + with tempfile.TemporaryDirectory() as directory: + root = Path(directory) + source, output = root / "source", root / "output" + source.mkdir() + output.mkdir() + path = source / "original.safetensors" + prefixed = {name + "." + key: value for key, value in state.items()} + save_file(prefixed, str(path), metadata={"provenance": "keep this exact original header"}) + original_bytes = path.read_bytes() + index = {key: path for key in prefixed} + info = write_layer_shard(output, "retained.safetensors", name, state, index, retain_source=True) + saved = output / "retained.safetensors" + self.assertTrue(info["source_file_copied_verbatim"]) + self.assertEqual(saved.read_bytes(), original_bytes) + self.assertEqual(path.read_bytes(), original_bytes) + self.assertEqual(info["sha256"], hashlib.sha256(original_bytes).hexdigest()) + with safe_open(str(saved), framework="pt", device="cpu") as handle: + self.assertEqual(handle.metadata()["provenance"], "keep this exact original header") + changed = {key: value + 1 for key, value in state.items()} + info = write_layer_shard(output, "candidate.safetensors", name, changed, index, retain_source=False) + self.assertFalse(info["source_file_copied_verbatim"]) + self.assertEqual(path.read_bytes(), original_bytes) + actual = load_file(str(output / "candidate.safetensors")) + for key, value in changed.items(): + self.assertTrue(torch.equal(actual[name + "." + key], value)) + + def test_shared_source_shard_does_not_copy_unrelated_layers(self): + from safetensors.torch import load_file, save_file + name = "transformer_blocks.0.attn.to_q" + other = "transformer_blocks.0.attn.to_k" + state = {"qweight": torch.ones(4, 4, dtype=torch.int8)} + with tempfile.TemporaryDirectory() as directory: + root = Path(directory) + output = root / "output" + output.mkdir() + path = root / "shared.safetensors" + payload = {name + ".qweight": state["qweight"], other + ".qweight": state["qweight"] * 2} + save_file(payload, str(path)) + original = path.read_bytes() + index = {key: path for key in payload} + info = write_layer_shard(output, "layer.safetensors", name, state, index, retain_source=True) + self.assertFalse(info["source_file_copied_verbatim"]) + self.assertEqual(path.read_bytes(), original) + actual = load_file(str(output / "layer.safetensors")) + self.assertEqual(set(actual), {name + ".qweight"}) + self.assertTrue(torch.equal(actual[name + ".qweight"], state["qweight"])) + + def test_dry_run_validates_recipes_without_gpu_or_output_writes(self): + from . import FORMAT + from . import refine_checkpoint_v3 as refinement + suffixes = ("attn.to_q", "attn.to_k", "attn.to_v", "attn.to_out.0", + "img_mlp.proj", "img_mlp.out", "img_mlp.gate_layer") + names = [f"transformer_blocks.{block}.{suffix}" for block in range(32) for suffix in suffixes] + class Reader: + fingerprint = "activation-sha" + def require(self, requested): + if set(requested) != set(names): + raise AssertionError("Dry run changed the expected layer set") + with tempfile.TemporaryDirectory() as directory: + root = Path(directory) + source, model, output = root / "source", root / "bf16", root / "new-output" + source.mkdir() + model.mkdir() + config = b'{"_class_name":"QwenImage21Transformer2DModel","num_layers":32}' + (model / "config.json").write_bytes(config) + (source / "config.json").write_bytes(config) + baseline = root / "calibration.safetensors" + baseline.write_bytes(b"Only hashed in dry run; tensor reads must never occur") + manifest = { + "backend_format": FORMAT, "complete": True, + "conversion_identity": { + "activation_fingerprint": "activation-sha", + "source_config_sha256": hashlib.sha256(config).hexdigest(), + "baseline_calibration_sha256": hashlib.sha256(baseline.read_bytes()).hexdigest(), + }, + "layers": {name: {"in_features": 128, "out_features": 128, "rank": 128, + "bias": False, "precision": "int4"} for name in names}, + "files": {"not-read-in-dry-run.safetensors": {"layer_stats": {name: source_stats() for name in names}}}, + } + (source / "manifest.json").write_text(json.dumps(manifest)) + (source / "model.safetensors.index.json").write_text(json.dumps({"weight_map": {}})) + before = {p.name: p.read_bytes() for p in source.iterdir()} + argv = ["refine_checkpoint_v3", "--source-checkpoint", str(source), + "--model-path", str(model), "--activations", str(root / "mock-activations"), + "--baseline-calibration", str(baseline), "--out", str(output), + "--device", "cuda:0", "--threads", "2", "--dry-run"] + captured = io.StringIO() + with patch("sys.argv", argv), patch.object(refinement, "ActivationReader", return_value=Reader()), \ + patch.object(refinement, "_actual_forward", side_effect=AssertionError("Unexpected kernel call")), \ + patch("torch.cuda._lazy_init", side_effect=AssertionError("Unexpected CUDA initialization")), \ + contextlib.redirect_stdout(captured): + refinement.main() + report = json.loads(captured.getvalue()) + self.assertTrue(report["dry_run"]) + self.assertFalse(report["gpu_work_executed"]) + self.assertEqual(report["replayed_layers"], 224) + self.assertEqual(report["iterations"], 100) + self.assertFalse(output.exists()) + self.assertEqual(before, {p.name: p.read_bytes() for p in source.iterdir()}) + + def test_baseline_recipe_requires_explicit_retention(self): + source = source_stats() + source["selected"]["family"] = "one_pass_baseline" + with self.assertRaises(NotImplementedError): + fixed_recipe(source) + + +if __name__ == "__main__": + unittest.main(verbosity=2) diff --git a/reproduction/nunchaku_backend/refinement100_ready.md b/reproduction/nunchaku_backend/refinement100_ready.md new file mode 100644 index 0000000000000000000000000000000000000000..ab30034b49805789c3ddf369e14ed6db8ef30ae5 --- /dev/null +++ b/reproduction/nunchaku_backend/refinement100_ready.md @@ -0,0 +1,27 @@ +# Full selected-recipe 100-iteration refinement — prepared, not launched; deferred after probe + +The completed [three-layer probe](iterations100_probe.md) did not show meaningful held-out improvement. The non-cap control was unchanged; attention-output validation MSE improved 0.548% while held-out MSE worsened 3.871%; Q validation and held-out MSE worsened 0.286% and 0.120%. Based on this evidence, the main agent decided not to run full refinement. This tooling remains available and CPU-validated, but no full refinement checkpoint has been produced and its GPU execution path remains untested. + +`refine_checkpoint_v3.py` replays only each immutable V3 layer's selected smoothing family/alpha, rank, weighting, factorization and seed, with a larger iteration limit. It does not reopen the broad smoothing/rank grid. Original early stopping remains active. This tooling is available if the three-layer probe justifies a full refinement; it does not queue or start one automatically. + +For each layer it loads the **actual packed V3 tensors** and evaluates them with Nunchaku against the same BF16 linear teacher and validation inputs. It independently evaluates the proposed packed tensors and adopts them only for strict validation improvement. An optional `--min-validation-improvement .001` requires at least0.1% lower validation MSE. Ties, regressions, unchanged tensor sets, incompatible recipes, and original one-pass-baseline selections retain V3. Heldout tensors are read only after the decision is frozen, and both source/candidate heldout results are reported. + +In the ordinary one-layer-per-shard V3 layout, fallback copies the source safetensors file **byte-for-byte**, including its header/metadata. BF16 boundary tensors are unchanged. The output is a separate checkpoint directory with the usual224-layer manifest and225 weight shards, compatible with the existing runtime. Extra per-layer JSON reports record proposed candidates and decisions. The source config, index, manifest, and all source weight shards are hashed and rechecked for immutability. Resuming checks output shard/report hashes and identical source/data/code/settings. + +GPU command, only if the main agent decides to proceed after its probe: + +```sh +python -m nunchaku_backend.refine_checkpoint_v3 \ + --source-checkpoint /cache/qwen-nunchaku-v3-r128 \ + --model-path /cache/huggingface/hub/models--Qwen--Qwen-Image-2.1/snapshots/b3179ad355be050328e483a9dfdd9e60cd62adfa \ + --activations /cache/qwen21-activation-v3-40 \ + --baseline-calibration /cache/qwen21-calibration \ + --out /cache/qwen-nunchaku-v3-refined100 \ + --device cuda:0 --iterations 100 +``` + +For initial GPU verification, append `--max-layers 1`; this writes an intentionally incomplete checkpoint. Re-run without that limit and with `--resume` to continue. The new directory must remain separate from the immutable source; same-directory, ancestor/descendant, and symlink-equivalent paths are rejected. This mode is independently prepared and has not yet received a GPU execution test. + +Append `--dry-run` for CPU-only recipe/provenance validation. This still accepts the planned `--device cuda:0` but never initializes CUDA or creates an output checkpoint. The actual host validation used `NVIDIA_VISIBLE_DEVICES=void` and empty `CUDA_VISIBLE_DEVICES`; it confirmed224 replayable recipes, zero unsupported recipes, and `cuda_initialized:false` against the source manifest SHA256 `cf69e83a646981d22ee6df8d5239b46a50df25d8eb73c9f0478feae87323e6cb`. + +Final target-image CPU verification passed all 30 tests in 0.389 seconds, followed by a successful dry-run against all 224 actual source recipes with CUDA uninitialized. CPU tests cover strict validation-only selection, heldout independence, exact tensor and file fallback, candidate schema consistency, malformed recipes, source/data provenance, isolated output paths, and a dry-run GPU-initialization tripwire. These tests do not establish image quality. If a candidate wins validation but worsens heldout or visual quality, that remains visible in the report and needs review before adoption. diff --git a/reproduction/nunchaku_backend/role_sweep_v3.py b/reproduction/nunchaku_backend/role_sweep_v3.py new file mode 100644 index 0000000000000000000000000000000000000000..c7e28279d7de2572d842102df2b2c15940fb75ce --- /dev/null +++ b/reproduction/nunchaku_backend/role_sweep_v3.py @@ -0,0 +1,273 @@ +"""Serial BF16-restoration sensitivity sweep on a fixed denoiser capture. + +Run only after the root agent releases the POC GPU:: + + python -m nunchaku_backend.role_sweep_v3 \ + --checkpoint /cache/qwen21-nunchaku-v3 \ + --source /cache/original-snapshot \ + --capture /cache/denoiser-probe-v3 \ + --out results/role-sweep-v3 --include-qk --include-qk-gate + +By default this measures an unmodified baseline plus each of seven projection +roles restored independently in all 32 blocks. Optional combined variants are +explicit. Each variant gets a fresh subprocess/model/CUDA context; the existing +denoiser comparison independently rebuilds its own prefix KV for every step. + +This is sensitivity on the development mug capture, not the expanded image +evaluation set and not proof of general image-quality improvement. No model +checkpoint, production service, or running GPU process is changed by this +driver. Import and --help perform no GPU work. +""" +from __future__ import annotations + +import argparse +import hashlib +import json +import math +from pathlib import Path +import re +import subprocess +import sys +import time + + +ROLES = ( + ("q", "attn.to_q"), ("k", "attn.to_k"), ("v", "attn.to_v"), + ("attention-output", "attn.to_out.0"), ("mlp-projection", "img_mlp.proj"), + ("mlp-gate", "img_mlp.gate_layer"), ("mlp-output", "img_mlp.out"), +) + + +def _write(path, data): + path = Path(path) + temporary = path.with_suffix(path.suffix + ".tmp") + temporary.write_text(json.dumps(data, indent=2, sort_keys=True) + "\n") + temporary.replace(path) + + +def _hash(path): + digest = hashlib.sha256() + with Path(path).open("rb") as handle: + for chunk in iter(lambda: handle.read(1024 * 1024), b""): + digest.update(chunk) + return digest.hexdigest() + + +def _variants(include_qk=False, include_qk_gate=False, combinations=()): + variants = [{"name": "baseline", "roles": []}] + variants.extend({"name": name, "roles": [role]} for name, role in ROLES) + if include_qk: + variants.append({"name": "qk", "roles": ["attn.to_q", "attn.to_k"]}) + if include_qk_gate: + variants.append({"name": "qk-gate", "roles": ["attn.to_q", "attn.to_k", "img_mlp.gate_layer"]}) + valid_roles = {role for _, role in ROLES} + for specification in combinations: + name, separator, value = specification.partition("=") + roles = value.split(",") + if not separator or not re.fullmatch(r"[a-z0-9][a-z0-9-]*", name): + raise ValueError("Combination must be lowercase-name=role,role") + if any(role not in valid_roles for role in roles) or len(set(roles)) != len(roles): + raise ValueError(f"Invalid or duplicate combination roles: {roles}") + if any(variant["name"] == name for variant in variants): + raise ValueError(f"Duplicate variant name: {name}") + variants.append({"name": name, "roles": roles}) + return variants + + +def _backend_metrics(report, name, capture_sha, checkpoint_sha): + if report.get("complete") is not True or report.get("capture_manifest_sha256") != capture_sha: + raise ValueError("Incomplete comparison or mismatched teacher capture") + matches = [backend for backend in report.get("backends", []) if backend.get("name") == name] + if len(matches) != 1: + raise ValueError(f"Expected one comparison backend named {name}") + backend = matches[0] + if backend.get("manifest_sha256") != checkpoint_sha: + raise ValueError("Compared checkpoint manifest changed") + probes = sorted(backend["probes"], key=lambda row: row["step_zero_based"]) + if len(probes) != 3 or probes[0]["step_zero_based"] != 0: + raise ValueError("Expected first/middle/late probe results") + errors, teacher_energy = [], 0.0 + for probe in probes: + metric = probe["metrics"] + if metric.get("finite") is not True: + raise ValueError("Non-finite prediction probe") + relative, norm = metric["relative_l2"], metric["teacher_l2_norm"] + if not math.isfinite(relative) or not math.isfinite(norm) or norm <= 0: + raise ValueError("Invalid relative error or teacher prediction norm") + errors.append(relative) + teacher_energy += norm * norm + weighted_error = sum((probe["metrics"]["relative_l2"] * probe["metrics"]["teacher_l2_norm"]) ** 2 + for probe in probes) + return { + "steps": [{"step_zero_based": probe["step_zero_based"], "timestep": probe["timestep"], + **probe["metrics"]} for probe in probes], + "mean_relative_l2": sum(errors) / len(errors), + "pooled_relative_l2": math.sqrt(weighted_error / teacher_energy), + "worst_relative_l2": max(errors), + "late_relative_l2": errors[-1], + "minimum_cosine_similarity": min(probe["metrics"]["cosine_similarity"] for probe in probes), + "peak_allocated_mib": max(probe["peak_allocated_mib"] for probe in probes), + "peak_reserved_mib": max(probe["peak_reserved_mib"] for probe in probes), + "hybrid_override": backend.get("hybrid_override"), + } + + +def _summary(run): + successful = [row for row in run["results"] if row.get("status") == "completed"] + baseline = next((row for row in successful if row["name"] == "baseline"), None) + rows = [] + for result in successful: + metric = result["metrics"] + row = {"name": result["name"], "roles": result["roles"], "report": result["report"], **metric} + if baseline: + base = baseline["metrics"] + if [step["step_zero_based"] for step in metric["steps"]] != [step["step_zero_based"] for step in base["steps"]]: + raise ValueError("Variant/baseline captured timesteps differ") + row["mean_relative_l2_ratio_to_baseline"] = metric["mean_relative_l2"] / max(base["mean_relative_l2"], 1e-30) + row["pooled_relative_l2_ratio_to_baseline"] = metric["pooled_relative_l2"] / max(base["pooled_relative_l2"], 1e-30) + row["per_step_ratio_to_baseline"] = [ + {"step_zero_based": step["step_zero_based"], + "ratio": step["relative_l2"] / max(ref["relative_l2"], 1e-30)} + for step, ref in zip(metric["steps"], base["steps"], strict=True) + ] + rows.append(row) + winners = [] + if rows: + for index, step in enumerate(rows[0]["steps"]): + best = min(rows, key=lambda row: row["steps"][index]["relative_l2"]) + winners.append({"step_zero_based": step["step_zero_based"], "name": best["name"], + "relative_l2": best["steps"][index]["relative_l2"]}) + return { + "format": "qwen21-role-sensitivity-summary-v3", "complete": run["complete"], + "identity": run["identity"], "job": run["job"], "successful_variants": len(rows), + "failed_variants": [{"name": row["name"], "error": row.get("error"), "log": row.get("log")} + for row in run["results"] if row.get("status") == "failed"], + "variants_by_mean_relative_l2": sorted(rows, key=lambda row: row["mean_relative_l2"]), + "ranking_by_pooled_relative_l2": [row["name"] for row in sorted(rows, key=lambda row: row["pooled_relative_l2"])], + "per_step_lowest_relative_l2": winners, + "interpretation": "Development-capture sensitivity only. Rankings do not select a production backend or establish image quality.", + "aggregate_definitions": { + "mean_relative_l2": "Unweighted mean of first/middle/late relative prediction errors", + "pooled_relative_l2": "sqrt(sum prediction-error L2 squared / sum teacher-prediction L2 squared)", + }, + } + + +def main(): + parser = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter) + parser.add_argument("--checkpoint", required=True) + parser.add_argument("--source", required=True) + parser.add_argument("--capture", required=True) + parser.add_argument("--out", required=True) + parser.add_argument("--include-qk", action="store_true") + parser.add_argument("--include-qk-gate", action="store_true") + parser.add_argument("--combo", action="append", default=[], + help="Additional role combination, e.g. v-output=attn.to_v,attn.to_out.0") + parser.add_argument("--resume", action="store_true") + parser.add_argument("--stop-on-error", action="store_true") + parser.add_argument("--baseline-report", help="Reuse a completed comparison on this exact checkpoint and teacher capture") + parser.add_argument("--baseline-name", default="v3", help="Backend name inside --baseline-report") + args = parser.parse_args() + checkpoint, source, capture, output = (Path(value).resolve() for value in (args.checkpoint, args.source, args.capture, args.out)) + source_config = source / "transformer" / "config.json" if (source / "transformer").is_dir() else source / "config.json" + capture_data = json.loads((capture / "capture.json").read_text()) + if not capture_data.get("complete") or capture_data.get("format") != "qwen21-denoiser-teacher-capture-v3": + raise ValueError("Need a complete bounded BF16 teacher capture") + variants = _variants(args.include_qk, args.include_qk_gate, args.combo) + identity = { + "checkpoint": str(checkpoint), "checkpoint_manifest_sha256": _hash(checkpoint / "manifest.json"), + "source": str(source), "source_config_sha256": _hash(source_config), + "capture": str(capture), "capture_manifest_sha256": _hash(capture / "capture.json"), + "variants": variants, + } + output.mkdir(parents=True, exist_ok=True) + manifest = output / "run.json" + if args.resume: + run = json.loads(manifest.read_text()) + if run["identity"] != identity: + raise ValueError("Resume source/checkpoint/capture/variants differ from original sweep") + else: + if any(output.iterdir()): + raise ValueError("Use a new output directory or --resume") + run = {"format": "qwen21-role-sweep-v3", "complete": False, "identity": identity, + "job": capture_data["job"], "results": [], "started_unix": time.time()} + run["complete"] = False + _write(manifest, run) + for variant in variants: + previous = next((row for row in run["results"] if row["name"] == variant["name"]), None) + if previous and previous.get("status") == "completed": + # Verify persisted results again rather than trusting a stale resume flag. + report = json.loads(Path(previous["report"]).read_text()) + _backend_metrics(report, previous.get("backend_name", variant["name"]), + identity["capture_manifest_sha256"], identity["checkpoint_manifest_sha256"]) + continue + run["results"] = [row for row in run["results"] if row["name"] != variant["name"]] + report_path, log_path = output / f"{variant['name']}.json", output / f"{variant['name']}.log" + row = {**variant, "report": str(report_path), "log": str(log_path), "status": "running", + "backend_name": variant["name"], "started_unix": time.time()} + run["results"].append(row) + _write(manifest, run) + print(json.dumps({"event": "role_sweep_variant_start", **variant}), flush=True) + try: + if variant["name"] == "baseline" and args.baseline_report: + report = json.loads(Path(args.baseline_report).read_text()) + row["backend_name"] = args.baseline_name + row["reused_report"] = str(Path(args.baseline_report).resolve()) + _write(report_path, report) + log_path.write_text("Reused explicitly supplied complete baseline report; no GPU call.\n") + else: + command = [sys.executable, "-m", "nunchaku_backend.denoiser_probe_v3", "compare", + "--capture", str(capture), "--checkpoint", f"{variant['name']}={checkpoint}", + "--out", str(report_path)] + if variant["roles"]: + command.extend(["--bf16-source", str(source)]) + for role in variant["roles"]: + command.extend(["--restore-role", role]) + row["command"] = command + _write(manifest, run) + with log_path.open("w") as log: + process = subprocess.run(command, cwd=Path(__file__).resolve().parents[1], + stdout=log, stderr=subprocess.STDOUT, check=False) + row["exit_code"] = process.returncode + if process.returncode: + # Preserve full logs; limit inline errors and continue so one + # too-large optional hybrid cannot hide the seven core roles. + tail = log_path.read_text(errors="replace")[-5000:] + raise RuntimeError(f"Variant process exited {process.returncode}: {tail}") + report = json.loads(report_path.read_text()) + metric = _backend_metrics(report, row["backend_name"], identity["capture_manifest_sha256"], + identity["checkpoint_manifest_sha256"]) + hybrid = metric["hybrid_override"] + if variant["roles"]: + if not hybrid or sorted(hybrid["roles"]) != sorted(variant["roles"]): + raise ValueError("Hybrid result did not restore the requested roles") + if hybrid["source_config_sha256"] != identity["source_config_sha256"]: + raise ValueError("Hybrid BF16 source differs from sweep source") + elif hybrid: + raise ValueError("Baseline must contain no BF16 hybrid overrides") + row.update(status="completed", metrics=metric) + print(json.dumps({"event": "role_sweep_variant_result", "name": variant["name"], + "mean_relative_l2": metric["mean_relative_l2"], "late_relative_l2": metric["late_relative_l2"], + "peak_allocated_mib": metric["peak_allocated_mib"]}), flush=True) + except Exception as error: + row.update(status="failed", error=f"{type(error).__name__}: {error}") + print(json.dumps({"event": "role_sweep_variant_failure", "name": variant["name"], "error": row["error"]}), flush=True) + if args.stop_on_error: + _write(manifest, run) + _write(output / "summary.json", _summary(run)) + raise + finally: + row["finished_unix"] = time.time() + _write(manifest, run) + _write(output / "summary.json", _summary(run)) + run["complete"] = True # All variants attempted; failures remain explicit. + run["finished_unix"] = time.time() + _write(manifest, run) + _write(output / "summary.json", _summary(run)) + print(json.dumps({"event": "role_sweep_finished", "summary": str(output / "summary.json"), + "completed": sum(row["status"] == "completed" for row in run["results"]), + "failed": sum(row["status"] == "failed" for row in run["results"])}), flush=True) + + +if __name__ == "__main__": + main() diff --git a/reproduction/nunchaku_backend/runner_hybrid_test.py b/reproduction/nunchaku_backend/runner_hybrid_test.py new file mode 100644 index 0000000000000000000000000000000000000000..d20f12157d16f2cb7464b0dc63faff22b6e49489 --- /dev/null +++ b/reproduction/nunchaku_backend/runner_hybrid_test.py @@ -0,0 +1,166 @@ +"""CPU-only configuration and initialization-order tests for Engine hybrids.""" +import argparse +import builtins +from contextlib import redirect_stdout +import io +import os +import sys +import types +import unittest +from unittest.mock import patch + +import runner + + +def restored_report(): + return { + "roles": ["img_mlp.proj"], "restored_names": [f"transformer_blocks.{i}.img_mlp.proj" for i in range(32)], + "restored_count": 32, "source_transformer": "/original/transformer", "source_config_sha256": "source-hash", + "bf16_state_bytes": 8192, "replaced_packed_state_bytes": 2048, "net_state_bytes_delta": 6144, + } + + +class RunnerHybridTests(unittest.TestCase): + def setUp(self): + self.environment = patch.dict(os.environ, {}, clear=True) + self.environment.start() + self.addCleanup(self.environment.stop) + + def test_defaults_leave_original_behavior(self): + args = runner.default_args() + self.assertEqual(args.backend, "nf4") + self.assertEqual(args.restore_roles, []) + self.assertIsNone(args.bf16_source) + self.assertTrue(args.offload and args.cache and args.stage_offload and args.release_kv) + self.assertFalse(args.compile or args.flex or args.tiling or args.bf16_transformer) + self.assertFalse(runner._validate_hybrid_args(args)) + with patch("nunchaku_backend.hybrid_v3.restore_bf16_projections") as restore: + self.assertIsNone(runner._apply_hybrid_override(object(), args)) + restore.assert_not_called() + + def test_legacy_namespace_without_new_fields_stays_disabled(self): + args = argparse.Namespace(backend="nunchaku") + self.assertFalse(runner._validate_hybrid_args(args)) + self.assertIsNone(runner._apply_hybrid_override(object(), args)) + + def test_environment_enables_api_engine_configuration(self): + with patch.dict(os.environ, {"QWEN_BACKEND": "nunchaku", "QWEN_BF16_SOURCE": "/original", + "QWEN_BF16_ROLES": " img_mlp.proj, img_mlp.out,img_mlp.proj "}): + args = runner.default_args() + self.assertTrue(runner._validate_hybrid_args(args)) + self.assertEqual(args.bf16_source, "/original") + self.assertEqual(args.restore_roles, ["img_mlp.out", "img_mlp.proj"]) + + def test_source_required_and_source_without_roles_rejected(self): + for source, roles in [(None, ["attn.to_q"]), ("/original", [])]: + with self.subTest(source=source, roles=roles): + args = argparse.Namespace(backend="nunchaku", bf16_source=source, restore_roles=roles) + with self.assertRaisesRegex(ValueError, "requires both"): + runner._validate_hybrid_args(args) + + def test_invalid_role_backend_and_save_combination_rejected(self): + cases = [ + argparse.Namespace(backend="nunchaku", bf16_source="/original", restore_roles=["img_mlp.typo"]), + argparse.Namespace(backend="nf4", bf16_source="/original", restore_roles=["img_mlp.proj"]), + argparse.Namespace(backend="nunchaku", bf16_source="/original", restore_roles=["img_mlp.proj"], save_prequant="/save"), + argparse.Namespace(backend="nunchaku", bf16_source="/original", restore_roles="img_mlp.proj"), + ] + for args in cases: + with self.subTest(args=args), self.assertRaises(ValueError): + runner._validate_hybrid_args(args) + + def test_engine_configuration_guard_precedes_torch_import(self): + args = runner.default_args() + args.backend = "nunchaku" + args.restore_roles = ["img_mlp.proj"] + real_import = builtins.__import__ + imports = [] + def guarded(name, *positional, **kwargs): + imports.append(name) + if name == "torch": + raise AssertionError("Torch imported before hybrid configuration validation") + return real_import(name, *positional, **kwargs) + with patch("builtins.__import__", side_effect=guarded): + with self.assertRaisesRegex(ValueError, "requires both"): + runner.Engine(args) + self.assertNotIn("torch", imports) + + def test_repeated_cli_roles_override_environment_roles(self): + with patch.dict(os.environ, {"QWEN_BF16_ROLES": "attn.to_q", "QWEN_BF16_SOURCE": "/env-original"}): + args = runner._parse_cli_args(["--backend", "nunchaku", "--restore-source", "/cli-original", + "--restore-role", "img_mlp.proj", "--restore-role", "img_mlp.out"]) + self.assertTrue(runner._validate_hybrid_args(args)) + self.assertEqual(args.bf16_source, "/cli-original") + self.assertEqual(args.restore_roles, ["img_mlp.out", "img_mlp.proj"]) + + def test_cli_env_fallback_and_default_unchanged(self): + args = runner._parse_cli_args([]) + self.assertEqual(args.backend, "nf4") + self.assertEqual(args.restore_roles, []) + self.assertIsNone(args.bf16_source) + with patch.dict(os.environ, {"QWEN_BF16_ROLES": "img_mlp.proj", "QWEN_BF16_SOURCE": "/env-original"}): + args = runner._parse_cli_args(["--backend", "nunchaku"]) + self.assertTrue(runner._validate_hybrid_args(args)) + self.assertEqual(args.restore_roles, ["img_mlp.proj"]) + + def test_metadata_describes_actual_restoration(self): + args = argparse.Namespace(restore_roles=["img_mlp.proj"], bf16_source="/original") + transformer = object() + report = restored_report() + with patch("nunchaku_backend.hybrid_v3.restore_bf16_projections", return_value=report) as restore: + self.assertIs(runner._apply_hybrid_override(transformer, args), report) + restore.assert_called_once_with(transformer, "/original", roles=["img_mlp.proj"]) + self.assertEqual(args.hybrid_roles, report["roles"]) + self.assertEqual(args.hybrid_names, report["restored_names"]) + self.assertEqual(args.hybrid_restored_count, 32) + self.assertEqual(args.hybrid_bf16_state_bytes, 8192) + self.assertEqual(args.hybrid_replaced_packed_state_bytes, 2048) + self.assertEqual(args.hybrid_state_bytes_delta, 6144) + + def test_restoration_happens_before_pipeline_and_offload_hooks(self): + events = [] + class ReachedPipeline(Exception): + pass + fake_torch = types.ModuleType("torch") + fake_torch.bfloat16 = "bf16" + fake_torch.set_num_threads = lambda count: None + fake_torch.cuda = types.SimpleNamespace(empty_cache=lambda: None) + encoder = types.SimpleNamespace( + config=types.SimpleNamespace(text_config=types.SimpleNamespace()), + get_memory_footprint=lambda: 0, to=lambda device: None, modules=lambda: [], + model=types.SimpleNamespace(visual=types.SimpleNamespace(modules=lambda: [])), + ) + transformer = types.SimpleNamespace(modules=lambda: []) + def load(*args, **kwargs): + self.assertEqual(kwargs["device"], "cpu") + events.append("load_cpu") + return transformer + def restore(model, source, roles): + self.assertIs(model, transformer) + events.append("restore_cpu") + return restored_report() + def pipeline(*args, **kwargs): + events.append("pipeline") + self.assertEqual(events, ["load_cpu", "restore_cpu", "pipeline"]) + raise ReachedPipeline() + fake_diffusers = types.ModuleType("diffusers") + fake_diffusers.QwenImage21Pipeline = types.SimpleNamespace(from_pretrained=pipeline) + fake_diffusers.QwenImage21Transformer2DModel = object + fake_diffusers.BitsAndBytesConfig = lambda **kwargs: kwargs + fake_transformers = types.ModuleType("transformers") + fake_transformers.Qwen3VLForConditionalGeneration = types.SimpleNamespace(from_pretrained=lambda *args, **kwargs: encoder) + fake_transformers.BitsAndBytesConfig = lambda **kwargs: kwargs + args = runner.default_args() + args.backend, args.nunchaku_checkpoint = "nunchaku", "/packed" + args.bf16_source, args.restore_roles = "/original", ["img_mlp.proj"] + with patch.dict(sys.modules, {"torch": fake_torch, "diffusers": fake_diffusers, "transformers": fake_transformers}), \ + patch("nunchaku_backend.runtime.load_transformer", side_effect=load), \ + patch("nunchaku_backend.hybrid_v3.restore_bf16_projections", side_effect=restore), \ + redirect_stdout(io.StringIO()): + with self.assertRaises(ReachedPipeline): + runner.Engine(args) + self.assertEqual(events, ["load_cpu", "restore_cpu", "pipeline"]) + + +if __name__ == "__main__": + unittest.main() diff --git a/reproduction/nunchaku_backend/runtime.py b/reproduction/nunchaku_backend/runtime.py new file mode 100644 index 0000000000000000000000000000000000000000..64a5c08d5cfd1b6e90ed227e1b0d9acee86b682a --- /dev/null +++ b/reproduction/nunchaku_backend/runtime.py @@ -0,0 +1,130 @@ +"""Load custom Qwen2.1 checkpoints using upstream generic Nunchaku linears. + +The original Diffusers model still owns positional encoding, shared modulation, +attention, prefix-KV caching, timestep handling and output projection. Only the +224 block projection matrices use SVDQW4A4Linear. This is not a port of the older +Qwen-Image architecture or an NF4 wrapper. +""" +import json +from pathlib import Path + +from . import FORMAT +from .layout import BLOCK_LINEAR_PATTERN + + +def probe_backend(): + """Import-only probe, safe in a container with NVIDIA_VISIBLE_DEVICES=void.""" + import inspect + import torch + import nunchaku + from importlib.metadata import version, PackageNotFoundError + from nunchaku.models.linear import SVDQW4A4Linear + parameters = inspect.signature(SVDQW4A4Linear.__init__).parameters + required = {"in_features", "out_features", "rank", "precision", "torch_dtype", "device"} + if not required.issubset(parameters): + raise RuntimeError(f"Unsupported Nunchaku SVDQW4A4Linear signature: {list(parameters)}") + try: + nunchaku_version = version("nunchaku") + except PackageNotFoundError: + nunchaku_version = getattr(nunchaku, "__version__", "unknown") + return {"torch": torch.__version__, "nunchaku": nunchaku_version, + "linear_class": SVDQW4A4Linear.__module__ + "." + SVDQW4A4Linear.__name__, + "precision": "int4", "weight_bits": 4, "activation_bits": 4} + + +def _read_manifest(directory): + manifest = json.loads((directory / "manifest.json").read_text()) + if manifest.get("backend_format") != FORMAT or manifest.get("complete") is not True: + raise ValueError("Not a compatible Qwen2.1 Nunchaku INT4 checkpoint") + layers = manifest.get("layers", {}) + if len(layers) != 224 or any(not BLOCK_LINEAR_PATTERN.fullmatch(n) for n in layers): + raise ValueError("Checkpoint must describe exactly224 Qwen2.1 block linears") + for name, info in layers.items(): + if info["in_features"] % 128 or info["out_features"] % 128 or info["rank"] % 16: + raise ValueError(f"Unsupported kernel dimensions: {name}: {info}") + if info.get("precision", "int4") != "int4": + raise ValueError(f"Unsupported precision for4070 Ti SUPER: {name}") + return manifest + + +def _state_files(directory): + index = directory / "model.safetensors.index.json" + if index.exists(): + weight_map = json.loads(index.read_text())["weight_map"] + filenames = sorted(set(weight_map.values())) + else: + filenames = ["model.safetensors"] + for filename in filenames: + path = directory / filename + if path.parent.resolve() != directory.resolve() or not path.is_file(): + raise ValueError(f"Invalid or missing checkpoint shard: {filename}") + yield path + + +def load_transformer(checkpoint, *, device="cpu", torch_dtype=None): + """Return a QwenImage21Transformer2DModel with real Nunchaku W4A4 blocks. + + Construction and checkpoint load happen on CPU; moving to CUDA is explicit. + `device='cuda:0'` means the GPU visible as0 in the isolated POC container. + Use normal Diffusers model CPU offload if desired. Load this checkpoint with + this function, not stock from_pretrained; its linears have a custom schema. + """ + import torch + from accelerate import init_empty_weights + from diffusers import QwenImage21Transformer2DModel + from nunchaku.models.linear import SVDQW4A4Linear + from safetensors.torch import load_file + + probe = probe_backend() + directory = Path(checkpoint) + manifest = _read_manifest(directory) + torch_dtype = torch.bfloat16 if torch_dtype is None else torch_dtype + if torch_dtype != torch.bfloat16: + raise ValueError("This initial checkpoint/runtime contract requires BF16 compute") + config = json.loads((directory / "config.json").read_text()) + with init_empty_weights(include_buffers=False): + model = QwenImage21Transformer2DModel.from_config(config) + for name, info in manifest["layers"].items(): + parent_name, child_name = name.rsplit(".", 1) + parent = model.get_submodule(parent_name) + original = parent.get_submodule(child_name) + if (original.in_features, original.out_features) != (info["in_features"], info["out_features"]): + raise ValueError(f"Model/checkpoint dimension mismatch: {name}") + if (original.bias is not None) != info["bias"]: + raise ValueError(f"Model/checkpoint bias mismatch: {name}") + parent._modules[child_name] = SVDQW4A4Linear( + info["in_features"], info["out_features"], rank=info["rank"], + bias=info["bias"], precision="int4", act_unsigned=False, + torch_dtype=torch_dtype, device="meta", + ) + expected = set(model.state_dict()) + seen = set() + for path in _state_files(directory): + state = load_file(str(path), device="cpu") + duplicate = seen.intersection(state) + unexpected = set(state).difference(expected) + if duplicate or unexpected: + raise ValueError(f"Invalid shard {path.name}: duplicate={sorted(duplicate)}, unexpected={sorted(unexpected)}") + model.load_state_dict(state, strict=False, assign=True) + seen.update(state) + del state + missing = expected - seen + if missing: + raise ValueError(f"Incomplete checkpoint; missing {sorted(missing)}") + meta = [n for n, p in model.named_parameters() if p.is_meta] + if meta: + raise ValueError(f"Uninitialized model parameters: {meta}") + for name, module in model.named_modules(): + if BLOCK_LINEAR_PATTERN.fullmatch(name): + if module.qweight.dtype != torch.int8 or module.wscales.dtype != torch_dtype: + raise ValueError(f"Wrong packed tensor dtypes: {name}") + model.eval().requires_grad_(False) + model._qwen21_backend = {"kind": "nunchaku-svdq-w4a4-int4", "checkpoint": str(directory), + "manifest": manifest, "runtime": probe} + # No dtype cast here: preserve integer kernel storage and BF16 scales. + model.to(device=device) + return model + + +if __name__ == "__main__": + print(json.dumps(probe_backend(), indent=2)) diff --git a/reproduction/nunchaku_backend/summarize_v3.py b/reproduction/nunchaku_backend/summarize_v3.py new file mode 100644 index 0000000000000000000000000000000000000000..1b384941afe9b2b98b179a5a13fe2a95a59b20e5 --- /dev/null +++ b/reproduction/nunchaku_backend/summarize_v3.py @@ -0,0 +1,57 @@ +"""Summarize recorded validation/heldout evidence without opening model weights.""" +import argparse +from collections import Counter, defaultdict +import json +from pathlib import Path +import statistics + + +def summarize(manifest): + layers = {name: stats for file in manifest.get("files", {}).values() + for name, stats in file.get("layer_stats", {}).items()} + if not layers: + raise ValueError("No converted-layer metrics in manifest") + groups = defaultdict(list) + for name, stats in layers.items(): + groups[".".join(name.split(".")[2:])].append(stats) + def metrics(rows): + ratios = [r["heldout_mse_ratio_to_baseline"] for r in rows] + return {"layers": len(rows), "heldout_improved": sum(r < 1 for r in ratios), + "heldout_mse_ratio_median": statistics.median(ratios), + "heldout_mse_ratio_min": min(ratios), "heldout_mse_ratio_max": max(ratios), + "heldout_relative_l2_median": statistics.median(r["heldout"]["relative_l2"] for r in rows), + "baseline_heldout_relative_l2_median": statistics.median(r["baseline"]["heldout"]["relative_l2"] for r in rows), + "a16_weight_proxy_relative_l2_median": statistics.median(r["heldout_weight_only_proxy"]["relative_l2"] for r in rows), + "baseline_a16_weight_proxy_relative_l2_median": statistics.median(r["baseline"]["heldout_weight_only_proxy"]["relative_l2"] for r in rows), + "gptq_accepted": sum(bool(r["selected"].get("gptq")) for r in rows), + "selected_iteration_cap_count": sum(r["selected"]["iteration"] == r["search"]["iterations"] - 1 for r in rows), + "rank_counts": dict(Counter(str(r["selected"]["rank"]) for r in rows)), + "iteration_counts": dict(Counter(str(r["selected"]["iteration"]) for r in rows)), + "smoothing_counts": dict(Counter(f'{r["selected"]["family"]}:{r["selected"]["alpha"]}' for r in rows)), + "output_correction_counts": dict(Counter(str(r["selected"]["output_correction"]) for r in rows)), + "summed_layer_seconds": sum(r["seconds"] for r in rows)} + output = {"complete": manifest.get("complete", False), "conversion_identity": manifest.get("conversion_identity"), + "aggregate": metrics(list(layers.values())), "by_projection": {k: metrics(v) for k, v in sorted(groups.items())}, + "limitations": "Unweighted per-layer metric summaries. A16 proxy also changes accumulator arithmetic and is not an exact attribution of A4 error. The regenerated one-pass rank32 baseline is not the existing rank128 v2 checkpoint. Per-linear error reductions do not establish end-to-end quality or losslessness."} + output["worst_heldout_layers"] = [ + {"name": name, "heldout_relative_l2": row["heldout"]["relative_l2"], + "heldout_mse_ratio": row["heldout_mse_ratio_to_baseline"], "selected": row["selected"]} + for name, row in sorted(layers.items(), key=lambda item: item[1]["heldout"]["relative_l2"], reverse=True)[:20]] + return output + + +def main(): + parser = argparse.ArgumentParser() + parser.add_argument("manifest", type=Path) + parser.add_argument("--out", type=Path) + args = parser.parse_args() + report = summarize(json.loads(args.manifest.read_text())) + text = json.dumps(report, indent=2, sort_keys=True) + "\n" + if args.out: + args.out.parent.mkdir(parents=True, exist_ok=True) + args.out.write_text(text) + print(text) + + +if __name__ == "__main__": + main() diff --git a/reproduction/nunchaku_backend/teacher_v3.py b/reproduction/nunchaku_backend/teacher_v3.py new file mode 100644 index 0000000000000000000000000000000000000000..2b83e66d75cfdbc906900a25fa843fd8d4e40acc --- /dev/null +++ b/reproduction/nunchaku_backend/teacher_v3.py @@ -0,0 +1,66 @@ +"""BF16 DiT teacher with group offload; NF4 encoder held fixed for DiT comparisons.""" +import argparse,json,time +from pathlib import Path + + +def enable_teacher_offload(engine, *, stream=True): + import torch + pipe=engine.pipe + pipe.remove_all_hooks() + pipe._all_hooks=[] + pipe.text_encoder.to('cpu');pipe.vae.to('cpu');pipe.transformer.to('cpu') + torch.cuda.empty_cache() + pipe.transformer.enable_group_offload(onload_device=torch.device('cuda:0'),offload_device=torch.device('cpu'), + offload_type='block_level',num_blocks_per_group=1 if stream else 4,use_stream=stream,record_stream=stream, + exclude_kwargs=['kv_cache']) + # The runner's existing stage-offload wrappers unload these components. + original_prompt=pipe._get_qwen_prompt_embeds + def prompt(*args,**kwargs): + pipe.text_encoder.to('cuda:0') + return original_prompt(*args,**kwargs) + pipe._get_qwen_prompt_embeds=prompt + original_encode=pipe._encode_vae_image + def encode(*args,**kwargs): + pipe.vae.to('cuda:0') + return original_encode(*args,**kwargs) + pipe._encode_vae_image=encode + original_decode=pipe.vae.decode + def decode(*args,**kwargs): + pipe.vae.to('cuda:0') + try:return original_decode(*args,**kwargs) + finally: + pipe.vae.to('cpu');torch.cuda.empty_cache() + pipe.vae.decode=decode + engine.args.teacher_offload='streamed-group1' if stream else 'synchronous-group4' + print(json.dumps({'event':'teacher_offload','mode':engine.args.teacher_offload,'execution_device':str(pipe._execution_device)}),flush=True) + return engine + + +def make_teacher(sample_dir=None,stream=True): + from runner import Engine,default_args + args=default_args();args.backend='nf4';args.bf16_transformer=True;args.prequant='/cache/qwen-nf4' + args.cache=False;args.compile=False;args.sample_dir=sample_dir + engine=Engine(args) + return enable_teacher_offload(engine,stream=stream) + + +def main(): + p=argparse.ArgumentParser();p.add_argument('--jobs');p.add_argument('--sample-dir',default='/poc/samples/fidelity-v3/bf16-dit') + p.add_argument('--parity',action='store_true');p.add_argument('--no-stream',action='store_true');args=p.parse_args() + if args.parity: + from runner import Engine,default_args + s=default_args();s.backend='nf4';s.bf16_transformer=True;s.prequant='/cache/qwen-nf4';s.cache=False;s.compile=False;s.sample_dir=args.sample_dir + engine=Engine(s) + job=dict(label='teacher-resident-parity',prompt='A golden retriever lying on a wooden porch beside a red watering can, realistic morning light.',width=512,height=512,steps=4,seed=808) + a=engine.generate(job) + enable_teacher_offload(engine,stream=not args.no_stream) + b=engine.generate({**job,'label':'teacher-streamed-parity'}) + same=Path(a['path']).read_bytes()==Path(b['path']).read_bytes() + report=dict(byte_identical=same,resident_seconds=a['seconds'],offloaded_seconds=b['seconds'],offload_mode=engine.args.teacher_offload) + Path('/poc/results/teacher-v3-parity.json').write_text(json.dumps(report,indent=2)+'\n');print(json.dumps(report),flush=True) + if not same:raise RuntimeError('Teacher offload parity failed') + else:engine=make_teacher(args.sample_dir,not args.no_stream) + if args.jobs: + for job in json.loads(Path(args.jobs).read_text()):engine.generate(job) + +if __name__=='__main__':main() diff --git a/reproduction/nunchaku_backend/upgrade_mlp_rank_v3.md b/reproduction/nunchaku_backend/upgrade_mlp_rank_v3.md new file mode 100644 index 0000000000000000000000000000000000000000..73bebdd4038a2a3a410f8fe3d2f1caf1a694385a --- /dev/null +++ b/reproduction/nunchaku_backend/upgrade_mlp_rank_v3.md @@ -0,0 +1,29 @@ +# Selective MLP projection rank upgrade — complete; CPU validation passed + +The [six-case rank probe](rank_probe_v3.md) completed with validation and held-out gains in all six cases. The main agent then completed staged first-layer verification and the full 32-layer upgrade. All 32 projections selected rank 512. An [independent CPU audit](../research/mlpproj-rank-upgrade-validation.md) verified the complete checkpoint, source immutability, all 193 untouched output shards, finite loaded state, and exact 4.71047 GiB model-state size. Median per-layer held-out MSE ratio is 0.64203 versus actual V3. Parent-MLP, whole-denoiser and image evaluation remain separate requirements before quality acceptance; the exporter does not automatically deploy its output. + +For each of 32 `img_mlp.proj` layers, the exporter tests rank 256 and 512 using that layer's exact V3 selected smoothing family/alpha, weighting, factorization, seed and solver settings. The default limit stays at 16 iterations. It compares actual packed Nunchaku outputs against the BF16 teacher on the original 40-step calibration validation inputs. The actual source rank-128 checkpoint is included in every comparison and wins ties or regressions. Candidates use training data for fitting; all source/rank decisions are frozen before held-out inputs are read. Held-out errors are then reported for the source and both candidates. + +Only the two low-rank parameter shapes may grow. Input/output dimensions, key sets, bias presence, INT4 residual layout, and BF16 scale/smoothing layouts must remain valid; all candidate floating tensors must be finite. The INT4 residual itself is refitted for each candidate. If the optimizer's old one-pass rank-32 guard wins, that result cannot replace the requested higher-rank candidate. All other 192 quantized layers and the BF16 boundary shard are copied and verified **byte-for-byte**. A target that retains rank 128 also retains its original shard byte-for-byte. + +The output keeps 224 quantized layer entries and 225 weight shards, using the existing runtime's per-layer rank metadata. The weight index keeps its tensor mapping and updates its total tensor-byte count. Per-layer JSON reports record actual-source validation/held-out comparisons, all candidate histories, selected rank, timing and low-rank storage cost. The original manifest, index, config, and all 225 source shards are hashed for provenance and rechecked after execution. No production or runner configuration is changed. + +GPU command, only if the main agent chooses to proceed and reserves GPU0: + +```sh +python -m nunchaku_backend.upgrade_mlp_rank_v3 \ + --source-checkpoint /cache/qwen-nunchaku-v3-r128 \ + --model-path /cache/huggingface/hub/models--Qwen--Qwen-Image-2.1/snapshots/b3179ad355be050328e483a9dfdd9e60cd62adfa \ + --activations /cache/qwen21-activation-v3-40 \ + --baseline-calibration /cache/qwen21-calibration \ + --out /cache/qwen-nunchaku-v3-mlpproj-upgrade \ + --device cuda:0 --ranks 256 512 --iterations 16 +``` + +Append `--dry-run` for CPU-only validation of all 32 targets, 64 recipes, source shard hashes and installed Nunchaku rank constructors. It creates no output directory and does not initialize CUDA. For initial GPU verification append `--max-layers 1`; the output remains incomplete and cannot be loaded by the normal runtime. Continue with identical settings and `--resume`, omitting the limit. Resume verifies exact source/data/code/settings plus output shard/report/index hashes. An interruption between replacing a shard and committing its manifest fails these checks rather than silently accepting a partial update. + +Rank candidates can be restricted with `--ranks 256`. Sorted ascending candidate ranks make equal-score candidate ties prefer the lower rank after the source tie preference. The output directory must be separate from the source and new unless resuming. Source files are copied, not hard-linked. The complete output is a new comparison checkpoint, never an automatic deployment. + +Uniform rank-256 selection adds 128 MiB of low-rank weights across the 32 projections; uniform rank-512 selection adds 384 MiB. Mixed selection reports its exact sum. These values exclude temporary low-rank activations and allocator effects. Per-linear validation improvement does not establish whole-denoiser or image quality; the original checkpoint and reports remain available for those comparisons. + +Preparation verification passed in the actual image with GPUs hidden. The [actual-source dry-run](../results/nunchaku-v3-mlpproj-upgrade-dryrun.json) validated 32 targets, 64 recipes, all 225 source shard hashes, and rank-256/512 constructors with `cuda_initialized:false`. All 54 combined CPU tests passed in 0.467 seconds, including strict source/tie selection, rank-dependent schema validation, held-out independence, and 225-shard dry-runs that block CUDA initialization/inference and assert no source or output mutation. GPU export subsequently completed under the main agent's coordination; the independent full CPU model load passed in 3.477 seconds. diff --git a/reproduction/nunchaku_backend/upgrade_mlp_rank_v3.py b/reproduction/nunchaku_backend/upgrade_mlp_rank_v3.py new file mode 100644 index 0000000000000000000000000000000000000000..376fb95160e5dc937014c31c92e3d5410f819f11 --- /dev/null +++ b/reproduction/nunchaku_backend/upgrade_mlp_rank_v3.py @@ -0,0 +1,297 @@ +"""Optional 32-layer MLP projection rank upgrade, guarded by actual V3. + +Prepared tooling only: no automatic execution, full-model launch or deployment. +Every other quantized layer and BF16 boundary shard is copied byte-for-byte. +""" +from __future__ import annotations + +import argparse +import copy +import json +from pathlib import Path +import shutil +import time + +from .compare_iterations_v3 import layer_stats +from .checkpoint_io import _json_write, _source_index, _read_tensor +from .export_v3 import ActivationReader, _hash_file +from .rank_probe_v3 import probe_recipe, lowrank_cost, cpu_constructor_probe +from .refine_checkpoint_v3 import assert_separate_output, require_matching_provenance, _actual_forward, _source_layer_state + + +def target_layers(manifest): + names = sorted((name for name in manifest["layers"] if name.endswith(".img_mlp.proj")), key=lambda n: int(n.split(".")[1])) + expected = [f"transformer_blocks.{block}.img_mlp.proj" for block in range(32)] + if names != expected: + raise ValueError("Source must contain exactly one MLP projection per block, 0 through31") + return names + + +def source_layout(source, manifest, names, index): + """Require isolated source layer shards so unmodified shards stay exact.""" + filenames = {} + for name in names: + keys = {key for key in index if key.startswith(name + ".")} + paths = {index[key] for key in keys} + if len(paths) != 1: + raise ValueError(f"Selected layer must occupy one isolated source shard: {name}") + path = next(iter(paths)) + if {key for key, value in index.items() if value == path} != keys: + raise ValueError(f"Selected source shard contains unrelated tensors: {name}") + if path.parent.resolve() != source.resolve() or path.name not in manifest["files"]: + raise ValueError("Invalid source shard path or manifest") + filenames[name] = path.name + if len(set(filenames.values())) != 32: + raise ValueError("Expected32 independent MLP projection source shards") + return filenames + + +def main(): + parser = argparse.ArgumentParser(description=__doc__) + parser.add_argument("--source-checkpoint", type=Path, required=True) + parser.add_argument("--model-path", type=Path, required=True) + parser.add_argument("--activations", type=Path, required=True) + parser.add_argument("--baseline-calibration", type=Path, required=True) + parser.add_argument("--out", type=Path, required=True) + parser.add_argument("--device", required=True) + parser.add_argument("--ranks", type=int, nargs="+", default=[256, 512]) + parser.add_argument("--iterations", type=int, default=16) + parser.add_argument("--threads", type=int, default=8) + parser.add_argument("--dry-run", action="store_true") + parser.add_argument("--resume", action="store_true") + parser.add_argument("--max-layers", type=int, help="Stop after this many new targets; output stays incomplete") + args = parser.parse_args() + assert_separate_output(args.source_checkpoint, args.out) + if args.max_layers is not None and args.max_layers < 1: + parser.error("--max-layers must be positive") + import torch + from safetensors.torch import load_file, save_file + from .runtime import _read_manifest, probe_backend + from .rank_upgrade_helpers import require_unique_ranks, validate_rank_state, select_validation_candidate + from .optimize_v3 import optimize_linear_weight, output_metrics + torch.set_num_threads(args.threads) + device = torch.device(args.device) + if device.type != "cuda" or device.index is None: + parser.error("Use an explicitly reserved indexed CUDA device") + source = args.source_checkpoint.resolve() + source_manifest = _read_manifest(source) + source_manifest_hash = _hash_file(source / "manifest.json") + names = target_layers(source_manifest) + stats = layer_stats(source_manifest) + if len(stats) != 224 or set(stats) != set(source_manifest["layers"]): + raise ValueError("Source must have complete224-layer V3 statistics") + for name in names: + info = source_manifest["layers"][name] + if info["rank"] != 128 or stats[name]["selected"]["rank"] != info["rank"]: + raise ValueError("This bounded upgrade requires actual V3 rank128 source MLP projections") + ranks = tuple(sorted(require_unique_ranks(args.ranks, source_rank=128))) + directory = args.model_path / "transformer" if (args.model_path / "transformer").is_dir() else args.model_path + baseline_path = args.baseline_calibration / "activation_stats.safetensors" if args.baseline_calibration.is_dir() else args.baseline_calibration + source_metadata_hashes = {filename: _hash_file(source / filename) for filename in ("config.json", "model.safetensors.index.json")} + if json.loads((source / "config.json").read_text()) != json.loads((directory / "config.json").read_text()): + raise ValueError("Source and BF16 teacher configs differ") + activations = ActivationReader(args.activations) + activations.require(names) + require_matching_provenance(source_manifest, activations.fingerprint, _hash_file(directory / "config.json"), _hash_file(baseline_path)) + recipes = {name: {rank: probe_recipe(stats[name], rank, args.iterations) for rank in ranks} for name in names} + source_index = _source_index(source) + filenames = source_layout(source, source_manifest, names, source_index) + source_index_document = json.loads((source / "model.safetensors.index.json").read_text()) + weight_map = source_index_document["weight_map"] + if set(weight_map) != set(source_index) or any(weight_map[key] != path.name for key, path in source_index.items()): + raise ValueError("Source weight map does not match checkpoint shards") + paths = sorted(set(source_index.values())) + if len(paths) != 225 or set(source_manifest["files"]) != {path.name for path in paths}: + raise ValueError("Expected the complete225-shard V3 source layout") + source_hashes = {path.name: _hash_file(path) for path in paths} + for filename, digest in source_hashes.items(): + if source_manifest["files"][filename].get("sha256", digest) != digest: + raise ValueError(f"Source shard differs from its saved manifest: {filename}") + constructor = cpu_constructor_probe(source_manifest["layers"][names[0]], ranks) + costs = {rank: lowrank_cost(4096, 12288, rank) for rank in ranks} + if args.dry_run: + print(json.dumps({"dry_run": True, "gpu_work_executed": False, "cuda_initialized": torch.cuda.is_initialized(), + "source_manifest_sha256": source_manifest_hash, "target_layers": names, + "target_count": len(names), "unchanged_quantized_layers": 192, + "source_weight_shards_verified": len(paths), "candidate_recipes": len(names) * len(ranks), + "candidate_ranks": ranks, "iterations": args.iterations, "cpu_constructor": constructor, + "uniform_upgrade_lowrank_costs": costs, "out": str(args.out)}, indent=2)) + return + model_index = _source_index(directory) + baseline = load_file(str(baseline_path), device="cpu") + if any(name + ".input_absmax" not in baseline for name in names): + raise ValueError("Original absmax baseline archive lacks target layers") + identity = {"algorithm": "selective_mlp_projection_rank_upgrade_actual_source_validation_guard", + "source_checkpoint": str(source), "source_manifest_sha256": source_manifest_hash, + "source_file_sha256": source_hashes, "source_metadata_sha256": source_metadata_hashes, + "activation_fingerprint": activations.fingerprint, "source_config_sha256": _hash_file(directory / "config.json"), + "baseline_calibration_sha256": _hash_file(baseline_path), "ranks": ranks, + "iterations": args.iterations, "device": str(device), "target_layers": names, + "code_sha256": {filename: _hash_file(Path(__file__).parent / filename) for filename in + ("upgrade_mlp_rank_v3.py", "rank_upgrade_helpers.py", "rank_probe_v3.py", "refine_checkpoint_v3.py", "optimize_v3.py", "gptq_v3.py", "baseline_candidate.py", "packing.py", "checkpoint_io.py", "layout.py", "runtime.py")}} + # JSON-normalize tuples for exact resume identity comparisons. + identity = json.loads(json.dumps(identity)) + manifest_path = args.out / "manifest.json" + if args.out.exists(): + if not args.resume or not manifest_path.is_file(): + raise ValueError("Use a new output directory, or --resume the same incomplete upgrade") + manifest = json.loads(manifest_path.read_text()) + if manifest.get("rank_upgrade_identity") != identity: + raise ValueError("Resume requires identical source/data/code/settings") + for filename, entry in manifest["files"].items(): + if _hash_file(args.out / filename) != entry["sha256"]: + raise ValueError(f"Modified or incomplete output shard: {filename}") + for entry in manifest["rank_upgrade_reports"].values(): + if _hash_file(args.out / entry["file"]) != entry["sha256"]: + raise ValueError("Modified output layer report") + if _hash_file(args.out / "config.json") != source_metadata_hashes["config.json"]: + raise ValueError("Modified output config") + if _hash_file(args.out / "model.safetensors.index.json") != manifest["output_index_sha256"]: + raise ValueError("Modified output index") + else: + args.out.mkdir(parents=True) + for path in paths: + shutil.copyfile(path, args.out / path.name) + for filename in source_metadata_hashes: + shutil.copyfile(source / filename, args.out / filename) + manifest = copy.deepcopy(source_manifest) + for filename, digest in source_hashes.items(): + manifest["files"][filename]["sha256"] = digest + manifest.update(complete=False, created_unix=time.time(), rank_upgrade_identity=identity, + rank_upgrade_reports={}, runtime=probe_backend(), + output_index_sha256=_hash_file(args.out / "model.safetensors.index.json"), + algorithm="V3 with selectively higher-rank MLP projection branches; actual source wins validation ties/regressions") + manifest["conversion_identity"] = {**copy.deepcopy(source_manifest["conversion_identity"]), + "rank_upgrade": identity} + manifest["limitations"] = "Per-linear validation selection, heldout reporting only. Variable rank does not imply whole-denoiser or image equivalence. Other192 quantized layers and BF16 boundaries are byte-identical to source." + _json_write(manifest_path, manifest) + completed_now = 0 + for name in names: + if name in manifest["rank_upgrade_reports"]: + continue + started = time.perf_counter() + info = source_manifest["layers"][name] + original = _source_layer_state(source_index, name) + weight = _read_tensor(model_index, name + ".weight") + bias = _read_tensor(model_index, name + ".bias") if name + ".bias" in model_index else None + if weight.dtype != torch.bfloat16 or tuple(weight.shape) != (info["out_features"], info["in_features"]): + raise ValueError(f"BF16 teacher weight mismatch: {name}") + train = activations.read(name, "train") + validation = activations.read(name, "validation").to(device) + train_absmax = activations.read(name, "input_absmax") + teacher_weight = weight.to(device) + teacher_bias = bias.to(device) if bias is not None else None + candidates = [] + with torch.inference_mode(): + teacher_validation = torch.nn.functional.linear(validation, teacher_weight, teacher_bias).float() + source_validation = output_metrics(_actual_forward(original, info, validation, device), teacher_validation) + for rank in ranks: + rank_started = time.perf_counter() + state, candidate_stats, reference = optimize_linear_weight( + weight, train, validation, None, bias, **recipes[name][rank], + train_absmax=train_absmax, baseline_absmax=baseline[name + ".input_absmax"], + conversion_device=device, objective_backend="nunchaku", + output_correction=source_manifest["conversion_identity"]["settings"].get("output_correction", True)) + selected = candidate_stats["selected"] + requested_recipe = (selected["rank"] == rank and selected["family"] == stats[name]["selected"]["family"] + and selected.get("alpha") == stats[name]["selected"].get("alpha")) + state = {key: value.detach().cpu().contiguous() for key, value in state.items()} + if requested_recipe: + candidate_info = validate_rank_state(original, state, info, rank) + metric = output_metrics(_actual_forward(state, candidate_info, validation, device), teacher_validation) + else: + candidate_info = {**info, "rank": selected["rank"]} + metric = candidate_stats["validation"] + # Keep only CPU packed states while fitting the next rank. + candidates.append({"rank": rank, "validation": metric if requested_recipe else {"mse": float("inf"), "finite": False}, + "actual_validation": metric, "state": state, "info": candidate_info, + "requested_recipe": requested_recipe, "optimizer": candidate_stats, + "fit_seconds": time.perf_counter() - rank_started}) + del state, candidate_stats, reference + selected_index = select_validation_candidate(source_validation, candidates) + # All rank/source selection is now frozen before heldout is read. + heldout = activations.read(name, "heldout").to(device) + teacher_heldout = torch.nn.functional.linear(heldout, teacher_weight, teacher_bias).float() + source_heldout = output_metrics(_actual_forward(original, info, heldout, device), teacher_heldout) + candidate_reports = [] + for candidate in candidates: + metric = output_metrics(_actual_forward(candidate["state"], candidate["info"], heldout, device), teacher_heldout) + candidate_reports.append({"requested_rank": candidate["rank"], "selected_recipe_matches_request": candidate["requested_recipe"], + "selected": candidate["optimizer"]["selected"], "validation": candidate["actual_validation"], "heldout": metric, + "validation_mse_ratio_to_actual_v3": candidate["actual_validation"]["mse"] / max(source_validation["mse"], 1e-30), + "heldout_mse_ratio_to_actual_v3": metric["mse"] / max(source_heldout["mse"], 1e-30), + "optimizer_stats": candidate["optimizer"], "fit_seconds": candidate["fit_seconds"]}) + if selected_index is None: + chosen_state, chosen_info = original, info + chosen_stats = copy.deepcopy(stats[name]) + selected_validation, selected_heldout = source_validation, source_heldout + else: + winner = candidates[selected_index] + chosen_state, chosen_info = winner["state"], winner["info"] + chosen_stats = copy.deepcopy(winner["optimizer"]) + selected_validation, selected_heldout = winner["actual_validation"], candidate_reports[selected_index]["heldout"] + # Existing summary tooling's rank32 comparison remains explicitly historical. + chosen_stats["baseline"] = copy.deepcopy(stats[name]["baseline"]) + chosen_stats["baseline_metrics_provenance"] = "Original V3 one-pass rank32 metrics on identical archive; actual V3 guard is recorded in rank_upgrade report" + chosen_stats.update(validation=selected_validation, heldout=selected_heldout) + chosen_stats["validation_mse_ratio_to_baseline"] = selected_validation["mse"] / max(stats[name]["baseline"]["validation"]["mse"], 1e-30) + chosen_stats["heldout_mse_ratio_to_baseline"] = selected_heldout["mse"] / max(stats[name]["baseline"]["heldout"]["mse"], 1e-30) + filename = filenames[name] + if selected_index is not None: + temporary = args.out / (filename + ".tmp") + save_file({name + "." + key: value for key, value in chosen_state.items()}, str(temporary)) + temporary.replace(args.out / filename) + else: + # Already copied source; assert exact fallback including metadata. + if _hash_file(args.out / filename) != source_hashes[filename]: + raise ValueError("Fallback shard was unexpectedly modified") + report = {"layer": name, "source_rank": info["rank"], "selected_rank": chosen_info["rank"], + "source_retained": selected_index is None, "selected_candidate_index": selected_index, + "source_actual_validation": source_validation, "source_actual_heldout": source_heldout, + "source_saved_validation": stats[name]["validation"], "candidates": candidate_reports, + "decision_uses_heldout": False, "seconds": time.perf_counter() - started, + "lowrank_cost": lowrank_cost(info["in_features"], info["out_features"], chosen_info["rank"], info["rank"])} + report_filename = f"rank-upgrade-block-{int(name.split('.')[1]):02d}.json" + _json_write(args.out / report_filename, report) + chosen_stats["rank_upgrade"] = {"report": report_filename, "selected_rank": chosen_info["rank"], "source_retained": selected_index is None} + manifest["layers"][name] = copy.deepcopy(chosen_info) + manifest["files"][filename] = {**copy.deepcopy(source_manifest["files"][filename]), + "sha256": _hash_file(args.out / filename), + "bytes": sum(t.numel() * t.element_size() for t in chosen_state.values()), + "source_file_copied_verbatim": selected_index is None, "layer_stats": {name: chosen_stats}} + manifest["rank_upgrade_reports"][name] = {"file": report_filename, "sha256": _hash_file(args.out / report_filename), + "selected_rank": chosen_info["rank"], "source_retained": selected_index is None} + output_index = copy.deepcopy(source_index_document) + output_index.setdefault("metadata", {})["total_size"] = sum(entry["bytes"] for entry in manifest["files"].values()) + _json_write(args.out / "model.safetensors.index.json", output_index) + manifest["output_index_sha256"] = _hash_file(args.out / "model.safetensors.index.json") + _json_write(manifest_path, manifest) + print(json.dumps({"event": "mlp_rank_selected", "layer": name, "selected_rank": chosen_info["rank"], + "validation_ratio_to_actual_v3": selected_validation["mse"] / max(source_validation["mse"], 1e-30), + "heldout_ratio_to_actual_v3": selected_heldout["mse"] / max(source_heldout["mse"], 1e-30), "seconds": report["seconds"]}), flush=True) + del original, weight, bias, train, validation, heldout, teacher_weight, teacher_bias, teacher_validation, teacher_heldout, candidates, chosen_state + completed_now += 1 + if args.max_layers and completed_now >= args.max_layers: + break + if _hash_file(source / "manifest.json") != source_manifest_hash: + raise ValueError("Source manifest changed during upgrade") + for filename, digest in {**source_hashes, **source_metadata_hashes}.items(): + if _hash_file(source / filename) != digest: + raise ValueError(f"Source changed during upgrade: {filename}") + untouched_files = set(source_hashes) - set(filenames.values()) + for filename in untouched_files: + if _hash_file(args.out / filename) != source_hashes[filename]: + raise ValueError(f"Unrelated output shard changed: {filename}") + manifest["complete"] = len(manifest["rank_upgrade_reports"]) == 32 + manifest["rank_upgrade_summary"] = {"processed": len(manifest["rank_upgrade_reports"]), + "rank_counts": {str(rank): sum(row["selected_rank"] == rank for row in manifest["rank_upgrade_reports"].values()) for rank in (128, *ranks)}, + "unmodified_quantized_layers": 192, "untouched_shards_verified": len(untouched_files), + "lowrank_extra_bytes": sum(2 * (source_manifest["layers"][name]["in_features"] + source_manifest["layers"][name]["out_features"]) * (row["selected_rank"] - 128) for name, row in manifest["rank_upgrade_reports"].items())} + if manifest["complete"]: + manifest.setdefault("completed_unix", time.time()) + _json_write(manifest_path, manifest) + print(json.dumps({"complete": manifest["complete"], **manifest["rank_upgrade_summary"]}), flush=True) + + +if __name__ == "__main__": + main() diff --git a/reproduction/nunchaku_backend/upgrade_mlp_rank_v3_test.py b/reproduction/nunchaku_backend/upgrade_mlp_rank_v3_test.py new file mode 100644 index 0000000000000000000000000000000000000000..38a2ef33e434c046a4ec8d24b3df057d11aa8c16 --- /dev/null +++ b/reproduction/nunchaku_backend/upgrade_mlp_rank_v3_test.py @@ -0,0 +1,184 @@ +"""CPU tests for bounded full-checkpoint rank-upgrade preparation. + +Tiny real safetensors exercise shard discovery/hashes; they are deliberately +not inference weights. The dry run mocks only activation metadata and the +installed Nunchaku constructor, while CUDA/inference are explicit tripwires. +""" +import contextlib +import copy +import hashlib +import io +import json +from pathlib import Path +import tempfile +import unittest +from unittest.mock import patch + +import torch +from safetensors.torch import save_file + +from . import FORMAT +from . import upgrade_mlp_rank_v3 as upgrade + + +def _stats(): + return {"selected": {"rank": 128, "family": "activation_only", "alpha": 0.5, + "weighting": "none", "factorization": "up_singular"}, + "search": {"seed": 1947, "niter": 4, "oversample": 16, "ridge": 0.01, + "gptq_damp": 0.01, "final_gptq": True}, + "baseline": {"conversion": {"rank": 32}}, + "validation": {"mse": 0.1, "relative_l2": 0.2, "finite": True}, + "heldout": {"mse": 0.2, "relative_l2": 0.3, "finite": True}} + + +def _manifest(): + layers = {} + for block in reversed(range(32)): + for role in ("attn.to_q", "attn.to_k", "attn.to_v", "attn.to_out.0", + "img_mlp.proj", "img_mlp.out", "img_mlp.gate_layer"): + layers[f"transformer_blocks.{block}.{role}"] = { + "in_features": 4096, "out_features": 12288 if role == "img_mlp.proj" else 4096, + "rank": 128, "bias": False, "precision": "int4"} + return {"backend_format": FORMAT, "complete": True, "layers": layers, "files": {}} + + +def _layout(source, manifest): + index = {} + for number, name in enumerate(manifest["layers"]): + filename = f"layer-{number:03d}.safetensors" + index[name + ".qweight"] = source / filename + manifest["files"][filename] = {"layer_stats": {name: _stats()}} + index["proj_out.weight"] = source / "boundary.safetensors" + manifest["files"]["boundary.safetensors"] = {} + return index + + +class UpgradeCheckpointTests(unittest.TestCase): + def test_targets_are_exactly_32_projections_in_depth_order(self): + manifest = _manifest() + before = copy.deepcopy(manifest) + expected = [f"transformer_blocks.{block}.img_mlp.proj" for block in range(32)] + self.assertEqual(upgrade.target_layers(manifest), expected) + self.assertEqual(manifest, before) + self.assertEqual(len(manifest["layers"]) - len(expected), 192) + + def test_targets_reject_missing_or_extra_projection_block(self): + for mode in ("missing", "extra"): + manifest = _manifest() + if mode == "missing": + del manifest["layers"]["transformer_blocks.17.img_mlp.proj"] + else: + manifest["layers"]["transformer_blocks.32.img_mlp.proj"] = {} + with self.subTest(mode=mode), self.assertRaises(ValueError): + upgrade.target_layers(manifest) + + def test_source_layout_selects_32_leaving_exactly_193_untouched_shards(self): + source = Path("/tmp/isolated-upgrade-source") + manifest = _manifest() + index = _layout(source, manifest) + result = upgrade.source_layout(source, manifest, upgrade.target_layers(manifest), index) + self.assertEqual(len(result), 32) + self.assertEqual(len(set(index.values())), 225) + self.assertEqual(len(set(manifest["files"]) - set(result.values())), 193) + + def test_source_layout_rejects_missing_split_shared_and_external_shards(self): + for mode in ("missing", "split", "shared", "external", "unlisted"): + source = Path("/tmp/isolated-upgrade-source") + manifest = _manifest() + index = _layout(source, manifest) + name = upgrade.target_layers(manifest)[0] + key = name + ".qweight" + original = index[key] + if mode == "missing": + del index[key] + elif mode == "split": + index[name + ".proj_up"] = source / "split.safetensors" + elif mode == "shared": + index["unrelated.weight"] = original + elif mode == "external": + index[key] = source.parent / original.name + else: + del manifest["files"][original.name] + with self.subTest(mode=mode), self.assertRaises(ValueError): + upgrade.source_layout(source, manifest, upgrade.target_layers(manifest), index) + + def test_dryrun_verifies_real_225_shards_without_cuda_or_output_mutation(self): + with tempfile.TemporaryDirectory() as directory: + root = Path(directory) + source, model, output = root / "source", root / "model", root / "output" + source.mkdir() + model.mkdir() + config = b'{"_class_name":"QwenImage21Transformer2DModel","num_layers":32}' + (source / "config.json").write_bytes(config) + (model / "config.json").write_bytes(config) + baseline = root / "baseline.safetensors" + baseline.write_bytes(b"Baseline is hashed, never loaded during dry run") + manifest = _manifest() + index = _layout(source, manifest) + for key, path in index.items(): + save_file({key: torch.zeros(1, dtype=torch.int8)}, str(path), metadata={"fixture": "CPU dryrun"}) + manifest["files"][path.name]["sha256"] = hashlib.sha256(path.read_bytes()).hexdigest() + manifest["conversion_identity"] = { + "activation_fingerprint": "test-activation-fingerprint", + "source_config_sha256": hashlib.sha256(config).hexdigest(), + "baseline_calibration_sha256": hashlib.sha256(baseline.read_bytes()).hexdigest()} + (source / "manifest.json").write_text(json.dumps(manifest)) + (source / "model.safetensors.index.json").write_text(json.dumps({ + "metadata": {"total_size": 225}, "weight_map": {key: path.name for key, path in index.items()}})) + before = {path.name: path.read_bytes() for path in source.iterdir()} + expected_names = upgrade.target_layers(manifest) + + class Reader: + fingerprint = "test-activation-fingerprint" + + def require(self, names): + if names != expected_names: + raise AssertionError("Dry run requested wrong activation layers") + + def read(self, *args): + raise AssertionError("Dry run must not read activation tensors") + + argv = ["upgrade_mlp_rank_v3", "--source-checkpoint", str(source), + "--model-path", str(model), "--activations", str(root / "activations"), + "--baseline-calibration", str(baseline), "--out", str(output), + "--device", "cuda:0", "--threads", "2", "--dry-run", "--ranks", "512", "256"] + captured = io.StringIO() + with patch("sys.argv", argv), patch.object(upgrade, "ActivationReader", return_value=Reader()), \ + patch.object(upgrade, "cpu_constructor_probe", return_value={"cuda_initialized": False}) as constructor, \ + patch.object(upgrade, "_actual_forward", side_effect=AssertionError("Unexpected inference")), \ + patch("torch.cuda._lazy_init", side_effect=AssertionError("Unexpected CUDA initialization")), \ + contextlib.redirect_stdout(captured): + upgrade.main() + report = json.loads(captured.getvalue()) + self.assertTrue(report["dry_run"]) + self.assertFalse(report["gpu_work_executed"]) + self.assertFalse(report["cuda_initialized"]) + self.assertEqual(report["target_count"], 32) + self.assertEqual(report["unchanged_quantized_layers"], 192) + self.assertEqual(report["source_weight_shards_verified"], 225) + self.assertEqual(report["candidate_recipes"], 64) + self.assertEqual(report["candidate_ranks"], [256, 512]) + self.assertEqual(tuple(constructor.call_args.args[1]), (256, 512)) + self.assertFalse(output.exists()) + self.assertEqual(before, {path.name: path.read_bytes() for path in source.iterdir()}) + + # The actual V3 source has no per-file hashes; preparation must + # compute fresh hashes without mutating or rejecting that source. + for entry in manifest["files"].values(): + del entry["sha256"] + (source / "manifest.json").write_text(json.dumps(manifest)) + before = {path.name: path.read_bytes() for path in source.iterdir()} + captured = io.StringIO() + with patch("sys.argv", argv), patch.object(upgrade, "ActivationReader", return_value=Reader()), \ + patch.object(upgrade, "cpu_constructor_probe", return_value={"cuda_initialized": False}), \ + patch.object(upgrade, "_actual_forward", side_effect=AssertionError("Unexpected inference")), \ + patch("torch.cuda._lazy_init", side_effect=AssertionError("Unexpected CUDA initialization")), \ + contextlib.redirect_stdout(captured): + upgrade.main() + self.assertEqual(json.loads(captured.getvalue())["source_weight_shards_verified"], 225) + self.assertFalse(output.exists()) + self.assertEqual(before, {path.name: path.read_bytes() for path in source.iterdir()}) + + +if __name__ == "__main__": + unittest.main(verbosity=2) diff --git a/reproduction/nunchaku_backend/upstream/deepcompressor-tree.json b/reproduction/nunchaku_backend/upstream/deepcompressor-tree.json new file mode 100644 index 0000000000000000000000000000000000000000..7d3d3a96925cc21eb7c3fc6dccd121f962a95383 --- /dev/null +++ b/reproduction/nunchaku_backend/upstream/deepcompressor-tree.json @@ -0,0 +1 @@ +{"sha": "69f3473f5e1c1504bae35cc50c7858ef900a9b17", "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/trees/69f3473f5e1c1504bae35cc50c7858ef900a9b17", "tree": [{"path": ".gitignore", "mode": "100644", "type": "blob", "sha": "04c27b4d8959c28ee9483f69f2f8eada6ef8b543", "size": 3182, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/04c27b4d8959c28ee9483f69f2f8eada6ef8b543"}, {"path": "LICENSE", "mode": "100644", "type": "blob", "sha": "228229d3a5fc1b67b4cc60c07753213614712676", "size": 11401, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/228229d3a5fc1b67b4cc60c07753213614712676"}, {"path": "README.md", "mode": "100644", "type": "blob", "sha": "f942a152918104f50b5255fecb15f4f18072bee9", "size": 18642, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/f942a152918104f50b5255fecb15f4f18072bee9"}, {"path": "assets", "mode": "040000", "type": "tree", "sha": "a9a1d7eadabb27f4a721fab35aecc436bdae22db", "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/trees/a9a1d7eadabb27f4a721fab35aecc436bdae22db"}, {"path": "assets/deepcompressor.png", "mode": "100644", "type": "blob", "sha": "8d4a1ef7d7a458f557e2bb05b2d3aacdf4649ac9", "size": 47764, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/8d4a1ef7d7a458f557e2bb05b2d3aacdf4649ac9"}, {"path": "assets/diffusion", "mode": "040000", "type": "tree", "sha": "67b05daf9ea1bfc7a2d16ff71d02bf87a185afc4", "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/trees/67b05daf9ea1bfc7a2d16ff71d02bf87a185afc4"}, {"path": "assets/diffusion/.gitkeep", "mode": "100644", "type": "blob", "sha": "e69de29bb2d1d6434b8b29ae775ad8c2e48c5391", "size": 0, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/e69de29bb2d1d6434b8b29ae775ad8c2e48c5391"}, {"path": "assets/diffusion/svdquant", "mode": "040000", "type": "tree", "sha": "2a551e5ef979799aa06cf3bee38b6866e7f3b693", "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/trees/2a551e5ef979799aa06cf3bee38b6866e7f3b693"}, {"path": "assets/diffusion/svdquant/svdquant.gif", "mode": "100644", "type": "blob", "sha": "d0b675d48490d9bf934c0ed6cf541042a30b9a22", "size": 294087, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/d0b675d48490d9bf934c0ed6cf541042a30b9a22"}, {"path": "assets/diffusion/svdquant/teaser.jpg", "mode": "100644", "type": "blob", "sha": "5623fc46278ee09acfec1a8dbc95cc3331827561", "size": 1766517, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/5623fc46278ee09acfec1a8dbc95cc3331827561"}, {"path": "assets/llm", "mode": "040000", "type": "tree", "sha": "524877b325394974201c9a6c47b0c3e4f106fa8e", "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/trees/524877b325394974201c9a6c47b0c3e4f106fa8e"}, {"path": "assets/llm/.gitkeep", "mode": "100644", "type": "blob", "sha": "e69de29bb2d1d6434b8b29ae775ad8c2e48c5391", "size": 0, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/e69de29bb2d1d6434b8b29ae775ad8c2e48c5391"}, {"path": "assets/llm/qoq", "mode": "040000", "type": "tree", "sha": "c3c3be8572616dc63fb79d68e93f24c0ed139a07", "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/trees/c3c3be8572616dc63fb79d68e93f24c0ed139a07"}, {"path": "assets/llm/qoq/qoq-qserve.png", "mode": "100644", "type": "blob", "sha": "b2dfc972605a2d39cd1f5d09fbabfd4d3cfcff1c", "size": 245133, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/b2dfc972605a2d39cd1f5d09fbabfd4d3cfcff1c"}, {"path": "assets/llm/qoq/qoq.png", "mode": "100644", "type": "blob", "sha": "f892e115bbea2c30965b4b0b3784aac15633bbf4", "size": 625316, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/f892e115bbea2c30965b4b0b3784aac15633bbf4"}, {"path": "deepcompressor", "mode": "040000", "type": "tree", "sha": "e908a2119e2b795d92cdcf5885794aa6497849a7", "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/trees/e908a2119e2b795d92cdcf5885794aa6497849a7"}, {"path": "deepcompressor/__init__.py", "mode": "100644", "type": "blob", "sha": "5c0cf285cb9b647150914087937f12440ab102b2", "size": 47, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/5c0cf285cb9b647150914087937f12440ab102b2"}, {"path": "deepcompressor/app", "mode": "040000", "type": "tree", "sha": "284e6b874f766141d5979e02fecf9fbbdd073ac9", "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/trees/284e6b874f766141d5979e02fecf9fbbdd073ac9"}, {"path": "deepcompressor/app/__init__.py", "mode": "100644", "type": "blob", "sha": "e69de29bb2d1d6434b8b29ae775ad8c2e48c5391", "size": 0, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/e69de29bb2d1d6434b8b29ae775ad8c2e48c5391"}, {"path": "deepcompressor/app/diffusion", "mode": "040000", "type": "tree", "sha": "6bd8887a9647e24b73e07e9b470c54c6605e9a60", "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/trees/6bd8887a9647e24b73e07e9b470c54c6605e9a60"}, {"path": "deepcompressor/app/diffusion/__init__.py", "mode": "100644", "type": "blob", "sha": "e69de29bb2d1d6434b8b29ae775ad8c2e48c5391", "size": 0, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/e69de29bb2d1d6434b8b29ae775ad8c2e48c5391"}, {"path": "deepcompressor/app/diffusion/cache", "mode": "040000", "type": "tree", "sha": "8d759c0e428e9d05778ac0b4c78c476449cece5d", "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/trees/8d759c0e428e9d05778ac0b4c78c476449cece5d"}, {"path": "deepcompressor/app/diffusion/cache/__init__.py", "mode": "100644", "type": "blob", "sha": "82c3e1f7234966cb1ceca31d06703eb0e5c34d8e", "size": 71, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/82c3e1f7234966cb1ceca31d06703eb0e5c34d8e"}, {"path": "deepcompressor/app/diffusion/cache/config.py", "mode": "100644", "type": "blob", "sha": "74b774dd716367be2ae1abd57f933b2b6e6718c3", "size": 2403, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/74b774dd716367be2ae1abd57f933b2b6e6718c3"}, {"path": "deepcompressor/app/diffusion/config.py", "mode": "100644", "type": "blob", "sha": "14d55fd87cd041694de57ae5acd832ed479439a6", "size": 8440, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/14d55fd87cd041694de57ae5acd832ed479439a6"}, {"path": "deepcompressor/app/diffusion/dataset", "mode": "040000", "type": "tree", "sha": "e867042d7b470df9d327e1ce2b6df1dcab81b4d5", "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/trees/e867042d7b470df9d327e1ce2b6df1dcab81b4d5"}, {"path": "deepcompressor/app/diffusion/dataset/__init__.py", "mode": "100644", "type": "blob", "sha": "4c7af45afefffe6f81474fe9db694c9af1cf299e", "size": 138, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/4c7af45afefffe6f81474fe9db694c9af1cf299e"}, {"path": "deepcompressor/app/diffusion/dataset/base.py", "mode": "100644", "type": "blob", "sha": "535b19172fd2e6344932dde3acb331a9c2d584a7", "size": 2623, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/535b19172fd2e6344932dde3acb331a9c2d584a7"}, {"path": "deepcompressor/app/diffusion/dataset/calib.py", "mode": "100644", "type": "blob", "sha": "f794d30830bce9fc15b7da12c619c9d011ef174b", "size": 15238, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/f794d30830bce9fc15b7da12c619c9d011ef174b"}, {"path": "deepcompressor/app/diffusion/dataset/collect", "mode": "040000", "type": "tree", "sha": "78a81649be8850b37950ce49ffc70a88bf1ffaa6", "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/trees/78a81649be8850b37950ce49ffc70a88bf1ffaa6"}, {"path": "deepcompressor/app/diffusion/dataset/collect/calib.py", "mode": "100644", "type": "blob", "sha": "5b0e99c03c7f471a238220fc5a8c573bcdd1f816", "size": 5657, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/5b0e99c03c7f471a238220fc5a8c573bcdd1f816"}, {"path": "deepcompressor/app/diffusion/dataset/collect/utils.py", "mode": "100644", "type": "blob", "sha": "dc4169d6a1931f4718a2d84b7d7ad4a10cad6788", "size": 3293, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/dc4169d6a1931f4718a2d84b7d7ad4a10cad6788"}, {"path": "deepcompressor/app/diffusion/dataset/data", "mode": "040000", "type": "tree", "sha": "c6f6b23a92cd901e26c8e871bc7523d8ba180989", "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/trees/c6f6b23a92cd901e26c8e871bc7523d8ba180989"}, {"path": "deepcompressor/app/diffusion/dataset/data/COCO", "mode": "040000", "type": "tree", "sha": "85422db737c9a226d539aae0bf45ebb9183e7bc2", "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/trees/85422db737c9a226d539aae0bf45ebb9183e7bc2"}, {"path": "deepcompressor/app/diffusion/dataset/data/COCO/COCO.py", "mode": "100644", "type": "blob", "sha": "2f155c752eedb5f0c4f8a4d8293e5700be3f5f23", "size": 7724, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/2f155c752eedb5f0c4f8a4d8293e5700be3f5f23"}, {"path": "deepcompressor/app/diffusion/dataset/data/COCO/__init__.py", "mode": "100644", "type": "blob", "sha": "e69de29bb2d1d6434b8b29ae775ad8c2e48c5391", "size": 0, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/e69de29bb2d1d6434b8b29ae775ad8c2e48c5391"}, {"path": "deepcompressor/app/diffusion/dataset/data/DCI", "mode": "040000", "type": "tree", "sha": "7ff18e460af148b17a28a55a3f1c4c4a01e9caa5", "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/trees/7ff18e460af148b17a28a55a3f1c4c4a01e9caa5"}, {"path": "deepcompressor/app/diffusion/dataset/data/DCI/DCI.py", "mode": "100644", "type": "blob", "sha": "a65454e95cb90646d386e95db9013566bcef7ff7", "size": 4032, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/a65454e95cb90646d386e95db9013566bcef7ff7"}, {"path": "deepcompressor/app/diffusion/dataset/data/DCI/__init__.py", "mode": "100644", "type": "blob", "sha": "e69de29bb2d1d6434b8b29ae775ad8c2e48c5391", "size": 0, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/e69de29bb2d1d6434b8b29ae775ad8c2e48c5391"}, {"path": "deepcompressor/app/diffusion/dataset/data/MJHQ", "mode": "040000", "type": "tree", "sha": "9848ca4d6775747d029e9ae0fbaaf520f5fe1725", "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/trees/9848ca4d6775747d029e9ae0fbaaf520f5fe1725"}, {"path": "deepcompressor/app/diffusion/dataset/data/MJHQ/MJHQ.py", "mode": "100644", "type": "blob", "sha": "8cfa715964c272bb324a9f7e7c71e550e0efcd90", "size": 3958, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/8cfa715964c272bb324a9f7e7c71e550e0efcd90"}, {"path": "deepcompressor/app/diffusion/dataset/data/MJHQ/__init__.py", "mode": "100644", "type": "blob", "sha": "e69de29bb2d1d6434b8b29ae775ad8c2e48c5391", "size": 0, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/e69de29bb2d1d6434b8b29ae775ad8c2e48c5391"}, {"path": "deepcompressor/app/diffusion/dataset/data/__init__.py", "mode": "100644", "type": "blob", "sha": "496f394b1e7d092025922694d2ef48c39422fd9e", "size": 2624, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/496f394b1e7d092025922694d2ef48c39422fd9e"}, {"path": "deepcompressor/app/diffusion/dataset/data/dump.py", "mode": "100644", "type": "blob", "sha": "4082fa3cb28bcb73c646af089aa4272a0f70e00b", "size": 3905, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/4082fa3cb28bcb73c646af089aa4272a0f70e00b"}, {"path": "deepcompressor/app/diffusion/eval", "mode": "040000", "type": "tree", "sha": "0755986d7972684148169f0ecbb254bbfc54e6d2", "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/trees/0755986d7972684148169f0ecbb254bbfc54e6d2"}, {"path": "deepcompressor/app/diffusion/eval/__init__.py", "mode": "100644", "type": "blob", "sha": "b0960a027e582038ec3efe5eb5676ddb2c4f0e08", "size": 65, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/b0960a027e582038ec3efe5eb5676ddb2c4f0e08"}, {"path": "deepcompressor/app/diffusion/eval/config.py", "mode": "100644", "type": "blob", "sha": "12de797185bc150a4f4fd377b49bcc9b4b15045d", "size": 9873, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/12de797185bc150a4f4fd377b49bcc9b4b15045d"}, {"path": "deepcompressor/app/diffusion/eval/metrics", "mode": "040000", "type": "tree", "sha": "582829074617b4fd0ca05974d918476fde62c2be", "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/trees/582829074617b4fd0ca05974d918476fde62c2be"}, {"path": "deepcompressor/app/diffusion/eval/metrics/__init__.py", "mode": "100644", "type": "blob", "sha": "bfe9aa285d185f09bc2fa3945fabd2757772a981", "size": 4543, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/bfe9aa285d185f09bc2fa3945fabd2757772a981"}, {"path": "deepcompressor/app/diffusion/eval/metrics/fid.py", "mode": "100644", "type": "blob", "sha": "8c054bf9850728c7c13521da3501d152fed74d07", "size": 4298, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/8c054bf9850728c7c13521da3501d152fed74d07"}, {"path": "deepcompressor/app/diffusion/eval/metrics/image_reward.py", "mode": "100644", "type": "blob", "sha": "f964431ae5a777a195b15601b07dc6f2e315e365", "size": 889, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/f964431ae5a777a195b15601b07dc6f2e315e365"}, {"path": "deepcompressor/app/diffusion/eval/metrics/multimodal.py", "mode": "100644", "type": "blob", "sha": "455aaf688066ad3e92f882adcfde7a6f11b03353", "size": 2573, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/455aaf688066ad3e92f882adcfde7a6f11b03353"}, {"path": "deepcompressor/app/diffusion/eval/metrics/run.py", "mode": "100644", "type": "blob", "sha": "782b79961c449da4ecf3ea534df5a3272155a3a8", "size": 940, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/782b79961c449da4ecf3ea534df5a3272155a3a8"}, {"path": "deepcompressor/app/diffusion/eval/metrics/similarity.py", "mode": "100644", "type": "blob", "sha": "b93adb82173b66be1924f59f2b86cd0cec360f4e", "size": 3943, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/b93adb82173b66be1924f59f2b86cd0cec360f4e"}, {"path": "deepcompressor/app/diffusion/nn", "mode": "040000", "type": "tree", "sha": "875e938e213de6372565d421fccba36413a08a1a", "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/trees/875e938e213de6372565d421fccba36413a08a1a"}, {"path": "deepcompressor/app/diffusion/nn/__init__.py", "mode": "100644", "type": "blob", "sha": "40a96afc6ff09d58a702b76e3f7dd412fe975e26", "size": 24, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/40a96afc6ff09d58a702b76e3f7dd412fe975e26"}, {"path": "deepcompressor/app/diffusion/nn/attention.py", "mode": "100644", "type": "blob", "sha": "6f29a9fab2ddcbab17d7f227d10582b6923159e6", "size": 7213, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/6f29a9fab2ddcbab17d7f227d10582b6923159e6"}, {"path": "deepcompressor/app/diffusion/nn/patch.py", "mode": "100644", "type": "blob", "sha": "a39ff406899565d9f7577b9ea2b2139589e92c8e", "size": 6069, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/a39ff406899565d9f7577b9ea2b2139589e92c8e"}, {"path": "deepcompressor/app/diffusion/nn/struct.py", "mode": "100644", "type": "blob", "sha": "a25c4f4d1438d502cc7899af94bde2c3394948e5", "size": 87134, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/a25c4f4d1438d502cc7899af94bde2c3394948e5"}, {"path": "deepcompressor/app/diffusion/pipeline", "mode": "040000", "type": "tree", "sha": "7a29d0eb36864e87fde1f73a57202b245f662cd9", "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/trees/7a29d0eb36864e87fde1f73a57202b245f662cd9"}, {"path": "deepcompressor/app/diffusion/pipeline/__init__.py", "mode": "100644", "type": "blob", "sha": "d667085cd8cfd9a393aac21d7d69a2b35767ce8a", "size": 69, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/d667085cd8cfd9a393aac21d7d69a2b35767ce8a"}, {"path": "deepcompressor/app/diffusion/pipeline/config.py", "mode": "100644", "type": "blob", "sha": "bc7cbe204016860835937ba9f317108c296a97c4", "size": 18410, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/bc7cbe204016860835937ba9f317108c296a97c4"}, {"path": "deepcompressor/app/diffusion/ptq.py", "mode": "100644", "type": "blob", "sha": "74c1c173739359b19dab978446ebf7b4cf17c483", "size": 17924, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/74c1c173739359b19dab978446ebf7b4cf17c483"}, {"path": "deepcompressor/app/diffusion/quant", "mode": "040000", "type": "tree", "sha": "e5f01d78bbbcbc9158623958cc24285dc7e2a881", "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/trees/e5f01d78bbbcbc9158623958cc24285dc7e2a881"}, {"path": "deepcompressor/app/diffusion/quant/__init__.py", "mode": "100644", "type": "blob", "sha": "3dda2a1e02ddc1aeab89067cd72a303b5e227ec0", "size": 382, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/3dda2a1e02ddc1aeab89067cd72a303b5e227ec0"}, {"path": "deepcompressor/app/diffusion/quant/activation.py", "mode": "100644", "type": "blob", "sha": "7e0dd27dc40ec8d5f66868c76954eccb102a8501", "size": 13170, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/7e0dd27dc40ec8d5f66868c76954eccb102a8501"}, {"path": "deepcompressor/app/diffusion/quant/config.py", "mode": "100644", "type": "blob", "sha": "54a51de7a03139c804523ce710c083966d47aacf", "size": 24884, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/54a51de7a03139c804523ce710c083966d47aacf"}, {"path": "deepcompressor/app/diffusion/quant/quantizer", "mode": "040000", "type": "tree", "sha": "6f786a57734c1425e470c341e0ec55c4eac3da16", "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/trees/6f786a57734c1425e470c341e0ec55c4eac3da16"}, {"path": "deepcompressor/app/diffusion/quant/quantizer/__init__.py", "mode": "100644", "type": "blob", "sha": "cc8153fa65eb683511cca05f21657be60741832c", "size": 154, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/cc8153fa65eb683511cca05f21657be60741832c"}, {"path": "deepcompressor/app/diffusion/quant/quantizer/config.py", "mode": "100644", "type": "blob", "sha": "3c3f7e25d1a1cfd33cebe3f1ff046f61ce321184", "size": 17573, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/3c3f7e25d1a1cfd33cebe3f1ff046f61ce321184"}, {"path": "deepcompressor/app/diffusion/quant/quantizer/quantizer.py", "mode": "100644", "type": "blob", "sha": "8103ca92798fa5269b8668623bdbbc6376cc0033", "size": 14274, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/8103ca92798fa5269b8668623bdbbc6376cc0033"}, {"path": "deepcompressor/app/diffusion/quant/rotate.py", "mode": "100644", "type": "blob", "sha": "718d38d3ac76a1862a6531c7fd6316885095adbf", "size": 6221, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/718d38d3ac76a1862a6531c7fd6316885095adbf"}, {"path": "deepcompressor/app/diffusion/quant/smooth.py", "mode": "100644", "type": "blob", "sha": "22f7b12477bdf80490768f9c7d640fa1483d7dad", "size": 30633, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/22f7b12477bdf80490768f9c7d640fa1483d7dad"}, {"path": "deepcompressor/app/diffusion/quant/utils.py", "mode": "100644", "type": "blob", "sha": "b2c329d4f77a7208714a8d95f1d4d075134c0c49", "size": 3895, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/b2c329d4f77a7208714a8d95f1d4d075134c0c49"}, {"path": "deepcompressor/app/diffusion/quant/weight.py", "mode": "100644", "type": "blob", "sha": "c9a046dec0af96dfe36d59d881423bba16ed37db", "size": 22087, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/c9a046dec0af96dfe36d59d881423bba16ed37db"}, {"path": "deepcompressor/app/diffusion/utils.py", "mode": "100644", "type": "blob", "sha": "dae2565eafede3782bc94f7746dbcc28d56e9288", "size": 6090, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/dae2565eafede3782bc94f7746dbcc28d56e9288"}, {"path": "deepcompressor/app/llm", "mode": "040000", "type": "tree", "sha": "8f5e41eab12613ed82aa87dddc0504bf6333990f", "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/trees/8f5e41eab12613ed82aa87dddc0504bf6333990f"}, {"path": "deepcompressor/app/llm/__init__.py", "mode": "100644", "type": "blob", "sha": "40a96afc6ff09d58a702b76e3f7dd412fe975e26", "size": 24, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/40a96afc6ff09d58a702b76e3f7dd412fe975e26"}, {"path": "deepcompressor/app/llm/cache", "mode": "040000", "type": "tree", "sha": "8a338e6207f1f9ad969b289f7c4fc4c70b5b66ce", "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/trees/8a338e6207f1f9ad969b289f7c4fc4c70b5b66ce"}, {"path": "deepcompressor/app/llm/cache/__init__.py", "mode": "100644", "type": "blob", "sha": "40a96afc6ff09d58a702b76e3f7dd412fe975e26", "size": 24, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/40a96afc6ff09d58a702b76e3f7dd412fe975e26"}, {"path": "deepcompressor/app/llm/cache/config.py", "mode": "100644", "type": "blob", "sha": "d3bff454d949d767eb3a7b52240aa6c9f36f740a", "size": 1689, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/d3bff454d949d767eb3a7b52240aa6c9f36f740a"}, {"path": "deepcompressor/app/llm/config.py", "mode": "100644", "type": "blob", "sha": "0cb6967ec20b10f5311ec0a6a18486eb44acfb4c", "size": 4728, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/0cb6967ec20b10f5311ec0a6a18486eb44acfb4c"}, {"path": "deepcompressor/app/llm/eval", "mode": "040000", "type": "tree", "sha": "ecdee5de27a58efcdc9b8c8c7fc923c092f01260", "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/trees/ecdee5de27a58efcdc9b8c8c7fc923c092f01260"}, {"path": "deepcompressor/app/llm/eval/__init__.py", "mode": "100644", "type": "blob", "sha": "40a96afc6ff09d58a702b76e3f7dd412fe975e26", "size": 24, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/40a96afc6ff09d58a702b76e3f7dd412fe975e26"}, {"path": "deepcompressor/app/llm/eval/base.py", "mode": "100644", "type": "blob", "sha": "43cb96eca3fc72caaf9d19cf196af5e88397c108", "size": 694, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/43cb96eca3fc72caaf9d19cf196af5e88397c108"}, {"path": "deepcompressor/app/llm/eval/config.py", "mode": "100644", "type": "blob", "sha": "5e0555b2821f8ca3a12d476e48be8615286960ec", "size": 9091, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/5e0555b2821f8ca3a12d476e48be8615286960ec"}, {"path": "deepcompressor/app/llm/eval/custom.py", "mode": "100644", "type": "blob", "sha": "9636df26377f99056b5394370a7886836a8eff3f", "size": 3516, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/9636df26377f99056b5394370a7886836a8eff3f"}, {"path": "deepcompressor/app/llm/eval/lm_eval.py", "mode": "100644", "type": "blob", "sha": "7732938d30d898ecd028b7a065ee660eac766dda", "size": 1829, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/7732938d30d898ecd028b7a065ee660eac766dda"}, {"path": "deepcompressor/app/llm/eval/longbench", "mode": "040000", "type": "tree", "sha": "adf1e1bf352675375c3defcd5363f651eecada10", "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/trees/adf1e1bf352675375c3defcd5363f651eecada10"}, {"path": "deepcompressor/app/llm/eval/longbench/__init__.py", "mode": "100644", "type": "blob", "sha": "4a7816a643bdf48a29d95634111c987f88254798", "size": 54, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/4a7816a643bdf48a29d95634111c987f88254798"}, {"path": "deepcompressor/app/llm/eval/longbench/eval.py", "mode": "100644", "type": "blob", "sha": "59592263eafb306a4b33c4f8a05a19cccf14a4a3", "size": 13172, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/59592263eafb306a4b33c4f8a05a19cccf14a4a3"}, {"path": "deepcompressor/app/llm/eval/longbench/metrics.py", "mode": "100644", "type": "blob", "sha": "7bc4afe284047beb64a4b5037001876c07b02e43", "size": 5231, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/7bc4afe284047beb64a4b5037001876c07b02e43"}, {"path": "deepcompressor/app/llm/eval/longbench/task2prompt.json", "mode": "100644", "type": "blob", "sha": "1c85f6bc0f0df4e42131aa49a867797ee7043ecf", "size": 5437, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/1c85f6bc0f0df4e42131aa49a867797ee7043ecf"}, {"path": "deepcompressor/app/llm/model", "mode": "040000", "type": "tree", "sha": "60baf19a3770d5bd78015559503af895bd52f732", "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/trees/60baf19a3770d5bd78015559503af895bd52f732"}, {"path": "deepcompressor/app/llm/model/__init__.py", "mode": "100644", "type": "blob", "sha": "40a96afc6ff09d58a702b76e3f7dd412fe975e26", "size": 24, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/40a96afc6ff09d58a702b76e3f7dd412fe975e26"}, {"path": "deepcompressor/app/llm/model/config.py", "mode": "100644", "type": "blob", "sha": "6b26ce230ad3317040e9dd9d64b92d0d8f90d0f5", "size": 6553, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/6b26ce230ad3317040e9dd9d64b92d0d8f90d0f5"}, {"path": "deepcompressor/app/llm/nn", "mode": "040000", "type": "tree", "sha": "3da9394d78f2edbd1f6b8515e43b8ab223c15216", "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/trees/3da9394d78f2edbd1f6b8515e43b8ab223c15216"}, {"path": "deepcompressor/app/llm/nn/__init__.py", "mode": "100644", "type": "blob", "sha": "41a9268ea12db3c36786d1b9669f8cb7dfe6aa0c", "size": 109, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/41a9268ea12db3c36786d1b9669f8cb7dfe6aa0c"}, {"path": "deepcompressor/app/llm/nn/patch.py", "mode": "100644", "type": "blob", "sha": "cd007c37bbcc8e9a9b8477b01e95d1b0b5f77a56", "size": 8993, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/cd007c37bbcc8e9a9b8477b01e95d1b0b5f77a56"}, {"path": "deepcompressor/app/llm/nn/struct.py", "mode": "100644", "type": "blob", "sha": "85643c2ed170fef8ea3546698ce3717652d937a5", "size": 37512, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/85643c2ed170fef8ea3546698ce3717652d937a5"}, {"path": "deepcompressor/app/llm/ptq.py", "mode": "100644", "type": "blob", "sha": "0aa3691624c9292c59c119b16cbcfc1f5a052059", "size": 17880, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/0aa3691624c9292c59c119b16cbcfc1f5a052059"}, {"path": "deepcompressor/app/llm/quant", "mode": "040000", "type": "tree", "sha": "8d43749de42fe20fe4b5b081f4276252bd3e8e69", "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/trees/8d43749de42fe20fe4b5b081f4276252bd3e8e69"}, {"path": "deepcompressor/app/llm/quant/__init__.py", "mode": "100644", "type": "blob", "sha": "eb0936a9f4b6191ec27676e0bd667a686da01cf9", "size": 332, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/eb0936a9f4b6191ec27676e0bd667a686da01cf9"}, {"path": "deepcompressor/app/llm/quant/activation.py", "mode": "100644", "type": "blob", "sha": "71a9a9b3cc28955dea33d914a5ffa8d96a2e6310", "size": 11352, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/71a9a9b3cc28955dea33d914a5ffa8d96a2e6310"}, {"path": "deepcompressor/app/llm/quant/config.py", "mode": "100644", "type": "blob", "sha": "9c433328ee66d0b8d943cb94823ad1aecf4168fa", "size": 18564, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/9c433328ee66d0b8d943cb94823ad1aecf4168fa"}, {"path": "deepcompressor/app/llm/quant/dataset.py", "mode": "100644", "type": "blob", "sha": "6dfce864abf1cba9369abae78c341bfad978f003", "size": 13498, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/6dfce864abf1cba9369abae78c341bfad978f003"}, {"path": "deepcompressor/app/llm/quant/quantizer", "mode": "040000", "type": "tree", "sha": "9a05b08aaa46be3c54bf3a9738561bc55b6a6e72", "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/trees/9a05b08aaa46be3c54bf3a9738561bc55b6a6e72"}, {"path": "deepcompressor/app/llm/quant/quantizer/__init__.py", "mode": "100644", "type": "blob", "sha": "5c5acaafaefdf8037bd00a46c315f3b46ae0dd81", "size": 136, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/5c5acaafaefdf8037bd00a46c315f3b46ae0dd81"}, {"path": "deepcompressor/app/llm/quant/quantizer/config.py", "mode": "100644", "type": "blob", "sha": "fbe241a9c59e7b1d5f220675363834e40b9ddcbe", "size": 10156, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/fbe241a9c59e7b1d5f220675363834e40b9ddcbe"}, {"path": "deepcompressor/app/llm/quant/quantizer/quantizer.py", "mode": "100644", "type": "blob", "sha": "92aab0562072d5fa0186e440e91232a7876e187c", "size": 12056, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/92aab0562072d5fa0186e440e91232a7876e187c"}, {"path": "deepcompressor/app/llm/quant/reorder.py", "mode": "100644", "type": "blob", "sha": "7714196865183a7ce6f69aa1b1615275592a7396", "size": 16797, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/7714196865183a7ce6f69aa1b1615275592a7396"}, {"path": "deepcompressor/app/llm/quant/rotate.py", "mode": "100644", "type": "blob", "sha": "19dd92e9eeb9a0997d7cf4058251fa1a7df6fde7", "size": 8660, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/19dd92e9eeb9a0997d7cf4058251fa1a7df6fde7"}, {"path": "deepcompressor/app/llm/quant/smooth.py", "mode": "100644", "type": "blob", "sha": "de9d17fc9b7e48b890e75260fd08ed334361d872", "size": 10462, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/de9d17fc9b7e48b890e75260fd08ed334361d872"}, {"path": "deepcompressor/app/llm/quant/utils.py", "mode": "100644", "type": "blob", "sha": "3b731d71cd5906924b67ff2a85c01b03e8b12082", "size": 3073, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/3b731d71cd5906924b67ff2a85c01b03e8b12082"}, {"path": "deepcompressor/app/llm/quant/weight.py", "mode": "100644", "type": "blob", "sha": "a663c2ab32af8058620b261f5649d1edc7044e1a", "size": 7685, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/a663c2ab32af8058620b261f5649d1edc7044e1a"}, {"path": "deepcompressor/backend", "mode": "040000", "type": "tree", "sha": "ee98a1992cf8b1b126f92b0ec6b6e235ed99e26b", "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/trees/ee98a1992cf8b1b126f92b0ec6b6e235ed99e26b"}, {"path": "deepcompressor/backend/__init__.py", "mode": "100644", "type": "blob", "sha": "e69de29bb2d1d6434b8b29ae775ad8c2e48c5391", "size": 0, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/e69de29bb2d1d6434b8b29ae775ad8c2e48c5391"}, {"path": "deepcompressor/backend/nunchaku", "mode": "040000", "type": "tree", "sha": "32314dfab8a3f52cde7b5b3d891daf2c7cfe4364", "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/trees/32314dfab8a3f52cde7b5b3d891daf2c7cfe4364"}, {"path": "deepcompressor/backend/nunchaku/__init__.py", "mode": "100644", "type": "blob", "sha": "e69de29bb2d1d6434b8b29ae775ad8c2e48c5391", "size": 0, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/e69de29bb2d1d6434b8b29ae775ad8c2e48c5391"}, {"path": "deepcompressor/backend/nunchaku/convert.py", "mode": "100644", "type": "blob", "sha": "f667f3bbabae68b42de0eee854c87f5b3ab21460", "size": 18740, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/f667f3bbabae68b42de0eee854c87f5b3ab21460"}, {"path": "deepcompressor/backend/nunchaku/convert_lora.py", "mode": "100644", "type": "blob", "sha": "a5fe24521a41a04e1828d16d1fd59ef0f2ec2957", "size": 14185, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/a5fe24521a41a04e1828d16d1fd59ef0f2ec2957"}, {"path": "deepcompressor/backend/nunchaku/utils.py", "mode": "100644", "type": "blob", "sha": "4022512dad4bcbcef3424b1bc7c027f87e3ccd86", "size": 19739, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/4022512dad4bcbcef3424b1bc7c027f87e3ccd86"}, {"path": "deepcompressor/backend/qserve", "mode": "040000", "type": "tree", "sha": "02b59ea338a23a99f6c8a3bcca6e4f6b5e26731d", "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/trees/02b59ea338a23a99f6c8a3bcca6e4f6b5e26731d"}, {"path": "deepcompressor/backend/qserve/__init__.py", "mode": "100644", "type": "blob", "sha": "e69de29bb2d1d6434b8b29ae775ad8c2e48c5391", "size": 0, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/e69de29bb2d1d6434b8b29ae775ad8c2e48c5391"}, {"path": "deepcompressor/backend/qserve/convert.py", "mode": "100644", "type": "blob", "sha": "630cfabb2f705d906e1f8ad9ba071eac1343a1ea", "size": 9066, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/630cfabb2f705d906e1f8ad9ba071eac1343a1ea"}, {"path": "deepcompressor/backend/qserve/utils.py", "mode": "100644", "type": "blob", "sha": "181a4ec5c6d127e28933e197950acd5fb95c63a3", "size": 10249, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/181a4ec5c6d127e28933e197950acd5fb95c63a3"}, {"path": "deepcompressor/backend/tinychat", "mode": "040000", "type": "tree", "sha": "6b8b00228bedd322e9b2ffeed95fd665818a7cd0", "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/trees/6b8b00228bedd322e9b2ffeed95fd665818a7cd0"}, {"path": "deepcompressor/backend/tinychat/__init__.py", "mode": "100644", "type": "blob", "sha": "e69de29bb2d1d6434b8b29ae775ad8c2e48c5391", "size": 0, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/e69de29bb2d1d6434b8b29ae775ad8c2e48c5391"}, {"path": "deepcompressor/backend/tinychat/convert.py", "mode": "100644", "type": "blob", "sha": "47a898c9071a2f96b4056faa26206007ef432112", "size": 6941, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/47a898c9071a2f96b4056faa26206007ef432112"}, {"path": "deepcompressor/backend/tinychat/csrc", "mode": "040000", "type": "tree", "sha": "cde48527e5f3c61f0ca5c976467653b7454b7508", "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/trees/cde48527e5f3c61f0ca5c976467653b7454b7508"}, {"path": "deepcompressor/backend/tinychat/csrc/load.py", "mode": "100644", "type": "blob", "sha": "142f81ec68f30dc0f08ae519a542b87682cf85db", "size": 1040, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/142f81ec68f30dc0f08ae519a542b87682cf85db"}, {"path": "deepcompressor/backend/tinychat/csrc/pybind.cpp", "mode": "100644", "type": "blob", "sha": "34b52b62044d56cc76e0ac6e4ebfc4d516b509db", "size": 368, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/34b52b62044d56cc76e0ac6e4ebfc4d516b509db"}, {"path": "deepcompressor/backend/tinychat/csrc/quantization", "mode": "040000", "type": "tree", "sha": "0d7a9367565dd8fccd6d7abd5248276953ecac47", "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/trees/0d7a9367565dd8fccd6d7abd5248276953ecac47"}, {"path": "deepcompressor/backend/tinychat/csrc/quantization/dequantize.cuh", "mode": "100644", "type": "blob", "sha": "1d3018cb1020942ca1e43afbb2542fffd0a6c4a8", "size": 6201, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/1d3018cb1020942ca1e43afbb2542fffd0a6c4a8"}, {"path": "deepcompressor/backend/tinychat/csrc/quantization/gemm", "mode": "040000", "type": "tree", "sha": "e50e4693d5719fedd2182f121dcf1917f1b56355", "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/trees/e50e4693d5719fedd2182f121dcf1917f1b56355"}, {"path": "deepcompressor/backend/tinychat/csrc/quantization/gemm/gemm_cuda.cu", "mode": "100644", "type": "blob", "sha": "851d24b705aa0b944bad7ba41279357cb54428db", "size": 55402, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/851d24b705aa0b944bad7ba41279357cb54428db"}, {"path": "deepcompressor/backend/tinychat/csrc/quantization/gemm/gemm_cuda.h", "mode": "100644", "type": "blob", "sha": "58b2625d27b4cfb9f13d6c2eaecf3d8626b394ca", "size": 177, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/58b2625d27b4cfb9f13d6c2eaecf3d8626b394ca"}, {"path": "deepcompressor/backend/tinychat/csrc/quantization/gemm/semaphore.h", "mode": "100644", "type": "blob", "sha": "acc636f745c53fec08521c7ec863b5d1baf675f4", "size": 3886, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/acc636f745c53fec08521c7ec863b5d1baf675f4"}, {"path": "deepcompressor/backend/tinychat/csrc/quantization/gemv", "mode": "040000", "type": "tree", "sha": "6e5df4e74f48cf9758a3a47499f6f746f41fdce2", "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/trees/6e5df4e74f48cf9758a3a47499f6f746f41fdce2"}, {"path": "deepcompressor/backend/tinychat/csrc/quantization/gemv/gemv_cuda.cu", "mode": "100644", "type": "blob", "sha": "0307321db6d8fa1730146cdc2d9fbd642a6903cc", "size": 11067, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/0307321db6d8fa1730146cdc2d9fbd642a6903cc"}, {"path": "deepcompressor/backend/tinychat/csrc/quantization/gemv/gemv_cuda.h", "mode": "100644", "type": "blob", "sha": "5600040525dde6865444bee13cceebb4e88b49d5", "size": 252, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/5600040525dde6865444bee13cceebb4e88b49d5"}, {"path": "deepcompressor/backend/tinychat/csrc/utils.cuh", "mode": "100644", "type": "blob", "sha": "8414d7c78269d1741744ed83425afc2680cdd914", "size": 11205, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/8414d7c78269d1741744ed83425afc2680cdd914"}, {"path": "deepcompressor/backend/tinychat/linear.py", "mode": "100644", "type": "blob", "sha": "719086d460444f2d417f69d40cb7df8142ab5c61", "size": 6593, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/719086d460444f2d417f69d40cb7df8142ab5c61"}, {"path": "deepcompressor/backend/tinychat/utils.py", "mode": "100644", "type": "blob", "sha": "0d03d4a93334dd266ab5610eb4707cf3067f0071", "size": 4367, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/0d03d4a93334dd266ab5610eb4707cf3067f0071"}, {"path": "deepcompressor/backend/utils.py", "mode": "100644", "type": "blob", "sha": "9c3fe12739ef7f85a5224ad750cfcd3e799fe21a", "size": 6038, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/9c3fe12739ef7f85a5224ad750cfcd3e799fe21a"}, {"path": "deepcompressor/calib", "mode": "040000", "type": "tree", "sha": "d5e4e996255b979ad7d890c7dc6399cb8f2ece67", "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/trees/d5e4e996255b979ad7d890c7dc6399cb8f2ece67"}, {"path": "deepcompressor/calib/__init__.py", "mode": "100644", "type": "blob", "sha": "40a96afc6ff09d58a702b76e3f7dd412fe975e26", "size": 24, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/40a96afc6ff09d58a702b76e3f7dd412fe975e26"}, {"path": "deepcompressor/calib/config", "mode": "040000", "type": "tree", "sha": "381628957f7b622befd5d0e59a975208f30e0f69", "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/trees/381628957f7b622befd5d0e59a975208f30e0f69"}, {"path": "deepcompressor/calib/config/__init__.py", "mode": "100644", "type": "blob", "sha": "fb9abf45c595c5b56ff13dc6a2057bd910ccb601", "size": 549, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/fb9abf45c595c5b56ff13dc6a2057bd910ccb601"}, {"path": "deepcompressor/calib/config/lowrank.py", "mode": "100644", "type": "blob", "sha": "4839e78477bb47203f12970dfae49d2fda217db8", "size": 4647, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/4839e78477bb47203f12970dfae49d2fda217db8"}, {"path": "deepcompressor/calib/config/range.py", "mode": "100644", "type": "blob", "sha": "70e5d7b1fc4fadfda8d07d5e95f8a7c0b1cdfb43", "size": 6681, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/70e5d7b1fc4fadfda8d07d5e95f8a7c0b1cdfb43"}, {"path": "deepcompressor/calib/config/reorder.py", "mode": "100644", "type": "blob", "sha": "fd64073a4c6f00e38fa541c018cf2ef1b3164dbb", "size": 5582, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/fd64073a4c6f00e38fa541c018cf2ef1b3164dbb"}, {"path": "deepcompressor/calib/config/rotation.py", "mode": "100644", "type": "blob", "sha": "ab6f911503c35151fe72d4a8bc9c0c1d9ebcd932", "size": 3214, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/ab6f911503c35151fe72d4a8bc9c0c1d9ebcd932"}, {"path": "deepcompressor/calib/config/search.py", "mode": "100644", "type": "blob", "sha": "a59cee7f25ab11c693a691355a13d45255174b32", "size": 4780, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/a59cee7f25ab11c693a691355a13d45255174b32"}, {"path": "deepcompressor/calib/config/smooth.py", "mode": "100644", "type": "blob", "sha": "6b8e2feed8bbcf1c7f7756e36a9bb3c5fafca7aa", "size": 17516, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/6b8e2feed8bbcf1c7f7756e36a9bb3c5fafca7aa"}, {"path": "deepcompressor/calib/lowrank.py", "mode": "100644", "type": "blob", "sha": "960423e8726e33694523d88f5d1a26e63e0e6df7", "size": 9693, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/960423e8726e33694523d88f5d1a26e63e0e6df7"}, {"path": "deepcompressor/calib/metric.py", "mode": "100644", "type": "blob", "sha": "5ab01eaa5c3bfe2cbc13fff61ce0fb1c06e56f71", "size": 7350, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/5ab01eaa5c3bfe2cbc13fff61ce0fb1c06e56f71"}, {"path": "deepcompressor/calib/range.py", "mode": "100644", "type": "blob", "sha": "5a395e7f04c347f6d6e5a115d74bfa2f94b3c151", "size": 21119, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/5a395e7f04c347f6d6e5a115d74bfa2f94b3c151"}, {"path": "deepcompressor/calib/reorder.py", "mode": "100644", "type": "blob", "sha": "9a0b0093dbc23e42e01a2a35df68f20e1f6ecc31", "size": 21619, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/9a0b0093dbc23e42e01a2a35df68f20e1f6ecc31"}, {"path": "deepcompressor/calib/rotate.py", "mode": "100644", "type": "blob", "sha": "2e587c0740eedfc9fe6b994e6dcae33aefb4ac6e", "size": 10238, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/2e587c0740eedfc9fe6b994e6dcae33aefb4ac6e"}, {"path": "deepcompressor/calib/search.py", "mode": "100644", "type": "blob", "sha": "28ea6219383681ec8d2b1d984304c51605ea8b94", "size": 51037, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/28ea6219383681ec8d2b1d984304c51605ea8b94"}, {"path": "deepcompressor/calib/smooth.py", "mode": "100644", "type": "blob", "sha": "876daa10a4d60b3cef707234c39a6106e517bb93", "size": 47439, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/876daa10a4d60b3cef707234c39a6106e517bb93"}, {"path": "deepcompressor/csrc", "mode": "040000", "type": "tree", "sha": "8aeab721f68a0bf62102446fc07c749c380cef7b", "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/trees/8aeab721f68a0bf62102446fc07c749c380cef7b"}, {"path": "deepcompressor/csrc/load.py", "mode": "100644", "type": "blob", "sha": "6aae3f1b64847ce8bba8786542957aca6a3bf2ed", "size": 951, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/6aae3f1b64847ce8bba8786542957aca6a3bf2ed"}, {"path": "deepcompressor/csrc/pybind.cpp", "mode": "100644", "type": "blob", "sha": "37b11320a104dfee8856b1b34deaa9281090ad90", "size": 356, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/37b11320a104dfee8856b1b34deaa9281090ad90"}, {"path": "deepcompressor/csrc/quantize", "mode": "040000", "type": "tree", "sha": "32a8572ed60e091d7a22b8a3aa1ec925dc9d18ba", "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/trees/32a8572ed60e091d7a22b8a3aa1ec925dc9d18ba"}, {"path": "deepcompressor/csrc/quantize/quantize.cu", "mode": "100644", "type": "blob", "sha": "d2331d25b07d33a909d0c6a4995a96470274eefd", "size": 4126, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/d2331d25b07d33a909d0c6a4995a96470274eefd"}, {"path": "deepcompressor/csrc/quantize/quantize.h", "mode": "100644", "type": "blob", "sha": "06869abbf80d876844615575af96a0167121e402", "size": 322, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/06869abbf80d876844615575af96a0167121e402"}, {"path": "deepcompressor/data", "mode": "040000", "type": "tree", "sha": "329951999682163729b32c8ca6ca2a1a153bb1ee", "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/trees/329951999682163729b32c8ca6ca2a1a153bb1ee"}, {"path": "deepcompressor/data/__init__.py", "mode": "100644", "type": "blob", "sha": "1b91235212261b8bd99c8ca5d674775012fcf405", "size": 199, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/1b91235212261b8bd99c8ca5d674775012fcf405"}, {"path": "deepcompressor/data/cache.py", "mode": "100644", "type": "blob", "sha": "79b5fd3be4da6296a19860bf95a8ee9b800bf086", "size": 9875, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/79b5fd3be4da6296a19860bf95a8ee9b800bf086"}, {"path": "deepcompressor/data/codebook.py", "mode": "100644", "type": "blob", "sha": "6c0d5a5e3469e7f050ecbe0a7fed4c4cc09360d9", "size": 7616, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/6c0d5a5e3469e7f050ecbe0a7fed4c4cc09360d9"}, {"path": "deepcompressor/data/common.py", "mode": "100644", "type": "blob", "sha": "747bc59077ba9d7c6dacd55e4aa59ab7b4fa3294", "size": 230, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/747bc59077ba9d7c6dacd55e4aa59ab7b4fa3294"}, {"path": "deepcompressor/data/dtype.py", "mode": "100644", "type": "blob", "sha": "8ab64e1bba00633956f8468c195f3c47c2bb726b", "size": 16225, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/8ab64e1bba00633956f8468c195f3c47c2bb726b"}, {"path": "deepcompressor/data/range.py", "mode": "100644", "type": "blob", "sha": "d0f0d5d75ea09615fd9993c9f30e29f4ba44c610", "size": 18548, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/d0f0d5d75ea09615fd9993c9f30e29f4ba44c610"}, {"path": "deepcompressor/data/scale.py", "mode": "100644", "type": "blob", "sha": "25e003cc6f4c59f282b461c6d91ce827ad009fc6", "size": 4272, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/25e003cc6f4c59f282b461c6d91ce827ad009fc6"}, {"path": "deepcompressor/data/tensor.py", "mode": "100644", "type": "blob", "sha": "d32607796bbd1c3d0f64687eac2c41b133353912", "size": 1252, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/d32607796bbd1c3d0f64687eac2c41b133353912"}, {"path": "deepcompressor/data/utils", "mode": "040000", "type": "tree", "sha": "02e84b568065dc61ca963908c7b388a2f8ab4819", "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/trees/02e84b568065dc61ca963908c7b388a2f8ab4819"}, {"path": "deepcompressor/data/utils/__init__.py", "mode": "100644", "type": "blob", "sha": "04b502bc23d54a45d2ccabe55267e5c8cee1d793", "size": 127, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/04b502bc23d54a45d2ccabe55267e5c8cee1d793"}, {"path": "deepcompressor/data/utils/dtype.py", "mode": "100644", "type": "blob", "sha": "8b9a0eb3c13035e2662ba6a16a693a7b0d10bfc7", "size": 3494, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/8b9a0eb3c13035e2662ba6a16a693a7b0d10bfc7"}, {"path": "deepcompressor/data/utils/reshape.py", "mode": "100644", "type": "blob", "sha": "24227ed9305e89728b9f1a7bc876f015aea62262", "size": 4821, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/24227ed9305e89728b9f1a7bc876f015aea62262"}, {"path": "deepcompressor/data/utils/scale.py", "mode": "100644", "type": "blob", "sha": "335ca0cdad88f942606f09897937b7280caa8ef8", "size": 1853, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/335ca0cdad88f942606f09897937b7280caa8ef8"}, {"path": "deepcompressor/data/utils/shape.py", "mode": "100644", "type": "blob", "sha": "4e43244eb5b3d37ccd0d48c0a5dca443aba43a4d", "size": 7996, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/4e43244eb5b3d37ccd0d48c0a5dca443aba43a4d"}, {"path": "deepcompressor/data/zero.py", "mode": "100644", "type": "blob", "sha": "b00d92ddeacb87daedffebc64317d2d0fc6f09b2", "size": 224, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/b00d92ddeacb87daedffebc64317d2d0fc6f09b2"}, {"path": "deepcompressor/dataset", "mode": "040000", "type": "tree", "sha": "b5d7bc6b57d423b1a872dc79157b5e2915555b91", "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/trees/b5d7bc6b57d423b1a872dc79157b5e2915555b91"}, {"path": "deepcompressor/dataset/__init__.py", "mode": "100644", "type": "blob", "sha": "653cd8d4b3d36f825fb9d594f49ebb26ce2154d6", "size": 157, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/653cd8d4b3d36f825fb9d594f49ebb26ce2154d6"}, {"path": "deepcompressor/dataset/action.py", "mode": "100644", "type": "blob", "sha": "298f649c0a618c86d0a30eafe5bc023d7ec71f2b", "size": 7987, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/298f649c0a618c86d0a30eafe5bc023d7ec71f2b"}, {"path": "deepcompressor/dataset/cache.py", "mode": "100644", "type": "blob", "sha": "7d63d48ec98dcb484e830703e58d8fec9906c34d", "size": 21085, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/7d63d48ec98dcb484e830703e58d8fec9906c34d"}, {"path": "deepcompressor/dataset/config.py", "mode": "100644", "type": "blob", "sha": "4d880bc66b341d4ab8e48defe975e7def6512fbc", "size": 1404, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/4d880bc66b341d4ab8e48defe975e7def6512fbc"}, {"path": "deepcompressor/nn", "mode": "040000", "type": "tree", "sha": "f8fcefef97ae24b9bd529088d9051e8526b43f85", "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/trees/f8fcefef97ae24b9bd529088d9051e8526b43f85"}, {"path": "deepcompressor/nn/__init__.py", "mode": "100644", "type": "blob", "sha": "40a96afc6ff09d58a702b76e3f7dd412fe975e26", "size": 24, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/40a96afc6ff09d58a702b76e3f7dd412fe975e26"}, {"path": "deepcompressor/nn/patch", "mode": "040000", "type": "tree", "sha": "696630f6468df6aff0b67452ad11954ee419316c", "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/trees/696630f6468df6aff0b67452ad11954ee419316c"}, {"path": "deepcompressor/nn/patch/__init__.py", "mode": "100644", "type": "blob", "sha": "acdbee736db54e6528da3cfa1c14dd080546754e", "size": 110, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/acdbee736db54e6528da3cfa1c14dd080546754e"}, {"path": "deepcompressor/nn/patch/conv.py", "mode": "100644", "type": "blob", "sha": "d094b4f1eeea07c6657db088131a35ff5089858d", "size": 6808, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/d094b4f1eeea07c6657db088131a35ff5089858d"}, {"path": "deepcompressor/nn/patch/linear.py", "mode": "100644", "type": "blob", "sha": "08570ea70657f257d2465b01eb72cec733899d88", "size": 4794, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/08570ea70657f257d2465b01eb72cec733899d88"}, {"path": "deepcompressor/nn/patch/lowrank.py", "mode": "100644", "type": "blob", "sha": "624774116d37b4a5d42299d4af85b04a24a32279", "size": 4095, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/624774116d37b4a5d42299d4af85b04a24a32279"}, {"path": "deepcompressor/nn/patch/sdpa.py", "mode": "100644", "type": "blob", "sha": "bf1ce83974c5bc76dc96e8332bd730cf716771af", "size": 691, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/bf1ce83974c5bc76dc96e8332bd730cf716771af"}, {"path": "deepcompressor/nn/struct", "mode": "040000", "type": "tree", "sha": "3de8d52147ef3d8922d12ae3a1b1e991079f8172", "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/trees/3de8d52147ef3d8922d12ae3a1b1e991079f8172"}, {"path": "deepcompressor/nn/struct/__init__.py", "mode": "100644", "type": "blob", "sha": "88078120ef8ec8210b3e1ca700eb3c2712bf2bac", "size": 65, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/88078120ef8ec8210b3e1ca700eb3c2712bf2bac"}, {"path": "deepcompressor/nn/struct/attn.py", "mode": "100644", "type": "blob", "sha": "0a63aa69f1f9876e197353588b9bf6bb6124f13d", "size": 32134, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/0a63aa69f1f9876e197353588b9bf6bb6124f13d"}, {"path": "deepcompressor/nn/struct/base.py", "mode": "100644", "type": "blob", "sha": "65526ecd13d06e467eb824c2302e82b9f4744891", "size": 5935, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/65526ecd13d06e467eb824c2302e82b9f4744891"}, {"path": "deepcompressor/quantizer", "mode": "040000", "type": "tree", "sha": "3b7c6d6cc38c4927963daa55a6061d90c17cd8fe", "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/trees/3b7c6d6cc38c4927963daa55a6061d90c17cd8fe"}, {"path": "deepcompressor/quantizer/__init__.py", "mode": "100644", "type": "blob", "sha": "feb560436cc31034e54adab84daf34bcd52aa0fb", "size": 58, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/feb560436cc31034e54adab84daf34bcd52aa0fb"}, {"path": "deepcompressor/quantizer/config", "mode": "040000", "type": "tree", "sha": "9652550e512386bdb6830ec7b21a4295e21aa8a7", "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/trees/9652550e512386bdb6830ec7b21a4295e21aa8a7"}, {"path": "deepcompressor/quantizer/config/__init__.py", "mode": "100644", "type": "blob", "sha": "7ad8d0f767ba77544a866cd2440a44dac9c17a80", "size": 249, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/7ad8d0f767ba77544a866cd2440a44dac9c17a80"}, {"path": "deepcompressor/quantizer/config/base.py", "mode": "100644", "type": "blob", "sha": "95cfb2f09f3a99e4c9386625c16d8abed05966c1", "size": 15332, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/95cfb2f09f3a99e4c9386625c16d8abed05966c1"}, {"path": "deepcompressor/quantizer/config/kernel.py", "mode": "100644", "type": "blob", "sha": "5db9d41d0a42138847348e40b699c2547e109165", "size": 5232, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/5db9d41d0a42138847348e40b699c2547e109165"}, {"path": "deepcompressor/quantizer/config/lowrank.py", "mode": "100644", "type": "blob", "sha": "81a459eca1965838970109824374144e5ba76a58", "size": 1358, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/81a459eca1965838970109824374144e5ba76a58"}, {"path": "deepcompressor/quantizer/impl", "mode": "040000", "type": "tree", "sha": "3ac20dc6e7456445af8a752e388f9d751ef25553", "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/trees/3ac20dc6e7456445af8a752e388f9d751ef25553"}, {"path": "deepcompressor/quantizer/impl/__init__.py", "mode": "100644", "type": "blob", "sha": "e69de29bb2d1d6434b8b29ae775ad8c2e48c5391", "size": 0, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/e69de29bb2d1d6434b8b29ae775ad8c2e48c5391"}, {"path": "deepcompressor/quantizer/impl/base.py", "mode": "100644", "type": "blob", "sha": "2b26f0333f980a58f8ef3db04d92ea6ccc13aee9", "size": 14507, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/2b26f0333f980a58f8ef3db04d92ea6ccc13aee9"}, {"path": "deepcompressor/quantizer/impl/info.py", "mode": "100644", "type": "blob", "sha": "d9207866484836ec4e23904ffc2dc7786cf1eb86", "size": 6728, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/d9207866484836ec4e23904ffc2dc7786cf1eb86"}, {"path": "deepcompressor/quantizer/impl/scale.py", "mode": "100644", "type": "blob", "sha": "cb09af080bd48282c66f1819b72d40541816922a", "size": 11502, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/cb09af080bd48282c66f1819b72d40541816922a"}, {"path": "deepcompressor/quantizer/impl/simple.py", "mode": "100644", "type": "blob", "sha": "4173ece3796d9fb31fd285a13fb0f147f263840f", "size": 2459, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/4173ece3796d9fb31fd285a13fb0f147f263840f"}, {"path": "deepcompressor/quantizer/impl/ste.py", "mode": "100644", "type": "blob", "sha": "9e1f39d8d9ee51223be07d8c9ddd5a19a6ae83e5", "size": 779, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/9e1f39d8d9ee51223be07d8c9ddd5a19a6ae83e5"}, {"path": "deepcompressor/quantizer/kernel", "mode": "040000", "type": "tree", "sha": "26761faa25e64fb11e4ac0b36ed6f81add20d8eb", "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/trees/26761faa25e64fb11e4ac0b36ed6f81add20d8eb"}, {"path": "deepcompressor/quantizer/kernel/__init__.py", "mode": "100644", "type": "blob", "sha": "f891180c692097d56455bdea6322937d983a1be5", "size": 137, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/f891180c692097d56455bdea6322937d983a1be5"}, {"path": "deepcompressor/quantizer/kernel/gptq.py", "mode": "100644", "type": "blob", "sha": "274a9badb2b54ab0cba6c3ecd5f4babdbbd9e294", "size": 12177, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/274a9badb2b54ab0cba6c3ecd5f4babdbbd9e294"}, {"path": "deepcompressor/quantizer/kernel/rtn.py", "mode": "100644", "type": "blob", "sha": "587b2f77c9cad2f640234d87b786054faf2f1264", "size": 3786, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/587b2f77c9cad2f640234d87b786054faf2f1264"}, {"path": "deepcompressor/quantizer/processor.py", "mode": "100644", "type": "blob", "sha": "36a098b0240ae3452bfe570310eaec513785ccb4", "size": 17147, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/36a098b0240ae3452bfe570310eaec513785ccb4"}, {"path": "deepcompressor/utils", "mode": "040000", "type": "tree", "sha": "156153c2efd5c16cd805a9e669c49f5bc06e1d71", "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/trees/156153c2efd5c16cd805a9e669c49f5bc06e1d71"}, {"path": "deepcompressor/utils/__init__.py", "mode": "100644", "type": "blob", "sha": "91bee6090d7570e3ae436afab268a6f7f3dd13a1", "size": 68, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/91bee6090d7570e3ae436afab268a6f7f3dd13a1"}, {"path": "deepcompressor/utils/common.py", "mode": "100644", "type": "blob", "sha": "8040a9af2746b3c3c23d1333afd6f585fc549291", "size": 6556, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/8040a9af2746b3c3c23d1333afd6f585fc549291"}, {"path": "deepcompressor/utils/config", "mode": "040000", "type": "tree", "sha": "e16ac626a97743fdc13461d19f1aaa1a36fddd29", "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/trees/e16ac626a97743fdc13461d19f1aaa1a36fddd29"}, {"path": "deepcompressor/utils/config/__init__.py", "mode": "100644", "type": "blob", "sha": "a10f95b67e0259dff0fa60c0b15d84be9e957188", "size": 85, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/a10f95b67e0259dff0fa60c0b15d84be9e957188"}, {"path": "deepcompressor/utils/config/base.py", "mode": "100644", "type": "blob", "sha": "7b29a2a2785a883da5506ab183a3a3c45eda1539", "size": 7204, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/7b29a2a2785a883da5506ab183a3a3c45eda1539"}, {"path": "deepcompressor/utils/config/model.py", "mode": "100644", "type": "blob", "sha": "62aad76eda75e9fd7660350edc037a195bf77647", "size": 1619, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/62aad76eda75e9fd7660350edc037a195bf77647"}, {"path": "deepcompressor/utils/config/output.py", "mode": "100644", "type": "blob", "sha": "e2384ed2e1e25274a1e12822c553628af803d93b", "size": 3874, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/e2384ed2e1e25274a1e12822c553628af803d93b"}, {"path": "deepcompressor/utils/config/path.py", "mode": "100644", "type": "blob", "sha": "66615a973dcd836b4cdd7d0ed37374f51b7ee600", "size": 2690, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/66615a973dcd836b4cdd7d0ed37374f51b7ee600"}, {"path": "deepcompressor/utils/dataclass.py", "mode": "100644", "type": "blob", "sha": "93fd47c9e499188d30fc5c566dd2d2c5e6cf7680", "size": 1066, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/93fd47c9e499188d30fc5c566dd2d2c5e6cf7680"}, {"path": "deepcompressor/utils/hooks", "mode": "040000", "type": "tree", "sha": "48454288b51f69c6ca0db00b55e9138742d09069", "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/trees/48454288b51f69c6ca0db00b55e9138742d09069"}, {"path": "deepcompressor/utils/hooks/__init__.py", "mode": "100644", "type": "blob", "sha": "b4d2b077ba28417a8cbb7623d2c498e01881c340", "size": 331, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/b4d2b077ba28417a8cbb7623d2c498e01881c340"}, {"path": "deepcompressor/utils/hooks/branch.py", "mode": "100644", "type": "blob", "sha": "274e1979bf879d118c67b0f4b43f0db1ec60fa87", "size": 2256, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/274e1979bf879d118c67b0f4b43f0db1ec60fa87"}, {"path": "deepcompressor/utils/hooks/hook.py", "mode": "100644", "type": "blob", "sha": "d58d89fd37150234b3a1c4ca00643ea03e3f4ca3", "size": 7187, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/d58d89fd37150234b3a1c4ca00643ea03e3f4ca3"}, {"path": "deepcompressor/utils/hooks/packager.py", "mode": "100644", "type": "blob", "sha": "3b07a1bb7fbc555e3c0ab449eabfc307c212c534", "size": 9943, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/3b07a1bb7fbc555e3c0ab449eabfc307c212c534"}, {"path": "deepcompressor/utils/hooks/processor.py", "mode": "100644", "type": "blob", "sha": "714187e781955f1b5c203bc7912942e92b2bf14d", "size": 3139, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/714187e781955f1b5c203bc7912942e92b2bf14d"}, {"path": "deepcompressor/utils/math", "mode": "040000", "type": "tree", "sha": "90f2868e4db934fb30de80a76b08944ba888bd73", "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/trees/90f2868e4db934fb30de80a76b08944ba888bd73"}, {"path": "deepcompressor/utils/math/__init__.py", "mode": "100644", "type": "blob", "sha": "3e46d620aaef30f2c4b303a562516c5c61dba884", "size": 75, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/3e46d620aaef30f2c4b303a562516c5c61dba884"}, {"path": "deepcompressor/utils/math/functional.py", "mode": "100644", "type": "blob", "sha": "562c19f53b8fcbab4ab44f1166fc463ee9db98ca", "size": 740, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/562c19f53b8fcbab4ab44f1166fc463ee9db98ca"}, {"path": "deepcompressor/utils/math/hadamard.py", "mode": "100644", "type": "blob", "sha": "f5624457acb6f877c3698d20b0da35d252ee08ba", "size": 2321799, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/f5624457acb6f877c3698d20b0da35d252ee08ba"}, {"path": "deepcompressor/utils/patch.py", "mode": "100644", "type": "blob", "sha": "d6f19472295ae6038b390e008528d1b6f0c92e4c", "size": 2229, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/d6f19472295ae6038b390e008528d1b6f0c92e4c"}, {"path": "deepcompressor/utils/tools", "mode": "040000", "type": "tree", "sha": "61bca22341b9755ddda55ee130c66ea6fb43f69b", "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/trees/61bca22341b9755ddda55ee130c66ea6fb43f69b"}, {"path": "deepcompressor/utils/tools/__init__.py", "mode": "100644", "type": "blob", "sha": "7c202f598949d1a6e8a1b91553566d08560e03b5", "size": 52, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/7c202f598949d1a6e8a1b91553566d08560e03b5"}, {"path": "deepcompressor/utils/tools/logging.py", "mode": "100644", "type": "blob", "sha": "e39f015c46dc72404df29371ecee8f6a1a7870f2", "size": 6658, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/e39f015c46dc72404df29371ecee8f6a1a7870f2"}, {"path": "deepcompressor/utils/tools/sys.py", "mode": "100644", "type": "blob", "sha": "9401c0cf567ff8d57bf8dbb7d9d7fc4a80d0812f", "size": 1212, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/9401c0cf567ff8d57bf8dbb7d9d7fc4a80d0812f"}, {"path": "deepcompressor/version.py", "mode": "100644", "type": "blob", "sha": "42d9553f192d8dba06cdd1c914ca250fc1ddf2fd", "size": 74, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/42d9553f192d8dba06cdd1c914ca250fc1ddf2fd"}, {"path": "environment.yml", "mode": "100644", "type": "blob", "sha": "caac1dab7e5270601e3178a0716f3604eaa069b2", "size": 85, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/caac1dab7e5270601e3178a0716f3604eaa069b2"}, {"path": "examples", "mode": "040000", "type": "tree", "sha": "a0e48834ac8ff12dccbb8ae59bb0268212a98c7e", "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/trees/a0e48834ac8ff12dccbb8ae59bb0268212a98c7e"}, {"path": "examples/diffusion", "mode": "040000", "type": "tree", "sha": "af0d722ce92a44f8c40ac7734d13a62cc26c87a6", "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/trees/af0d722ce92a44f8c40ac7734d13a62cc26c87a6"}, {"path": "examples/diffusion/.gitignore", "mode": "100644", "type": "blob", "sha": "c52c4ac7569eb5483ad787c1a341c204aacd43b5", "size": 126, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/c52c4ac7569eb5483ad787c1a341c204aacd43b5"}, {"path": "examples/diffusion/README.md", "mode": "100644", "type": "blob", "sha": "bfd51960d5073a38a7b110fbb818b60bad007161", "size": 12907, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/bfd51960d5073a38a7b110fbb818b60bad007161"}, {"path": "examples/diffusion/configs", "mode": "040000", "type": "tree", "sha": "5195900d0cbcb18d443332e870bdf59c2e4e2306", "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/trees/5195900d0cbcb18d443332e870bdf59c2e4e2306"}, {"path": "examples/diffusion/configs/__default__.yaml", "mode": "100644", "type": "blob", "sha": "c0ad56cd22c6e4f2766f09a2fb8475503a292c91", "size": 2439, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/c0ad56cd22c6e4f2766f09a2fb8475503a292c91"}, {"path": "examples/diffusion/configs/collect", "mode": "040000", "type": "tree", "sha": "036d5f9f587b6be172b6d64c3d9fd278c55b2ce0", "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/trees/036d5f9f587b6be172b6d64c3d9fd278c55b2ce0"}, {"path": "examples/diffusion/configs/collect/qdiff.yaml", "mode": "100644", "type": "blob", "sha": "31426107cdcdeb15675d2ab0451ec689959eb6a5", "size": 98, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/31426107cdcdeb15675d2ab0451ec689959eb6a5"}, {"path": "examples/diffusion/configs/lora", "mode": "040000", "type": "tree", "sha": "89844e67b61f27d36827b5c2b7b282375fe81a9b", "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/trees/89844e67b61f27d36827b5c2b7b282375fe81a9b"}, {"path": "examples/diffusion/configs/lora/__default__.yaml", "mode": "100644", "type": "blob", "sha": "c33ebbe90b78ab32814f544868b9d5d91e1e0413", "size": 46, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/c33ebbe90b78ab32814f544868b9d5d91e1e0413"}, {"path": "examples/diffusion/configs/lora/flux.1-dev", "mode": "040000", "type": "tree", "sha": "a0bc7aec2f84161c74629ab21c18a1b331ee84f1", "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/trees/a0bc7aec2f84161c74629ab21c18a1b331ee84f1"}, {"path": "examples/diffusion/configs/lora/flux.1-dev/anime.yaml", "mode": "100644", "type": "blob", "sha": "d44abb2454c3c87899a22c1039f655345aee44a4", "size": 326, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/d44abb2454c3c87899a22c1039f655345aee44a4"}, {"path": "examples/diffusion/configs/lora/flux.1-dev/ghibsky.yaml", "mode": "100644", "type": "blob", "sha": "77519d990606bda5a197726151357a1d80df7eae", "size": 334, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/77519d990606bda5a197726151357a1d80df7eae"}, {"path": "examples/diffusion/configs/lora/flux.1-dev/realism.yaml", "mode": "100644", "type": "blob", "sha": "487fa5fced22d64984c116db25bf26b7257055cc", "size": 333, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/487fa5fced22d64984c116db25bf26b7257055cc"}, {"path": "examples/diffusion/configs/lora/flux.1-dev/sketch.yaml", "mode": "100644", "type": "blob", "sha": "eaf27ddb59bf06d5977852415be0957e2536bd08", "size": 410, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/eaf27ddb59bf06d5977852415be0957e2536bd08"}, {"path": "examples/diffusion/configs/lora/flux.1-dev/yarn.yaml", "mode": "100644", "type": "blob", "sha": "1c473c526ff72b24c4f0938cb59ba93d70c99b90", "size": 337, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/1c473c526ff72b24c4f0938cb59ba93d70c99b90"}, {"path": "examples/diffusion/configs/model", "mode": "040000", "type": "tree", "sha": "b8302e73b408087247e85461e6c1247a0d1ba3dd", "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/trees/b8302e73b408087247e85461e6c1247a0d1ba3dd"}, {"path": "examples/diffusion/configs/model/flux.1-dev.yaml", "mode": "100644", "type": "blob", "sha": "0a0ae8b24785e558c9436dc0264fb9b2a3c2144f", "size": 1181, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/0a0ae8b24785e558c9436dc0264fb9b2a3c2144f"}, {"path": "examples/diffusion/configs/model/flux.1-schnell.yaml", "mode": "100644", "type": "blob", "sha": "9355dca7d3e141c4c8121efae90f99fe62e0942b", "size": 1182, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/9355dca7d3e141c4c8121efae90f99fe62e0942b"}, {"path": "examples/diffusion/configs/model/pixart-sigma.yaml", "mode": "100644", "type": "blob", "sha": "b094c4b7509f3a41eecd53e0f62c1bfcec8336eb", "size": 921, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/b094c4b7509f3a41eecd53e0f62c1bfcec8336eb"}, {"path": "examples/diffusion/configs/model/sana-1.6b.yaml", "mode": "100644", "type": "blob", "sha": "129e8f869c8278cd89027be6cedf45f356fe2e03", "size": 1368, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/129e8f869c8278cd89027be6cedf45f356fe2e03"}, {"path": "examples/diffusion/configs/svdquant", "mode": "040000", "type": "tree", "sha": "a1b97da3d19e68f40d2dfa225f560de0da37db0e", "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/trees/a1b97da3d19e68f40d2dfa225f560de0da37db0e"}, {"path": "examples/diffusion/configs/svdquant/__default__.yaml", "mode": "100644", "type": "blob", "sha": "ec39459138662de8f2700d953b61c0084f6ebcad", "size": 849, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/ec39459138662de8f2700d953b61c0084f6ebcad"}, {"path": "examples/diffusion/configs/svdquant/fast.yaml", "mode": "100644", "type": "blob", "sha": "236b892a7e039fb6e4418a03dc97fee1e9f52994", "size": 75, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/236b892a7e039fb6e4418a03dc97fee1e9f52994"}, {"path": "examples/diffusion/configs/svdquant/gptq.yaml", "mode": "100644", "type": "blob", "sha": "bc9b7c63ed30d8f47f3ba305cdb0d7b87cb4b5e4", "size": 165, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/bc9b7c63ed30d8f47f3ba305cdb0d7b87cb4b5e4"}, {"path": "examples/diffusion/configs/svdquant/int4.yaml", "mode": "100644", "type": "blob", "sha": "626635b0145e6e938d29901f01f2691c5b2f0618", "size": 331, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/626635b0145e6e938d29901f01f2691c5b2f0618"}, {"path": "examples/diffusion/configs/svdquant/nvfp4.yaml", "mode": "100644", "type": "blob", "sha": "c839b98661a33106983239fb81f92cb7c73a0cb4", "size": 557, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/c839b98661a33106983239fb81f92cb7c73a0cb4"}, {"path": "examples/diffusion/configs/text", "mode": "040000", "type": "tree", "sha": "b60922a547a2a82a447a6e545fc7cec2965d1dbb", "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/trees/b60922a547a2a82a447a6e545fc7cec2965d1dbb"}, {"path": "examples/diffusion/configs/text/__default__.yaml", "mode": "100644", "type": "blob", "sha": "ef87b92782542c2897e58c0d53653d0ba8c4b723", "size": 3967, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/ef87b92782542c2897e58c0d53653d0ba8c4b723"}, {"path": "examples/diffusion/configs/text/awq.yaml", "mode": "100644", "type": "blob", "sha": "2f6f3049e966c097ab4078a9afd0a958f583a166", "size": 996, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/2f6f3049e966c097ab4078a9afd0a958f583a166"}, {"path": "examples/diffusion/prompts", "mode": "040000", "type": "tree", "sha": "84c4172d7bd28350c2fcb83fda2751c443e05085", "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/trees/84c4172d7bd28350c2fcb83fda2751c443e05085"}, {"path": "examples/diffusion/prompts/lora", "mode": "040000", "type": "tree", "sha": "a7ee172cfbfe164aa35d796fcb6b03330ffec32d", "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/trees/a7ee172cfbfe164aa35d796fcb6b03330ffec32d"}, {"path": "examples/diffusion/prompts/lora/anime.yaml", "mode": "100644", "type": "blob", "sha": "15830f40b60fd10b0471a1c1df30b6d972abe897", "size": 5664, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/15830f40b60fd10b0471a1c1df30b6d972abe897"}, {"path": "examples/diffusion/prompts/lora/ghibsky.yaml", "mode": "100644", "type": "blob", "sha": "c55905f6e6c498a875d6a20712c1029607aab2a5", "size": 13573, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/c55905f6e6c498a875d6a20712c1029607aab2a5"}, {"path": "examples/diffusion/prompts/lora/realism.yaml", "mode": "100644", "type": "blob", "sha": "ce7209dee287378cdbc52f63e68eb43e9be2f153", "size": 9995, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/ce7209dee287378cdbc52f63e68eb43e9be2f153"}, {"path": "examples/diffusion/prompts/lora/sketch.yaml", "mode": "100644", "type": "blob", "sha": "7c3313172387f6e58f9dec79347483641243723a", "size": 10469, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/7c3313172387f6e58f9dec79347483641243723a"}, {"path": "examples/diffusion/prompts/lora/yarn.yaml", "mode": "100644", "type": "blob", "sha": "d3b437aa67702c5d7b0794d51c9014bdda0759da", "size": 5429, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/d3b437aa67702c5d7b0794d51c9014bdda0759da"}, {"path": "examples/diffusion/prompts/qdiff.yaml", "mode": "100644", "type": "blob", "sha": "e8e0a3905b113e3de75f754fe88f98315cb5ff66", "size": 62384, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/e8e0a3905b113e3de75f754fe88f98315cb5ff66"}, {"path": "examples/diffusion/scripts", "mode": "040000", "type": "tree", "sha": "033088b10d75cf01ca975447f0eb23c799edaed0", "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/trees/033088b10d75cf01ca975447f0eb23c799edaed0"}, {"path": "examples/diffusion/scripts/svdquant.sh", "mode": "100644", "type": "blob", "sha": "566ca2eb9469a7c50d17122d8cb9b55734a15e99", "size": 103, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/566ca2eb9469a7c50d17122d8cb9b55734a15e99"}, {"path": "examples/llm", "mode": "040000", "type": "tree", "sha": "595413222a0781f6c0faa59853d3925b8fd78768", "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/trees/595413222a0781f6c0faa59853d3925b8fd78768"}, {"path": "examples/llm/.gitignore", "mode": "100644", "type": "blob", "sha": "8bd1bd091ce8bfb01aca93648834cb03fa2e1dc8", "size": 11, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/8bd1bd091ce8bfb01aca93648834cb03fa2e1dc8"}, {"path": "examples/llm/README.md", "mode": "100644", "type": "blob", "sha": "034dd0d9d35182fe7f89b9d08f96d6f0b7cea3f9", "size": 12410, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/034dd0d9d35182fe7f89b9d08f96d6f0b7cea3f9"}, {"path": "examples/llm/configs", "mode": "040000", "type": "tree", "sha": "af5ca973d5370b1dcb545b55db8418f293881ddd", "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/trees/af5ca973d5370b1dcb545b55db8418f293881ddd"}, {"path": "examples/llm/configs/__default__.yaml", "mode": "100644", "type": "blob", "sha": "960fcc01ad826c20323a2ca653f0a0866e5e8525", "size": 3354, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/960fcc01ad826c20323a2ca653f0a0866e5e8525"}, {"path": "examples/llm/configs/awq.yaml", "mode": "100644", "type": "blob", "sha": "b6f9a408032503f5b10c7cda649dc5131e170787", "size": 975, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/b6f9a408032503f5b10c7cda649dc5131e170787"}, {"path": "examples/llm/configs/gptq.yaml", "mode": "100644", "type": "blob", "sha": "b3e45792031469b09a6b036ea3db2487336184d9", "size": 985, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/b3e45792031469b09a6b036ea3db2487336184d9"}, {"path": "examples/llm/configs/ooo.yaml", "mode": "100644", "type": "blob", "sha": "9a6098fc1b62d86f30b82ed2b58bd8b6a68f1386", "size": 1640, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/9a6098fc1b62d86f30b82ed2b58bd8b6a68f1386"}, {"path": "examples/llm/configs/qoq-g128.yaml", "mode": "100644", "type": "blob", "sha": "67900543baab21f92a14054d8382c1d5cc4b132f", "size": 1512, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/67900543baab21f92a14054d8382c1d5cc4b132f"}, {"path": "examples/llm/configs/qoq-gchn.yaml", "mode": "100644", "type": "blob", "sha": "b5e804996d08dd56d11f44da6bf8c8feadeedf8b", "size": 1425, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/b5e804996d08dd56d11f44da6bf8c8feadeedf8b"}, {"path": "examples/llm/configs/smoothquant-dynamic.yaml", "mode": "100644", "type": "blob", "sha": "7432542a6ff81fb52932a9e66999cd02ae2c02d4", "size": 710, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/7432542a6ff81fb52932a9e66999cd02ae2c02d4"}, {"path": "examples/llm/configs/smoothquant-static.yaml", "mode": "100644", "type": "blob", "sha": "ffcebf419e0f19faae4d1d34887bcf8ed78fcebb", "size": 1002, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/ffcebf419e0f19faae4d1d34887bcf8ed78fcebb"}, {"path": "examples/llm/scripts", "mode": "040000", "type": "tree", "sha": "dd97f4de684bcac091a1a097d17b46610c5d242b", "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/trees/dd97f4de684bcac091a1a097d17b46610c5d242b"}, {"path": "examples/llm/scripts/awq.sh", "mode": "100644", "type": "blob", "sha": "59133fb1d439861f786cbca55b9a705efd87467c", "size": 534, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/59133fb1d439861f786cbca55b9a705efd87467c"}, {"path": "examples/llm/scripts/gptq.sh", "mode": "100644", "type": "blob", "sha": "8d4a14eeaa409da3dd174aede508cf05f133b16f", "size": 554, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/8d4a14eeaa409da3dd174aede508cf05f133b16f"}, {"path": "examples/llm/scripts/qoq.sh", "mode": "100644", "type": "blob", "sha": "b37207eafa3e3cbcf25076c7fca5b150971058c0", "size": 6057, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/b37207eafa3e3cbcf25076c7fca5b150971058c0"}, {"path": "examples/llm/scripts/smoothquant.sh", "mode": "100644", "type": "blob", "sha": "766f7b08f2cd4d2fe8293da00f32bb556ecbee3f", "size": 1740, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/766f7b08f2cd4d2fe8293da00f32bb556ecbee3f"}, {"path": "pyproject.toml", "mode": "100644", "type": "blob", "sha": "55eb2efef6703dfe624036231126393d271ccf0c", "size": 1757, "url": "https://api.github.com/repos/nunchux-ai/deepcompressor/git/blobs/55eb2efef6703dfe624036231126393d271ccf0c"}], "truncated": false} \ No newline at end of file diff --git a/reproduction/nunchaku_backend/upstream/nunchaku-tree.json b/reproduction/nunchaku_backend/upstream/nunchaku-tree.json new file mode 100644 index 0000000000000000000000000000000000000000..e16bd70e17b25d007f7b3599653c116f48731caa --- /dev/null +++ b/reproduction/nunchaku_backend/upstream/nunchaku-tree.json @@ -0,0 +1 @@ +{"sha": "302e0e97024ebd68688fe890e5df83731edf7b54", "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/trees/302e0e97024ebd68688fe890e5df83731edf7b54", "tree": [{"path": ".clang-format", "mode": "100644", "type": "blob", "sha": "f2886ffbbf9c5e803ff0dc758540869cd790ade1", "size": 1130, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/f2886ffbbf9c5e803ff0dc758540869cd790ade1"}, {"path": ".clang-format-ignore", "mode": "100644", "type": "blob", "sha": "5eb22646923e16e2565bb7a5244037d63cd565d0", "size": 14, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/5eb22646923e16e2565bb7a5244037d63cd565d0"}, {"path": ".github", "mode": "040000", "type": "tree", "sha": "06c0e601e723476d6fec0312b5fdbefb2b8eff3d", "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/trees/06c0e601e723476d6fec0312b5fdbefb2b8eff3d"}, {"path": ".github/CODEOWNERS", "mode": "100644", "type": "blob", "sha": "5e4c72b129c4bab470e4e05bf810c35cc6ba21b8", "size": 17, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/5e4c72b129c4bab470e4e05bf810c35cc6ba21b8"}, {"path": ".github/ISSUE_TEMPLATE", "mode": "040000", "type": "tree", "sha": "0f2a88f6b1bc8c05bc46665498ace5132bafc91b", "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/trees/0f2a88f6b1bc8c05bc46665498ace5132bafc91b"}, {"path": ".github/ISSUE_TEMPLATE/1-bug-report.yml", "mode": "100644", "type": "blob", "sha": "55dbda70abe02386836b3ee01ad4ab31c5c571c5", "size": 1924, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/55dbda70abe02386836b3ee01ad4ab31c5c571c5"}, {"path": ".github/ISSUE_TEMPLATE/2-feature-request.yml", "mode": "100644", "type": "blob", "sha": "23e078cdf8ba4c57ebac757bec433ca893cd2638", "size": 981, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/23e078cdf8ba4c57ebac757bec433ca893cd2638"}, {"path": ".github/pull_request_template.md", "mode": "100644", "type": "blob", "sha": "bb04f4f10d9ae21fb4f1df4b6be63cebce718578", "size": 1393, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/bb04f4f10d9ae21fb4f1df4b6be63cebce718578"}, {"path": ".github/workflows", "mode": "040000", "type": "tree", "sha": "5af2959c2e580ff890b9ae3d71c876d4a7d521fb", "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/trees/5af2959c2e580ff890b9ae3d71c876d4a7d521fb"}, {"path": ".github/workflows/build-docs.yaml", "mode": "100644", "type": "blob", "sha": "67174415b09c8420ead04f7ef361bb8143003a9a", "size": 2044, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/67174415b09c8420ead04f7ef361bb8143003a9a"}, {"path": ".github/workflows/clean-nightly-releases.yaml", "mode": "100644", "type": "blob", "sha": "f315c17467e0cadc412025b02a0c1e12767f16b5", "size": 1306, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/f315c17467e0cadc412025b02a0c1e12767f16b5"}, {"path": ".github/workflows/close-inactive-issues.yaml", "mode": "100644", "type": "blob", "sha": "22496809bf47fdde6aba909a51ac890737c20a6e", "size": 3738, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/22496809bf47fdde6aba909a51ac890737c20a6e"}, {"path": ".github/workflows/lint.yaml", "mode": "100644", "type": "blob", "sha": "82aee9e29ad46085a8d144da67fc3d7bcc8409ad", "size": 613, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/82aee9e29ad46085a8d144da67fc3d7bcc8409ad"}, {"path": ".github/workflows/nightly-build.yaml", "mode": "100644", "type": "blob", "sha": "76e8a830407fdc345542ee6b6c223d5dc0fdac11", "size": 5561, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/76e8a830407fdc345542ee6b6c223d5dc0fdac11"}, {"path": ".github/workflows/pr-test.yaml", "mode": "100644", "type": "blob", "sha": "84bfc6d8bcc6f3042c49dfccb68cfc64d24bbc51", "size": 3721, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/84bfc6d8bcc6f3042c49dfccb68cfc64d24bbc51"}, {"path": ".github/workflows/release-build.yaml", "mode": "100644", "type": "blob", "sha": "a14bff9fe9f1bba9599c82e925cb73af1649b20b", "size": 5168, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/a14bff9fe9f1bba9599c82e925cb73af1649b20b"}, {"path": ".github/workflows/reopen-issues.yaml", "mode": "100644", "type": "blob", "sha": "1932275d2f870ebf18e9a63495f79ccadc209965", "size": 2688, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/1932275d2f870ebf18e9a63495f79ccadc209965"}, {"path": ".github/workflows/run_all_tests.py", "mode": "100644", "type": "blob", "sha": "53655c03fb17c3ae46a3147f7af6801f08c61fd5", "size": 1333, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/53655c03fb17c3ae46a3147f7af6801f08c61fd5"}, {"path": ".github/workflows/sync-to-private.yaml", "mode": "100644", "type": "blob", "sha": "5d52a0d46c9c0a3882d13fd3c0bfbef145a083c7", "size": 3155, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/5d52a0d46c9c0a3882d13fd3c0bfbef145a083c7"}, {"path": ".gitignore", "mode": "100644", "type": "blob", "sha": "cdfb47107e2758ac638c2288ab13273b71502bc2", "size": 3593, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/cdfb47107e2758ac638c2288ab13273b71502bc2"}, {"path": ".gitmodules", "mode": "100644", "type": "blob", "sha": "de59083eb287c7894c9cee5b4d9159b3d25b3fdf", "size": 579, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/de59083eb287c7894c9cee5b4d9159b3d25b3fdf"}, {"path": ".pre-commit-config.yaml", "mode": "100644", "type": "blob", "sha": "c0bf5ba79801416f5aa75ba0589eab37efc979b0", "size": 2343, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/c0bf5ba79801416f5aa75ba0589eab37efc979b0"}, {"path": "LICENCE.txt", "mode": "100644", "type": "blob", "sha": "656a072054e4c42a9c3b0ca512b240f9c6d99602", "size": 11345, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/656a072054e4c42a9c3b0ca512b240f9c6d99602"}, {"path": "MANIFEST.in", "mode": "100644", "type": "blob", "sha": "7c018dcf8b0995b080eee84d116679441cff7575", "size": 760, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/7c018dcf8b0995b080eee84d116679441cff7575"}, {"path": "README.md", "mode": "100644", "type": "blob", "sha": "fab0e4a064220e7e1ff4931a0dcec54c0a5cce1e", "size": 17394, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/fab0e4a064220e7e1ff4931a0dcec54c0a5cce1e"}, {"path": "app", "mode": "040000", "type": "tree", "sha": "db1f63e18e7174842ec6b7112308178b1a8f67dc", "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/trees/db1f63e18e7174842ec6b7112308178b1a8f67dc"}, {"path": "app/flux.1", "mode": "040000", "type": "tree", "sha": "510e926e6ab9807b7c91389c2fcceea5d818f59b", "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/trees/510e926e6ab9807b7c91389c2fcceea5d818f59b"}, {"path": "app/flux.1/depth_canny", "mode": "040000", "type": "tree", "sha": "cb4d1cb2dfde6f84a4e83852a6f09ddb7ee3821b", "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/trees/cb4d1cb2dfde6f84a4e83852a6f09ddb7ee3821b"}, {"path": "app/flux.1/depth_canny/README.md", "mode": "100644", "type": "blob", "sha": "5177d5629de25ffafe2baef1d78a9a0d0c0383cb", "size": 1254, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/5177d5629de25ffafe2baef1d78a9a0d0c0383cb"}, {"path": "app/flux.1/depth_canny/assets", "mode": "040000", "type": "tree", "sha": "0d7bb1867e0b0f7f91c7da25593428d3cd6c5742", "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/trees/0d7bb1867e0b0f7f91c7da25593428d3cd6c5742"}, {"path": "app/flux.1/depth_canny/assets/description.html", "mode": "100644", "type": "blob", "sha": "38a370b049c7650758873680abc9e9e1b8d81942", "size": 1179, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/38a370b049c7650758873680abc9e9e1b8d81942"}, {"path": "app/flux.1/depth_canny/assets/style.css", "mode": "100644", "type": "blob", "sha": "6df54d33fe823c2dcb49ead8b33fb24fc0037bc1", "size": 718, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/6df54d33fe823c2dcb49ead8b33fb24fc0037bc1"}, {"path": "app/flux.1/depth_canny/run_gradio.py", "mode": "100644", "type": "blob", "sha": "c31bdd81d5e27a8003695f5677ea953d01e9a98f", "size": 10200, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/c31bdd81d5e27a8003695f5677ea953d01e9a98f"}, {"path": "app/flux.1/depth_canny/utils.py", "mode": "100644", "type": "blob", "sha": "9fa4980d803e829f28869aa700d1c0d6cd55caba", "size": 812, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/9fa4980d803e829f28869aa700d1c0d6cd55caba"}, {"path": "app/flux.1/depth_canny/vars.py", "mode": "100644", "type": "blob", "sha": "4b48646a1a23679a3651054c617d4c49890f0e64", "size": 3254, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/4b48646a1a23679a3651054c617d4c49890f0e64"}, {"path": "app/flux.1/fill", "mode": "040000", "type": "tree", "sha": "3556eccc967558b3e998a3d9095a4039666b3178", "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/trees/3556eccc967558b3e998a3d9095a4039666b3178"}, {"path": "app/flux.1/fill/README.md", "mode": "100644", "type": "blob", "sha": "8c4a77bb93fe4ea58163de9050c5540e467dd95c", "size": 744, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/8c4a77bb93fe4ea58163de9050c5540e467dd95c"}, {"path": "app/flux.1/fill/assets", "mode": "040000", "type": "tree", "sha": "6879492fe7ef7182a4c70dfc403ec2f146319388", "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/trees/6879492fe7ef7182a4c70dfc403ec2f146319388"}, {"path": "app/flux.1/fill/assets/description.html", "mode": "100644", "type": "blob", "sha": "d5421d7bfb1130b52eb0420414777fa9ee045f74", "size": 1194, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/d5421d7bfb1130b52eb0420414777fa9ee045f74"}, {"path": "app/flux.1/fill/assets/style.css", "mode": "100644", "type": "blob", "sha": "6df54d33fe823c2dcb49ead8b33fb24fc0037bc1", "size": 718, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/6df54d33fe823c2dcb49ead8b33fb24fc0037bc1"}, {"path": "app/flux.1/fill/run_gradio.py", "mode": "100644", "type": "blob", "sha": "93259decf70acb4408bf273e531cab556bab240c", "size": 9091, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/93259decf70acb4408bf273e531cab556bab240c"}, {"path": "app/flux.1/fill/utils.py", "mode": "100644", "type": "blob", "sha": "35956ee7688adc6fd052930dabc9cbfa95eae889", "size": 668, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/35956ee7688adc6fd052930dabc9cbfa95eae889"}, {"path": "app/flux.1/fill/vars.py", "mode": "100644", "type": "blob", "sha": "b28d69490aec97a8a04d6fed547cfd51e3f793c7", "size": 1723, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/b28d69490aec97a8a04d6fed547cfd51e3f793c7"}, {"path": "app/flux.1/kontext", "mode": "040000", "type": "tree", "sha": "948b9a4896f0ba2abedef3d258897b13b81cf52b", "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/trees/948b9a4896f0ba2abedef3d258897b13b81cf52b"}, {"path": "app/flux.1/kontext/README.md", "mode": "100644", "type": "blob", "sha": "f97be1005826aa11a86e1da103f150e8169f70e5", "size": 465, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/f97be1005826aa11a86e1da103f150e8169f70e5"}, {"path": "app/flux.1/kontext/assets", "mode": "040000", "type": "tree", "sha": "8b6224e380db276c0dc0aed4926f82bbd8d5eb10", "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/trees/8b6224e380db276c0dc0aed4926f82bbd8d5eb10"}, {"path": "app/flux.1/kontext/assets/description.html", "mode": "100644", "type": "blob", "sha": "7601df355012efdac011a41a5cc47355938e6f42", "size": 1035, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/7601df355012efdac011a41a5cc47355938e6f42"}, {"path": "app/flux.1/kontext/assets/style.css", "mode": "100644", "type": "blob", "sha": "6df54d33fe823c2dcb49ead8b33fb24fc0037bc1", "size": 718, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/6df54d33fe823c2dcb49ead8b33fb24fc0037bc1"}, {"path": "app/flux.1/kontext/run_gradio.py", "mode": "100644", "type": "blob", "sha": "5955fd693c1e07cfbacf21c7c9e8b083b33bdd8d", "size": 7421, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/5955fd693c1e07cfbacf21c7c9e8b083b33bdd8d"}, {"path": "app/flux.1/kontext/utils.py", "mode": "100644", "type": "blob", "sha": "35956ee7688adc6fd052930dabc9cbfa95eae889", "size": 668, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/35956ee7688adc6fd052930dabc9cbfa95eae889"}, {"path": "app/flux.1/kontext/vars.py", "mode": "100644", "type": "blob", "sha": "421fd72f50f79ef405d6026de9707712fa26a955", "size": 1674, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/421fd72f50f79ef405d6026de9707712fa26a955"}, {"path": "app/flux.1/redux", "mode": "040000", "type": "tree", "sha": "e4f2f2d9c6f9fb1c099d149f6f8c61443f7719d4", "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/trees/e4f2f2d9c6f9fb1c099d149f6f8c61443f7719d4"}, {"path": "app/flux.1/redux/README.md", "mode": "100644", "type": "blob", "sha": "ba6c3fb466c8cebeec107a58c42f45c70124313d", "size": 624, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/ba6c3fb466c8cebeec107a58c42f45c70124313d"}, {"path": "app/flux.1/redux/assets", "mode": "040000", "type": "tree", "sha": "570d778304750a6466424d399fb64576f63b5433", "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/trees/570d778304750a6466424d399fb64576f63b5433"}, {"path": "app/flux.1/redux/assets/description.html", "mode": "100644", "type": "blob", "sha": "f010c77c7c6dff9edf4070a8fc98140bfd853255", "size": 1051, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/f010c77c7c6dff9edf4070a8fc98140bfd853255"}, {"path": "app/flux.1/redux/assets/style.css", "mode": "100644", "type": "blob", "sha": "6df54d33fe823c2dcb49ead8b33fb24fc0037bc1", "size": 718, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/6df54d33fe823c2dcb49ead8b33fb24fc0037bc1"}, {"path": "app/flux.1/redux/run_gradio.py", "mode": "100644", "type": "blob", "sha": "dab5415cb4857a56fd7d11a2d781461136e4f9d6", "size": 7609, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/dab5415cb4857a56fd7d11a2d781461136e4f9d6"}, {"path": "app/flux.1/redux/utils.py", "mode": "100644", "type": "blob", "sha": "bf11be9b10bd9e1443668bb61aca69ac23373bc3", "size": 505, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/bf11be9b10bd9e1443668bb61aca69ac23373bc3"}, {"path": "app/flux.1/redux/vars.py", "mode": "100644", "type": "blob", "sha": "84e272bf6050c07b26e8511ad60facb98378dfc8", "size": 449, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/84e272bf6050c07b26e8511ad60facb98378dfc8"}, {"path": "app/flux.1/sketch", "mode": "040000", "type": "tree", "sha": "ec387e7707071c5f993663296e4ee6315ebe6d41", "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/trees/ec387e7707071c5f993663296e4ee6315ebe6d41"}, {"path": "app/flux.1/sketch/README.md", "mode": "100644", "type": "blob", "sha": "915760d84dd07424533148f0926da6e620d132c6", "size": 835, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/915760d84dd07424533148f0926da6e620d132c6"}, {"path": "app/flux.1/sketch/assets", "mode": "040000", "type": "tree", "sha": "fa322f538f2960ad51f7c45171e1190cc70ecf52", "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/trees/fa322f538f2960ad51f7c45171e1190cc70ecf52"}, {"path": "app/flux.1/sketch/assets/description.html", "mode": "100644", "type": "blob", "sha": "1be1f01ad9b45808496cf48b9793b1f8051b8327", "size": 1186, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/1be1f01ad9b45808496cf48b9793b1f8051b8327"}, {"path": "app/flux.1/sketch/assets/style.css", "mode": "100644", "type": "blob", "sha": "6df54d33fe823c2dcb49ead8b33fb24fc0037bc1", "size": 718, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/6df54d33fe823c2dcb49ead8b33fb24fc0037bc1"}, {"path": "app/flux.1/sketch/convert_ckpt.py", "mode": "100644", "type": "blob", "sha": "fe045df9434c15987fc284edfc0e028ad3295ec6", "size": 4506, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/fe045df9434c15987fc284edfc0e028ad3295ec6"}, {"path": "app/flux.1/sketch/flux_pix2pix_pipeline.py", "mode": "100644", "type": "blob", "sha": "406d71385ba5e03b46d48e01d2bc6fbe77f7d31b", "size": 6987, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/406d71385ba5e03b46d48e01d2bc6fbe77f7d31b"}, {"path": "app/flux.1/sketch/run.py", "mode": "100644", "type": "blob", "sha": "d63713757d51caa149fe3c31f1a7a6fb48510dbf", "size": 1288, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/d63713757d51caa149fe3c31f1a7a6fb48510dbf"}, {"path": "app/flux.1/sketch/run_gradio.py", "mode": "100644", "type": "blob", "sha": "5f96fc0a4410db7a3ad22d60d755ca196365f2e0", "size": 9410, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/5f96fc0a4410db7a3ad22d60d755ca196365f2e0"}, {"path": "app/flux.1/sketch/utils.py", "mode": "100644", "type": "blob", "sha": "35956ee7688adc6fd052930dabc9cbfa95eae889", "size": 668, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/35956ee7688adc6fd052930dabc9cbfa95eae889"}, {"path": "app/flux.1/sketch/vars.py", "mode": "100644", "type": "blob", "sha": "adc4e52c0c58efa11b0e433ce426ee31c32d8979", "size": 1385, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/adc4e52c0c58efa11b0e433ce426ee31c32d8979"}, {"path": "app/flux.1/t2i", "mode": "040000", "type": "tree", "sha": "2286b24d3a6b2f733be32ef88d85942758559323", "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/trees/2286b24d3a6b2f733be32ef88d85942758559323"}, {"path": "app/flux.1/t2i/README.md", "mode": "100644", "type": "blob", "sha": "ca7beccc99a01964fac206046d17fae06c3990e2", "size": 5473, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/ca7beccc99a01964fac206046d17fae06c3990e2"}, {"path": "app/flux.1/t2i/assets", "mode": "040000", "type": "tree", "sha": "033a5296ad0c5a996ea8a4cd588a7808a6f70486", "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/trees/033a5296ad0c5a996ea8a4cd588a7808a6f70486"}, {"path": "app/flux.1/t2i/assets/common.css", "mode": "100644", "type": "blob", "sha": "736f2df8119d4aae64882e183a91de1feaf1f65a", "size": 210, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/736f2df8119d4aae64882e183a91de1feaf1f65a"}, {"path": "app/flux.1/t2i/assets/demo.jpg", "mode": "100644", "type": "blob", "sha": "e97a3d4e60c38b2bf0aedd66b5ff33b03fbd0baf", "size": 543832, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/e97a3d4e60c38b2bf0aedd66b5ff33b03fbd0baf"}, {"path": "app/flux.1/t2i/assets/description.html", "mode": "100644", "type": "blob", "sha": "676ea01b37f0d8a29c1e68c8e7d644c22be66012", "size": 1265, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/676ea01b37f0d8a29c1e68c8e7d644c22be66012"}, {"path": "app/flux.1/t2i/assets/frame1.css", "mode": "100644", "type": "blob", "sha": "edbef3e6cb71dc9be453013ea67c4f8c920020d6", "size": 112, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/edbef3e6cb71dc9be453013ea67c4f8c920020d6"}, {"path": "app/flux.1/t2i/assets/frame2.css", "mode": "100644", "type": "blob", "sha": "1e4e19ab75b80ea5830bcf965173636b972b8114", "size": 113, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/1e4e19ab75b80ea5830bcf965173636b972b8114"}, {"path": "app/flux.1/t2i/data", "mode": "040000", "type": "tree", "sha": "a85246c55664799b9762f14fcbb119c8091e5ff2", "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/trees/a85246c55664799b9762f14fcbb119c8091e5ff2"}, {"path": "app/flux.1/t2i/data/DCI", "mode": "040000", "type": "tree", "sha": "c4cf98f463985ac9c7e57a155d0501c539759d96", "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/trees/c4cf98f463985ac9c7e57a155d0501c539759d96"}, {"path": "app/flux.1/t2i/data/DCI/DCI.py", "mode": "100644", "type": "blob", "sha": "e35761f6c78ba0c561cef18aaefd096466e75000", "size": 3935, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/e35761f6c78ba0c561cef18aaefd096466e75000"}, {"path": "app/flux.1/t2i/data/MJHQ", "mode": "040000", "type": "tree", "sha": "509feaa9fb966dc9cec76833f0e98ae914717fd3", "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/trees/509feaa9fb966dc9cec76833f0e98ae914717fd3"}, {"path": "app/flux.1/t2i/data/MJHQ/MJHQ.py", "mode": "100644", "type": "blob", "sha": "c7c0c88bfe11581add115aff0c605ba768e4cece", "size": 3873, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/c7c0c88bfe11581add115aff0c605ba768e4cece"}, {"path": "app/flux.1/t2i/data/__init__.py", "mode": "100644", "type": "blob", "sha": "b79aeee5f79c68eebd4c678d1465b4d2ba18af58", "size": 789, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/b79aeee5f79c68eebd4c678d1465b4d2ba18af58"}, {"path": "app/flux.1/t2i/evaluate.py", "mode": "100644", "type": "blob", "sha": "99f78a96050a646181364aa9ffabfc04e7c09ef5", "size": 3052, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/99f78a96050a646181364aa9ffabfc04e7c09ef5"}, {"path": "app/flux.1/t2i/generate.py", "mode": "100644", "type": "blob", "sha": "4b95149a87c85c444dff48656c0f52e95456b146", "size": 2565, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/4b95149a87c85c444dff48656c0f52e95456b146"}, {"path": "app/flux.1/t2i/get_metrics.py", "mode": "100644", "type": "blob", "sha": "b24550adb7f42f7c331bf7c15e3fb59c074e8037", "size": 2764, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/b24550adb7f42f7c331bf7c15e3fb59c074e8037"}, {"path": "app/flux.1/t2i/latency.py", "mode": "100644", "type": "blob", "sha": "25679701bd26004f9bc994c02b002d2abd9f4c50", "size": 3830, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/25679701bd26004f9bc994c02b002d2abd9f4c50"}, {"path": "app/flux.1/t2i/metrics", "mode": "040000", "type": "tree", "sha": "20b3a84412ac368e9d249e22d4b7b9b93fa92fc1", "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/trees/20b3a84412ac368e9d249e22d4b7b9b93fa92fc1"}, {"path": "app/flux.1/t2i/metrics/__init__.py", "mode": "100644", "type": "blob", "sha": "e69de29bb2d1d6434b8b29ae775ad8c2e48c5391", "size": 0, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/e69de29bb2d1d6434b8b29ae775ad8c2e48c5391"}, {"path": "app/flux.1/t2i/metrics/fid.py", "mode": "100644", "type": "blob", "sha": "3c890e50a584075be3eb77b52cc6adeaecdc427c", "size": 4271, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/3c890e50a584075be3eb77b52cc6adeaecdc427c"}, {"path": "app/flux.1/t2i/metrics/image_reward.py", "mode": "100644", "type": "blob", "sha": "2960b71924fd548c197f3db630cf2108ed41e7d1", "size": 783, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/2960b71924fd548c197f3db630cf2108ed41e7d1"}, {"path": "app/flux.1/t2i/metrics/multimodal.py", "mode": "100644", "type": "blob", "sha": "3808ed35cef0a20028c84683c8236fe5df59c678", "size": 2539, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/3808ed35cef0a20028c84683c8236fe5df59c678"}, {"path": "app/flux.1/t2i/metrics/similarity.py", "mode": "100644", "type": "blob", "sha": "72a01fc2a7335540c0ff1a177377afb0dbbc5fee", "size": 3909, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/72a01fc2a7335540c0ff1a177377afb0dbbc5fee"}, {"path": "app/flux.1/t2i/run_gradio.py", "mode": "100644", "type": "blob", "sha": "260afa5b2945a9a8b36059146f826b214c1a57db", "size": 11499, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/260afa5b2945a9a8b36059146f826b214c1a57db"}, {"path": "app/flux.1/t2i/utils.py", "mode": "100644", "type": "blob", "sha": "e7976d702aab10178f902fc5a08f35a0627ff67e", "size": 5033, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/e7976d702aab10178f902fc5a08f35a0627ff67e"}, {"path": "app/flux.1/t2i/vars.py", "mode": "100644", "type": "blob", "sha": "175ac9a02ceb25f5e2d9734a5b32e2c34bb044b7", "size": 4804, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/175ac9a02ceb25f5e2d9734a5b32e2c34bb044b7"}, {"path": "app/sana", "mode": "040000", "type": "tree", "sha": "799c7aaab7fda2d54e79a3b26cbfeea2fb2c8973", "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/trees/799c7aaab7fda2d54e79a3b26cbfeea2fb2c8973"}, {"path": "app/sana/t2i", "mode": "040000", "type": "tree", "sha": "a68beb50ff28ee6c8480873c176b27cb21b9bd8d", "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/trees/a68beb50ff28ee6c8480873c176b27cb21b9bd8d"}, {"path": "app/sana/t2i/README.md", "mode": "100644", "type": "blob", "sha": "419eee10d807f7d3fe15ad5e87e404c4d17b6561", "size": 2254, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/419eee10d807f7d3fe15ad5e87e404c4d17b6561"}, {"path": "app/sana/t2i/assets", "mode": "040000", "type": "tree", "sha": "6f9afa5172c22cfc0cdebcf3197708049d56a25b", "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/trees/6f9afa5172c22cfc0cdebcf3197708049d56a25b"}, {"path": "app/sana/t2i/assets/common.css", "mode": "100644", "type": "blob", "sha": "736f2df8119d4aae64882e183a91de1feaf1f65a", "size": 210, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/736f2df8119d4aae64882e183a91de1feaf1f65a"}, {"path": "app/sana/t2i/assets/description.html", "mode": "100644", "type": "blob", "sha": "063ebd40c665a1c95ec976497808485f4cf458de", "size": 2635, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/063ebd40c665a1c95ec976497808485f4cf458de"}, {"path": "app/sana/t2i/assets/frame1.css", "mode": "100644", "type": "blob", "sha": "edbef3e6cb71dc9be453013ea67c4f8c920020d6", "size": 112, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/edbef3e6cb71dc9be453013ea67c4f8c920020d6"}, {"path": "app/sana/t2i/assets/frame2.css", "mode": "100644", "type": "blob", "sha": "1e4e19ab75b80ea5830bcf965173636b972b8114", "size": 113, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/1e4e19ab75b80ea5830bcf965173636b972b8114"}, {"path": "app/sana/t2i/generate.py", "mode": "100644", "type": "blob", "sha": "895ec7672f415d86dc4ff3d69809c220adc99947", "size": 1912, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/895ec7672f415d86dc4ff3d69809c220adc99947"}, {"path": "app/sana/t2i/latency.py", "mode": "100644", "type": "blob", "sha": "f554c867a50fcd216c4e3d1825d2aab16759d839", "size": 3933, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/f554c867a50fcd216c4e3d1825d2aab16759d839"}, {"path": "app/sana/t2i/run_gradio.py", "mode": "100644", "type": "blob", "sha": "e5c8efe8e67161c193071086bc6f1e5a50da14b8", "size": 7661, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/e5c8efe8e67161c193071086bc6f1e5a50da14b8"}, {"path": "app/sana/t2i/utils.py", "mode": "100644", "type": "blob", "sha": "356eaca826a2b96b59caec7a5bb3d71a57c3aa7f", "size": 1412, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/356eaca826a2b96b59caec7a5bb3d71a57c3aa7f"}, {"path": "app/sana/t2i/vars.py", "mode": "100644", "type": "blob", "sha": "f63bed2f6d1883133269e356d47a2f4f41197806", "size": 2374, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/f63bed2f6d1883133269e356d47a2f4f41197806"}, {"path": "assets", "mode": "040000", "type": "tree", "sha": "cb63d6f624c0be7fbaf6ab331b307afe536342ad", "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/trees/cb63d6f624c0be7fbaf6ab331b307afe536342ad"}, {"path": "assets/nunchaku.svg", "mode": "100644", "type": "blob", "sha": "503380cebad8b1b1a0e0240b52e341ade81e8aec", "size": 9862, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/503380cebad8b1b1a0e0240b52e341ade81e8aec"}, {"path": "assets/svdquant.svg", "mode": "100644", "type": "blob", "sha": "8e7318195312440b397f18653ad927ed0988411c", "size": 3382, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/8e7318195312440b397f18653ad927ed0988411c"}, {"path": "docs", "mode": "040000", "type": "tree", "sha": "adf7207400cfc85e756cc9cef62f23631d21d363", "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/trees/adf7207400cfc85e756cc9cef62f23631d21d363"}, {"path": "docs/Makefile", "mode": "100644", "type": "blob", "sha": "d0c3cbf1020d5c292abdedf27627c6abe25e2293", "size": 638, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/d0c3cbf1020d5c292abdedf27627c6abe25e2293"}, {"path": "docs/doxyfile", "mode": "100644", "type": "blob", "sha": "3bd55df068794ccc233d7193e02348fdab79c3c2", "size": 132052, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/3bd55df068794ccc233d7193e02348fdab79c3c2"}, {"path": "docs/source", "mode": "040000", "type": "tree", "sha": "df23f81651a9e3edff2f4e67c60d2d9eeda9efc0", "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/trees/df23f81651a9e3edff2f4e67c60d2d9eeda9efc0"}, {"path": "docs/source/_static", "mode": "040000", "type": "tree", "sha": "e711a5b97635c3855eacd20e83fcd751eb474ecb", "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/trees/e711a5b97635c3855eacd20e83fcd751eb474ecb"}, {"path": "docs/source/_static/nunchaku.ico", "mode": "100644", "type": "blob", "sha": "ed48d652ca71ce932088ddcf9b37d443e5943f8b", "size": 65534, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/ed48d652ca71ce932088ddcf9b37d443e5943f8b"}, {"path": "docs/source/conf.py", "mode": "100644", "type": "blob", "sha": "e937ae09dce0babfd4bf0d8e9d1d13000c90a7d9", "size": 2699, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/e937ae09dce0babfd4bf0d8e9d1d13000c90a7d9"}, {"path": "docs/source/developer", "mode": "040000", "type": "tree", "sha": "c65eeb764b110178f9d3274e65fc63d2b31b3e97", "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/trees/c65eeb764b110178f9d3274e65fc63d2b31b3e97"}, {"path": "docs/source/developer/build_docs.rst", "mode": "100644", "type": "blob", "sha": "1d4e802e8e5fed76fcc50573c34e06677e9c09ae", "size": 1239, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/1d4e802e8e5fed76fcc50573c34e06677e9c09ae"}, {"path": "docs/source/developer/contribution_guide.rst", "mode": "100644", "type": "blob", "sha": "454ff1a324a78d590a93d6a6540ec4bd46597af2", "size": 4588, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/454ff1a324a78d590a93d6a6540ec4bd46597af2"}, {"path": "docs/source/developer/docstring.rst", "mode": "100644", "type": "blob", "sha": "9903113fd2aa18fd9bec49f9cb9dfe73e065048b", "size": 3775, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/9903113fd2aa18fd9bec49f9cb9dfe73e065048b"}, {"path": "docs/source/faq", "mode": "040000", "type": "tree", "sha": "6b77d79678ac56eae95145bf2b7bc7431d8ba4d9", "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/trees/6b77d79678ac56eae95145bf2b7bc7431d8ba4d9"}, {"path": "docs/source/faq/faq.rst", "mode": "100644", "type": "blob", "sha": "8778afc61d0a026d998231f1925a40e76d8a212c", "size": 140, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/8778afc61d0a026d998231f1925a40e76d8a212c"}, {"path": "docs/source/faq/model.rst", "mode": "100644", "type": "blob", "sha": "8b953161da65ea033bfdb1332f6eb654eddc8fc4", "size": 264, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/8b953161da65ea033bfdb1332f6eb654eddc8fc4"}, {"path": "docs/source/faq/usage.rst", "mode": "100644", "type": "blob", "sha": "4a0e8c7c5ed8b4e4602863b337da47c5575c8b36", "size": 1134, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/4a0e8c7c5ed8b4e4602863b337da47c5575c8b36"}, {"path": "docs/source/index.rst", "mode": "100644", "type": "blob", "sha": "908e4300f71047d552f78aa6735515d5a795bdc7", "size": 1388, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/908e4300f71047d552f78aa6735515d5a795bdc7"}, {"path": "docs/source/installation", "mode": "040000", "type": "tree", "sha": "4a668f4f47f64ab66a9369bd2cbe9674ba2bfbe6", "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/trees/4a668f4f47f64ab66a9369bd2cbe9674ba2bfbe6"}, {"path": "docs/source/installation/installation.rst", "mode": "100644", "type": "blob", "sha": "d37bd673b8b4878a44a79d78ac47378a40c5e927", "size": 5507, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/d37bd673b8b4878a44a79d78ac47378a40c5e927"}, {"path": "docs/source/installation/setup_windows.rst", "mode": "100644", "type": "blob", "sha": "3414cb34b93fb037ee42c635c42fecb8086d1d9b", "size": 7557, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/3414cb34b93fb037ee42c635c42fecb8086d1d9b"}, {"path": "docs/source/links", "mode": "040000", "type": "tree", "sha": "3c66df737294dc0ac6abebed48ce40ffbffa9652", "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/trees/3c66df737294dc0ac6abebed48ce40ffbffa9652"}, {"path": "docs/source/links/blog.txt", "mode": "100644", "type": "blob", "sha": "742467db935ee231106d97a10b52241a2d819cc3", "size": 149, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/742467db935ee231106d97a10b52241a2d819cc3"}, {"path": "docs/source/links/cdn.txt", "mode": "100644", "type": "blob", "sha": "e69de29bb2d1d6434b8b29ae775ad8c2e48c5391", "size": 0, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/e69de29bb2d1d6434b8b29ae775ad8c2e48c5391"}, {"path": "docs/source/links/github.txt", "mode": "100644", "type": "blob", "sha": "da44aedf2385f54f7e087ca5e60f6e13cdb9e548", "size": 1090, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/da44aedf2385f54f7e087ca5e60f6e13cdb9e548"}, {"path": "docs/source/links/huggingface.txt", "mode": "100644", "type": "blob", "sha": "aceadfc92050112a8e47b5410e7b96057839becd", "size": 1262, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/aceadfc92050112a8e47b5410e7b96057839becd"}, {"path": "docs/source/links/misc.txt", "mode": "100644", "type": "blob", "sha": "61feee504e3e16cb439466a9c11a0baedfdc89b0", "size": 528, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/61feee504e3e16cb439466a9c11a0baedfdc89b0"}, {"path": "docs/source/links/modelscope.txt", "mode": "100644", "type": "blob", "sha": "fc9358b17fe8505baee22bb41a52dffc5cd3595f", "size": 142, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/fc9358b17fe8505baee22bb41a52dffc5cd3595f"}, {"path": "docs/source/links/paper.txt", "mode": "100644", "type": "blob", "sha": "7fba8199e4ca2fe1ebb2eef54557889e963dbff1", "size": 150, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/7fba8199e4ca2fe1ebb2eef54557889e963dbff1"}, {"path": "docs/source/python_api", "mode": "040000", "type": "tree", "sha": "e96a8e28a0bf1181b0cb5efefd752ed539dec56b", "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/trees/e96a8e28a0bf1181b0cb5efefd752ed539dec56b"}, {"path": "docs/source/python_api/nunchaku.caching.diffusers_adapters.flux.rst", "mode": "100644", "type": "blob", "sha": "1751f656851824f79f99128b89806e2891c22be4", "size": 196, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/1751f656851824f79f99128b89806e2891c22be4"}, {"path": "docs/source/python_api/nunchaku.caching.diffusers_adapters.rst", "mode": "100644", "type": "blob", "sha": "a540f427dac656b9db4692f51d0f9700b4404350", "size": 300, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/a540f427dac656b9db4692f51d0f9700b4404350"}, {"path": "docs/source/python_api/nunchaku.caching.diffusers_adapters.sana.rst", "mode": "100644", "type": "blob", "sha": "bc7cb04b180db1e3a4de94e51d48df6d824fed6f", "size": 196, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/bc7cb04b180db1e3a4de94e51d48df6d824fed6f"}, {"path": "docs/source/python_api/nunchaku.caching.rst", "mode": "100644", "type": "blob", "sha": "56e3d3f6d5fe4ca840c79efe81171fc4a0d64400", "size": 130, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/56e3d3f6d5fe4ca840c79efe81171fc4a0d64400"}, {"path": "docs/source/python_api/nunchaku.caching.utils.rst", "mode": "100644", "type": "blob", "sha": "6308a59f5b85425f48e5a0b76e72decd0ca85407", "size": 121, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/6308a59f5b85425f48e5a0b76e72decd0ca85407"}, {"path": "docs/source/python_api/nunchaku.lora.flux.compose.rst", "mode": "100644", "type": "blob", "sha": "81b6daa485de766ee76d8a04b2453b5177fd9584", "size": 152, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/81b6daa485de766ee76d8a04b2453b5177fd9584"}, {"path": "docs/source/python_api/nunchaku.lora.flux.convert.rst", "mode": "100644", "type": "blob", "sha": "09941e35188065ec9ed5e5aeb3f4a269e4457b01", "size": 152, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/09941e35188065ec9ed5e5aeb3f4a269e4457b01"}, {"path": "docs/source/python_api/nunchaku.lora.flux.diffusers_converter.rst", "mode": "100644", "type": "blob", "sha": "58384d98af8df17084ab2cb40dd5e9c60f1a8e8b", "size": 190, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/58384d98af8df17084ab2cb40dd5e9c60f1a8e8b"}, {"path": "docs/source/python_api/nunchaku.lora.flux.nunchaku_converter.rst", "mode": "100644", "type": "blob", "sha": "891843ac485555ccd3c3f9e1ce67c8808f856138", "size": 187, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/891843ac485555ccd3c3f9e1ce67c8808f856138"}, {"path": "docs/source/python_api/nunchaku.lora.flux.packer.rst", "mode": "100644", "type": "blob", "sha": "2e4e907eb4bb03939c19e8c5d5a7ff06fd6bbb7b", "size": 130, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/2e4e907eb4bb03939c19e8c5d5a7ff06fd6bbb7b"}, {"path": "docs/source/python_api/nunchaku.lora.flux.rst", "mode": "100644", "type": "blob", "sha": "244f45892e328f450eec3679c18c37cb12956747", "size": 269, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/244f45892e328f450eec3679c18c37cb12956747"}, {"path": "docs/source/python_api/nunchaku.lora.flux.utils.rst", "mode": "100644", "type": "blob", "sha": "37748e95db7a0476f262eb0163d435da3dabb469", "size": 146, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/37748e95db7a0476f262eb0163d435da3dabb469"}, {"path": "docs/source/python_api/nunchaku.lora.rst", "mode": "100644", "type": "blob", "sha": "270d4d467d91ce05475c3a2bffe24d88013637f5", "size": 81, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/270d4d467d91ce05475c3a2bffe24d88013637f5"}, {"path": "docs/source/python_api/nunchaku.merge_safetensors.rst", "mode": "100644", "type": "blob", "sha": "a535a61a003604662cb3832ccb47c2511f1a46c4", "size": 154, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/a535a61a003604662cb3832ccb47c2511f1a46c4"}, {"path": "docs/source/python_api/nunchaku.models.attention.rst", "mode": "100644", "type": "blob", "sha": "6f5c24422a229c5671709d63632c0625363c8d06", "size": 174, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/6f5c24422a229c5671709d63632c0625363c8d06"}, {"path": "docs/source/python_api/nunchaku.models.attention_processors.flux.rst", "mode": "100644", "type": "blob", "sha": "4dc13f2e04382873e8495742ff47fb399c645140", "size": 197, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/4dc13f2e04382873e8495742ff47fb399c645140"}, {"path": "docs/source/python_api/nunchaku.models.attention_processors.qwenimage.rst", "mode": "100644", "type": "blob", "sha": "8d7681b5445e0e0c8deda9781a96b83a9ff3e166", "size": 212, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/8d7681b5445e0e0c8deda9781a96b83a9ff3e166"}, {"path": "docs/source/python_api/nunchaku.models.attention_processors.rst", "mode": "100644", "type": "blob", "sha": "1f43b394b391fd960023c181d9f4fadb466d0fef", "size": 247, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/1f43b394b391fd960023c181d9f4fadb466d0fef"}, {"path": "docs/source/python_api/nunchaku.models.attention_processors.zimage.rst", "mode": "100644", "type": "blob", "sha": "9d919490ab47ee3df33c1925c3ce9ffea3ebb34b", "size": 203, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/9d919490ab47ee3df33c1925c3ce9ffea3ebb34b"}, {"path": "docs/source/python_api/nunchaku.models.embeddings.rst", "mode": "100644", "type": "blob", "sha": "53778a69f5bf378186e033ea5d07a9fde8e00349", "size": 173, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/53778a69f5bf378186e033ea5d07a9fde8e00349"}, {"path": "docs/source/python_api/nunchaku.models.ip_adapter.diffusers_adapters.flux.rst", "mode": "100644", "type": "blob", "sha": "bcc8a53c2ba81b5ac0ac70dca4f432c338db9b64", "size": 224, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/bcc8a53c2ba81b5ac0ac70dca4f432c338db9b64"}, {"path": "docs/source/python_api/nunchaku.models.ip_adapter.diffusers_adapters.rst", "mode": "100644", "type": "blob", "sha": "2b3afcf77063fa32de0fbadfe9bc18a74a017e5b", "size": 294, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/2b3afcf77063fa32de0fbadfe9bc18a74a017e5b"}, {"path": "docs/source/python_api/nunchaku.models.ip_adapter.rst", "mode": "100644", "type": "blob", "sha": "dc1c893544fc6a370b88f6569e94cb5bc282e1da", "size": 170, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/dc1c893544fc6a370b88f6569e94cb5bc282e1da"}, {"path": "docs/source/python_api/nunchaku.models.ip_adapter.utils.rst", "mode": "100644", "type": "blob", "sha": "2a292122cb087303d58fc2d9004be3b5dc264673", "size": 170, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/2a292122cb087303d58fc2d9004be3b5dc264673"}, {"path": "docs/source/python_api/nunchaku.models.linear.rst", "mode": "100644", "type": "blob", "sha": "12c424b634daf111c56b7bec791f554ed32608a8", "size": 161, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/12c424b634daf111c56b7bec791f554ed32608a8"}, {"path": "docs/source/python_api/nunchaku.models.normalization.rst", "mode": "100644", "type": "blob", "sha": "d0783273bab4d077b62bb37b39fa92373788da17", "size": 182, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/d0783273bab4d077b62bb37b39fa92373788da17"}, {"path": "docs/source/python_api/nunchaku.models.pulid.encoders_transformer.rst", "mode": "100644", "type": "blob", "sha": "46843029ec5dee51afbaa99d85d21d95b587017d", "size": 202, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/46843029ec5dee51afbaa99d85d21d95b587017d"}, {"path": "docs/source/python_api/nunchaku.models.pulid.pulid_forward.rst", "mode": "100644", "type": "blob", "sha": "22ad81d3548b4bb435265348747653468b166a80", "size": 181, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/22ad81d3548b4bb435265348747653468b166a80"}, {"path": "docs/source/python_api/nunchaku.models.pulid.rst", "mode": "100644", "type": "blob", "sha": "1591e073d1462fc82f93ba284c18536f755c146b", "size": 191, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/1591e073d1462fc82f93ba284c18536f755c146b"}, {"path": "docs/source/python_api/nunchaku.models.pulid.utils.rst", "mode": "100644", "type": "blob", "sha": "51e748de4a2e89173255a0359c448b53f82b49e1", "size": 155, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/51e748de4a2e89173255a0359c448b53f82b49e1"}, {"path": "docs/source/python_api/nunchaku.models.rst", "mode": "100644", "type": "blob", "sha": "03192d8c0689d2de49e2375c75904bc2499b979a", "size": 425, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/03192d8c0689d2de49e2375c75904bc2499b979a"}, {"path": "docs/source/python_api/nunchaku.models.safety_checker.rst", "mode": "100644", "type": "blob", "sha": "3d2d23b5c4bece4d9a7e2aba8cb4ec6881f3942e", "size": 166, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/3d2d23b5c4bece4d9a7e2aba8cb4ec6881f3942e"}, {"path": "docs/source/python_api/nunchaku.models.text_encoders.linear.rst", "mode": "100644", "type": "blob", "sha": "3f7bb5e352ad763bbe91586527d3ed94ed678af0", "size": 184, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/3f7bb5e352ad763bbe91586527d3ed94ed678af0"}, {"path": "docs/source/python_api/nunchaku.models.text_encoders.rst", "mode": "100644", "type": "blob", "sha": "15609ba012084fbb2fdb1b50434eed3321847582", "size": 225, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/15609ba012084fbb2fdb1b50434eed3321847582"}, {"path": "docs/source/python_api/nunchaku.models.text_encoders.t5_encoder.rst", "mode": "100644", "type": "blob", "sha": "c9c120eda8bd650cb636c2752c6376dd0a962d67", "size": 198, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/c9c120eda8bd650cb636c2752c6376dd0a962d67"}, {"path": "docs/source/python_api/nunchaku.models.text_encoders.tinychat_utils.rst", "mode": "100644", "type": "blob", "sha": "0a248f44bdc866779ac55080581a4cc09b83bc92", "size": 210, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/0a248f44bdc866779ac55080581a4cc09b83bc92"}, {"path": "docs/source/python_api/nunchaku.models.transformers.rst", "mode": "100644", "type": "blob", "sha": "38b74a1f9e729a68e79da0fae6f7072a67616c8a", "size": 382, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/38b74a1f9e729a68e79da0fae6f7072a67616c8a"}, {"path": "docs/source/python_api/nunchaku.models.transformers.transformer_flux.rst", "mode": "100644", "type": "blob", "sha": "4f6ca1bcc90916212727cdef55b81aab24dc3e73", "size": 211, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/4f6ca1bcc90916212727cdef55b81aab24dc3e73"}, {"path": "docs/source/python_api/nunchaku.models.transformers.transformer_flux_v2.rst", "mode": "100644", "type": "blob", "sha": "3720f1645998ebfa1388cf81837231432d3f7069", "size": 222, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/3720f1645998ebfa1388cf81837231432d3f7069"}, {"path": "docs/source/python_api/nunchaku.models.transformers.transformer_qwenimage.rst", "mode": "100644", "type": "blob", "sha": "2045213cd0f15c75d82393c799dfad5d11e73f12", "size": 226, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/2045213cd0f15c75d82393c799dfad5d11e73f12"}, {"path": "docs/source/python_api/nunchaku.models.transformers.transformer_sana.rst", "mode": "100644", "type": "blob", "sha": "b4c4860de85515b05ecb82374e27f9c237b18d20", "size": 211, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/b4c4860de85515b05ecb82374e27f9c237b18d20"}, {"path": "docs/source/python_api/nunchaku.models.transformers.transformer_zimage.rst", "mode": "100644", "type": "blob", "sha": "e7a06355599e6e7fe3cba3d0e0812cc45f502c14", "size": 217, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/e7a06355599e6e7fe3cba3d0e0812cc45f502c14"}, {"path": "docs/source/python_api/nunchaku.models.transformers.utils.rst", "mode": "100644", "type": "blob", "sha": "7bac4b3dc9f00b0e23e7453cf0c45187e4caadb1", "size": 197, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/7bac4b3dc9f00b0e23e7453cf0c45187e4caadb1"}, {"path": "docs/source/python_api/nunchaku.models.unets.rst", "mode": "100644", "type": "blob", "sha": "3393e20e185507231cab2d46b7a67233232e13a1", "size": 110, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/3393e20e185507231cab2d46b7a67233232e13a1"}, {"path": "docs/source/python_api/nunchaku.models.unets.unet_sdxl.rst", "mode": "100644", "type": "blob", "sha": "201b223bc5abaa0aad5689e1dca6a26954dbea79", "size": 169, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/201b223bc5abaa0aad5689e1dca6a26954dbea79"}, {"path": "docs/source/python_api/nunchaku.models.utils.rst", "mode": "100644", "type": "blob", "sha": "ce48709e2624473711c3f0c7cea341076aeb37e6", "size": 137, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/ce48709e2624473711c3f0c7cea341076aeb37e6"}, {"path": "docs/source/python_api/nunchaku.ops.fused.rst", "mode": "100644", "type": "blob", "sha": "6010c5a56a1d63fd946c4f28a22df1b8aa5d7083", "size": 149, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/6010c5a56a1d63fd946c4f28a22df1b8aa5d7083"}, {"path": "docs/source/python_api/nunchaku.ops.gemm.rst", "mode": "100644", "type": "blob", "sha": "2a85da673be8798b2639dfbc4d3339e97e7d31db", "size": 146, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/2a85da673be8798b2639dfbc4d3339e97e7d31db"}, {"path": "docs/source/python_api/nunchaku.ops.gemv.rst", "mode": "100644", "type": "blob", "sha": "949c4ccb197b4713959eb24c00a9a5fd6183974b", "size": 146, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/949c4ccb197b4713959eb24c00a9a5fd6183974b"}, {"path": "docs/source/python_api/nunchaku.ops.quantize.rst", "mode": "100644", "type": "blob", "sha": "74aac60d6308c09e06eaffe9b5ff1b6e1e757e04", "size": 158, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/74aac60d6308c09e06eaffe9b5ff1b6e1e757e04"}, {"path": "docs/source/python_api/nunchaku.ops.rst", "mode": "100644", "type": "blob", "sha": "96eca139bc8993fc6d8d401f1c4097212121ce27", "size": 146, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/96eca139bc8993fc6d8d401f1c4097212121ce27"}, {"path": "docs/source/python_api/nunchaku.pipeline.pipeline_flux_pulid.rst", "mode": "100644", "type": "blob", "sha": "8665691a836e514b4163f9e9b78af51e25f8c7f9", "size": 170, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/8665691a836e514b4163f9e9b78af51e25f8c7f9"}, {"path": "docs/source/python_api/nunchaku.pipeline.rst", "mode": "100644", "type": "blob", "sha": "b3138c6cf8b8c18f096caa5bdcaeaa1e7e3b811a", "size": 108, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/b3138c6cf8b8c18f096caa5bdcaeaa1e7e3b811a"}, {"path": "docs/source/python_api/nunchaku.rst", "mode": "100644", "type": "blob", "sha": "f4deed583722156d107be57c3c14ec3d11de9dea", "size": 297, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/f4deed583722156d107be57c3c14ec3d11de9dea"}, {"path": "docs/source/python_api/nunchaku.test.rst", "mode": "100644", "type": "blob", "sha": "4218fe05a1f0480a19ca84df233008896e717fba", "size": 113, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/4218fe05a1f0480a19ca84df233008896e717fba"}, {"path": "docs/source/python_api/nunchaku.utils.rst", "mode": "100644", "type": "blob", "sha": "4c0147d6743bf1e2eabf00dabad52ed827dac540", "size": 116, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/4c0147d6743bf1e2eabf00dabad52ed827dac540"}, {"path": "docs/source/usage", "mode": "040000", "type": "tree", "sha": "2f961731466ae98c9001338f5816f97302298b51", "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/trees/2f961731466ae98c9001338f5816f97302298b51"}, {"path": "docs/source/usage/attention.rst", "mode": "100644", "type": "blob", "sha": "c57ac7c24adffb527e738ef03175fdba47594d49", "size": 953, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/c57ac7c24adffb527e738ef03175fdba47594d49"}, {"path": "docs/source/usage/basic_usage.rst", "mode": "100644", "type": "blob", "sha": "0fcf6f28b33fa03a76c7a78f9538bfe31812136a", "size": 1910, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/0fcf6f28b33fa03a76c7a78f9538bfe31812136a"}, {"path": "docs/source/usage/cache.rst", "mode": "100644", "type": "blob", "sha": "4ba3214be297a426157c5fc8b625b9dfd5e3d139", "size": 1868, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/4ba3214be297a426157c5fc8b625b9dfd5e3d139"}, {"path": "docs/source/usage/controlnet.rst", "mode": "100644", "type": "blob", "sha": "d612d2b67afc317a60153e7925a2c3674c6ef248", "size": 4189, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/d612d2b67afc317a60153e7925a2c3674c6ef248"}, {"path": "docs/source/usage/ip_adapter.rst", "mode": "100644", "type": "blob", "sha": "a963b0039e5154660f5b982b9592d919b9f79a33", "size": 2075, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/a963b0039e5154660f5b982b9592d919b9f79a33"}, {"path": "docs/source/usage/kontext.rst", "mode": "100644", "type": "blob", "sha": "22506841a588f79d927554ab868f9056e0b3136d", "size": 687, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/22506841a588f79d927554ab868f9056e0b3136d"}, {"path": "docs/source/usage/lora.rst", "mode": "100644", "type": "blob", "sha": "06aebaa78d402f948a88520c5dc52796ac4bce71", "size": 7486, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/06aebaa78d402f948a88520c5dc52796ac4bce71"}, {"path": "docs/source/usage/offload.rst", "mode": "100644", "type": "blob", "sha": "53427995f0d5d302d49d9e29c0632342c747f295", "size": 1275, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/53427995f0d5d302d49d9e29c0632342c747f295"}, {"path": "docs/source/usage/pulid.rst", "mode": "100644", "type": "blob", "sha": "f76a1be3ce9f41a697c41b7594cdcc704843d4bc", "size": 1958, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/f76a1be3ce9f41a697c41b7594cdcc704843d4bc"}, {"path": "docs/source/usage/qencoder.rst", "mode": "100644", "type": "blob", "sha": "46dd33ef546bad8820868573eff7e40ac0d98a77", "size": 1005, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/46dd33ef546bad8820868573eff7e40ac0d98a77"}, {"path": "docs/source/usage/qwen-image-edit.rst", "mode": "100644", "type": "blob", "sha": "56e00bcfc7bae62a1625177e1f1a88aa8e191b08", "size": 3818, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/56e00bcfc7bae62a1625177e1f1a88aa8e191b08"}, {"path": "docs/source/usage/qwen-image.rst", "mode": "100644", "type": "blob", "sha": "f647504e6a7bab09ec17bbca4cc07df54dc29f05", "size": 2543, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/f647504e6a7bab09ec17bbca4cc07df54dc29f05"}, {"path": "docs/source/usage/sdxl.rst", "mode": "100644", "type": "blob", "sha": "803b0a715296bb3747772ca113a7fddaff201eff", "size": 807, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/803b0a715296bb3747772ca113a7fddaff201eff"}, {"path": "docs/source/usage/zimage.rst", "mode": "100644", "type": "blob", "sha": "2c51d70bfd1f41bf58e729ed8495b1d4318d9dcb", "size": 530, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/2c51d70bfd1f41bf58e729ed8495b1d4318d9dcb"}, {"path": "examples", "mode": "040000", "type": "tree", "sha": "3d671a76cfa96ec6b286e7648a67dd8f67a1b279", "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/trees/3d671a76cfa96ec6b286e7648a67dd8f67a1b279"}, {"path": "examples/flux.1-canny-dev-lora.py", "mode": "100644", "type": "blob", "sha": "97376f42d4a7f6c5b38eab7b9052b2c080bcaf12", "size": 1568, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/97376f42d4a7f6c5b38eab7b9052b2c080bcaf12"}, {"path": "examples/flux.1-canny-dev.py", "mode": "100644", "type": "blob", "sha": "4fc0eac115e4e8be8ec1fdebb7a60be96ef30865", "size": 1263, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/4fc0eac115e4e8be8ec1fdebb7a60be96ef30865"}, {"path": "examples/flux.1-depth-dev-lora.py", "mode": "100644", "type": "blob", "sha": "6eb13c917c4c33c426b5ae7aac0ffb841be0de22", "size": 1608, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/6eb13c917c4c33c426b5ae7aac0ffb841be0de22"}, {"path": "examples/flux.1-depth-dev.py", "mode": "100644", "type": "blob", "sha": "1571b6450a5314f4d0acee8028c01b574dcbad54", "size": 1258, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/1571b6450a5314f4d0acee8028c01b574dcbad54"}, {"path": "examples/flux.1-dev-IP-adapter.py", "mode": "100644", "type": "blob", "sha": "1a5ab5039d01f8de2c5535a7a1c8989534da74c3", "size": 1463, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/1a5ab5039d01f8de2c5535a7a1c8989534da74c3"}, {"path": "examples/flux.1-dev-cache.py", "mode": "100644", "type": "blob", "sha": "8a8bca821dad785a37cca61f0c7a363a92226fe6", "size": 910, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/8a8bca821dad785a37cca61f0c7a363a92226fe6"}, {"path": "examples/flux.1-dev-colossus.py", "mode": "100644", "type": "blob", "sha": "7371b2c838056a1c1f786c0c0c65a8c0023c6545", "size": 790, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/7371b2c838056a1c1f786c0c0c65a8c0023c6545"}, {"path": "examples/flux.1-dev-controlnet-union-pro.py", "mode": "100644", "type": "blob", "sha": "b227bbfcafbd444fed8863533be2ae26e4ab94b9", "size": 2021, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/b227bbfcafbd444fed8863533be2ae26e4ab94b9"}, {"path": "examples/flux.1-dev-double_cache.py", "mode": "100644", "type": "blob", "sha": "03b1ccc0f2afb62c3363154db29d6aacd16a1a18", "size": 830, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/03b1ccc0f2afb62c3363154db29d6aacd16a1a18"}, {"path": "examples/flux.1-dev-double_cache_offloading.py", "mode": "100644", "type": "blob", "sha": "4066c909a98ade2cf333bf396666bf1a601c9537", "size": 849, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/4066c909a98ade2cf333bf396666bf1a601c9537"}, {"path": "examples/flux.1-dev-fp16attn.py", "mode": "100644", "type": "blob", "sha": "a5ab987095685b8b79dd37540afe1e2e7c9d6612", "size": 764, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/a5ab987095685b8b79dd37540afe1e2e7c9d6612"}, {"path": "examples/flux.1-dev-lora.py", "mode": "100644", "type": "blob", "sha": "1f255543db8966958979e2b16c7faa89a4b0baa4", "size": 1115, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/1f255543db8966958979e2b16c7faa89a4b0baa4"}, {"path": "examples/flux.1-dev-multiple-lora.py", "mode": "100644", "type": "blob", "sha": "3d58d359f3c38c73c6c356363107f2c56c043d33", "size": 1251, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/3d58d359f3c38c73c6c356363107f2c56c043d33"}, {"path": "examples/flux.1-dev-offload.py", "mode": "100644", "type": "blob", "sha": "7e0e77caaabe3de009dd5b293cb88d642b79a50d", "size": 849, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/7e0e77caaabe3de009dd5b293cb88d642b79a50d"}, {"path": "examples/flux.1-dev-pulid.py", "mode": "100644", "type": "blob", "sha": "5488a7d7e2ae83dba9232f1a401d77266b7e7196", "size": 1090, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/5488a7d7e2ae83dba9232f1a401d77266b7e7196"}, {"path": "examples/flux.1-dev-qencoder.py", "mode": "100644", "type": "blob", "sha": "266c5d67b39c154d6732313ca14e8bb026047378", "size": 873, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/266c5d67b39c154d6732313ca14e8bb026047378"}, {"path": "examples/flux.1-dev-teacache-batch.py", "mode": "100644", "type": "blob", "sha": "448c0970cc7bf36787af1c419e1cbe4caf22a192", "size": 1506, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/448c0970cc7bf36787af1c419e1cbe4caf22a192"}, {"path": "examples/flux.1-dev-teacache.py", "mode": "100644", "type": "blob", "sha": "53eed08814db5ec4bcf729befc0e7811b98aed10", "size": 1106, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/53eed08814db5ec4bcf729befc0e7811b98aed10"}, {"path": "examples/flux.1-dev-turing.py", "mode": "100644", "type": "blob", "sha": "903db4efdc76548c877015b68c12e4586e91469c", "size": 1018, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/903db4efdc76548c877015b68c12e4586e91469c"}, {"path": "examples/flux.1-dev.py", "mode": "100644", "type": "blob", "sha": "78e05c635edb366c7cdfb14c829de9f607786f29", "size": 688, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/78e05c635edb366c7cdfb14c829de9f607786f29"}, {"path": "examples/flux.1-fill-dev.py", "mode": "100644", "type": "blob", "sha": "0b5a6844f56cab2d91266e74b7a4f1035322a0e4", "size": 1062, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/0b5a6844f56cab2d91266e74b7a4f1035322a0e4"}, {"path": "examples/flux.1-kontext-FALAI_lora.py", "mode": "100644", "type": "blob", "sha": "8cf9d30808845fc0799983cc8be83c0b3a98f744", "size": 1273, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/8cf9d30808845fc0799983cc8be83c0b3a98f744"}, {"path": "examples/flux.1-kontext-dev-teacache.py", "mode": "100644", "type": "blob", "sha": "f9f40c61677bf522231b3edb498e9606a641897a", "size": 1191, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/f9f40c61677bf522231b3edb498e9606a641897a"}, {"path": "examples/flux.1-kontext-dev.py", "mode": "100644", "type": "blob", "sha": "2868d582f158d49c16d27ee39a8b0b40d1b3db6d", "size": 892, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/2868d582f158d49c16d27ee39a8b0b40d1b3db6d"}, {"path": "examples/flux.1-krea-dev.py", "mode": "100644", "type": "blob", "sha": "60496edfc7a91b817427d7f320cf75113aa0c30c", "size": 856, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/60496edfc7a91b817427d7f320cf75113aa0c30c"}, {"path": "examples/flux.1-redux-dev.py", "mode": "100644", "type": "blob", "sha": "7c4859f262b7e46f8b22445c56114628fc92fbb8", "size": 1012, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/7c4859f262b7e46f8b22445c56114628fc92fbb8"}, {"path": "examples/flux.1-schnell.py", "mode": "100644", "type": "blob", "sha": "be22bc225873614b855a17bb2beab03b831a883a", "size": 732, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/be22bc225873614b855a17bb2beab03b831a883a"}, {"path": "examples/sana1.6b.py", "mode": "100644", "type": "blob", "sha": "502e56616d3de67173db2ebebdf191bdc99e64df", "size": 760, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/502e56616d3de67173db2ebebdf191bdc99e64df"}, {"path": "examples/sana1.6b_pag.py", "mode": "100644", "type": "blob", "sha": "f4a8f8b2cdbf200ad018761b7f5a6a15c57decaf", "size": 845, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/f4a8f8b2cdbf200ad018761b7f5a6a15c57decaf"}, {"path": "examples/v1", "mode": "040000", "type": "tree", "sha": "69badac7af4dfdc273b7ad30f920dfefdfd20459", "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/trees/69badac7af4dfdc273b7ad30f920dfefdfd20459"}, {"path": "examples/v1/flux.1-canny-dev.py", "mode": "100644", "type": "blob", "sha": "dbe575163a2bde94089c726bfa06c3671437fe33", "size": 1267, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/dbe575163a2bde94089c726bfa06c3671437fe33"}, {"path": "examples/v1/flux.1-depth-dev.py", "mode": "100644", "type": "blob", "sha": "c86dc33be5775c353542299d50e10c6fe400188a", "size": 1262, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/c86dc33be5775c353542299d50e10c6fe400188a"}, {"path": "examples/v1/flux.1-dev-cache-dit.py", "mode": "100644", "type": "blob", "sha": "2475ad367aa6092f994af15daf2c9b341429e021", "size": 1150, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/2475ad367aa6092f994af15daf2c9b341429e021"}, {"path": "examples/v1/flux.1-dev-cache.py", "mode": "100644", "type": "blob", "sha": "70fbbd0f0f2541ac4b03afae46751cccccd52373", "size": 965, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/70fbbd0f0f2541ac4b03afae46751cccccd52373"}, {"path": "examples/v1/flux.1-dev.py", "mode": "100644", "type": "blob", "sha": "4391dcb394c9809c2766b490b7895bc37972a267", "size": 692, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/4391dcb394c9809c2766b490b7895bc37972a267"}, {"path": "examples/v1/flux.1-fill-dev.py", "mode": "100644", "type": "blob", "sha": "f6c49685591ad36153443ee2234747bb1f9700dc", "size": 1066, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/f6c49685591ad36153443ee2234747bb1f9700dc"}, {"path": "examples/v1/flux.1-kontext-dev.py", "mode": "100644", "type": "blob", "sha": "9383bc07c5ce9fa052ff1292d6c0056881abf270", "size": 920, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/9383bc07c5ce9fa052ff1292d6c0056881abf270"}, {"path": "examples/v1/flux.1-krea-dev.py", "mode": "100644", "type": "blob", "sha": "ead3d3e47ec99e06acf154ab9c8b33801bcf5681", "size": 884, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/ead3d3e47ec99e06acf154ab9c8b33801bcf5681"}, {"path": "examples/v1/flux.1-redux-dev.py", "mode": "100644", "type": "blob", "sha": "e6ace47b0cdfca5e3b2844efe79dbe59ae0cbd77", "size": 1016, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/e6ace47b0cdfca5e3b2844efe79dbe59ae0cbd77"}, {"path": "examples/v1/flux.1-schnell.py", "mode": "100644", "type": "blob", "sha": "6403d0b7e583612726e238c9e9ba400638b34a53", "size": 753, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/6403d0b7e583612726e238c9e9ba400638b34a53"}, {"path": "examples/v1/qwen-image-cache-dit.py", "mode": "100644", "type": "blob", "sha": "3f9ec8b6a19a67df3c15444acf27b89158e04756", "size": 2350, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/3f9ec8b6a19a67df3c15444acf27b89158e04756"}, {"path": "examples/v1/qwen-image-controlnet.py", "mode": "100644", "type": "blob", "sha": "5f4629a3475481f0c55af6a733405a5cdd323890", "size": 2079, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/5f4629a3475481f0c55af6a733405a5cdd323890"}, {"path": "examples/v1/qwen-image-edit-2509-lightning.py", "mode": "100644", "type": "blob", "sha": "2156bfd170ab68cc75e1a72019c7fd98d0d28dce", "size": 2874, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/2156bfd170ab68cc75e1a72019c7fd98d0d28dce"}, {"path": "examples/v1/qwen-image-edit-2509.py", "mode": "100644", "type": "blob", "sha": "70a5e7087adaf58be57f7da12463590c85906b94", "size": 1843, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/70a5e7087adaf58be57f7da12463590c85906b94"}, {"path": "examples/v1/qwen-image-edit-lightning.py", "mode": "100644", "type": "blob", "sha": "23ea82adee06a3521f1d082717acae28f535e2b1", "size": 2627, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/23ea82adee06a3521f1d082717acae28f535e2b1"}, {"path": "examples/v1/qwen-image-edit.py", "mode": "100644", "type": "blob", "sha": "0c223144d068183b4d785fa5858ecaa86bcf3230", "size": 1472, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/0c223144d068183b4d785fa5858ecaa86bcf3230"}, {"path": "examples/v1/qwen-image-lightning.py", "mode": "100644", "type": "blob", "sha": "a114e2dc135a969b1ef66de8e7c0cfe179f9dce1", "size": 2715, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/a114e2dc135a969b1ef66de8e7c0cfe179f9dce1"}, {"path": "examples/v1/qwen-image.py", "mode": "100644", "type": "blob", "sha": "fad061c362b5207d98b5c9f791def756dd7df077", "size": 1965, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/fad061c362b5207d98b5c9f791def756dd7df077"}, {"path": "examples/v1/sdxl-turbo.py", "mode": "100644", "type": "blob", "sha": "9fd3b4e89756e45bb28d273a2749fa2a8d4c9fbf", "size": 693, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/9fd3b4e89756e45bb28d273a2749fa2a8d4c9fbf"}, {"path": "examples/v1/sdxl.py", "mode": "100644", "type": "blob", "sha": "dd9d30d30845873b0cd7b2432ce895aa7274f41a", "size": 749, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/dd9d30d30845873b0cd7b2432ce895aa7274f41a"}, {"path": "examples/v1/z-image-turbo.py", "mode": "100644", "type": "blob", "sha": "e0a42c360ec8542022674324f8117c8d76e15ea7", "size": 1506, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/e0a42c360ec8542022674324f8117c8d76e15ea7"}, {"path": "nunchaku", "mode": "040000", "type": "tree", "sha": "02ffc7f37d0067e1e57f478cb56f182a4422eecd", "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/trees/02ffc7f37d0067e1e57f478cb56f182a4422eecd"}, {"path": "nunchaku/__init__.py", "mode": "100644", "type": "blob", "sha": "cb5eb8729600111634fcd0ac225c468b5d8a9804", "size": 485, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/cb5eb8729600111634fcd0ac225c468b5d8a9804"}, {"path": "nunchaku/__version__.py", "mode": "100644", "type": "blob", "sha": "81c3bbff17d964975b308f88baacea926bef6959", "size": 25, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/81c3bbff17d964975b308f88baacea926bef6959"}, {"path": "nunchaku/caching", "mode": "040000", "type": "tree", "sha": "4023c9554d573ef8e7a1be27aa1a0fa5f2d2c947", "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/trees/4023c9554d573ef8e7a1be27aa1a0fa5f2d2c947"}, {"path": "nunchaku/caching/__init__.py", "mode": "100644", "type": "blob", "sha": "e69de29bb2d1d6434b8b29ae775ad8c2e48c5391", "size": 0, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/e69de29bb2d1d6434b8b29ae775ad8c2e48c5391"}, {"path": "nunchaku/caching/diffusers_adapters", "mode": "040000", "type": "tree", "sha": "d7cd4772aef014001777a33ea4efef1e2fa441b5", "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/trees/d7cd4772aef014001777a33ea4efef1e2fa441b5"}, {"path": "nunchaku/caching/diffusers_adapters/__init__.py", "mode": "100644", "type": "blob", "sha": "1331889aa5af4f34a1376964910847d561390f8a", "size": 2594, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/1331889aa5af4f34a1376964910847d561390f8a"}, {"path": "nunchaku/caching/diffusers_adapters/flux.py", "mode": "100644", "type": "blob", "sha": "916dbefb85e93f6bfffd94ef539bf054ff2be542", "size": 5518, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/916dbefb85e93f6bfffd94ef539bf054ff2be542"}, {"path": "nunchaku/caching/diffusers_adapters/flux_v2.py", "mode": "100644", "type": "blob", "sha": "05441fc88b8926450be60e7e396d3a77d3904687", "size": 3006, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/05441fc88b8926450be60e7e396d3a77d3904687"}, {"path": "nunchaku/caching/diffusers_adapters/sana.py", "mode": "100644", "type": "blob", "sha": "002bc82bbba79a27d9aed1cca3301d27a5762fcb", "size": 3295, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/002bc82bbba79a27d9aed1cca3301d27a5762fcb"}, {"path": "nunchaku/caching/fbcache.py", "mode": "100644", "type": "blob", "sha": "0cc57aaf6661897924d1921a14bc90b4d9ff8835", "size": 14729, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/0cc57aaf6661897924d1921a14bc90b4d9ff8835"}, {"path": "nunchaku/caching/teacache.py", "mode": "100644", "type": "blob", "sha": "f2047a065653c6a8da77cc54a896a3084acf343c", "size": 18630, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/f2047a065653c6a8da77cc54a896a3084acf343c"}, {"path": "nunchaku/caching/utils.py", "mode": "100644", "type": "blob", "sha": "c2d24d41943b47005f31f26c306a43639bd3bf3c", "size": 26524, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/c2d24d41943b47005f31f26c306a43639bd3bf3c"}, {"path": "nunchaku/caching/utils_v2.py", "mode": "100644", "type": "blob", "sha": "02423c8676c5ae1d3536920ffd6a6f84fe504463", "size": 12157, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/02423c8676c5ae1d3536920ffd6a6f84fe504463"}, {"path": "nunchaku/csrc", "mode": "040000", "type": "tree", "sha": "d3d3bbdd9ecdd1021dcb7c99e209701da64d435c", "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/trees/d3d3bbdd9ecdd1021dcb7c99e209701da64d435c"}, {"path": "nunchaku/csrc/flux.h", "mode": "100644", "type": "blob", "sha": "132ad4e4bb0f98b7a9d4c91eae184f33ce801ae4", "size": 10931, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/132ad4e4bb0f98b7a9d4c91eae184f33ce801ae4"}, {"path": "nunchaku/csrc/gemm.h", "mode": "100644", "type": "blob", "sha": "74d1fb4b5ea15200d643fa98e406ef836ccfd307", "size": 3814, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/74d1fb4b5ea15200d643fa98e406ef836ccfd307"}, {"path": "nunchaku/csrc/gemm88.h", "mode": "100644", "type": "blob", "sha": "aa3dd3175f0a4e89c42c4ff4643a339d49b9e504", "size": 1045, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/aa3dd3175f0a4e89c42c4ff4643a339d49b9e504"}, {"path": "nunchaku/csrc/module.h", "mode": "100644", "type": "blob", "sha": "812c82720203f3abf5eb5cd3cda2bf6fc9747bce", "size": 2115, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/812c82720203f3abf5eb5cd3cda2bf6fc9747bce"}, {"path": "nunchaku/csrc/ops.h", "mode": "100644", "type": "blob", "sha": "dbd15fb22ce89492e520c77775311e612c7e18ed", "size": 8383, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/dbd15fb22ce89492e520c77775311e612c7e18ed"}, {"path": "nunchaku/csrc/pybind.cpp", "mode": "100644", "type": "blob", "sha": "74fc37f0f04aec50dc73072234d978e365de7d37", "size": 5811, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/74fc37f0f04aec50dc73072234d978e365de7d37"}, {"path": "nunchaku/csrc/sana.h", "mode": "100644", "type": "blob", "sha": "8d6c75ffba160cc6125772a5b583b99e7c7de541", "size": 4533, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/8d6c75ffba160cc6125772a5b583b99e7c7de541"}, {"path": "nunchaku/csrc/utils.h", "mode": "100644", "type": "blob", "sha": "967f87772901aec5f66d4fc0860f8e642e51e60b", "size": 1091, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/967f87772901aec5f66d4fc0860f8e642e51e60b"}, {"path": "nunchaku/lora", "mode": "040000", "type": "tree", "sha": "68953928f118f4e0471cc3a4b779cc9fb2865370", "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/trees/68953928f118f4e0471cc3a4b779cc9fb2865370"}, {"path": "nunchaku/lora/__init__.py", "mode": "100644", "type": "blob", "sha": "e69de29bb2d1d6434b8b29ae775ad8c2e48c5391", "size": 0, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/e69de29bb2d1d6434b8b29ae775ad8c2e48c5391"}, {"path": "nunchaku/lora/flux", "mode": "040000", "type": "tree", "sha": "88bedcf1787e9619afdb026755cd3615f3bd5103", "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/trees/88bedcf1787e9619afdb026755cd3615f3bd5103"}, {"path": "nunchaku/lora/flux/__init__.py", "mode": "100644", "type": "blob", "sha": "8d49b5f570dcbe27f81f934f7fb73f5dd7c5456c", "size": 273, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/8d49b5f570dcbe27f81f934f7fb73f5dd7c5456c"}, {"path": "nunchaku/lora/flux/compose.py", "mode": "100644", "type": "blob", "sha": "a44a8aff4726c026f0dc2859c06a054453539c3c", "size": 9877, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/a44a8aff4726c026f0dc2859c06a054453539c3c"}, {"path": "nunchaku/lora/flux/convert.py", "mode": "100644", "type": "blob", "sha": "ff11f40accb842f67197469a8c7cc8be0f3e102e", "size": 2730, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/ff11f40accb842f67197469a8c7cc8be0f3e102e"}, {"path": "nunchaku/lora/flux/diffusers_converter.py", "mode": "100644", "type": "blob", "sha": "e0c30134ef718ef79509d1d2d22182963cd97651", "size": 9052, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/e0c30134ef718ef79509d1d2d22182963cd97651"}, {"path": "nunchaku/lora/flux/nunchaku_converter.py", "mode": "100644", "type": "blob", "sha": "b10309c3dee19ba10192a0d3c237bc7a40858cb0", "size": 41194, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/b10309c3dee19ba10192a0d3c237bc7a40858cb0"}, {"path": "nunchaku/lora/flux/packer.py", "mode": "100644", "type": "blob", "sha": "7ab18f6ea4099da6bd74a71d3b0788832fe4c91a", "size": 22145, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/7ab18f6ea4099da6bd74a71d3b0788832fe4c91a"}, {"path": "nunchaku/lora/flux/utils.py", "mode": "100644", "type": "blob", "sha": "65b86ea815e365d11a8d91634f0f677e31d17e84", "size": 2655, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/65b86ea815e365d11a8d91634f0f677e31d17e84"}, {"path": "nunchaku/merge_safetensors.py", "mode": "100644", "type": "blob", "sha": "c89ecf63dcaa0653a4587611af1108d12dc7b2ec", "size": 7331, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/c89ecf63dcaa0653a4587611af1108d12dc7b2ec"}, {"path": "nunchaku/models", "mode": "040000", "type": "tree", "sha": "83065ccfb913006872253288b55f43f39329ab57", "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/trees/83065ccfb913006872253288b55f43f39329ab57"}, {"path": "nunchaku/models/__init__.py", "mode": "100644", "type": "blob", "sha": "70a25eebb47ac32f5906e56c628dc3f632787d12", "size": 524, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/70a25eebb47ac32f5906e56c628dc3f632787d12"}, {"path": "nunchaku/models/attention.py", "mode": "100644", "type": "blob", "sha": "9a9fdbce2c3cdc2ed81a4a47caa91111bbbf3369", "size": 3769, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/9a9fdbce2c3cdc2ed81a4a47caa91111bbbf3369"}, {"path": "nunchaku/models/attention_processors", "mode": "040000", "type": "tree", "sha": "3cf4e4b026664045272214a4147a0edfa6f68310", "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/trees/3cf4e4b026664045272214a4147a0edfa6f68310"}, {"path": "nunchaku/models/attention_processors/flux.py", "mode": "100644", "type": "blob", "sha": "1c2b4f8144336714f30767da4c4d27b13b7de248", "size": 10313, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/1c2b4f8144336714f30767da4c4d27b13b7de248"}, {"path": "nunchaku/models/attention_processors/qwenimage.py", "mode": "100644", "type": "blob", "sha": "d5ab3d71fd7b45dd7367f5a81ba67a92722f435d", "size": 5278, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/d5ab3d71fd7b45dd7367f5a81ba67a92722f435d"}, {"path": "nunchaku/models/attention_processors/sdxl.py", "mode": "100644", "type": "blob", "sha": "39e73b83b4b35759e06ab462cc8104127e0cb7d1", "size": 4388, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/39e73b83b4b35759e06ab462cc8104127e0cb7d1"}, {"path": "nunchaku/models/attention_processors/zimage.py", "mode": "100644", "type": "blob", "sha": "809a3e8fdecf63d4872fce209ec20770f1c14e4e", "size": 2407, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/809a3e8fdecf63d4872fce209ec20770f1c14e4e"}, {"path": "nunchaku/models/embeddings.py", "mode": "100644", "type": "blob", "sha": "12b50966200dc321c6b99679ac7cadfb07f8e594", "size": 4117, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/12b50966200dc321c6b99679ac7cadfb07f8e594"}, {"path": "nunchaku/models/ip_adapter", "mode": "040000", "type": "tree", "sha": "dc9f3af8d3f2f1393d1945533daf7b6c2b5ead00", "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/trees/dc9f3af8d3f2f1393d1945533daf7b6c2b5ead00"}, {"path": "nunchaku/models/ip_adapter/__init__.py", "mode": "100644", "type": "blob", "sha": "e69de29bb2d1d6434b8b29ae775ad8c2e48c5391", "size": 0, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/e69de29bb2d1d6434b8b29ae775ad8c2e48c5391"}, {"path": "nunchaku/models/ip_adapter/diffusers_adapters", "mode": "040000", "type": "tree", "sha": "85a938f64f0c6d6757f47ac3c69ebd778bfbd10f", "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/trees/85a938f64f0c6d6757f47ac3c69ebd778bfbd10f"}, {"path": "nunchaku/models/ip_adapter/diffusers_adapters/__init__.py", "mode": "100644", "type": "blob", "sha": "3d2af4236bb7c0a24fa7cf161da8a78530203485", "size": 1270, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/3d2af4236bb7c0a24fa7cf161da8a78530203485"}, {"path": "nunchaku/models/ip_adapter/diffusers_adapters/flux.py", "mode": "100644", "type": "blob", "sha": "a4e4a5500f483a0cbebf5d733f6bbc49426284b3", "size": 4236, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/a4e4a5500f483a0cbebf5d733f6bbc49426284b3"}, {"path": "nunchaku/models/ip_adapter/utils.py", "mode": "100644", "type": "blob", "sha": "502031ddd5b9aae6768e740d552ebabdc12f59a6", "size": 19111, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/502031ddd5b9aae6768e740d552ebabdc12f59a6"}, {"path": "nunchaku/models/linear.py", "mode": "100644", "type": "blob", "sha": "f083cc923804176dbc0f8488bebb639e99cf91b1", "size": 13977, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/f083cc923804176dbc0f8488bebb639e99cf91b1"}, {"path": "nunchaku/models/normalization.py", "mode": "100644", "type": "blob", "sha": "eede8c73bbd1291137ab931351b3ed8d31ba6b29", "size": 5566, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/eede8c73bbd1291137ab931351b3ed8d31ba6b29"}, {"path": "nunchaku/models/pulid", "mode": "040000", "type": "tree", "sha": "8e5c51375865a3561dc487aa95a1e51e47279cb1", "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/trees/8e5c51375865a3561dc487aa95a1e51e47279cb1"}, {"path": "nunchaku/models/pulid/__init__.py", "mode": "100644", "type": "blob", "sha": "e69de29bb2d1d6434b8b29ae775ad8c2e48c5391", "size": 0, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/e69de29bb2d1d6434b8b29ae775ad8c2e48c5391"}, {"path": "nunchaku/models/pulid/encoders_transformer.py", "mode": "100644", "type": "blob", "sha": "01e37d9bbd13337344b234d53cf98ba88ca3d6ae", "size": 9727, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/01e37d9bbd13337344b234d53cf98ba88ca3d6ae"}, {"path": "nunchaku/models/pulid/eva_clip", "mode": "040000", "type": "tree", "sha": "1d6ceebce32f04d809b11a532a4d8a0f31168a07", "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/trees/1d6ceebce32f04d809b11a532a4d8a0f31168a07"}, {"path": "nunchaku/models/pulid/eva_clip/__init__.py", "mode": "100644", "type": "blob", "sha": "7e3d0996ebb88b7a34a09751b79078273c362bc1", "size": 200, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/7e3d0996ebb88b7a34a09751b79078273c362bc1"}, {"path": "nunchaku/models/pulid/eva_clip/constants.py", "mode": "100644", "type": "blob", "sha": "a670bb3fab442baeb9af53b91c312e6982af57ee", "size": 116, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/a670bb3fab442baeb9af53b91c312e6982af57ee"}, {"path": "nunchaku/models/pulid/eva_clip/eva_vit_model.py", "mode": "100644", "type": "blob", "sha": "81db6039bf4c8634160942fee947ba15f195936b", "size": 23254, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/81db6039bf4c8634160942fee947ba15f195936b"}, {"path": "nunchaku/models/pulid/eva_clip/factory.py", "mode": "100644", "type": "blob", "sha": "fd29d4b63ad338257f4920d6590256afd4e1e840", "size": 17353, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/fd29d4b63ad338257f4920d6590256afd4e1e840"}, {"path": "nunchaku/models/pulid/eva_clip/hf_configs.py", "mode": "100644", "type": "blob", "sha": "ddd2c672fdcc895ce2e437f6da5ee439c09840da", "size": 2075, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/ddd2c672fdcc895ce2e437f6da5ee439c09840da"}, {"path": "nunchaku/models/pulid/eva_clip/hf_model.py", "mode": "100644", "type": "blob", "sha": "b6734ab0d84b17445fe87ca42cb3f242853da914", "size": 5281, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/b6734ab0d84b17445fe87ca42cb3f242853da914"}, {"path": "nunchaku/models/pulid/eva_clip/model.py", "mode": "100644", "type": "blob", "sha": "f7754bb585680b8f4daea52f90361b95e5aaa3af", "size": 10979, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/f7754bb585680b8f4daea52f90361b95e5aaa3af"}, {"path": "nunchaku/models/pulid/eva_clip/model_configs", "mode": "040000", "type": "tree", "sha": "fcb212fd3c4e6d27b5e690ac99964de3583c39f5", "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/trees/fcb212fd3c4e6d27b5e690ac99964de3583c39f5"}, {"path": "nunchaku/models/pulid/eva_clip/model_configs/EVA02-CLIP-L-14-336.json", "mode": "100644", "type": "blob", "sha": "5feb657558795a41b99a0e135d05f5962a438638", "size": 655, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/5feb657558795a41b99a0e135d05f5962a438638"}, {"path": "nunchaku/models/pulid/eva_clip/modified_resnet.py", "mode": "100644", "type": "blob", "sha": "8781c5f799d22982666bcfc320bed345bc6e4810", "size": 6802, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/8781c5f799d22982666bcfc320bed345bc6e4810"}, {"path": "nunchaku/models/pulid/eva_clip/pretrained.py", "mode": "100644", "type": "blob", "sha": "bee820cbc4c202e294b2d17a3937a1a1cbe7c1e4", "size": 10993, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/bee820cbc4c202e294b2d17a3937a1a1cbe7c1e4"}, {"path": "nunchaku/models/pulid/eva_clip/rope.py", "mode": "100644", "type": "blob", "sha": "fb2b4bec289f07b0f487137a12a9b1483b52b8cc", "size": 3510, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/fb2b4bec289f07b0f487137a12a9b1483b52b8cc"}, {"path": "nunchaku/models/pulid/eva_clip/transform.py", "mode": "100644", "type": "blob", "sha": "73396e3ef172d9226af15104c483b65c43e70c68", "size": 2875, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/73396e3ef172d9226af15104c483b65c43e70c68"}, {"path": "nunchaku/models/pulid/eva_clip/transformer.py", "mode": "100644", "type": "blob", "sha": "eb8a9e598e6d475978c78e909cef841c46a6e31d", "size": 16094, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/eb8a9e598e6d475978c78e909cef841c46a6e31d"}, {"path": "nunchaku/models/pulid/eva_clip/utils.py", "mode": "100644", "type": "blob", "sha": "b924b9043c95feabb99416428710ab37292dd699", "size": 8853, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/b924b9043c95feabb99416428710ab37292dd699"}, {"path": "nunchaku/models/pulid/pulid_forward.py", "mode": "100644", "type": "blob", "sha": "6e8d84111cc51f03535b10d9f033961384def40d", "size": 5943, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/6e8d84111cc51f03535b10d9f033961384def40d"}, {"path": "nunchaku/models/pulid/utils.py", "mode": "100644", "type": "blob", "sha": "bea2a8aeb0da7dead54a94400c3bc652be3b6d95", "size": 5490, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/bea2a8aeb0da7dead54a94400c3bc652be3b6d95"}, {"path": "nunchaku/models/safety_checker.py", "mode": "100644", "type": "blob", "sha": "e53974003b67e8d1c0de8c92de534a0e03bcb0b7", "size": 3474, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/e53974003b67e8d1c0de8c92de534a0e03bcb0b7"}, {"path": "nunchaku/models/text_encoders", "mode": "040000", "type": "tree", "sha": "f3a76752a25aa24821eb129c051a600b5f77f675", "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/trees/f3a76752a25aa24821eb129c051a600b5f77f675"}, {"path": "nunchaku/models/text_encoders/__init__.py", "mode": "100644", "type": "blob", "sha": "e69de29bb2d1d6434b8b29ae775ad8c2e48c5391", "size": 0, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/e69de29bb2d1d6434b8b29ae775ad8c2e48c5391"}, {"path": "nunchaku/models/text_encoders/linear.py", "mode": "100644", "type": "blob", "sha": "a878c13c634ec67d3e9fb57ab67ee121df238c20", "size": 8105, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/a878c13c634ec67d3e9fb57ab67ee121df238c20"}, {"path": "nunchaku/models/text_encoders/t5_encoder.py", "mode": "100644", "type": "blob", "sha": "6223437978d92948a5355af3b525b0d13e0e1513", "size": 4608, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/6223437978d92948a5355af3b525b0d13e0e1513"}, {"path": "nunchaku/models/text_encoders/tinychat_utils.py", "mode": "100644", "type": "blob", "sha": "03cea464e7371cd3e966aefef3a9842e9ba2caca", "size": 6987, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/03cea464e7371cd3e966aefef3a9842e9ba2caca"}, {"path": "nunchaku/models/transformers", "mode": "040000", "type": "tree", "sha": "1fed6edf147413102117f69f517de4568daa8677", "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/trees/1fed6edf147413102117f69f517de4568daa8677"}, {"path": "nunchaku/models/transformers/__init__.py", "mode": "100644", "type": "blob", "sha": "6233f23b60cbea7bfc2d2f22260a0035f47c8c21", "size": 538, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/6233f23b60cbea7bfc2d2f22260a0035f47c8c21"}, {"path": "nunchaku/models/transformers/transformer_flux.py", "mode": "100644", "type": "blob", "sha": "10f5864460a111de04a90da5cac4005d484bdc64", "size": 39956, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/10f5864460a111de04a90da5cac4005d484bdc64"}, {"path": "nunchaku/models/transformers/transformer_flux_v2.py", "mode": "100644", "type": "blob", "sha": "0ce7ae2cb253e08467c6a6c8442d5d43199b865c", "size": 25576, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/0ce7ae2cb253e08467c6a6c8442d5d43199b865c"}, {"path": "nunchaku/models/transformers/transformer_qwenimage.py", "mode": "100644", "type": "blob", "sha": "8df6f5b25597ef5410fe4374848de94f116c2d99", "size": 23350, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/8df6f5b25597ef5410fe4374848de94f116c2d99"}, {"path": "nunchaku/models/transformers/transformer_sana.py", "mode": "100644", "type": "blob", "sha": "50320a125f56779fbc37ae49b31a329de57f618a", "size": 13664, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/50320a125f56779fbc37ae49b31a329de57f618a"}, {"path": "nunchaku/models/transformers/transformer_zimage.py", "mode": "100644", "type": "blob", "sha": "9b77e2f351e174597efaea50b5d4097252c91a3b", "size": 15483, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/9b77e2f351e174597efaea50b5d4097252c91a3b"}, {"path": "nunchaku/models/transformers/utils.py", "mode": "100644", "type": "blob", "sha": "5782ed3df3d05c9c9577b4ab353cb2ca730691e1", "size": 7652, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/5782ed3df3d05c9c9577b4ab353cb2ca730691e1"}, {"path": "nunchaku/models/unets", "mode": "040000", "type": "tree", "sha": "7f805aa0eb3eb61b90ebcce7b84ab395b49add3d", "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/trees/7f805aa0eb3eb61b90ebcce7b84ab395b49add3d"}, {"path": "nunchaku/models/unets/__init__.py", "mode": "100644", "type": "blob", "sha": "c5c8da0d3cbaea47d0dac94c992facdd20671cf6", "size": 417, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/c5c8da0d3cbaea47d0dac94c992facdd20671cf6"}, {"path": "nunchaku/models/unets/unet_sdxl.py", "mode": "100644", "type": "blob", "sha": "4d099a301499d3eb49ed8e6453deacd3d3104d79", "size": 20175, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/4d099a301499d3eb49ed8e6453deacd3d3104d79"}, {"path": "nunchaku/models/utils.py", "mode": "100644", "type": "blob", "sha": "090540ee37539ddaf229980d3331fec5aab591fd", "size": 9565, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/090540ee37539ddaf229980d3331fec5aab591fd"}, {"path": "nunchaku/ops", "mode": "040000", "type": "tree", "sha": "bcfdb1a496a5672f724076fdbb14b6b58f69162d", "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/trees/bcfdb1a496a5672f724076fdbb14b6b58f69162d"}, {"path": "nunchaku/ops/fused.py", "mode": "100644", "type": "blob", "sha": "dc9346ab4a1f45c18e1f036aa4d1366c26473d08", "size": 6616, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/dc9346ab4a1f45c18e1f036aa4d1366c26473d08"}, {"path": "nunchaku/ops/gemm.py", "mode": "100644", "type": "blob", "sha": "be0b39f9fd18fd7c9a8b2b198f568aa6c61fc9bd", "size": 6116, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/be0b39f9fd18fd7c9a8b2b198f568aa6c61fc9bd"}, {"path": "nunchaku/ops/gemv.py", "mode": "100644", "type": "blob", "sha": "b0658ba9264ae668ed460ade8ff4987b668161bc", "size": 1533, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/b0658ba9264ae668ed460ade8ff4987b668161bc"}, {"path": "nunchaku/ops/quantize.py", "mode": "100644", "type": "blob", "sha": "fdebb68f8af73af6bcecc7f72b463d067ab97a3f", "size": 3264, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/fdebb68f8af73af6bcecc7f72b463d067ab97a3f"}, {"path": "nunchaku/pipeline", "mode": "040000", "type": "tree", "sha": "f86a2a8ab1c7ba7a5df55ec3010a05f571090fde", "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/trees/f86a2a8ab1c7ba7a5df55ec3010a05f571090fde"}, {"path": "nunchaku/pipeline/__init__.py", "mode": "100644", "type": "blob", "sha": "661eef456bf915703015e7bb94367901a33626fd", "size": 142, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/661eef456bf915703015e7bb94367901a33626fd"}, {"path": "nunchaku/pipeline/pipeline_flux_pulid.py", "mode": "100644", "type": "blob", "sha": "36319686c75deb45dee510ad84c4495a5012eb01", "size": 32035, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/36319686c75deb45dee510ad84c4495a5012eb01"}, {"path": "nunchaku/test.py", "mode": "100644", "type": "blob", "sha": "9e13fe3f2c9e88a7fdee5fea1b17f598fb9eb2af", "size": 1349, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/9e13fe3f2c9e88a7fdee5fea1b17f598fb9eb2af"}, {"path": "nunchaku/utils.py", "mode": "100644", "type": "blob", "sha": "6be37f4bb6fcf5391925e8f4a61aa484f8e0cab7", "size": 11621, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/6be37f4bb6fcf5391925e8f4a61aa484f8e0cab7"}, {"path": "pyproject.toml", "mode": "100644", "type": "blob", "sha": "5ef1e353b468cfebfbf8727990155d1d82f5932c", "size": 2575, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/5ef1e353b468cfebfbf8727990155d1d82f5932c"}, {"path": "scripts", "mode": "040000", "type": "tree", "sha": "119231b54b2adc5b34e00e7b8b3f3f845489864e", "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/trees/119231b54b2adc5b34e00e7b8b3f3f845489864e"}, {"path": "scripts/build_all_linux_wheels.sh", "mode": "100644", "type": "blob", "sha": "591e7d3bd907f27976cb79fd8737a593e935d51c", "size": 1039, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/591e7d3bd907f27976cb79fd8737a593e935d51c"}, {"path": "scripts/build_all_windows_wheels.cmd", "mode": "100644", "type": "blob", "sha": "0ec01d54f7c6373036829bccba4c4963ffc6267b", "size": 1125, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/0ec01d54f7c6373036829bccba4c4963ffc6267b"}, {"path": "scripts/build_docker.sh", "mode": "100644", "type": "blob", "sha": "2b448a675929ecd226885a9601f568cac2b8f32a", "size": 1724, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/2b448a675929ecd226885a9601f568cac2b8f32a"}, {"path": "scripts/build_docker_torch27.sh", "mode": "100644", "type": "blob", "sha": "ba4dd138da03579356a74a461339e434594cd2b4", "size": 1156, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/ba4dd138da03579356a74a461339e434594cd2b4"}, {"path": "scripts/build_docker_torch28.sh", "mode": "100644", "type": "blob", "sha": "363f4d74f33d2b7f5f9d92bb072ff2b434209867", "size": 1732, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/363f4d74f33d2b7f5f9d92bb072ff2b434209867"}, {"path": "scripts/build_linux_wheel.sh", "mode": "100644", "type": "blob", "sha": "6f7de50bed08d5f84120ddd21095c862680e503c", "size": 1509, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/6f7de50bed08d5f84120ddd21095c862680e503c"}, {"path": "scripts/build_linux_wheel_torch_nightly.sh", "mode": "100644", "type": "blob", "sha": "caa26f393f742e4344a23f7a2bc5f46a471eb2a8", "size": 1202, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/caa26f393f742e4344a23f7a2bc5f46a471eb2a8"}, {"path": "scripts/build_windows_wheel.cmd", "mode": "100644", "type": "blob", "sha": "029f863e76788985ad7db1faf751593fc6c76737", "size": 1953, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/029f863e76788985ad7db1faf751593fc6c76737"}, {"path": "scripts/build_windows_wheel_torch_nightly.cmd", "mode": "100644", "type": "blob", "sha": "1d75c054bc459bfa0df3dbba56bc2b8f77de2fe2", "size": 1382, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/1d75c054bc459bfa0df3dbba56bc2b8f77de2fe2"}, {"path": "scripts/linux_cleanup.sh", "mode": "100644", "type": "blob", "sha": "c7c9bacd36af81cc1270dbfdbd8c17dde1f37d4e", "size": 151, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/c7c9bacd36af81cc1270dbfdbd8c17dde1f37d4e"}, {"path": "setup.py", "mode": "100644", "type": "blob", "sha": "337b11abdbe05372436b67e440a8378f16405bec", "size": 7890, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/337b11abdbe05372436b67e440a8378f16405bec"}, {"path": "src", "mode": "040000", "type": "tree", "sha": "a6b66b5855ce40f69bd8816d640de7656f22b294", "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/trees/a6b66b5855ce40f69bd8816d640de7656f22b294"}, {"path": "src/FluxModel.cpp", "mode": "100644", "type": "blob", "sha": "56b5119b4a1082b482b823b0ffda0eca8174b153", "size": 65376, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/56b5119b4a1082b482b823b0ffda0eca8174b153"}, {"path": "src/FluxModel.h", "mode": "100644", "type": "blob", "sha": "9e516cb86c6d2c26b19a320e4a5ce5ff4fc2fe6e", "size": 7353, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/9e516cb86c6d2c26b19a320e4a5ce5ff4fc2fe6e"}, {"path": "src/Linear.cpp", "mode": "100644", "type": "blob", "sha": "495d8d8bdcddb0ba266e34dd811f15c38f38d839", "size": 20011, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/495d8d8bdcddb0ba266e34dd811f15c38f38d839"}, {"path": "src/Linear.h", "mode": "100644", "type": "blob", "sha": "024ebec14be224e2ec4065a7b172846ac80e100a", "size": 3666, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/024ebec14be224e2ec4065a7b172846ac80e100a"}, {"path": "src/Module.cpp", "mode": "100644", "type": "blob", "sha": "7aa11fb759446dce2bd11c6f85632a05648f0e02", "size": 558, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/7aa11fb759446dce2bd11c6f85632a05648f0e02"}, {"path": "src/Module.h", "mode": "100644", "type": "blob", "sha": "7ff22a4a453818673a35d647cf3d6f5276dd669d", "size": 10249, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/7ff22a4a453818673a35d647cf3d6f5276dd669d"}, {"path": "src/SanaModel.cpp", "mode": "100644", "type": "blob", "sha": "6a29e2b9b701e38e9ca32eb7b5797dfb0964528d", "size": 14049, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/6a29e2b9b701e38e9ca32eb7b5797dfb0964528d"}, {"path": "src/SanaModel.h", "mode": "100644", "type": "blob", "sha": "c737a7cbe86f3d9fa7f5d04d6950a5c1e855482b", "size": 3091, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/c737a7cbe86f3d9fa7f5d04d6950a5c1e855482b"}, {"path": "src/Serialization.cpp", "mode": "100644", "type": "blob", "sha": "7c18861308d1aaa53f6477bf93f14c5775687286", "size": 7892, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/7c18861308d1aaa53f6477bf93f14c5775687286"}, {"path": "src/Serialization.h", "mode": "100644", "type": "blob", "sha": "224197d845aed371e11fc4379ff7305c01a91935", "size": 1676, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/224197d845aed371e11fc4379ff7305c01a91935"}, {"path": "src/Tensor.h", "mode": "100644", "type": "blob", "sha": "a241b37f27c2c8b47c65f3ddc14ebdcfe191298f", "size": 18544, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/a241b37f27c2c8b47c65f3ddc14ebdcfe191298f"}, {"path": "src/activation.cpp", "mode": "100644", "type": "blob", "sha": "b8b1f1907bb5a6689eea66be0af1ae4b92b8a1d7", "size": 1194, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/b8b1f1907bb5a6689eea66be0af1ae4b92b8a1d7"}, {"path": "src/activation.h", "mode": "100644", "type": "blob", "sha": "834170040042d6d7b86e9cfbec409f3c86b5de09", "size": 1071, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/834170040042d6d7b86e9cfbec409f3c86b5de09"}, {"path": "src/common.h", "mode": "100644", "type": "blob", "sha": "ec997b34ae9692a73d888d52bae8f7e9081565c6", "size": 6825, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/ec997b34ae9692a73d888d52bae8f7e9081565c6"}, {"path": "src/debug.h", "mode": "100644", "type": "blob", "sha": "2e8c430ade3f169c4f655bb803d99503d67fa6ca", "size": 401, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/2e8c430ade3f169c4f655bb803d99503d67fa6ca"}, {"path": "src/interop", "mode": "040000", "type": "tree", "sha": "af796ca4fda5bcf62ef49246cc6bbf474e27e79d", "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/trees/af796ca4fda5bcf62ef49246cc6bbf474e27e79d"}, {"path": "src/interop/torch.cpp", "mode": "100644", "type": "blob", "sha": "6bcc6aa9e013ea1e04f57322a131a0829040156d", "size": 3011, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/6bcc6aa9e013ea1e04f57322a131a0829040156d"}, {"path": "src/interop/torch.h", "mode": "100644", "type": "blob", "sha": "1ac97ef847e019052976d39b2b7fe12097bdb922", "size": 1468, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/1ac97ef847e019052976d39b2b7fe12097bdb922"}, {"path": "src/kernels", "mode": "040000", "type": "tree", "sha": "6723a23c11689500f5e4e3c30d348c3bbc6f326a", "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/trees/6723a23c11689500f5e4e3c30d348c3bbc6f326a"}, {"path": "src/kernels/activation_kernels.cu", "mode": "100644", "type": "blob", "sha": "093c2e10c9c516e8caf292ba05f798fc2dddb032", "size": 4528, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/093c2e10c9c516e8caf292ba05f798fc2dddb032"}, {"path": "src/kernels/activation_kernels.h", "mode": "100644", "type": "blob", "sha": "878f6524b6d9e01511db1478894c0cbeea25c44c", "size": 1041, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/878f6524b6d9e01511db1478894c0cbeea25c44c"}, {"path": "src/kernels/activation_kernels_impl.cuh", "mode": "100644", "type": "blob", "sha": "9d24d88d9e1606094d437a6dd07e1664243809a8", "size": 4272, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/9d24d88d9e1606094d437a6dd07e1664243809a8"}, {"path": "src/kernels/awq", "mode": "040000", "type": "tree", "sha": "3f454e97106d1d717d3ef69c0fa24e1873fd4bac", "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/trees/3f454e97106d1d717d3ef69c0fa24e1873fd4bac"}, {"path": "src/kernels/awq/dequantize.cuh", "mode": "100644", "type": "blob", "sha": "7bda3764417c9299b244aae57d481f8e68d2efb6", "size": 6861, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/7bda3764417c9299b244aae57d481f8e68d2efb6"}, {"path": "src/kernels/awq/gemm_awq.cu", "mode": "100644", "type": "blob", "sha": "a96fe3418cb8b99b8920f92026e91d68c6c845c0", "size": 72246, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/a96fe3418cb8b99b8920f92026e91d68c6c845c0"}, {"path": "src/kernels/awq/gemm_awq.h", "mode": "100644", "type": "blob", "sha": "ef78e10eda138122b9812b14cd8cacf312107193", "size": 150, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/ef78e10eda138122b9812b14cd8cacf312107193"}, {"path": "src/kernels/awq/gemv_awq.cu", "mode": "100644", "type": "blob", "sha": "694fd78d1e2899e8eee506fd15a782d46fbf6ff3", "size": 12076, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/694fd78d1e2899e8eee506fd15a782d46fbf6ff3"}, {"path": "src/kernels/awq/gemv_awq.h", "mode": "100644", "type": "blob", "sha": "0b99de5376e776bdecde46f203099f733af025fa", "size": 183, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/0b99de5376e776bdecde46f203099f733af025fa"}, {"path": "src/kernels/awq/semaphore.h", "mode": "100644", "type": "blob", "sha": "e061c5500504e59db7e14254cf2f09f671d60b8d", "size": 3847, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/e061c5500504e59db7e14254cf2f09f671d60b8d"}, {"path": "src/kernels/dispatch_cutlass.h", "mode": "100644", "type": "blob", "sha": "870f35fcd1a5d77e87522d637d330c3e7bec331a", "size": 448, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/870f35fcd1a5d77e87522d637d330c3e7bec331a"}, {"path": "src/kernels/dispatch_utils.h", "mode": "100644", "type": "blob", "sha": "bfecb48a319a7492eccf2b36a305332a40e24cca", "size": 2574, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/bfecb48a319a7492eccf2b36a305332a40e24cca"}, {"path": "src/kernels/dwconv.cu", "mode": "100644", "type": "blob", "sha": "71db1131eb34ca12b0672631829f20ecdd33d79a", "size": 13123, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/71db1131eb34ca12b0672631829f20ecdd33d79a"}, {"path": "src/kernels/dwconv.h", "mode": "100644", "type": "blob", "sha": "5c67f8cefd35a1d54cb84e8e2d2ef6ceae65e7bf", "size": 184, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/5c67f8cefd35a1d54cb84e8e2d2ef6ceae65e7bf"}, {"path": "src/kernels/gemm_batched.cu", "mode": "100644", "type": "blob", "sha": "33921cf370b38dc60dffbf41d48ab30d788e3162", "size": 3736, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/33921cf370b38dc60dffbf41d48ab30d788e3162"}, {"path": "src/kernels/gemm_batched.h", "mode": "100644", "type": "blob", "sha": "5716905c749b95f11e620517a584cfc586261e12", "size": 292, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/5716905c749b95f11e620517a584cfc586261e12"}, {"path": "src/kernels/gemm_f16.cu", "mode": "100644", "type": "blob", "sha": "7e1c06aa637276c6c8d6e46e2fcb8ec73985453b", "size": 6181, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/7e1c06aa637276c6c8d6e46e2fcb8ec73985453b"}, {"path": "src/kernels/gemm_f16.h", "mode": "100644", "type": "blob", "sha": "7b68388ee4e158468969750f82e4f293dd6ab2bf", "size": 231, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/7b68388ee4e158468969750f82e4f293dd6ab2bf"}, {"path": "src/kernels/gemm_w8a8.cu", "mode": "100644", "type": "blob", "sha": "e36d4c1ab932610f5ac96339bd88d5c7c4114fda", "size": 7117, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/e36d4c1ab932610f5ac96339bd88d5c7c4114fda"}, {"path": "src/kernels/gemm_w8a8.h", "mode": "100644", "type": "blob", "sha": "77c61cd3d54636c9e87eca661660b292dfd4382a", "size": 247, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/77c61cd3d54636c9e87eca661660b292dfd4382a"}, {"path": "src/kernels/layernorm_kernels.cu", "mode": "100644", "type": "blob", "sha": "94269225b1ab4b4ebc8eb7df29f569af0b730b6a", "size": 11474, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/94269225b1ab4b4ebc8eb7df29f569af0b730b6a"}, {"path": "src/kernels/layernorm_kernels.h", "mode": "100644", "type": "blob", "sha": "d28674ba52dba5a3280d44b6904359d4941d01f2", "size": 2140, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/d28674ba52dba5a3280d44b6904359d4941d01f2"}, {"path": "src/kernels/layernorm_kernels_impl.cuh", "mode": "100644", "type": "blob", "sha": "6f4c25ba8ebd735e4bb26565041ae9b68e070c89", "size": 14830, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/6f4c25ba8ebd735e4bb26565041ae9b68e070c89"}, {"path": "src/kernels/misc_kernels.cu", "mode": "100644", "type": "blob", "sha": "a0db9f4323ec071c5b5cccdf3e4fcacb81afa7e2", "size": 12423, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/a0db9f4323ec071c5b5cccdf3e4fcacb81afa7e2"}, {"path": "src/kernels/misc_kernels.h", "mode": "100644", "type": "blob", "sha": "a1b7300609b667d4fa003ae6cc2482030ae732ba", "size": 698, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/a1b7300609b667d4fa003ae6cc2482030ae732ba"}, {"path": "src/kernels/misc_kernels_impl.cuh", "mode": "100644", "type": "blob", "sha": "59fc4dc5adc4bbb0d22db20d2813cfca9b9cf4c7", "size": 8090, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/59fc4dc5adc4bbb0d22db20d2813cfca9b9cf4c7"}, {"path": "src/kernels/reduction_utils.cuh", "mode": "100644", "type": "blob", "sha": "e0389d115529a79760908a7522b418ddcd7689cd", "size": 4628, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/e0389d115529a79760908a7522b418ddcd7689cd"}, {"path": "src/kernels/utils.cuh", "mode": "100644", "type": "blob", "sha": "b46d37000531416f89bcbdb064c7208b0bbe51ce", "size": 11420, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/b46d37000531416f89bcbdb064c7208b0bbe51ce"}, {"path": "src/kernels/zgemm", "mode": "040000", "type": "tree", "sha": "de3a95fe2f26dadd12f207056d6260db7d5e0e1d", "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/trees/de3a95fe2f26dadd12f207056d6260db7d5e0e1d"}, {"path": "src/kernels/zgemm/attention.cu", "mode": "100644", "type": "blob", "sha": "c3595f04804e54c296163cdcefaa9d70b0af5b3a", "size": 4374, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/c3595f04804e54c296163cdcefaa9d70b0af5b3a"}, {"path": "src/kernels/zgemm/attention.cuh", "mode": "100644", "type": "blob", "sha": "01bb569409c32071de8cf8f55802531d1ffad590", "size": 26482, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/01bb569409c32071de8cf8f55802531d1ffad590"}, {"path": "src/kernels/zgemm/epilogues.cuh", "mode": "100644", "type": "blob", "sha": "92d03eecb544c6668eed3f579efbe73eea0cec1f", "size": 37194, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/92d03eecb544c6668eed3f579efbe73eea0cec1f"}, {"path": "src/kernels/zgemm/gemm_base.cuh", "mode": "100644", "type": "blob", "sha": "13412840948f427bb3ba166aca8f3d3a4485e88c", "size": 42382, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/13412840948f427bb3ba166aca8f3d3a4485e88c"}, {"path": "src/kernels/zgemm/gemm_utils.cuh", "mode": "100644", "type": "blob", "sha": "36ea0013513736f54db20830238eb5ddb77a0e44", "size": 15051, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/36ea0013513736f54db20830238eb5ddb77a0e44"}, {"path": "src/kernels/zgemm/gemm_w4a4.cu", "mode": "100644", "type": "blob", "sha": "4e76afcf5297331cbcf7010dfe4f5b71b6420a0b", "size": 6975, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/4e76afcf5297331cbcf7010dfe4f5b71b6420a0b"}, {"path": "src/kernels/zgemm/gemm_w4a4.cuh", "mode": "100644", "type": "blob", "sha": "f49ba7b636756f4dbebcdf663b5d9d36cb21db0c", "size": 47781, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/f49ba7b636756f4dbebcdf663b5d9d36cb21db0c"}, {"path": "src/kernels/zgemm/gemm_w4a4_launch.cuh", "mode": "100644", "type": "blob", "sha": "d7753605a8dd778092c6d6d800f3627c31f80f42", "size": 3532, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/d7753605a8dd778092c6d6d800f3627c31f80f42"}, {"path": "src/kernels/zgemm/gemm_w4a4_launch_bf16_fp4.cu", "mode": "100644", "type": "blob", "sha": "493ae32c49916c0107cf533de85fddfbb5a9eec9", "size": 132, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/493ae32c49916c0107cf533de85fddfbb5a9eec9"}, {"path": "src/kernels/zgemm/gemm_w4a4_launch_bf16_int4.cu", "mode": "100644", "type": "blob", "sha": "d0058cb5b2fa4eb44920649af6225ed405508149", "size": 133, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/d0058cb5b2fa4eb44920649af6225ed405508149"}, {"path": "src/kernels/zgemm/gemm_w4a4_launch_fp16_fp4.cu", "mode": "100644", "type": "blob", "sha": "9f378b265c2616e9b1bd9a01b7480077a8d0eb23", "size": 132, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/9f378b265c2616e9b1bd9a01b7480077a8d0eb23"}, {"path": "src/kernels/zgemm/gemm_w4a4_launch_fp16_int4.cu", "mode": "100644", "type": "blob", "sha": "c3eaf0350bd64d8df20dcf539e7dc8c9102182fa", "size": 133, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/c3eaf0350bd64d8df20dcf539e7dc8c9102182fa"}, {"path": "src/kernels/zgemm/gemm_w4a4_launch_fp16_int4_fasteri2f.cu", "mode": "100644", "type": "blob", "sha": "04fac2d6569f18166fb998d4e58fffbf3541466a", "size": 143, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/04fac2d6569f18166fb998d4e58fffbf3541466a"}, {"path": "src/kernels/zgemm/gemm_w4a4_launch_impl.cuh", "mode": "100644", "type": "blob", "sha": "567fca19b9cf4147526eca5b1c284d25ce1d8a0d", "size": 25729, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/567fca19b9cf4147526eca5b1c284d25ce1d8a0d"}, {"path": "src/kernels/zgemm/gemm_w4a4_test.cu", "mode": "100644", "type": "blob", "sha": "f9260a623fe522a913911ccc243741ce68d7e1dd", "size": 3884, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/f9260a623fe522a913911ccc243741ce68d7e1dd"}, {"path": "src/kernels/zgemm/gemm_w8a8.cu", "mode": "100644", "type": "blob", "sha": "d2f86e0927e806dad91790ab4323b997b5dc0c3b", "size": 6963, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/d2f86e0927e806dad91790ab4323b997b5dc0c3b"}, {"path": "src/kernels/zgemm/gemm_w8a8.cuh", "mode": "100644", "type": "blob", "sha": "9f2c09f1d0ae41221869846065bbdf8eec13ec9e", "size": 20349, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/9f2c09f1d0ae41221869846065bbdf8eec13ec9e"}, {"path": "src/kernels/zgemm/lora.cuh", "mode": "100644", "type": "blob", "sha": "f62043566a3176cf75874e53f6529fc0fc3402f9", "size": 13599, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/f62043566a3176cf75874e53f6529fc0fc3402f9"}, {"path": "src/kernels/zgemm/mma.cuh", "mode": "100644", "type": "blob", "sha": "8f3dc0683b51a7deef6599aafbaf2b87795a761c", "size": 5804, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/8f3dc0683b51a7deef6599aafbaf2b87795a761c"}, {"path": "src/kernels/zgemm/mma_earlycuda.cuh", "mode": "100644", "type": "blob", "sha": "2f2ebd046c157946283af3b11e0b6763f48eb70e", "size": 7833, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/2f2ebd046c157946283af3b11e0b6763f48eb70e"}, {"path": "src/kernels/zgemm/zgemm.h", "mode": "100644", "type": "blob", "sha": "84c7862574ad7b82bcc0a9a70878ffd01ec0ea46", "size": 3669, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/84c7862574ad7b82bcc0a9a70878ffd01ec0ea46"}, {"path": "src/layernorm.cpp", "mode": "100644", "type": "blob", "sha": "2b8282d47ebfc799830f827f63c3d0527ac8ec92", "size": 1910, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/2b8282d47ebfc799830f827f63c3d0527ac8ec92"}, {"path": "src/layernorm.h", "mode": "100644", "type": "blob", "sha": "5b51fe1edb3f3ff91d1f5cb4285d22cff9add8e1", "size": 2186, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/5b51fe1edb3f3ff91d1f5cb4285d22cff9add8e1"}, {"path": "src/pytorch_compat.h", "mode": "100644", "type": "blob", "sha": "eaa12de3454fd09bc30fe15f11f78e5c82b0d9b4", "size": 2385, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/eaa12de3454fd09bc30fe15f11f78e5c82b0d9b4"}, {"path": "tests", "mode": "040000", "type": "tree", "sha": "ed7315baddc6144bac504181e903c0d6e73bfabf", "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/trees/ed7315baddc6144bac504181e903c0d6e73bfabf"}, {"path": "tests/README.md", "mode": "100644", "type": "blob", "sha": "9b75d989f932af1d80e6d67cc4f2a2a4c9b8ada4", "size": 2642, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/9b75d989f932af1d80e6d67cc4f2a2a4c9b8ada4"}, {"path": "tests/__init__.py", "mode": "100644", "type": "blob", "sha": "e69de29bb2d1d6434b8b29ae775ad8c2e48c5391", "size": 0, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/e69de29bb2d1d6434b8b29ae775ad8c2e48c5391"}, {"path": "tests/data", "mode": "040000", "type": "tree", "sha": "af4961d931331b21d0a3fd130f1f4ac044bf61f9", "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/trees/af4961d931331b21d0a3fd130f1f4ac044bf61f9"}, {"path": "tests/data/MJHQ", "mode": "040000", "type": "tree", "sha": "40fd20ee693556277fcc4912971cde728aa0873a", "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/trees/40fd20ee693556277fcc4912971cde728aa0873a"}, {"path": "tests/data/MJHQ/MJHQ.py", "mode": "100644", "type": "blob", "sha": "a58572e72cac265cbe115cc1126c30e36cacf7bd", "size": 6438, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/a58572e72cac265cbe115cc1126c30e36cacf7bd"}, {"path": "tests/data/__init__.py", "mode": "100644", "type": "blob", "sha": "b148eb8125438391e0e29d5872dd9e043d84d76e", "size": 2077, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/b148eb8125438391e0e29d5872dd9e043d84d76e"}, {"path": "tests/flux", "mode": "040000", "type": "tree", "sha": "a2d1b6ab0b7ffecf4e10c5cf89c699f6c4d8bdea", "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/trees/a2d1b6ab0b7ffecf4e10c5cf89c699f6c4d8bdea"}, {"path": "tests/flux/__init__.py", "mode": "100644", "type": "blob", "sha": "e69de29bb2d1d6434b8b29ae775ad8c2e48c5391", "size": 0, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/e69de29bb2d1d6434b8b29ae775ad8c2e48c5391"}, {"path": "tests/flux/test_device_id.py", "mode": "100644", "type": "blob", "sha": "eb8a011a7e47d4febd36c2ff06a70185beefdfe4", "size": 1020, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/eb8a011a7e47d4febd36c2ff06a70185beefdfe4"}, {"path": "tests/flux/test_flux_cache.py", "mode": "100644", "type": "blob", "sha": "b631b6b60d9313b69930274eeb4d1b7b4fad23a8", "size": 1070, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/b631b6b60d9313b69930274eeb4d1b7b4fad23a8"}, {"path": "tests/flux/test_flux_dev.py", "mode": "100644", "type": "blob", "sha": "59e12a1bdcc822f1071b9eb77d7ab337b2758c92", "size": 1010, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/59e12a1bdcc822f1071b9eb77d7ab337b2758c92"}, {"path": "tests/flux/test_flux_dev_IPA.py", "mode": "100644", "type": "blob", "sha": "dbc100181e446ee69fac671a90c7e59259ffba1f", "size": 2960, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/dbc100181e446ee69fac671a90c7e59259ffba1f"}, {"path": "tests/flux/test_flux_dev_loras.py", "mode": "100644", "type": "blob", "sha": "4a555867c25714d84b4f4be9381ebc06f5cbfce9", "size": 4303, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/4a555867c25714d84b4f4be9381ebc06f5cbfce9"}, {"path": "tests/flux/test_flux_dev_pulid.py", "mode": "100644", "type": "blob", "sha": "c24181497c869c65ed783c56dea041b446008b85", "size": 2121, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/c24181497c869c65ed783c56dea041b446008b85"}, {"path": "tests/flux/test_flux_double_fb_cache.py", "mode": "100644", "type": "blob", "sha": "9396480d566d87419b0f8314a8582b8257208494", "size": 1491, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/9396480d566d87419b0f8314a8582b8257208494"}, {"path": "tests/flux/test_flux_double_fb_cache_v2.py", "mode": "100644", "type": "blob", "sha": "7251a8793a9d5393a16985e5cb60ebe49b4aa96b", "size": 7189, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/7251a8793a9d5393a16985e5cb60ebe49b4aa96b"}, {"path": "tests/flux/test_flux_examples.py", "mode": "100644", "type": "blob", "sha": "9eeeda099c9d05d0086d64ace7bd9354d707681c", "size": 560, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/9eeeda099c9d05d0086d64ace7bd9354d707681c"}, {"path": "tests/flux/test_flux_kontext.py", "mode": "100644", "type": "blob", "sha": "557ca260c09c328cdeb59af8064dfe0a4364f6c8", "size": 3704, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/557ca260c09c328cdeb59af8064dfe0a4364f6c8"}, {"path": "tests/flux/test_flux_kontext_lora.py", "mode": "100644", "type": "blob", "sha": "6a81370b077236114be4589b0822f42e8d3cecde", "size": 8094, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/6a81370b077236114be4589b0822f42e8d3cecde"}, {"path": "tests/flux/test_flux_memory.py", "mode": "100644", "type": "blob", "sha": "d8b9b00aaa65f13b83971b64a3e096ea8c766555", "size": 1695, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/d8b9b00aaa65f13b83971b64a3e096ea8c766555"}, {"path": "tests/flux/test_flux_qencoder.py", "mode": "100644", "type": "blob", "sha": "4b75a2effc841c9f42418a3e49ab6db605c39fd0", "size": 561, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/4b75a2effc841c9f42418a3e49ab6db605c39fd0"}, {"path": "tests/flux/test_flux_schnell.py", "mode": "100644", "type": "blob", "sha": "da2842fae472d6875f82b3bfd4fdaeca0c69dc52", "size": 969, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/da2842fae472d6875f82b3bfd4fdaeca0c69dc52"}, {"path": "tests/flux/test_flux_speed.py", "mode": "100644", "type": "blob", "sha": "c186d2444616f83251ff5afe46141c2392cbc715", "size": 2332, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/c186d2444616f83251ff5afe46141c2392cbc715"}, {"path": "tests/flux/test_flux_teacache.py", "mode": "100644", "type": "blob", "sha": "56f67d1da751b19b59297292e2896a6d4ba350bc", "size": 4766, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/56f67d1da751b19b59297292e2896a6d4ba350bc"}, {"path": "tests/flux/test_flux_tools.py", "mode": "100644", "type": "blob", "sha": "027b6c2d622a9a17b2d92ee01b7b6cc85db6f89d", "size": 3592, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/027b6c2d622a9a17b2d92ee01b7b6cc85db6f89d"}, {"path": "tests/flux/test_flux_txt2img_cache_controlnet.py", "mode": "100644", "type": "blob", "sha": "8f13b770f3db0f9aa369d42381594b6e20ec5862", "size": 3940, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/8f13b770f3db0f9aa369d42381594b6e20ec5862"}, {"path": "tests/flux/test_lora_reset.py", "mode": "100644", "type": "blob", "sha": "7c8629f771305a47e274eaee8eb467b918da1229", "size": 1844, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/7c8629f771305a47e274eaee8eb467b918da1229"}, {"path": "tests/flux/test_multiple_batch.py", "mode": "100644", "type": "blob", "sha": "8fb364801df17803574388bcc9d5f462e7a62fde", "size": 871, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/8fb364801df17803574388bcc9d5f462e7a62fde"}, {"path": "tests/flux/test_shuttle_jaguar.py", "mode": "100644", "type": "blob", "sha": "6f043cdab06c536760ad7a663f6f1df4316d214f", "size": 718, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/6f043cdab06c536760ad7a663f6f1df4316d214f"}, {"path": "tests/flux/test_turing.py", "mode": "100644", "type": "blob", "sha": "19de1a98fc6e964174f47f1a16ae1c55ddf6c8fe", "size": 833, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/19de1a98fc6e964174f47f1a16ae1c55ddf6c8fe"}, {"path": "tests/flux/utils.py", "mode": "100644", "type": "blob", "sha": "bf0c6e9dd244b3d022acafcb21f03c7eee9ebd34", "size": 14532, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/bf0c6e9dd244b3d022acafcb21f03c7eee9ebd34"}, {"path": "tests/sana", "mode": "040000", "type": "tree", "sha": "e022155021ba7cfdc9cb33cddd4a7739106ecd68", "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/trees/e022155021ba7cfdc9cb33cddd4a7739106ecd68"}, {"path": "tests/sana/__init__.py", "mode": "100644", "type": "blob", "sha": "e69de29bb2d1d6434b8b29ae775ad8c2e48c5391", "size": 0, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/e69de29bb2d1d6434b8b29ae775ad8c2e48c5391"}, {"path": "tests/sana/test_examples.py", "mode": "100644", "type": "blob", "sha": "d56b0672532ea963f2e67de1d2875710f4c7d50d", "size": 762, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/d56b0672532ea963f2e67de1d2875710f4c7d50d"}, {"path": "tests/utils.py", "mode": "100644", "type": "blob", "sha": "e75526fe1c2c9f495c874300bdbb9b7ac4fc853f", "size": 3772, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/e75526fe1c2c9f495c874300bdbb9b7ac4fc853f"}, {"path": "tests/v1", "mode": "040000", "type": "tree", "sha": "c88c2c8c98698cab4c77e8ce819f3d37b26b1881", "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/trees/c88c2c8c98698cab4c77e8ce819f3d37b26b1881"}, {"path": "tests/v1/__init__.py", "mode": "100644", "type": "blob", "sha": "e69de29bb2d1d6434b8b29ae775ad8c2e48c5391", "size": 0, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/e69de29bb2d1d6434b8b29ae775ad8c2e48c5391"}, {"path": "tests/v1/flux", "mode": "040000", "type": "tree", "sha": "a61f34b1eabde098d76f24179def9d2f34ba405c", "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/trees/a61f34b1eabde098d76f24179def9d2f34ba405c"}, {"path": "tests/v1/flux/__init__.py", "mode": "100644", "type": "blob", "sha": "e69de29bb2d1d6434b8b29ae775ad8c2e48c5391", "size": 0, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/e69de29bb2d1d6434b8b29ae775ad8c2e48c5391"}, {"path": "tests/v1/flux/test_flux1_canny_dev.py", "mode": "100644", "type": "blob", "sha": "73227ef483844d6d3fdac7a1cf4e46b1ccbb2ff7", "size": 6049, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/73227ef483844d6d3fdac7a1cf4e46b1ccbb2ff7"}, {"path": "tests/v1/flux/test_flux1_depth_dev.py", "mode": "100644", "type": "blob", "sha": "9ea9ea1491befa3712fd518ccbf87b5a05c41508", "size": 6050, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/9ea9ea1491befa3712fd518ccbf87b5a05c41508"}, {"path": "tests/v1/flux/test_flux1_dev.py", "mode": "100644", "type": "blob", "sha": "ad2c4eedc0a94e5ad8c439665bd45f2f8608704b", "size": 4719, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/ad2c4eedc0a94e5ad8c439665bd45f2f8608704b"}, {"path": "tests/v1/flux/test_flux1_fill_dev.py", "mode": "100644", "type": "blob", "sha": "b79ba70078e7f25f208558d11afead5b5b989721", "size": 6820, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/b79ba70078e7f25f208558d11afead5b5b989721"}, {"path": "tests/v1/flux/test_flux1_kontext_dev.py", "mode": "100644", "type": "blob", "sha": "e834e72ca41eadcbef9901b22e5a9af3f8193b03", "size": 4489, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/e834e72ca41eadcbef9901b22e5a9af3f8193b03"}, {"path": "tests/v1/flux/test_flux1_krea_dev.py", "mode": "100644", "type": "blob", "sha": "c38630c76b1d312d2d12d746a6c17b356c424dc8", "size": 4636, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/c38630c76b1d312d2d12d746a6c17b356c424dc8"}, {"path": "tests/v1/flux/test_flux1_schnell.py", "mode": "100644", "type": "blob", "sha": "16e565e54245eefda36313fdae705fddc31a263b", "size": 4707, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/16e565e54245eefda36313fdae705fddc31a263b"}, {"path": "tests/v1/qwenimage", "mode": "040000", "type": "tree", "sha": "8858d1044e3f535a8f6e166fedb01ae18b8cb1e3", "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/trees/8858d1044e3f535a8f6e166fedb01ae18b8cb1e3"}, {"path": "tests/v1/qwenimage/__init__.py", "mode": "100644", "type": "blob", "sha": "e69de29bb2d1d6434b8b29ae775ad8c2e48c5391", "size": 0, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/e69de29bb2d1d6434b8b29ae775ad8c2e48c5391"}, {"path": "tests/v1/qwenimage/test_qwenimage.py", "mode": "100644", "type": "blob", "sha": "5f01750b24de8873f0c5f000e6da7b640e718a34", "size": 6031, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/5f01750b24de8873f0c5f000e6da7b640e718a34"}, {"path": "tests/v1/qwenimage/test_qwenimage_controlnet.py", "mode": "100644", "type": "blob", "sha": "2487478a1d22c6d0a4bb7ad378d548c247277f41", "size": 6476, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/2487478a1d22c6d0a4bb7ad378d548c247277f41"}, {"path": "tests/v1/qwenimage/test_qwenimage_edit.py", "mode": "100644", "type": "blob", "sha": "04bdb6f952e1273650fc482809c8136f640ccb63", "size": 4507, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/04bdb6f952e1273650fc482809c8136f640ccb63"}, {"path": "tests/v1/qwenimage/test_qwenimage_edit_2509.py", "mode": "100644", "type": "blob", "sha": "7eb14e9b7820b45277ce37385cc1e1251543ade1", "size": 4498, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/7eb14e9b7820b45277ce37385cc1e1251543ade1"}, {"path": "tests/v1/qwenimage/test_qwenimage_edit_2509_lightning.py", "mode": "100644", "type": "blob", "sha": "c6116f0752674e63cda317f96d2eb9508acdb315", "size": 7043, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/c6116f0752674e63cda317f96d2eb9508acdb315"}, {"path": "tests/v1/qwenimage/test_qwenimage_edit_lightning.py", "mode": "100644", "type": "blob", "sha": "19997e1c2f9333002fe76531be964a546955e1fe", "size": 7020, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/19997e1c2f9333002fe76531be964a546955e1fe"}, {"path": "tests/v1/qwenimage/test_qwenimage_lightning.py", "mode": "100644", "type": "blob", "sha": "0c183792c31ea48f2192aea6bcc683b989f224e2", "size": 9036, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/0c183792c31ea48f2192aea6bcc683b989f224e2"}, {"path": "tests/v1/sdxl", "mode": "040000", "type": "tree", "sha": "bc4f8b07a98820eacdd87546568106c258006550", "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/trees/bc4f8b07a98820eacdd87546568106c258006550"}, {"path": "tests/v1/sdxl/__init__.py", "mode": "100644", "type": "blob", "sha": "e69de29bb2d1d6434b8b29ae775ad8c2e48c5391", "size": 0, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/e69de29bb2d1d6434b8b29ae775ad8c2e48c5391"}, {"path": "tests/v1/sdxl/test_sdxl.py", "mode": "100644", "type": "blob", "sha": "6a4055f9d451519253a02256f92c75ff7a260d6e", "size": 5229, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/6a4055f9d451519253a02256f92c75ff7a260d6e"}, {"path": "tests/v1/sdxl/test_sdxl_turbo.py", "mode": "100644", "type": "blob", "sha": "77ab8a291642ab7ff8df2f835d590b854c256ec1", "size": 7677, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/77ab8a291642ab7ff8df2f835d590b854c256ec1"}, {"path": "tests/v1/test_examples.py", "mode": "100644", "type": "blob", "sha": "3f5f14934faea50c9c48f43def2c8dc946eaccfa", "size": 669, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/3f5f14934faea50c9c48f43def2c8dc946eaccfa"}, {"path": "tests/v1/utils.py", "mode": "100644", "type": "blob", "sha": "671503a587efc01a974f107a0fac3640b365d765", "size": 1539, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/671503a587efc01a974f107a0fac3640b365d765"}, {"path": "tests/v1/z_image", "mode": "040000", "type": "tree", "sha": "7c5eed32db1d8c7f988089689f04ef69a5fcccfc", "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/trees/7c5eed32db1d8c7f988089689f04ef69a5fcccfc"}, {"path": "tests/v1/z_image/__init__.py", "mode": "100644", "type": "blob", "sha": "e69de29bb2d1d6434b8b29ae775ad8c2e48c5391", "size": 0, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/e69de29bb2d1d6434b8b29ae775ad8c2e48c5391"}, {"path": "tests/v1/z_image/test_z_image_turbo.py", "mode": "100644", "type": "blob", "sha": "5f6a0abf57a80039577cda749728c5ab2ef7345d", "size": 7070, "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/blobs/5f6a0abf57a80039577cda749728c5ab2ef7345d"}, {"path": "third_party", "mode": "040000", "type": "tree", "sha": "92a1e0fd864adb49cccc1aca493f8c92c9b5880d", "url": "https://api.github.com/repos/nunchux-ai/nunchaku/git/trees/92a1e0fd864adb49cccc1aca493f8c92c9b5880d"}, {"path": "third_party/Block-Sparse-Attention", "mode": "160000", "type": "commit", "sha": "99511c34554a13ffaa81321834faf66389ffcb30"}, {"path": "third_party/cutlass", "mode": "160000", "type": "commit", "sha": "a75b4ac483166189a45290783cb0a18af5ff0ea5"}, {"path": "third_party/json", "mode": "160000", "type": "commit", "sha": "63258397761b3dd96dd171e5a5ad5aa915834c35"}, {"path": "third_party/mio", "mode": "160000", "type": "commit", "sha": "8b6b7d878c89e81614d05edca7936de41ccdd2da"}, {"path": "third_party/spdlog", "mode": "160000", "type": "commit", "sha": "27cb4c76708608465c413f6d0e6b8d99a4d84302"}], "truncated": false} \ No newline at end of file diff --git a/reproduction/nunchaku_backend/upstream/nunchaku__models__linear.py b/reproduction/nunchaku_backend/upstream/nunchaku__models__linear.py new file mode 100644 index 0000000000000000000000000000000000000000..f083cc923804176dbc0f8488bebb639e99cf91b1 --- /dev/null +++ b/reproduction/nunchaku_backend/upstream/nunchaku__models__linear.py @@ -0,0 +1,415 @@ +""" +Quantized linear layers for Nunchaku. +""" + +import torch +from torch import nn + +from ..ops.gemm import svdq_gemm_w4a4_cuda +from ..ops.gemv import awq_gemv_w4a16_cuda +from ..ops.quantize import svdq_quantize_w4a4_act_fuse_lora_cuda + + +class SVDQW4A4Linear(nn.Module): + """ + `SVDQuant `_ W4A4 quantized linear layer. + + Parameters + ---------- + in_features : int + Input feature dimension. + out_features : int + Output feature dimension. + rank : int, optional + SVD low-rank dimension. Default is 32. + bias : bool, optional + If True, adds a learnable bias. Default is True. + precision : {'int4', 'nvfp4'}, optional + Quantization precision data type ('int4' or 'nvfp4'). Default is 'int4'. + act_unsigned : bool, optional + If True, use unsigned activation quantization (int4 only). Default is False. + torch_dtype : torch.dtype, optional + Parameter dtype. Default is torch.bfloat16. + device : str or torch.device or None, optional + Device for parameters. Default is CPU. + + Attributes + ---------- + in_features : int + out_features : int + rank : int + precision : str + 'int4' or 'nvfp4'. + group_size : int + 64 for int4, 16 for nvfp4. + qweight : nn.Parameter + Packed quantized weights, shape (out_features, in_features // 2), dtype int8. + bias : nn.Parameter or None + Bias tensor. + wscales : nn.Parameter + Weight scales, shape (in_features // group_size, out_features). + Dtype: bfloat16/float16 (int4), float8_e4m3fn (nvfp4). + smooth_factor : nn.Parameter + Smoothing factors, shape (in_features,). + smooth_factor_orig : nn.Parameter + Original smoothing factors, shape (in_features,). (Unused) + proj_down : nn.Parameter + Packed low-rank down projection, shape (in_features, rank), dtype bfloat16/float16. + proj_up : nn.Parameter + Packed low-rank up projection, shape (out_features, rank), dtype bfloat16/float16. + wtscale : float or None + Global weight scale (nvfp4 only). + wcscales : nn.Parameter or None + Channel-wise weight scale (nvfp4 only), shape (out_features,), dtype float8_e4m3fn. + act_unsigned : bool + If True, input activations are unsigned (int4 only). + """ + + def __init__( + self, + in_features: int, + out_features: int, + rank: int = 32, + bias: bool = True, + precision: str = "int4", + act_unsigned: bool = False, + torch_dtype: torch.dtype = torch.bfloat16, + device: str | torch.device | None = None, + ): + super(SVDQW4A4Linear, self).__init__() + if device is None: + device = torch.device("cpu") + self.in_features = in_features + self.out_features = out_features + self.rank = rank + + self.precision = precision + self.torch_dtype = torch_dtype + + if precision == "nvfp4": + self.group_size = 16 + elif precision == "int4": + self.group_size = 64 + else: + raise ValueError(f"Invalid precision: {precision}") + + self.qweight = nn.Parameter( + torch.empty(out_features, in_features // 2, dtype=torch.int8, device=device), requires_grad=False + ) + self.bias = ( + nn.Parameter(torch.empty(out_features, dtype=torch_dtype, device=device), requires_grad=True) + if bias + else None + ) + + self.wscales = nn.Parameter( + torch.empty( + in_features // self.group_size, + out_features, + dtype=torch_dtype if precision == "int4" else torch.float8_e4m3fn, + device=device, + ), + requires_grad=False, + ) + self.smooth_factor = nn.Parameter( + torch.empty(in_features, dtype=torch_dtype, device=device), requires_grad=False + ) + self.smooth_factor_orig = nn.Parameter( + torch.empty(in_features, dtype=torch_dtype, device=device), requires_grad=False + ) + + self.proj_down = nn.Parameter(torch.empty(in_features, rank, dtype=torch_dtype, device=device)) + self.proj_up = nn.Parameter(torch.empty(out_features, rank, dtype=torch_dtype, device=device)) + + if precision == "nvfp4": + self.wcscales = nn.Parameter( + torch.ones(out_features, dtype=torch_dtype, device=device), requires_grad=False + ) + self.wtscale = 1.0 + else: + self.wtscale = None + self.wcscales = None + + self.act_unsigned = act_unsigned + + @classmethod + def from_linear(cls, linear: nn.Linear, **kwargs): + """ + Create an SVDQW4A4Linear from a standard nn.Linear. The weight and bias are dummy tensors. + + Parameters + ---------- + linear : nn.Linear + Source linear layer. + **kwargs + Additional init arguments. + + Returns + ------- + SVDQW4A4Linear + """ + in_features = kwargs.pop("in_features", linear.in_features) + torch_dtype = kwargs.pop("torch_dtype", linear.weight.dtype) + return cls( + in_features=in_features, + out_features=linear.out_features, + bias=linear.bias is not None, + torch_dtype=torch_dtype, + device=linear.weight.device, + **kwargs, + ) + + def forward(self, x: torch.Tensor, output: torch.Tensor | None = None) -> torch.Tensor: + """ + Forward pass with 16-bit input. It will call :meth:`quantize` and :meth:`forward_quant`. + + Parameters + ---------- + x : torch.Tensor, shape (B, S, in_features), dtype float16 or bfloat16 + Input tensor. + output : torch.Tensor or None, optional + Optional output buffer. + + Returns + ------- + torch.Tensor, shape (B, S, out_features) + Output tensor. + + Notes + ----- + B: batch size, S: sequence length + """ + batch_size, seq_len, channels = x.shape + x = x.reshape(batch_size * seq_len, channels) + if output is None: + output = torch.empty(batch_size * seq_len, self.out_features, dtype=x.dtype, device=x.device) + quantized_x, ascales, lora_act_out = self.quantize(x) + output = self.forward_quant(quantized_x, ascales, lora_act_out, output) + output = output.reshape(batch_size, seq_len, -1) + return output + + def quantize(self, x: torch.Tensor, pad_size: int = 256) -> tuple[torch.Tensor, torch.Tensor, torch.Tensor]: + """ + Quantize input to 4-bit and compute low-rank hidden states. It will call :func:`~nunchaku.ops.quantize.svdq_quantize_w4a4_act_fuse_lora_cuda`. + + Parameters + ---------- + x : torch.Tensor, shape (N, in_features), dtype float16 or bfloat16 + Input tensor. + pad_size : int, optional + Batch padding size. Default is 256. + + Returns + ------- + quantized_x : torch.Tensor + Quantized input, shape (pad_size * ceil(N / pad_size), in_features // 2), dtype uint8. + ascales : torch.Tensor + Activation scales, shape (in_features // group_size,), dtype float8_e4m3fn for nvfp4 and input dtype for int4. + lora_act_out : torch.Tensor + Low-rank hidden states, shape (pad_size * ceil(N / pad_size), rank), dtype float32. + + Notes + ----- + N: batch size + """ + quantized_x, ascales, lora_act_out = svdq_quantize_w4a4_act_fuse_lora_cuda( + x, lora_down=self.proj_down, smooth=self.smooth_factor, fp4=self.precision == "nvfp4", pad_size=pad_size + ) + return quantized_x, ascales, lora_act_out + + def forward_quant( + self, + quantized_x: torch.Tensor, + ascales: torch.Tensor, + lora_act: torch.Tensor, + output: torch.Tensor | None = None, + ) -> torch.Tensor: + """ + Forward pass with pre-quantized input. It will call :func:`~nunchaku.ops.gemm.svdq_gemm_w4a4_cuda`. + + Parameters + ---------- + quantized_x : torch.Tensor + Quantized input, shape (N, in_features // 2), dtype uint8. + ascales : torch.Tensor + Activation scales, shape (in_features // group_size,), dtype float8_e4m3fn for nvfp4 and input dtype for int4. + lora_act : torch.Tensor + Low-rank hidden states, shape (N, rank), dtype float32. + output : torch.Tensor or None, optional + Optional output buffer. + + Returns + ------- + torch.Tensor + Output tensor, shape (N, out_features), dtype bfloat16/float16 for int4 and float8_e4m3fn for nvfp4. + + Notes + ----- + N: batch size + """ + if output is None: + output = torch.empty( + quantized_x.shape[0], self.out_features, dtype=self.proj_up.dtype, device=quantized_x.device + ) + + svdq_gemm_w4a4_cuda( + act=quantized_x, + wgt=self.qweight, + out=output, + ascales=ascales, + wscales=self.wscales, + lora_act_in=lora_act, + lora_up=self.proj_up, + bias=self.bias, + fp4=self.precision == "nvfp4", + alpha=self.wtscale, + wcscales=self.wcscales, + act_unsigned=self.act_unsigned, + ) + return output + + def __repr__(self): + return ( + f"SVDQW4A4Linear(in_features={self.in_features}, out_features={self.out_features}, " + f"rank={self.rank}, precision={self.precision}, act_unsigned={self.act_unsigned})" + ) + + +class AWQW4A16Linear(nn.Module): + """ + `AWQ `_ W4A16 quantized linear layer. + + Parameters + ---------- + in_features : int + Input feature dimension. + out_features : int + Output feature dimension. + bias : bool, optional + If True, adds learnable bias. Default is True. + group_size : int, optional + Quantization group size. Default is 64. + torch_dtype : torch.dtype, optional + Parameter dtype. Default is torch.bfloat16. + device : str or torch.device or None, optional + Device for parameters. Default is CPU. + + Attributes + ---------- + in_features : int + out_features : int + group_size : int + qweight : nn.Parameter + Packed quantized weights, shape (out_features // 4, in_features // 2), dtype int32. + bias : nn.Parameter or None + Bias tensor. + wscales : nn.Parameter + Weight scales, shape (in_features // group_size, out_features), dtype float16 or bfloat16. + wzeros : nn.Parameter + Weight zero points, shape (in_features // group_size, out_features), dtype float16 or bfloat16. + """ + + def __init__( + self, + in_features: int, + out_features: int, + bias: bool = True, + group_size: int = 64, + torch_dtype: torch.dtype = torch.bfloat16, + device: str | torch.device | None = None, + ): + super(AWQW4A16Linear, self).__init__() + if device is None: + device = torch.device("cpu") + self.in_features = in_features + self.out_features = out_features + self.group_size = group_size + + self.qweight = nn.Parameter( + torch.empty(out_features // 4, in_features // 2, dtype=torch.int32, device=device), requires_grad=False + ) + self.bias = ( + nn.Parameter(torch.empty(out_features, dtype=torch_dtype, device=device), requires_grad=True) + if bias + else None + ) + self.wscales = nn.Parameter( + torch.empty(in_features // self.group_size, out_features, dtype=torch_dtype, device=device), + requires_grad=False, + ) + self.wzeros = nn.Parameter( + torch.empty(in_features // self.group_size, out_features, dtype=torch_dtype, device=device), + requires_grad=False, + ) + + def forward(self, x: torch.Tensor) -> torch.Tensor: + """ + Forward pass for AWQW4A16Linear. + + Parameters + ---------- + x : torch.Tensor, shape (N, in_features) + Input tensor. + + Returns + ------- + torch.Tensor, shape (N, out_features) + Output tensor. + + Notes + ----- + N: batch size + """ + output = awq_gemv_w4a16_cuda( + in_feats=x, + kernel=self.qweight, + scaling_factors=self.wscales, + zeros=self.wzeros, + m=x.shape[0], + n=self.out_features, + k=self.in_features, + group_size=self.group_size, + ) + if self.bias is not None: + view_shape = [1] * (output.ndim - 1) + [-1] + output.add_(self.bias.view(view_shape)) + return output + + @classmethod + def from_linear( + cls, + linear: nn.Linear, + group_size: int = 64, + torch_dtype: torch.dtype = torch.bfloat16, + device: str = "cpu", + **kwargs, + ): + """ + Create an uninitialized AWQW4A16Linear from a standard nn.Linear. + + Parameters + ---------- + linear : nn.Linear + Source linear layer. + group_size : int, optional + Quantization group size. + torch_dtype : torch.dtype, optional + Parameter dtype. + device : str, optional + Device for parameters. + + Returns + ------- + AWQW4A16Linear + """ + return cls( + in_features=linear.in_features, + out_features=linear.out_features, + bias=linear.bias is not None, + group_size=group_size, + torch_dtype=torch_dtype, + device=device, + ) + + def __repr__(self): + return f"AWQW4A16Linear(in_features={self.in_features}, out_features={self.out_features}, group_size={self.group_size})" diff --git a/reproduction/nunchaku_backend/upstream/nunchaku__models__transformers__transformer_qwenimage.py b/reproduction/nunchaku_backend/upstream/nunchaku__models__transformers__transformer_qwenimage.py new file mode 100644 index 0000000000000000000000000000000000000000..8df6f5b25597ef5410fe4374848de94f116c2d99 --- /dev/null +++ b/reproduction/nunchaku_backend/upstream/nunchaku__models__transformers__transformer_qwenimage.py @@ -0,0 +1,617 @@ +""" +This module provides implementations of NunchakuQwenImageTransformer2DModel and its building blocks. +""" + +import gc +import json +import os +from pathlib import Path +from typing import Any, Dict, List, Optional, Tuple, Union +from warnings import warn + +import numpy as np +import torch +from diffusers.models.attention_processor import Attention +from diffusers.models.modeling_outputs import Transformer2DModelOutput +from diffusers.models.transformers.transformer_qwenimage import ( + QwenEmbedRope, + QwenImageTransformer2DModel, + QwenImageTransformerBlock, +) +from diffusers.utils import logging as diffusers_logging +from huggingface_hub import utils + +from ...utils import get_precision +from ..attention import NunchakuBaseAttention, NunchakuFeedForward +from ..attention_processors.qwenimage import NunchakuQwenImageNaiveFA2Processor +from ..linear import AWQW4A16Linear, SVDQW4A4Linear +from ..utils import CPUOffloadManager, fuse_linears +from .utils import NunchakuModelLoaderMixin, patch_scale_key + +logger = diffusers_logging.get_logger(__name__) + + +class NunchakuQwenAttention(NunchakuBaseAttention): + """ + Nunchaku-optimized quantized attention module for QwenImage. + + Parameters + ---------- + other : Attention + The original QwenImage Attention module to wrap and quantize. + processor : str, default="flashattn2" + The attention processor to use. + **kwargs + Additional arguments for quantization. + """ + + def __init__(self, other: Attention, processor: str = "flashattn2", **kwargs): + super(NunchakuQwenAttention, self).__init__(processor) + self.inner_dim = other.inner_dim + self.inner_kv_dim = other.inner_kv_dim + self.query_dim = other.query_dim + self.use_bias = other.use_bias + self.is_cross_attention = other.is_cross_attention + self.cross_attention_dim = other.cross_attention_dim + self.upcast_attention = other.upcast_attention + self.upcast_softmax = other.upcast_softmax + self.rescale_output_factor = other.rescale_output_factor + self.residual_connection = other.residual_connection + self.dropout = other.dropout + self.fused_projections = other.fused_projections + self.out_dim = other.out_dim + self.out_context_dim = other.out_context_dim + self.context_pre_only = other.context_pre_only + self.pre_only = other.pre_only + self.is_causal = other.is_causal + self.scale_qk = other.scale_qk + self.scale = other.scale + self.heads = other.heads + self.sliceable_head_dim = other.sliceable_head_dim + self.added_kv_proj_dim = other.added_kv_proj_dim + self.only_cross_attention = other.only_cross_attention + self.group_norm = other.group_norm + self.spatial_norm = other.spatial_norm + + self.norm_cross = other.norm_cross + + self.norm_q = other.norm_q + self.norm_k = other.norm_k + self.norm_added_q = other.norm_added_q + self.norm_added_k = other.norm_added_k + + # Fuse the QKV projections for quantization + with torch.device("meta"): + to_qkv = fuse_linears([other.to_q, other.to_k, other.to_v]) + self.to_qkv = SVDQW4A4Linear.from_linear(to_qkv, **kwargs) + self.to_out = other.to_out + self.to_out[0] = SVDQW4A4Linear.from_linear(self.to_out[0], **kwargs) + + assert self.added_kv_proj_dim is not None + # Fuse the additional QKV projections + with torch.device("meta"): + add_qkv_proj = fuse_linears([other.add_q_proj, other.add_k_proj, other.add_v_proj]) + self.add_qkv_proj = SVDQW4A4Linear.from_linear(add_qkv_proj, **kwargs) + self.to_add_out = SVDQW4A4Linear.from_linear(other.to_add_out, **kwargs) + + def forward( + self, + hidden_states: torch.FloatTensor, + encoder_hidden_states: torch.FloatTensor = None, + encoder_hidden_states_mask: torch.FloatTensor = None, + attention_mask: Optional[torch.FloatTensor] = None, + image_rotary_emb: Optional[torch.Tensor] = None, + **kwargs, + ): + """ + Forward pass for NunchakuQwenAttention. + + Parameters + ---------- + hidden_states : torch.FloatTensor + Image stream input. + encoder_hidden_states : torch.FloatTensor, optional + Text stream input. + encoder_hidden_states_mask : torch.FloatTensor, optional + Mask for encoder hidden states. + attention_mask : torch.FloatTensor, optional + Attention mask. + image_rotary_emb : torch.Tensor, optional + Rotary embedding for images. + **kwargs + Additional arguments. + + Returns + ------- + tuple + Attention outputs for image and text streams. + """ + return self.processor( + self, + hidden_states, + encoder_hidden_states, + encoder_hidden_states_mask, + attention_mask, + image_rotary_emb, + **kwargs, + ) + + def set_processor(self, processor: str): + """ + Set the attention processor. + + Parameters + ---------- + processor : str + Name of the processor to use. Only "flashattn2" is supported for now. See :class:`~nunchaku.models.attention_processors.qwenimage.NunchakuQwenImageNaiveFA2Processor`. + + Raises + ------ + ValueError + If the processor is not supported. + """ + if processor == "flashattn2": + self.processor = NunchakuQwenImageNaiveFA2Processor() + else: + raise ValueError(f"Processor {processor} is not supported") + + +class NunchakuQwenImageTransformerBlock(QwenImageTransformerBlock): + """ + Quantized QwenImage Transformer Block. + + This block supports quantized linear layers and joint attention for image and text streams. + + Parameters + ---------- + other : QwenImageTransformerBlock + The original transformer block to wrap and quantize. + scale_shift : float, default=1.0 + Value to add to scale parameters. Default is 1.0. + Nunchaku may have already fused the scale_shift into the linear weights, so you may want to set it to 0. + **kwargs + Additional arguments for quantization. + """ + + def __init__(self, other: QwenImageTransformerBlock, scale_shift: float = 1.0, **kwargs): + super(QwenImageTransformerBlock, self).__init__() + + self.dim = other.dim + self.img_mod = other.img_mod + self.img_mod[1] = AWQW4A16Linear.from_linear(other.img_mod[1], **kwargs) + self.img_norm1 = other.img_norm1 + self.attn = NunchakuQwenAttention(other.attn, **kwargs) + self.img_norm2 = other.img_norm2 + self.img_mlp = NunchakuFeedForward(other.img_mlp, **kwargs) + + # Text processing modules + self.txt_mod = other.txt_mod + self.txt_mod[1] = AWQW4A16Linear.from_linear(other.txt_mod[1], **kwargs) + self.txt_norm1 = other.txt_norm1 + # Text doesn't need separate attention - it's handled by img_attn joint computation + self.txt_norm2 = other.txt_norm2 + self.txt_mlp = NunchakuFeedForward(other.txt_mlp, **kwargs) + + self.scale_shift = scale_shift + + def _modulate(self, x: torch.Tensor, mod_params: torch.Tensor) -> Tuple[torch.Tensor, torch.Tensor]: + """ + Apply modulation to input tensor. + + Parameters + ---------- + x : torch.Tensor + Input tensor. + mod_params : torch.Tensor + Modulation parameters. + + Returns + ------- + tuple + Modulated tensor and gate tensor. + """ + shift, scale, gate = mod_params.chunk(3, dim=-1) + if self.scale_shift != 0: + scale.add_(self.scale_shift) + return x * scale.unsqueeze(1) + shift.unsqueeze(1), gate.unsqueeze(1) + + def forward( + self, + hidden_states: torch.Tensor, + encoder_hidden_states: torch.Tensor, + encoder_hidden_states_mask: torch.Tensor, + temb: torch.Tensor, + image_rotary_emb: Optional[Tuple[torch.Tensor, torch.Tensor]] = None, + joint_attention_kwargs: Optional[Dict[str, Any]] = None, + ) -> Tuple[torch.Tensor, torch.Tensor]: + """ + Forward pass for NunchakuQwenImageTransformerBlock. + + Parameters + ---------- + hidden_states : torch.Tensor + Image stream input. + encoder_hidden_states : torch.Tensor + Text stream input. + encoder_hidden_states_mask : torch.Tensor + Mask for encoder hidden states. + temb : torch.Tensor + Temporal embedding. + image_rotary_emb : tuple of torch.Tensor, optional + Rotary embedding for images. + joint_attention_kwargs : dict, optional + Additional arguments for joint attention. + + Returns + ------- + tuple + Updated encoder_hidden_states and hidden_states. + """ + # Get modulation parameters for both streams + img_mod_params = self.img_mod(temb) # [B, 6*dim] + txt_mod_params = self.txt_mod(temb) # [B, 6*dim] + + # nunchaku's mod_params is [B, 6*dim] instead of [B, dim*6] + img_mod_params = ( + img_mod_params.view(img_mod_params.shape[0], -1, 6).transpose(1, 2).reshape(img_mod_params.shape[0], -1) + ) + txt_mod_params = ( + txt_mod_params.view(txt_mod_params.shape[0], -1, 6).transpose(1, 2).reshape(txt_mod_params.shape[0], -1) + ) + + img_mod1, img_mod2 = img_mod_params.chunk(2, dim=-1) # Each [B, 3*dim] + txt_mod1, txt_mod2 = txt_mod_params.chunk(2, dim=-1) # Each [B, 3*dim] + + # Process image stream - norm1 + modulation + img_normed = self.img_norm1(hidden_states) + img_modulated, img_gate1 = self._modulate(img_normed, img_mod1) + + # Process text stream - norm1 + modulation + txt_normed = self.txt_norm1(encoder_hidden_states) + txt_modulated, txt_gate1 = self._modulate(txt_normed, txt_mod1) + + joint_attention_kwargs = joint_attention_kwargs or {} + attn_output = self.attn( + hidden_states=img_modulated, + encoder_hidden_states=txt_modulated, + encoder_hidden_states_mask=encoder_hidden_states_mask, + image_rotary_emb=image_rotary_emb, + **joint_attention_kwargs, + ) + + # QwenAttnProcessor2_0 returns (img_output, txt_output) when encoder_hidden_states is provided + img_attn_output, txt_attn_output = attn_output + + # Apply attention gates and add residual (like in Megatron) + hidden_states = hidden_states + img_gate1 * img_attn_output + encoder_hidden_states = encoder_hidden_states + txt_gate1 * txt_attn_output + + # Process image stream - norm2 + MLP + img_normed2 = self.img_norm2(hidden_states) + img_modulated2, img_gate2 = self._modulate(img_normed2, img_mod2) + img_mlp_output = self.img_mlp(img_modulated2) + hidden_states = hidden_states + img_gate2 * img_mlp_output + + # Process text stream - norm2 + MLP + txt_normed2 = self.txt_norm2(encoder_hidden_states) + txt_modulated2, txt_gate2 = self._modulate(txt_normed2, txt_mod2) + txt_mlp_output = self.txt_mlp(txt_modulated2) + encoder_hidden_states = encoder_hidden_states + txt_gate2 * txt_mlp_output + + # Clip to prevent overflow for fp16 + if encoder_hidden_states.dtype == torch.float16: + encoder_hidden_states = encoder_hidden_states.clip(-65504, 65504) + if hidden_states.dtype == torch.float16: + hidden_states = hidden_states.clip(-65504, 65504) + + return encoder_hidden_states, hidden_states + + +class NunchakuQwenImageTransformer2DModel(QwenImageTransformer2DModel, NunchakuModelLoaderMixin): + """ + Quantized QwenImage Transformer2DModel. + + This model supports quantized transformer blocks and optional CPU offloading for memory efficiency. + + Parameters + ---------- + *args + Positional arguments for the base model. + **kwargs + Keyword arguments for the base model and quantization. + + Attributes + ---------- + offload : bool + Whether CPU offloading is enabled. + offload_manager : CPUOffloadManager or None + Manager for offloading transformer blocks. + _is_initialized : bool + Whether the model has been patched for quantization. + """ + + def __init__(self, *args, **kwargs): + self.offload = kwargs.pop("offload", False) + self.offload_manager = None + self._is_initialized = False + super().__init__(*args, **kwargs) + + def _patch_model(self, **kwargs): + """ + Patch the transformer blocks for quantization. + + Parameters + ---------- + **kwargs + Additional arguments for quantization. + + Returns + ------- + self + """ + for i, block in enumerate(self.transformer_blocks): + self.transformer_blocks[i] = NunchakuQwenImageTransformerBlock(block, scale_shift=0, **kwargs) + self._is_initialized = True + return self + + @classmethod + @utils.validate_hf_hub_args + def from_pretrained(cls, pretrained_model_name_or_path: str | os.PathLike[str], **kwargs): + """ + Load a quantized model from a pretrained checkpoint. + + Parameters + ---------- + pretrained_model_name_or_path : str or os.PathLike + Path to the pretrained model checkpoint. It can be a local file or a remote HuggingFace path. + **kwargs + Additional arguments for loading and quantization. + + Returns + ------- + NunchakuQwenImageTransformer2DModel + The loaded and quantized model. + + Raises + ------ + AssertionError + If the checkpoint is not a safetensors file. + """ + device = kwargs.get("device", "cpu") + offload = kwargs.get("offload", False) + + torch_dtype = kwargs.get("torch_dtype", torch.bfloat16) + + if isinstance(pretrained_model_name_or_path, str): + pretrained_model_name_or_path = Path(pretrained_model_name_or_path) + + assert pretrained_model_name_or_path.is_file() or pretrained_model_name_or_path.name.endswith( + (".safetensors", ".sft") + ), "Only safetensors are supported" + transformer, model_state_dict, metadata = cls._build_model(pretrained_model_name_or_path, **kwargs) + quantization_config = json.loads(metadata.get("quantization_config", "{}")) + config = json.loads(metadata.get("config", "{}")) + rank = quantization_config.get("rank", 32) + transformer = transformer.to(torch_dtype) + + precision = get_precision() + if precision == "fp4": + precision = "nvfp4" + transformer._patch_model(precision=precision, rank=rank) + + transformer = transformer.to_empty(device=device) + # need to re-init the pos_embed as to_empty does not work on it + transformer.pos_embed = QwenEmbedRope( + theta=10000, axes_dim=list(config.get("axes_dims_rope", [16, 56, 56])), scale_rope=True + ) + + patch_scale_key(transformer, model_state_dict) + + transformer.load_state_dict(model_state_dict) + transformer.set_offload(offload) + + return transformer + + def set_offload(self, offload: bool, **kwargs): + """ + Enable or disable asynchronous CPU offloading for transformer blocks. + + Parameters + ---------- + offload : bool + Whether to enable offloading. + **kwargs + Additional arguments for offload manager. + + See Also + -------- + :class:`~nunchaku.models.utils.CPUOffloadManager` + """ + if offload == self.offload: + # nothing changed, just return + return + self.offload = offload + if offload: + self.offload_manager = CPUOffloadManager( + self.transformer_blocks, + use_pin_memory=kwargs.get("use_pin_memory", True), + on_gpu_modules=[ + self.img_in, + self.txt_in, + self.txt_norm, + self.time_text_embed, + self.norm_out, + self.proj_out, + ], + num_blocks_on_gpu=kwargs.get("num_blocks_on_gpu", 1), + ) + else: + self.offload_manager = None + gc.collect() + torch.cuda.empty_cache() + + def forward( + self, + hidden_states: torch.Tensor, + encoder_hidden_states: torch.Tensor = None, + encoder_hidden_states_mask: torch.Tensor = None, + timestep: torch.LongTensor = None, + img_shapes: Optional[List[Tuple[int, int, int]]] = None, + txt_seq_lens: Optional[List[int]] = None, + guidance: torch.Tensor = None, + attention_kwargs: Optional[Dict[str, Any]] = None, + controlnet_block_samples=None, + return_dict: bool = True, + ) -> Union[torch.Tensor, Transformer2DModelOutput]: + """ + Forward pass for the Nunchaku QwenImage transformer model with ControlNet support. + + Parameters + ---------- + hidden_states : torch.Tensor + Image stream input of shape `(batch_size, image_sequence_length, in_channels)`. + encoder_hidden_states : torch.Tensor, optional + Text stream input of shape `(batch_size, text_sequence_length, joint_attention_dim)`. + encoder_hidden_states_mask : torch.Tensor, optional + Mask for encoder hidden states of shape `(batch_size, text_sequence_length)`. + timestep : torch.LongTensor, optional + Timestep for temporal embedding. + img_shapes : list of tuple, optional + Image shapes for rotary embedding. + txt_seq_lens : list of int, optional + Text sequence lengths. + guidance : torch.Tensor, optional + Guidance tensor (for classifier-free guidance). + attention_kwargs : dict, optional + Additional attention arguments. A kwargs dictionary that if specified is passed along to the `AttentionProcessor`. + controlnet_block_samples : optional + ControlNet block samples for residual connections. + return_dict : bool, default=True + Whether to return a dict or tuple. + + Returns + ------- + torch.Tensor or Transformer2DModelOutput + If `return_dict` is True, an [`~models.transformer_2d.Transformer2DModelOutput`] is returned, otherwise a + `tuple` where the first element is the sample tensor. + """ + device = hidden_states.device + if self.offload: + self.offload_manager.set_device(device) + + hidden_states = self.img_in(hidden_states) + + timestep = timestep.to(hidden_states.dtype) + encoder_hidden_states = self.txt_norm(encoder_hidden_states) + encoder_hidden_states = self.txt_in(encoder_hidden_states) + + if guidance is not None: + guidance = guidance.to(hidden_states.dtype) * 1000 + + temb = ( + self.time_text_embed(timestep, hidden_states) + if guidance is None + else self.time_text_embed(timestep, guidance, hidden_states) + ) + + image_rotary_emb = self.pos_embed(img_shapes, txt_seq_lens, device=hidden_states.device) + + compute_stream = torch.cuda.current_stream() + if self.offload: + self.offload_manager.initialize(compute_stream) + for block_idx, block in enumerate(self.transformer_blocks): + with torch.cuda.stream(compute_stream): + if self.offload: + block = self.offload_manager.get_block(block_idx) + + if torch.is_grad_enabled() and self.gradient_checkpointing: + encoder_hidden_states, hidden_states = self._gradient_checkpointing_func( + block, + hidden_states, + encoder_hidden_states, + encoder_hidden_states_mask, + temb, + image_rotary_emb, + ) + else: + encoder_hidden_states, hidden_states = block( + hidden_states=hidden_states, + encoder_hidden_states=encoder_hidden_states, + encoder_hidden_states_mask=encoder_hidden_states_mask, + temb=temb, + image_rotary_emb=image_rotary_emb, + joint_attention_kwargs=attention_kwargs, + ) + + # controlnet residual - same logic as in diffusers QwenImageTransformer2DModel + if controlnet_block_samples is not None: + interval_control = len(self.transformer_blocks) / len(controlnet_block_samples) + interval_control = int(np.ceil(interval_control)) + hidden_states = hidden_states + controlnet_block_samples[block_idx // interval_control] + + if self.offload: + self.offload_manager.step(compute_stream) + + hidden_states = self.norm_out(hidden_states, temb) + output = self.proj_out(hidden_states) + + if self.offload: + torch.cuda.empty_cache() + + if not return_dict: + return (output,) + + return Transformer2DModelOutput(sample=output) + + def to(self, *args, **kwargs): + """ + Override the default ``.to()`` method. + + If offload is enabled, prevents moving the model to GPU. + Prevents changing dtype after quantization. + + Parameters + ---------- + *args + Positional arguments for ``.to()``. + **kwargs + Keyword arguments for ``.to()``. + + Returns + ------- + self + + Raises + ------ + ValueError + If attempting to change dtype after quantization. + """ + device_arg_or_kwarg_present = any(isinstance(arg, torch.device) for arg in args) or "device" in kwargs + dtype_present_in_args = "dtype" in kwargs + + # Try converting arguments to torch.device in case they are passed as strings + for arg in args: + if not isinstance(arg, str): + continue + try: + torch.device(arg) + device_arg_or_kwarg_present = True + except RuntimeError: + pass + + if not dtype_present_in_args: + for arg in args: + if isinstance(arg, torch.dtype): + dtype_present_in_args = True + break + + if dtype_present_in_args and self._is_initialized: + raise ValueError( + "Casting a quantized model to a new `dtype` is unsupported. To set the dtype of unquantized layers, please " + "use the `torch_dtype` argument when loading the model using `from_pretrained` or `from_single_file`." + ) + if self.offload: + if device_arg_or_kwarg_present: + warn("Skipping moving the model to GPU as offload is enabled", UserWarning) + return self + return super(type(self), self).to(*args, **kwargs) diff --git a/reproduction/nunchaku_backend/upstream/nunchaku__ops__quantize.py b/reproduction/nunchaku_backend/upstream/nunchaku__ops__quantize.py new file mode 100644 index 0000000000000000000000000000000000000000..fdebb68f8af73af6bcecc7f72b463d067ab97a3f --- /dev/null +++ b/reproduction/nunchaku_backend/upstream/nunchaku__ops__quantize.py @@ -0,0 +1,81 @@ +""" +This module provides Python wrappers for Nunchaku's high-performance SVDQuant quantization CUDA kernels. +""" + +import torch + +from .._C import ops +from ..utils import ceil_divide + + +def svdq_quantize_w4a4_act_fuse_lora_cuda( + input: torch.Tensor, + output: torch.Tensor | None = None, + oscales: torch.Tensor | None = None, + lora_down: torch.Tensor | None = None, + lora_act_out: torch.Tensor | None = None, + smooth: torch.Tensor | None = None, + fuse_glu: bool = False, + fp4: bool = False, + pad_size: int = 256, +) -> tuple[torch.Tensor, torch.Tensor, torch.Tensor]: + """ + Quantizes activations and computes LoRA down-projection using SVDQuant W4A4 CUDA kernel. + + Parameters + ---------- + input : torch.Tensor, shape (M, K), dtype bfloat16/float16 + Input activations. + output : torch.Tensor or None, shape (M_pad, K // 2), dtype uint8, optional + Packed output tensor for quantized activations. Allocated if None. + oscales : torch.Tensor or None, shape (K // G, M_pad), dtype float8_e4m3fn for NVFP4 or input dtype for INT4, optional + Output scales tensor. Allocated if None. + lora_down : torch.Tensor or None, shape (K, R), dtype bfloat16/float16, optional + Packed LoRA down-projection weights. + lora_act_out : torch.Tensor or None, shape (M_pad, R), dtype float32, optional + Packed output tensor for LoRA activations. Allocated if None. + smooth : torch.Tensor or None, optional, dtype bfloat16/float16 + Smoothing factor for quantization. + fuse_glu : bool, default=False + If True, fuse GLU activation. + fp4 : bool, default=False + If True, use NVFP4 quantization; else INT4. + pad_size : int, default=256 + Pad batch size to a multiple of this value for efficient CUDA execution. + + Returns + ------- + output : torch.Tensor, shape (M_pad, K // 2), dtype uint8 + Packed quantized activations. + oscales : torch.Tensor, shape (K // G, M_pad), dtype float8_e4m3fn for NVFP4 or input dtype for INT4 + Output scales. + lora_act_out : torch.Tensor, shape (M_pad, R), dtype float32 + Packed LoRA activation output. + + Notes + ----- + Notations: + + - M: batch size + - K: input channels + - R: LoRA rank + - G: group size (64 for INT4, 16 for NVFP4) + - M_pad: padded batch size = ceil(M / pad_size) * pad_size + """ + batch_size, channels = input.shape + rank = lora_down.shape[1] + batch_size_pad = ceil_divide(batch_size, pad_size) * pad_size + if output is None: + output = torch.empty(batch_size_pad, channels // 2, dtype=torch.uint8, device=input.device) + if oscales is None: + if fp4: + assert channels % 16 == 0 + oscales = torch.empty(channels // 16, batch_size_pad, dtype=torch.float8_e4m3fn, device=input.device) + else: + assert channels % 64 == 0 + oscales = torch.empty(channels // 64, batch_size_pad, dtype=input.dtype, device=input.device) + if lora_act_out is None: + lora_act_out = torch.empty(batch_size_pad, rank, dtype=torch.float32, device=input.device) + + ops.quantize_w4a4_act_fuse_lora(input, output, oscales, lora_down, lora_act_out, smooth, fuse_glu, fp4) + return output, oscales, lora_act_out diff --git a/reproduction/nunchaku_backend/v3_notes.md b/reproduction/nunchaku_backend/v3_notes.md new file mode 100644 index 0000000000000000000000000000000000000000..ca14450d5b217901dc810140e88ebac6d55944b6 --- /dev/null +++ b/reproduction/nunchaku_backend/v3_notes.md @@ -0,0 +1,105 @@ +# Activation-output calibrated V3 + +This is a separate experimental conversion path. `convert.py`, existing checkpoints, and `runtime.py` are unchanged. V3 improves the calibration objective and search; it does not change inference steps. Lower linear output error does not establish nearly lossless image quality. + +## Source-supported changes + +The pinned [DeepCompressor low-rank calibrator](https://github.com/mit-han-lab/deepcompressor/blob/69f3473f5e1c1504bae35cc50c7858ef900a9b17/deepcompressor/calib/lowrank.py) repeatedly computes `L = SVD_rank(W - Q_previous)` then `Q = INT4(W - L)`. Its output-error objective includes activation quantization, retains the best candidate, and can stop at the first worsening. The [default recipe](https://github.com/mit-han-lab/deepcompressor/blob/69f3473f5e1c1504bae35cc50c7858ef900a9b17/examples/diffusion/configs/svdquant/__default__.yaml) allows up to 100 iterations. V3 now implements that recurrence, instead of the original one-pass fit. + +Smoothing searches identity, activation-only `amax(X)^alpha`, and SmoothQuant `amax(X)^alpha / amax(W)^(1-alpha)`. The upstream default explores 39 candidates. The initial V3 probe uses a smaller explicit grid to measure cost and sensitivity before a full export. The actual shipped Klein checkpoint has rank32 and identity smoothing; rank32/identity is therefore included in the V3 grid. + +Additional experimental candidates use diagonal activation RMS weighting in SVD and a damped least-squares correction to the BF16 low-rank up projection. Corrections fit training data only, with validation rejecting overfitting. These extensions are not claimed to reproduce the official recipe. + +## Data and selection + +Each layer receives three independent BF16 input matrices: training, validation, and heldout. The intended collection uses different prompts/reference images for these sets and samples across early, middle, and late denoising calls. Preserve provenance about cached conditioning prefixes versus target-image tokens. Adjacent tokens from the same denoising call do not establish independent heldout coverage. + +Training supplies smoothing statistics, RMS weights, and output corrections. Validation chooses alpha, rank, iteration, weighting, and correction strength by MSE against the original BF16 linear output. Heldout inputs are evaluated only after the winner is frozen. The original one-pass conversion remains a candidate; validation cannot worsen relative to that candidate, but unseen outputs can. + +CUDA selection uses the **actual Nunchaku W4A4 kernel**, including activation rounding and groupwise accumulation. CPU tests use the independently validated groupwise reference. An additional A16 residual proxy helps diagnose weight versus activation sensitivity, but it also changes accumulator arithmetic and is not an exact decomposition of activation error or a deployed W4A16 backend. + +The baseline is regenerated using the original absmax archive, seed, rank, and randomized-SVD settings when supplied. A CPU-to-CUDA change in SVD computation can alter the regenerated factors; this is not necessarily byte-identical to an existing exported baseline. Baseline-only CPU behavior is tested byte-for-byte against `convert.py`. + +## API and files + +- `optimize_v3.py`: `optimize_linear_weight(weight, train, validation, heldout, bias, ...) -> (packed_state, stats, reference)`. +- `export_v3.py`: streaming, resumable exporter; one layer per shard and rank chosen per layer. The stable runtime supports this shard layout. +- `v3_test.py`: CPU tests for meaningful synthetic improvement, heldout independence, exact baseline preservation, input/RNG preservation, zero inputs/bias, and invalid data. + +Input format supported by `--activations`: files named `.safetensors`, each with `train`, `validation`, and `heldout` keys, plus a provenance manifest. Also supports combined archives with `.` keys or separate split directories with `.inputs` keys. Archives are lazily read per layer and hashed for resume identity. + +Example bounded first probe (GPU must first be reserved by the main agent): + +```sh +python -m nunchaku_backend.export_v3 \ + --model-path /cache/huggingface/hub/models--Qwen--Qwen-Image-2.1/snapshots/b3179ad355be050328e483a9dfdd9e60cd62adfa \ + --activations /cache/qwen21-activation-v3 \ + --baseline-calibration /cache/qwen21-calibration \ + --out /cache/qwen-v3-probe --device cuda:0 \ + --ranks 32 64 --alphas .5 --families activation_only smoothquant \ + --weighting none --iterations 8 --verbose-candidates \ + --layers transformer_blocks.0.attn.to_q \ + transformer_blocks.15.img_mlp.gate_layer \ + transformer_blocks.31.attn.to_out.0 +``` + +`--layers` produces an intentionally incomplete checkpoint suitable for diagnosis, not loading the whole model. Use a new output directory for a different recipe. A complete export requires all 224 linears and all three splits. No artifact is called lossless from these metrics alone. + +## Remaining differences from upstream + +Upstream can group Q/K/V for a shared low-rank branch and score Q/K candidates on attention output; V3 currently optimizes separate linear outputs. Attention softmax and prefix-KV sensitivity can amplify errors hidden by pooled linear MSE. Randomized FP32 SVD replaces the upstream full FP64 SVD; V3 retains the same projection seed during recurrence to reduce stochastic early-stop noise. We use BF16 factors and the same signed group64 packing already kernel-validated. + +Optional final GPTQ residual fitting is implemented in `gptq_v3.py`, activated with `--final-gptq`. It uses a full damped training Hessian, preserves original input-channel order/group64, and is retained only if actual-kernel validation improves. It does not directly optimize A4 activation rounding. GPTQ need not leave +/-7 extrema in each group, so the V3 CPU reference uses explicit stored scales. Stable reference files remain unchanged. + +`--factorization up_singular` supports the upstream `down=Vh`, `up=U*Sigma` placement for an explicit ablation against balanced square-root factors; exact arithmetic is equivalent, BF16 rounding may differ. No unsigned activation toggle is introduced: the generic signed quantizer does not become an unsigned quantizer by changing the GEMM flag. + +## Initial real-Qwen probe, September 21 + +These measurements use actual RTX4070TiSUPER Nunchaku execution and BF16-teacher inputs from independent training/validation/heldout jobs, with a **16-step calibration schedule**. They are preliminary: the next dataset uses40 steps and conditioning-token stratification. Editing calibration and heldout instructions differ but share the archived reference image. Final image evaluation uses different references. + +| Layer | One-pass rank32 heldout relative L2 | V3 heldout relative L2 | Heldout MSE / baseline | Selected | +|---|---:|---:|---:|---| +| block0 Q | 2.55% | 1.18% | 0.215 | activation-only alpha.5, rank128, iterative fit, output correction, GPTQ | +| block15 MLP gate | 9.18% | 7.19% | 0.614 | activation-only alpha.5, rank128, iterative fit, output correction; GPTQ rejected | +| block31 attention output | 6.58% | 4.67% | 0.505 | SmoothQuant alpha.5, rank128, RMS weighting, output correction, GPTQ | + +The large search tested rank32/64/128, both weighting choices, identity and the two alpha.5 smoothing families, up to8 iterations. All three winners reached iteration7 (zero-indexed), motivating a larger iteration budget. RMS weighting improved the last layer's validation MSE by only0.48% relative to its best unweighted candidate, and worsened the other two by about7%; omitting it in the full search frees time for additional alphas/iterations. Search time was4.8–10.2 seconds per layer, including actual-kernel candidate evaluation and GPTQ. + +Evidence: `results/nunchaku-v3-probe-manifest.json` and `results/nunchaku-v3-probe-gptq-manifest.json` preserve every candidate and complete source/calibration fingerprints. These are **per-linear measurements, not image-quality or losslessness results**. + +The improvement is not solely higher rank: holding rank128, smoothing family and weighting fixed, the final validation MSE divided by iteration0/no-output-correction MSE is0.621 for Q,0.877 for the gate, and0.691 for attention output. This is a validation-set ablation of the additional fitting steps; an independent heldout ablation for every intermediate candidate was intentionally not used for selection. + +Proposed full40-step-data export, after the short factorization ablation (record the chosen `--factorization` in the manifest): + +```sh +python -m nunchaku_backend.export_v3 \ + --model-path /cache/huggingface/hub/models--Qwen--Qwen-Image-2.1/snapshots/b3179ad355be050328e483a9dfdd9e60cd62adfa \ + --activations /cache/qwen21-activation-v3-40 \ + --baseline-calibration /cache/qwen21-calibration \ + --out /cache/qwen-nunchaku-v3-r128 --device cuda:0 \ + --ranks 128 --alphas .25 .5 .75 \ + --families activation_only smoothquant --weighting none \ + --iterations 16 --final-gptq --factorization balanced +``` + +The output name describes the searched rank. The rank32 baseline can still win individual layers; inspect manifest layer ranks rather than assuming every layer is rank128. + +## Completed full40-step-data export + +The full export completed September21 in1656.1 seconds (27.60 minutes, manifest creation to completion), producing225 shards at `/cache/qwen-nunchaku-v3-r128`. All224 selected rank128. CPU loading took3.58 seconds; all1417 state entries loaded, every parameter was on CPU, and all floating parameters were finite. Parameter storage totals4,655,177,728 bytes (4.335GiB); this is **not peak inference VRAM**. Runtime was Torch2.8.0+cu128 and Nunchaku1.2.1+cu12.8torch2.8. + +All224 layers improved heldout MSE versus the regenerated one-pass rank32 baseline. The unweighted median layer MSE ratio is0.6064, range0.1960–0.7612. GPTQ won25 layers and was rejected for199. Smoothing choices were activation-only180, SmoothQuant36, identity8. Best candidates hit the16-iteration cap in186 layers; this suggests additional fitting remains possible, but no end-to-end benefit from further iterations has yet been measured. + +| Projection role | Baseline median relative L2 | V3 median relative L2 | A16 proxy relative L2 | Median MSE ratio | +|---|---:|---:|---:|---:| +| Q | 6.55% | 4.39% | 2.60% | 0.448 | +| K | 6.52% | 4.13% | 2.36% | 0.406 | +| V | 12.87% | 10.83% | 6.50% | 0.708 | +| Attention output | 13.12% | 10.03% | 6.18% | 0.578 | +| MLP gate | 9.81% | 7.56% | 4.52% | 0.626 | +| MLP projection | 10.85% | 8.97% | 5.26% | 0.658 | +| MLP output | 12.18% | 10.13% | 5.69% | 0.653 | + +These40-step, stratified-input metrics use a different sampling distribution from the initial16-step probe; their raw errors cannot be compared as a before/after change in quality. The A16 proxy changes both activation quantization and accumulator behavior. Higher linear error suggests V, attention output and MLP output as additional mixed-precision candidates; gate sensitivity is also plausible because it feeds a nonlinearity. Actual denoiser and image comparisons must choose the final configuration. + +Complete artifacts are `results/nunchaku-v3-manifest.json.gz`, `results/nunchaku-v3-summary.json`, `results/nunchaku-v3-cpu-validation.json`, `results/nunchaku-v3-source-sha256.txt`, and `results/nunchaku-v3-export-container.txt`. The manifest retains every candidate and fingerprints all calibration files. The main agent owns all GPU evaluation after export; no further GPU work was launched by this subtask. diff --git a/reproduction/nunchaku_backend/v3_test.py b/reproduction/nunchaku_backend/v3_test.py new file mode 100644 index 0000000000000000000000000000000000000000..fd4102ef631c6b49e3cf77508d47f33b9c2f36d3 --- /dev/null +++ b/reproduction/nunchaku_backend/v3_test.py @@ -0,0 +1,212 @@ +"""CPU behavioral checks for experimental activation-output calibration. + +Run: PYTHONPATH=/opt/deepcompressor: python -m nunchaku_backend.v3_test +These synthetic checks do not establish final Qwen image quality. +""" + +import json +import unittest + +import torch + +from .baseline_candidate import convert_linear_weight +from .convert_reference import reference_forward_groupwise +from .optimize_v3 import Candidate, _reference_explicit_scales, optimize_linear_weight, pack_candidate + + +def _candidate_from_reference(reference): + return Candidate( + reference["residual_dequant"], reference["weight_scales"], + reference["down_unpacked"], reference["up_unpacked"], reference["smooth"], {}, + ) + + +def _unpack_weight_scales(packed, output_features): + groups = packed.numel() // output_features + return packed.reshape(output_features // 128, groups, 1, 8, 4, 2, 2).permute( + 0, 2, 3, 5, 4, 6, 1, + ).contiguous().reshape(output_features, groups) + + +class V3Tests(unittest.TestCase): + def setUp(self): + torch.set_num_threads(2) + + def test_fixed_smoothing_excludes_other_families(self): + weight, (train, validation, heldout) = self.fixture() + _, stats, _ = optimize_linear_weight( + weight, train, validation, heldout, ranks=(16,), baseline_rank=16, + alphas=(0.5,), smoothing_families=("activation_only",), + fixed_smoothing="activation_only", weighting=("none",), iterations=2, + ) + families = {row["family"] for row in stats["history"]} + self.assertEqual(families, {"one_pass_baseline", "activation_only"}) + self.assertEqual(stats["search"]["fixed_smoothing"], "activation_only") + with self.assertRaises(ValueError): + optimize_linear_weight(weight, train, validation, ranks=(16,), baseline_rank=16, + fixed_smoothing="activation_only", alphas=(0.25, 0.5)) + + def fixture(self): + generator = torch.Generator().manual_seed(1947) + weight = ( + torch.randn(128, 12, generator=generator) @ torch.randn(12, 128, generator=generator) * 0.07 + + torch.randn(128, 128, generator=generator) * 0.025 + ).bfloat16() + channel_scale = torch.ones(128) + channel_scale[::17] = 15 + inputs = [(torch.randn(n, 128, generator=generator) * channel_scale).bfloat16() for n in (80, 64, 72)] + return weight, inputs + + def test_lowrank_outliers_improve_independent_validation_and_test(self): + weight, (train, validation, heldout) = self.fixture() + state, stats, reference = optimize_linear_weight( + weight, train, validation, heldout, ranks=(16, 32), alphas=(0.5,), + iterations=3, weighting=("none", "rms"), baseline_rank=16, + output_correction=True, + ) + # Fixed independent draws verify a substantive improvement rather + # than accepting a train-set-only fit or just any finite checkpoint. + self.assertLess(stats["validation_mse_ratio_to_baseline"], 0.8) + self.assertLess(stats["heldout_mse_ratio_to_baseline"], 0.8) + best_reported = min(row["validation"]["mse"] for row in stats["history"]) + self.assertEqual(stats["validation"]["mse"], best_reported) + teacher = torch.nn.functional.linear(heldout, weight).float() + actual = reference_forward_groupwise(heldout, reference) + measured_mse = (actual.double() - teacher.double()).square().mean().item() + self.assertEqual(measured_mse, stats["heldout"]["mse"]) + self.assertEqual(state["proj_down"].shape[1], stats["selected"]["rank"]) + self.assertTrue(any(row["output_correction"] > 0 for row in stats["history"])) + json.dumps(stats, allow_nan=False) + + def test_heldout_never_changes_selection(self): + weight, (train, validation, heldout) = self.fixture() + opts = dict(ranks=(16, 32), alphas=(0.5,), iterations=2, + weighting=("none",), baseline_rank=16, output_correction=True) + # Deliberately change the test distribution, including its length. + other_test = torch.zeros(11, 128, dtype=torch.bfloat16) + other_test[:, 9] = 40 + for final_gptq in (False, True): + with self.subTest(final_gptq=final_gptq): + state_a, stats_a, _ = optimize_linear_weight(weight, train, validation, heldout, final_gptq=final_gptq, **opts) + state_b, stats_b, _ = optimize_linear_weight(weight, train, validation, other_test, final_gptq=final_gptq, **opts) + self.assertEqual(stats_a["selected"], stats_b["selected"]) + self.assertEqual(stats_a["history"], stats_b["history"]) + self.assertEqual(state_a.keys(), state_b.keys()) + for key in state_a: + self.assertTrue(torch.equal(state_a[key], state_b[key]), key) + self.assertNotEqual(stats_a["heldout"]["mse"], stats_b["heldout"]["mse"]) + + def test_final_gptq_selected_only_by_validation_and_returned_scales_match(self): + weight, (train, validation, heldout) = self.fixture() + bias = torch.linspace(-0.25, 0.25, 128).bfloat16() + state, stats, reference = optimize_linear_weight( + weight, train, validation, heldout, bias, ranks=(16, 32), alphas=(0.5,), + iterations=2, weighting=("none",), baseline_rank=16, + output_correction=True, final_gptq=True, + ) + self.assertTrue(stats["search"]["final_gptq"]) + self.assertIsNotNone(stats["gptq_diagnostics"]) + gptq_rows = [row for row in stats["history"] if row.get("event") == "final_gptq_candidate"] + self.assertEqual(len(gptq_rows), 1) + self.assertTrue(gptq_rows[0]["validation"]["finite"]) + all_mses = [row["validation"]["mse"] for row in stats["history"]] + self.assertEqual(stats["validation"]["mse"], min(all_mses)) + prior_best = min(row["validation"]["mse"] for row in stats["history"] if row.get("event") != "final_gptq_candidate") + if stats["selected"].get("gptq", False): + self.assertLess(gptq_rows[0]["validation"]["mse"], prior_best) + else: + self.assertGreaterEqual(gptq_rows[0]["validation"]["mse"], prior_best) + self.assertTrue(torch.equal(_unpack_weight_scales(state["wscales"], 128), reference["weight_scales"])) + candidate = _candidate_from_reference(reference) + for label, values in (("validation", validation), ("heldout", heldout)): + measured = _reference_explicit_scales(candidate, values, reference.get("bias")) + teacher = torch.nn.functional.linear(values, weight, bias).float() + mse = (measured.double() - teacher.double()).square().mean().item() + self.assertTrue(torch.isfinite(measured).all()) + self.assertEqual(mse, stats[label]["mse"]) + json.dumps(stats, allow_nan=False) + + def test_explicit_scales_support_groups_without_seven_extremum(self): + # GPTQ can leave an entire group at |q|<=1 while its original stored + # scale remains 0.5. Inferring absmax(residual)/7 silently changes it. + residual = torch.zeros(128, 128, dtype=torch.float32) + residual[:, 0] = 0.5 + scales = torch.empty(128, 2, dtype=torch.bfloat16) + scales[:, 0] = 0.5 + scales[:, 1] = 2 + candidate = Candidate( + residual, scales, torch.zeros(16, 128, dtype=torch.bfloat16), + torch.zeros(128, 16, dtype=torch.bfloat16), + torch.ones(128, dtype=torch.bfloat16), {"gptq": True}, + ) + x = torch.zeros(3, 128, dtype=torch.bfloat16) + x[:, 0] = torch.tensor([7, 14, 21], dtype=torch.bfloat16) + expected = torch.tensor([3.5, 7, 10.5]).unsqueeze(1).expand(3, 128) + actual = _reference_explicit_scales(candidate, x, None) + self.assertTrue(torch.equal(actual, expected)) + packed = pack_candidate(candidate) + self.assertTrue(torch.equal(_unpack_weight_scales(packed["wscales"], 128), scales)) + integer = (residual.reshape(128, 2, 64) / scales.float().unsqueeze(-1)).round() + self.assertEqual(integer.abs().max().item(), 1) + # The old RTN-only helper correctly rejects this inferred-scale case; + # the explicit-scale path above must continue to handle it. + with self.assertRaises(ValueError): + reference_forward_groupwise(x, candidate.reference()) + + def test_baseline_only_preserves_exact_checkpoint_and_inputs(self): + weight, (train, validation, heldout) = self.fixture() + bias = torch.linspace(-0.25, 0.25, 128).bfloat16() + originals = [t.clone() for t in (weight, train, validation, heldout, bias)] + seed = 31 + rng = torch.random.get_rng_state().clone() + state, stats, _ = optimize_linear_weight( + weight, train, validation, heldout, bias, ranks=(), smoothing_families=(), + iterations=1, baseline_rank=16, seed=seed, + ) + expected, _ = convert_linear_weight(weight, bias, rank=16, input_absmax=train.float().abs().amax(0), seed=seed) + self.assertTrue(torch.equal(rng, torch.random.get_rng_state())) + for value, original in zip((weight, train, validation, heldout, bias), originals): + self.assertTrue(torch.equal(value, original)) + self.assertEqual(stats["selected"]["family"], "one_pass_baseline") + self.assertEqual(stats["candidate_count"], 1) + for key in expected: + self.assertTrue(torch.equal(state[key], expected[key]), key) + + def test_zero_weight_bias_and_zero_error_metrics(self): + weight = torch.zeros(128, 128, dtype=torch.bfloat16) + train = torch.zeros(8, 128, dtype=torch.bfloat16) + validation = torch.ones(9, 128, dtype=torch.bfloat16) + heldout = -torch.ones(10, 128, dtype=torch.bfloat16) + bias = torch.linspace(-1, 1, 128).bfloat16() + _, stats, ref = optimize_linear_weight( + weight, train, validation, heldout, bias, ranks=(16,), alphas=(), + smoothing_families=(), iterations=2, weighting=("none",), baseline_rank=16, + final_gptq=True, + ) + self.assertEqual(stats["validation"]["mse"], 0) + self.assertEqual(stats["heldout"]["mse"], 0) + self.assertFalse(stats["selected"].get("gptq", False)) + self.assertEqual(stats["history"][-1]["event"], "final_gptq_candidate") + self.assertEqual(stats["history"][-1]["validation"]["mse"], 0) + self.assertTrue(torch.equal(reference_forward_groupwise(heldout, ref), bias.float().expand(10, 128))) + json.dumps(stats, allow_nan=False) + + def test_fail_closed_invalid_data_and_configuration(self): + weight, (train, validation, _) = self.fixture() + cases = [ + (weight, train, validation, {"ranks": (17,)}), + (weight, train[:, :127], validation, {}), + (weight, train, validation[:1], {}), + (weight, train, validation, {"iterations": 0}), + (weight, train, validation, {"ridge": 0}), + (weight, train, validation, {"objective_backend": "nunchaku"}), + (weight, train, torch.full_like(validation, float("nan")), {}), + ] + for w, t, v, kwargs in cases: + with self.subTest(kwargs=kwargs, train_shape=t.shape, validation_shape=v.shape): + with self.assertRaises(ValueError): + optimize_linear_weight(w, t, v, **kwargs) + + +if __name__ == "__main__": + unittest.main(verbosity=2) diff --git a/reproduction/nunchaku_backend/validate_rank_upgrade_cpu.py b/reproduction/nunchaku_backend/validate_rank_upgrade_cpu.py new file mode 100644 index 0000000000000000000000000000000000000000..42c91f67304fd06e249e2419c42e706529fb5819 --- /dev/null +++ b/reproduction/nunchaku_backend/validate_rank_upgrade_cpu.py @@ -0,0 +1,137 @@ +"""Independent CPU integrity/load audit of a completed selective rank upgrade.""" +import argparse +from collections import Counter +import json +from pathlib import Path +import statistics +import time + +from .export_v3 import _hash_file +from .checkpoint_io import _source_index +from .runtime import _read_manifest + + +def distribution(values): + return {"minimum": min(values), "median": statistics.median(values), "maximum": max(values), + "mean": statistics.mean(values)} + + +def main(): + parser = argparse.ArgumentParser(description=__doc__) + parser.add_argument("checkpoint", type=Path) + parser.add_argument("--source", type=Path, required=True) + parser.add_argument("--out", type=Path, required=True) + args = parser.parse_args() + import torch + from .runtime import load_transformer, probe_backend + from .layout import BLOCK_LINEAR_PATTERN + torch.set_num_threads(8) + if torch.cuda.is_available(): + raise RuntimeError("Run with NVIDIA_VISIBLE_DEVICES=void and CUDA_VISIBLE_DEVICES empty") + started = time.perf_counter() + checkpoint, source = args.checkpoint.resolve(), args.source.resolve() + manifest = _read_manifest(checkpoint) + original_manifest = _read_manifest(source) + identity = manifest["rank_upgrade_identity"] + if identity["source_manifest_sha256"] != _hash_file(source / "manifest.json"): + raise ValueError("Original source manifest changed") + for filename, digest in {**identity["source_file_sha256"], **identity["source_metadata_sha256"]}.items(): + if _hash_file(source / filename) != digest: + raise ValueError(f"Original source changed: {filename}") + output_index = json.loads((checkpoint / "model.safetensors.index.json").read_text()) + source_index = json.loads((source / "model.safetensors.index.json").read_text()) + if output_index["weight_map"] != source_index["weight_map"]: + raise ValueError("Output tensor-to-shard mapping changed") + if _hash_file(checkpoint / "model.safetensors.index.json") != manifest["output_index_sha256"]: + raise ValueError("Output index hash mismatch") + if _hash_file(checkpoint / "config.json") != identity["source_metadata_sha256"]["config.json"]: + raise ValueError("Output architecture config changed") + if len(manifest["files"]) != 225 or len(set(output_index["weight_map"].values())) != 225: + raise ValueError("Expected225 output weight shards") + index = _source_index(checkpoint) + if set(index) != set(output_index["weight_map"]): + raise ValueError("Output shard tensors disagree with index") + for filename, entry in manifest["files"].items(): + if _hash_file(checkpoint / filename) != entry["sha256"]: + raise ValueError(f"Output shard hash mismatch: {filename}") + targets = {f"transformer_blocks.{block}.img_mlp.proj" for block in range(32)} + if set(manifest["rank_upgrade_reports"]) != targets: + raise ValueError("Missing or unexpected upgraded layers") + target_files = {filename for key, filename in source_index["weight_map"].items() if any(key.startswith(name + ".") for name in targets)} + untouched = set(manifest["files"]) - target_files + if len(untouched) != 193: + raise ValueError("Expected193 untouched shards") + for filename in untouched: + if _hash_file(checkpoint / filename) != identity["source_file_sha256"][filename]: + raise ValueError(f"Unrelated shard changed: {filename}") + for name, info in manifest["layers"].items(): + if name not in targets and info != original_manifest["layers"][name]: + raise ValueError(f"Unrelated layer metadata changed: {name}") + rows = [] + for name in sorted(targets, key=lambda n: int(n.split(".")[1])): + entry = manifest["rank_upgrade_reports"][name] + report_path = checkpoint / entry["file"] + if _hash_file(report_path) != entry["sha256"]: + raise ValueError(f"Changed layer report: {name}") + report = json.loads(report_path.read_text()) + if report["decision_uses_heldout"] is not False or report["selected_rank"] != manifest["layers"][name]["rank"]: + raise ValueError(f"Report selection/rank mismatch: {name}") + selected = report["selected_candidate_index"] + if selected is None: + validation, heldout = report["source_actual_validation"], report["source_actual_heldout"] + else: + validation, heldout = report["candidates"][selected]["validation"], report["candidates"][selected]["heldout"] + source_validation, source_heldout = report["source_actual_validation"], report["source_actual_heldout"] + if not all(metric["finite"] for metric in (validation, heldout, source_validation, source_heldout)): + raise ValueError(f"Nonfinite reported outputs: {name}") + if validation["mse"] > source_validation["mse"]: + raise ValueError(f"Actual-source validation guard violated: {name}") + rows.append({"layer": name, "selected_rank": report["selected_rank"], + "validation_mse_ratio": validation["mse"] / source_validation["mse"], + "heldout_mse_ratio": heldout["mse"] / source_heldout["mse"], + "source_heldout_relative_l2": source_heldout["relative_l2"], + "selected_heldout_relative_l2": heldout["relative_l2"], "seconds": report["seconds"]}) + load_started = time.perf_counter() + model = load_transformer(checkpoint, device="cpu") + load_seconds = time.perf_counter() - load_started + state = model.state_dict() + if any(t.device.type != "cpu" or t.is_meta for t in state.values()): + raise ValueError("Non-CPU or meta state after full load") + invalid = [name for name, tensor in state.items() if tensor.is_floating_point() and not torch.isfinite(tensor).all()] + if invalid: + raise ValueError(f"Nonfinite output state: {invalid}") + linears = {name: module for name, module in model.named_modules() if BLOCK_LINEAR_PATTERN.fullmatch(name)} + if len(linears) != 224 or any(module.rank != manifest["layers"][name]["rank"] for name, module in linears.items()): + raise ValueError("Loaded model quantized ranks/count disagree with manifest") + state_bytes = sum(t.numel() * t.element_size() for t in state.values()) + parameter_bytes = sum(t.numel() * t.element_size() for t in model.parameters()) + if state_bytes != sum(entry["bytes"] for entry in manifest["files"].values()) or state_bytes != output_index["metadata"]["total_size"]: + raise ValueError("Actual state bytes disagree with manifest/index metadata") + source_bytes = sum(entry["bytes"] for entry in original_manifest["files"].values()) + result = {"passed": True, "gpu_work_executed": False, "cuda_initialized": torch.cuda.is_initialized(), + "checkpoint": str(checkpoint), "source_checkpoint": str(source), + "manifest_sha256": _hash_file(checkpoint / "manifest.json"), + "source_manifest_sha256": identity["source_manifest_sha256"], + "source_shards_unchanged": len(identity["source_file_sha256"]), "output_shards_hash_verified": len(manifest["files"]), + "untouched_shards_byte_identical": len(untouched), "unmodified_quantized_layers": 192, + "quantized_linears": len(linears), "all_rank_counts": dict(Counter(module.rank for module in linears.values())), + "upgraded_rank_counts": dict(Counter(row["selected_rank"] for row in rows)), + "state_entries": len(state), "state_bytes": state_bytes, "parameter_bytes": parameter_bytes, + "source_state_bytes": source_bytes, "extra_state_bytes": state_bytes - source_bytes, + "state_gib": state_bytes / 2**30, "all_state_cpu_and_finite": True, + "validation_mse_ratio": distribution([row["validation_mse_ratio"] for row in rows]), + "heldout_mse_ratio": distribution([row["heldout_mse_ratio"] for row in rows]), + "heldout_relative_l2_before": distribution([row["source_heldout_relative_l2"] for row in rows]), + "heldout_relative_l2_after": distribution([row["selected_heldout_relative_l2"] for row in rows]), + "heldout_improved_layers": sum(row["heldout_mse_ratio"] < 1 for row in rows), + "layer_processing_seconds_sum": sum(row["seconds"] for row in rows), + "load_seconds": load_seconds, "total_validation_seconds": time.perf_counter() - started, + "runtime": probe_backend(), "layers": rows, + "limitations": "CPU integrity/load validation only; no GPU execution. Model-state bytes are not measured VRAM. Linear error improvements do not establish parent-MLP, denoiser or image equivalence."} + args.out.parent.mkdir(parents=True, exist_ok=True) + args.out.write_text(json.dumps(result, indent=2) + "\n") + print(json.dumps({key: value for key, value in result.items() if key != "layers"}), flush=True) + + +if __name__ == "__main__": + main() diff --git a/reproduction/nunchaku_backend/validate_v3_cpu.py b/reproduction/nunchaku_backend/validate_v3_cpu.py new file mode 100644 index 0000000000000000000000000000000000000000..0e65bb2f0a0fbe643760449a3e953cf8da99e3ef --- /dev/null +++ b/reproduction/nunchaku_backend/validate_v3_cpu.py @@ -0,0 +1,46 @@ +"""Full checkpoint load check; launch with NVIDIA_VISIBLE_DEVICES=void. + +This intentionally refuses a CUDA-visible environment. It cannot start a GPU +workload while the main agent runs image evaluation on the reserved device. +""" +import argparse +import json +from pathlib import Path +import time + + +def main(): + parser = argparse.ArgumentParser() + parser.add_argument("checkpoint", type=Path) + parser.add_argument("--out", type=Path, required=True) + args = parser.parse_args() + import torch + from .runtime import load_transformer, probe_backend + from .layout import BLOCK_LINEAR_PATTERN + torch.set_num_threads(8) + if torch.cuda.is_available(): + raise RuntimeError("Run this CPU check with NVIDIA_VISIBLE_DEVICES=void and CUDA_VISIBLE_DEVICES empty") + started = time.perf_counter() + model = load_transformer(args.checkpoint, device="cpu") + load_seconds = time.perf_counter() - started + parameters = dict(model.named_parameters()) + if any(p.device.type != "cpu" or p.is_meta for p in parameters.values()): + raise ValueError("Checkpoint has non-CPU or meta parameters") + invalid = [name for name, p in parameters.items() if p.is_floating_point() and not torch.isfinite(p).all()] + if invalid: + raise ValueError(f"Nonfinite checkpoint parameters: {invalid}") + linears = {name: module for name, module in model.named_modules() if BLOCK_LINEAR_PATTERN.fullmatch(name)} + if len(linears) != 224: + raise ValueError(f"Expected224 quantized linears, found{len(linears)}") + result = {"passed": True, "cuda_available": False, "checkpoint": str(args.checkpoint), + "load_seconds": load_seconds, "total_seconds": time.perf_counter() - started, + "quantized_linears": len(linears), "parameter_bytes": sum(p.numel() * p.element_size() for p in parameters.values()), + "state_entries": len(model.state_dict()), "runtime": probe_backend(), + "all_parameters_cpu": True, "all_floating_parameters_finite": True} + args.out.parent.mkdir(parents=True, exist_ok=True) + args.out.write_text(json.dumps(result, indent=2) + "\n") + print(json.dumps(result), flush=True) + + +if __name__ == "__main__": + main() diff --git a/reproduction/repeatability.md b/reproduction/repeatability.md new file mode 100644 index 0000000000000000000000000000000000000000..1993a3aecf222cec71fd09bbcb26890f54c78bdb --- /dev/null +++ b/reproduction/repeatability.md @@ -0,0 +1,62 @@ +# API cache-repeat diagnosis + +The original repeat failure is a real pixel difference, not PNG compression. The completed four-run diagnostic now rules out prompt caching as the sole cause: both no-cache runs had identical prompt tensors and first-transformer inputs but different first-transformer outputs. Cached CPU tensors remained unchanged. Divergence is observed inside the first transformer forward; the exact kernel or state mechanism has not been isolated. Strict bitwise repeatability has failed and is not being relabeled as a pass. + +The first engine PNG and saved first API response are byte-for-byte identical. The second engine PNG differs visibly in some program-column positions, footer spacing and bridge strokes while retaining the same broad poster content. + +## Preserved initial evidence + +Both calls used seed 21044, the same bilingual-festival prompt, 25 steps, CFG 1.0, 1024×1024, no references, rank-128 Nunchaku transformer, NF4 encoder, eager execution and KV caching. Parsed measurement dictionaries match exactly after excluding the output label; engine settings also match. The first call took 18.125 seconds, the second 14.339 seconds. The second logged `prompt_cache_hits=1`; reference-VAE caching did not participate. + +| Comparison | Result | +|---|---:| +| First engine vs first saved API response | Bytes and decoded RGBA pixels identical | +| First vs second engine RGB mean absolute error | 6.5767 / 255 | +| RGB RMSE | 30.3013 / 255 | +| RGB PSNR | 18.5016 dB | +| Pixels with at least one differing RGB channel | 905,010 / 1,048,576 = 86.308% | +| Absolute RGB-channel error, median / 99th percentile | 1 / 191 | +| Differing alpha pixels | 12,387 | + +The large tail reflects displaced high-contrast lettering, not merely one-level background rounding. The initial smoke harness asserts equal encoded hashes before saving the second response or its metrics, so its JSON contains only the first request. The second engine artifact and complete engine measurement were recovered separately; the original failed second HTTP body was not persisted by the harness. + +Stable snapshots: + +| Artifact | SHA256 | +|---|---| +| [First engine PNG](../samples/fidelity-v3/api-cache-diagnosis-initial/api-ddf5ba02cd2b454492864b8f7b748493.png) | `fba4800fce6254a8e3a5c7bfe550a4b10b94332c442837d9263722eae35f7cb0` | +| [Second engine PNG](../samples/fidelity-v3/api-cache-diagnosis-initial/api-6a4fd749d36e4989af73d9fd5948c60f.png) | `ee64855f2fff0d02f5693683070128723b96063c19feec75112d5a52429ebae4` | +| [First saved API PNG](../samples/fidelity-v3/api-cache-diagnosis-initial/api-v3-expanded-bilingual-festival-25-0.png) | `fba4800fce6254a8e3a5c7bfe550a4b10b94332c442837d9263722eae35f7cb0` | + +[Pixel metrics](../results/api-cache-v3-pixel-diff.json) include decoded-pixel hashes. [Both engine measurements](../results/api-cache-v3-measurements.jsonl) preserve the exact prompt, seed, settings and timings. + +## Source findings + +`server.execute` forwards the explicit integer seed to `Engine.generate`; the engine constructs a fresh CUDA `torch.Generator` and calls `manual_seed` for each request. The pipeline uses that generator for initial latent noise, creates new prefix-KV objects each call, and resets the scheduler begin index. Request UUID/output labels do not affect sampling. + +The prompt cache stores detached CPU copies of the CUDA prompt tensors and returns device copies on hits. In the pinned pipeline, prompt embeddings are repeated into new storage before denoising, and image masks are extended using concatenation. The transformer constructs a new joint tensor before assigning image-token data. No specific in-place write into cached CPU prompt tensors was found along this text-only path. This source inspection does not replace tensor fingerprints at runtime. + +There is a concrete source-supported nondeterminism candidate: Nunchaku v1.2.1's [low-rank reduction](https://github.com/nunchaku-tech/nunchaku/blob/v1.2.1/src/kernels/zgemm/lora.cuh#L73) combines partial values through [FP32 global reduction instructions](https://github.com/nunchaku-tech/nunchaku/blob/v1.2.1/src/kernels/zgemm/gemm_utils.cuh#L321). Different floating-point accumulation orders can change rounding, so a fixed random seed alone does not prove bitwise repeatability of that kernel. This is a plausible mechanism, not a demonstrated attribution of the observed image differences. + +## Controlled diagnostic completed + +[`cache_repeat_diagnostic.py`](../nunchaku_backend/cache_repeat_diagnostic.py) loads the same rank-128 engine via `_parse_cli_args` and runs the same festival job four times: no-cache 1, no-cache 2, cache miss, cache hit. It records all three prompt output tensors' hashes/shapes/strides/dtypes; cached CPU contents before/after requests; first transformer latent, prompt, mask and timestep fingerprints; the first transformer output; and final PNG/pixel differences. Every run is saved before comparisons. Images go under `results/repeatability-diagnostic/samples`, outside quality-gallery discovery. Existing output directories are rejected. + +The main agent executed all four runs, and the complete [diagnostic report](../results/repeatability-diagnostic/diagnostic.json) and four PNGs have been synced locally. The report SHA256 is `384efadf773a862b08fa6986a2e37a506133caef22136fb52162d871f6603fa9`; all four local PNG hashes were verified against it. + +All four runs have exactly equal recorded prompt-tensor fingerprints, including shape, stride, dtype, device and content hash. Their first-transformer input fingerprints are also exactly equal for `hidden_states` (initial latents), `encoder_hidden_states`, `encoder_hidden_states_mask`, `img_mask` and `timestep`; non-tensor `img_shapes` and `kv_cache_mode` match too. **All four first-transformer output hashes differ.** The no-cache pair therefore diverges without prompt-cache reuse. The cache-miss stored tensors exactly match the cache-hit starting tensors, and the hit leaves every stored CPU tensor unchanged. + +| Pair | Prompt / first-transformer inputs | First-transformer output | Final RGB MAE / 255 | Final RGB RMSE / 255 | +|---|---|---|---:|---:| +| no-cache 1 → no-cache 2 | Identical | Different | 5.0577 | 25.5306 | +| no-cache 2 → cache miss | Identical | Different | 5.4012 | 26.3205 | +| cache miss → cache hit | Identical | Different | 6.4759 | 30.5930 | +| no-cache 1 → cache hit | Identical | Different | 5.2533 | 27.1202 | + +Every pair differs in encoded bytes and decoded pixels. These measurements support a numerical repeatability problem inside the transformer execution path, rather than a demonstrated prompt-cache mutation. They do not identify which operation causes it, establish a tolerance, or prove reference-VAE cache correctness; the diagnostic uses no references. The source-supported global-reduction mechanism above remains a candidate, not an isolated cause. Instrumentation adds transfers/synchronization, so these runs' elapsed times are not inference benchmarks. No GPU diagnostic was launched by this reviewing agent. + +## Minimal next changes + +The confirmed harness defect is loss of failure evidence: save each decoded image and response metadata before checking repeat equality, record encoded equality separately from pixel equality, and preserve the failed strict-repeat result. That can be fixed independently of the numerical diagnosis. The main agent plans the final four API requests in explicit `--record-repeat-differences` mode so serving, generation and reference-cache checks can complete while the strict bitwise pass field remains false. This is diagnostic recording, not an invented numerical tolerance or a repeatability pass. + +No runtime prompt-cache change is justified by the measured tensor evidence. If exact repeatability becomes a requirement, the next investigation must isolate the first divergent transformer operation and its reduction/state behavior; a fixed seed or unchanged embeddings alone cannot establish that guarantee. No runtime cache change or API tolerance relaxation has been applied by this reviewing agent. diff --git a/reproduction/runner.py b/reproduction/runner.py new file mode 100644 index 0000000000000000000000000000000000000000..728e85a790e0f108571fa4371db3ce6ee51f674a --- /dev/null +++ b/reproduction/runner.py @@ -0,0 +1,372 @@ +"""Pinned Qwen Image 2.1 experiments. One process, one GPU, serialized jobs.""" +import argparse +from collections import OrderedDict +import functools +import hashlib +import json +import os +from pathlib import Path +import time +import traceback + +MODEL = 'Qwen/Qwen-Image-2.1' +REVISION = 'b3179ad355be050328e483a9dfdd9e60cd62adfa' +ROOT = Path(__file__).resolve().parent + + +def _env_bf16_roles(): + """Comma-separated projection roles; empty environment means no override.""" + return [role.strip() for role in os.getenv('QWEN_BF16_ROLES', '').split(',') if role.strip()] + + +def default_args(): + return argparse.Namespace(quant='nf4', sample_dir=os.getenv('QWEN_SAMPLE_DIR'), backend=os.getenv('QWEN_BACKEND','nf4'), nunchaku_checkpoint=os.getenv('QWEN_NUNCHAKU_CHECKPOINT'), offload=True, compile=os.getenv('QWEN_COMPILE','0')=='1', flex=False, + cache=True, prequant=os.getenv('QWEN_PREQUANT'), save_prequant=None, tiling=False, lean_encoder=True, verify_lean=False, bf16_transformer=False, selective_nf4=False, bf16_vision=False, stage_offload=True, release_kv=True, + bf16_source=os.getenv('QWEN_BF16_SOURCE') or None, restore_roles=_env_bf16_roles()) + + +def _validate_hybrid_args(args): + """Validate opt-in configuration before importing Torch or loading weights.""" + roles = getattr(args, 'restore_roles', None) or [] + source = getattr(args, 'bf16_source', None) + if not roles and not source: + return False + if not isinstance(roles, (list, tuple)) or any(not isinstance(role, str) for role in roles): + raise ValueError('restore_roles must be a list of exact projection-role names') + if bool(roles) != bool(source): + raise ValueError('BF16 restoration requires both --bf16-source/QWEN_BF16_SOURCE and --restore-role/QWEN_BF16_ROLES') + if getattr(args, 'backend', 'nf4') != 'nunchaku': + raise ValueError('BF16 projection restoration requires the Nunchaku backend') + from nunchaku_backend.hybrid_v3 import _selected_names + _selected_names(roles, []) + if getattr(args, 'save_prequant', None): + raise ValueError('Hybrid restoration is a runtime override; do not combine it with --save-prequant') + args.restore_roles = sorted(set(roles)) + args.bf16_source = str(source) + return True + + +def _apply_hybrid_override(transformer, args): + """Restore CPU modules before the pipeline installs device/offload hooks.""" + if not getattr(args, 'restore_roles', None): + return None + from nunchaku_backend.hybrid_v3 import restore_bf16_projections + report = restore_bf16_projections(transformer, args.bf16_source, roles=args.restore_roles) + # Measurements serialize args, so record what was actually restored rather + # than only the requested selectors or the original checkpoint precision. + args.hybrid_roles = report['roles'] + args.hybrid_names = report['restored_names'] + args.hybrid_restored_count = report['restored_count'] + args.hybrid_source = report['source_transformer'] + args.hybrid_source_config_sha256 = report['source_config_sha256'] + args.hybrid_bf16_state_bytes = report['bf16_state_bytes'] + args.hybrid_replaced_packed_state_bytes = report['replaced_packed_state_bytes'] + args.hybrid_state_bytes_delta = report['net_state_bytes_delta'] + args.hybrid_precision = 'Original BF16 restored projections; Nunchaku W4A4 remaining projections' + return report + + +class Engine: + def __init__(self, args=None): + engine_init_started=time.perf_counter() + self.args = args or default_args() + args = self.args + _validate_hybrid_args(args) + import torch + from diffusers import QwenImage21Pipeline, QwenImage21Transformer2DModel, BitsAndBytesConfig + from transformers import Qwen3VLForConditionalGeneration, BitsAndBytesConfig as TBits + self.torch = torch + torch.set_num_threads(8) + self.metrics = {} + self.prompt_cache, self.vae_cache = OrderedDict(), OrderedDict() + started = time.perf_counter() + source = args.prequant or MODEL + kwargs = {} if args.prequant else {'revision': REVISION} + q = dict(load_in_4bit=True, bnb_4bit_quant_type='nf4', bnb_4bit_use_double_quant=True, + bnb_4bit_compute_dtype=torch.bfloat16) + te = Qwen3VLForConditionalGeneration.from_pretrained( + MODEL if args.bf16_vision else source, subfolder='text_encoder', dtype=torch.bfloat16, + quantization_config=TBits(**q, **({'llm_int8_skip_modules':['lm_head','model.visual']} if args.bf16_vision else {})) if args.quant == 'nf4' and (not args.prequant or args.bf16_vision) else None, + device_map={'': 0} if args.quant == 'nf4' else {'': 'cpu'}, + **({'revision':REVISION} if args.bf16_vision else kwargs)) + te.config.use_cache = False + te.config.text_config.use_cache = False + print(json.dumps({'event': 'text_encoder_loaded', 'seconds': time.perf_counter()-started, + 'footprint_bytes': te.get_memory_footprint()}), flush=True) + if args.offload: + te.to('cpu') + torch.cuda.empty_cache() + if getattr(args,'backend','nf4') == 'nunchaku': + if not getattr(args,'nunchaku_checkpoint',None): + raise ValueError('--nunchaku-checkpoint is required for the Nunchaku backend') + from nunchaku_backend.runtime import load_transformer + transformer = load_transformer(args.nunchaku_checkpoint,device='cpu',torch_dtype=torch.bfloat16) + else: + transformer = QwenImage21Transformer2DModel.from_pretrained( + MODEL if args.bf16_transformer or args.selective_nf4 else source, subfolder='transformer', torch_dtype=torch.bfloat16, + quantization_config=BitsAndBytesConfig(**q, **({'llm_int8_skip_modules':['time_text_embed','txt_in','img_in','modulation','norm_out','proj_out']} if args.selective_nf4 else {})) if args.quant == 'nf4' and (not args.prequant or args.selective_nf4) and not args.bf16_transformer else None, + **({'revision':REVISION} if args.bf16_transformer or args.selective_nf4 else kwargs)) + self.hybrid_metadata = _apply_hybrid_override(transformer, args) + print(json.dumps({'event':'quantization_scope', + 'encoder_linear4bit':sum(type(m).__name__=='Linear4bit' for m in te.modules()), + 'vision_linear4bit':sum(type(m).__name__=='Linear4bit' for m in te.model.visual.modules()), + 'transformer_linear4bit':sum(type(m).__name__=='Linear4bit' for m in transformer.modules()), + 'transformer_nunchaku_w4a4':sum(type(m).__name__=='SVDQW4A4Linear' for m in transformer.modules()), + **({'hybrid_roles':args.hybrid_roles, 'hybrid_restored_count':args.hybrid_restored_count, + 'hybrid_bf16_state_bytes':args.hybrid_bf16_state_bytes, + 'hybrid_state_bytes_delta':args.hybrid_state_bytes_delta} if self.hybrid_metadata else {})}),flush=True) + self.pipe = QwenImage21Pipeline.from_pretrained( + source, text_encoder=te, transformer=transformer, torch_dtype=torch.bfloat16, **kwargs) + if args.save_prequant: + if args.prequant: + # Transformers 5.17 cannot reserialize its already-loaded bnb + # conversion graph. Copy unchanged components; save only fresh ones. + import shutil + src,dst=Path(args.prequant).resolve(),Path(args.save_prequant).resolve() + if dst==src or dst.is_relative_to(src): + raise ValueError('Save checkpoint to a separate sibling directory') + shutil.copytree(src,dst,dirs_exist_ok=True) + if args.selective_nf4 or args.bf16_transformer: + shutil.rmtree(dst/'transformer') + transformer.save_pretrained(dst/'transformer',safe_serialization=True) + if args.bf16_vision: + shutil.rmtree(dst/'text_encoder') + te.save_pretrained(dst/'text_encoder',safe_serialization=True) + else: + self.pipe.save_pretrained(args.save_prequant, safe_serialization=True) + if args.tiling: + self.pipe.vae.enable_tiling() + if args.offload: + self.pipe.enable_model_cpu_offload(gpu_id=0) + else: + self.pipe.to('cuda') + if args.verify_lean: + from lean_encoder import verify_encoder_parity + from PIL import Image + reports=[verify_encoder_parity(te,self.pipe.processor,prompt='A red ceramic coffee mug'), + verify_encoder_parity(te,self.pipe.processor,prompt='Change the mug to blue', + images=[Image.new('RGB',(512,512),(150,50,40))])] + print(json.dumps({'event':'encoder_parity','reports':reports}),flush=True) + if not all(r['passed'] for r in reports): + raise RuntimeError('Lean encoder parity failed') + self.pipe.maybe_free_model_hooks() + if args.lean_encoder: + from lean_encoder import enable_lean_encoder + print(json.dumps({'event':'lean_encoder',**enable_lean_encoder(te)}),flush=True) + if args.flex: + if not args.compile: + raise ValueError('Flex attention requires compilation') + from diffusers.models.transformers.transformer_qwenimage21 import QwenImage21FlexAttnProcessor + self.pipe.transformer.set_attn_processor(QwenImage21FlexAttnProcessor()) + if args.compile: + self.pipe.transformer.compile(fullgraph=False, mode='default') + self._install_instrumentation() + torch.cuda.synchronize() + self.load_seconds = time.perf_counter() - started + self.startup_seconds=time.perf_counter()-engine_init_started + (ROOT/'results').mkdir(exist_ok=True) + (ROOT/'samples').mkdir(exist_ok=True) + print(json.dumps({'event': 'ready', 'load_seconds': self.load_seconds, 'startup_seconds':self.startup_seconds, 'settings':vars(args)}),flush=True) + + def sync(self): + self.torch.cuda.synchronize() + + def timed(self, name, function): + @functools.wraps(function) + def wrapper(*args, **kwargs): + self.sync() + started = time.perf_counter() + try: + return function(*args, **kwargs) + finally: + self.sync() + self.metrics[name] = self.metrics.get(name, 0) + time.perf_counter()-started + return wrapper + + @staticmethod + def _pil_key(image): + return (image.mode, image.size, hashlib.sha256(image.tobytes()).hexdigest()) + + def _install_instrumentation(self): + pipe = self.pipe + original = pipe._get_qwen_prompt_embeds + def compute_prompt(prompt=None,image=None,device=None): + result=original(prompt=prompt,image=image,device=device) + if self.args.offload and self.args.stage_offload: + pipe.text_encoder.to('cpu') + self.torch.cuda.empty_cache() + return result + def cached_prompt(prompt=None, image=None, device=None): + if not self.args.cache: + return compute_prompt(prompt=prompt, image=image, device=device) + key = (json.dumps(prompt), tuple(self._pil_key(i) for i in (image or []))) + target = device or pipe._execution_device + if self.args.cache and key in self.prompt_cache: + self.metrics['prompt_cache_hits'] = self.metrics.get('prompt_cache_hits', 0)+1 + self.prompt_cache.move_to_end(key) + return tuple(t.to(target) for t in self.prompt_cache[key]) + result = compute_prompt(prompt=prompt, image=image, device=device) + if self.args.cache: + self.prompt_cache[key] = tuple(t.detach().cpu() for t in result) + while len(self.prompt_cache)>8: + self.prompt_cache.popitem(last=False) + return result + pipe._get_qwen_prompt_embeds = self.timed('encode_prompt_seconds', cached_prompt) + original_vae = pipe._encode_vae_image + def compute_vae(image,generator): + result=original_vae(image,generator) + if self.args.offload and self.args.stage_offload: + pipe.vae.to('cpu') + self.torch.cuda.empty_cache() + return result + def cached_vae(image, generator): + if not self.args.cache: + return compute_vae(image,generator) + # Upstream uses argmax, so cached reference latents do not depend on RNG seed. + raw=image.detach().cpu().contiguous() + key=(tuple(raw.shape),str(raw.dtype),hashlib.sha256(raw.view(self.torch.uint8).numpy().tobytes()).hexdigest()) + if self.args.cache and key in self.vae_cache: + self.metrics['vae_cache_hits'] = self.metrics.get('vae_cache_hits', 0)+1 + self.vae_cache.move_to_end(key) + return self.vae_cache[key].to(image.device) + result=compute_vae(image,generator) + if self.args.cache: + self.vae_cache[key]=result.detach().cpu() + while len(self.vae_cache)>8: + self.vae_cache.popitem(last=False) + return result + pipe._encode_vae_image=self.timed('encode_reference_seconds',cached_vae) + pipe.vae.decode=self.timed('decode_seconds',pipe.vae.decode) + # Module hooks survive Accelerate rebuilding forward/offload hooks after each job. + # Keep timing calls outside Dynamo graphs. + @self.torch.compiler.disable + def before_transformer(module, inputs, kwargs): + cache=kwargs.get('kv_cache') + if cache is not None: + self._kv_caches[id(cache)]=cache + self.sync() + self._transformer_start=time.perf_counter() + @self.torch.compiler.disable + def after_transformer(module, inputs, output): + self.sync() + self.metrics['transformer_seconds']=self.metrics.get('transformer_seconds',0)+time.perf_counter()-self._transformer_start + pipe.transformer.register_forward_pre_hook(before_transformer,with_kwargs=True) + pipe.transformer.register_forward_hook(after_transformer) + + def generate(self, job): + try: + return self._generate(job) + except Exception: + # Restore stage hooks after an interrupted/OOM request before a later retry. + self.pipe.maybe_free_model_hooks() + self.torch.cuda.empty_cache() + raise + + def _generate(self, job): + job={**job,'kv_cache':job.get('kv_cache',True)} + from PIL import Image + torch=self.torch + refs=[Image.open(p).copy() for p in job.get('images', [])] + self.metrics={} + self._kv_caches={} + torch.cuda.reset_peak_memory_stats() + self.sync() + started=time.perf_counter() + started_unix=time.time() + last=started + steps=[] + def callback(pipe,index,timestep,kwargs): + nonlocal last + self.sync() + now=time.perf_counter() + steps.append(now-last) + last=now + # Prefix KV is dead after the final sampler step. Upstream otherwise + # retains it while the large VAE decode runs (about 2 GiB for one 1K ref). + if self.args.release_kv and index+1==pipe.num_timesteps: + released=0 + for cache in self._kv_caches.values(): + for layer in cache.layer_caches: + for name in ('k','v'): + value=getattr(layer,name) + if value is not None: + released+=value.numel()*value.element_size() + setattr(layer,name,None) + self._kv_caches.clear() + self.metrics['released_kv_mib']=released/2**20 + torch.cuda.empty_cache() + if index == 0 or (index+1)%10 == 0: + print(json.dumps({'event':'step','label':job.get('label'),'step':index+1,'elapsed':now-started}),flush=True) + return kwargs + result=self.pipe( + prompt=job['prompt'], image=refs or None, + width=job.get('width',1024), height=job.get('height',1024), + output_resolution=max(job.get('width',1024),job.get('height',1024)), + num_inference_steps=job.get('steps',40), true_cfg_scale=job.get('cfg',1.0), + generator=torch.Generator(device='cuda').manual_seed(job.get('seed',42)), + use_kv_cache=job.get('kv_cache',True), callback_on_step_end=callback).images[0] + self.sync() + elapsed=time.perf_counter()-started + label=job.get('label',str(time.time_ns())) + if not label or any(c not in 'abcdefghijklmnopqrstuvwxyzABCDEFGHIJKLMNOPQRSTUVWXYZ0123456789_-' for c in label): + raise ValueError('label must contain only letters, numbers, underscores or hyphens') + sample_dir=Path(getattr(self.args,'sample_dir',None) or ROOT/'samples') + sample_dir.mkdir(parents=True,exist_ok=True) + path=sample_dir/f'{label}.png' + result.save(path) + metrics={**self.metrics, 'path':str(path),'seconds':elapsed,'started_unix':started_unix,'finished_unix':time.time(), + 'peak_allocated_mib':torch.cuda.max_memory_allocated()/2**20, + 'peak_reserved_mib':torch.cuda.max_memory_reserved()/2**20, + 'step_seconds':steps,'job':job,'settings':vars(self.args),'load_seconds':self.load_seconds,'startup_seconds':self.startup_seconds, + 'output_size':result.size,'output_mode':result.mode} + with (ROOT/'results'/'measurements.jsonl').open('a') as f: + f.write(json.dumps(metrics)+'\n') + print(json.dumps({'event':'result',**metrics}),flush=True) + return metrics + + +def _parse_cli_args(argv=None): + parser=argparse.ArgumentParser() + parser.add_argument('--jobs') + parser.add_argument('--sample-dir',help='Separate output directory for a quality experiment') + parser.add_argument('--backend',choices=['nf4','nunchaku'],default='nf4') + parser.add_argument('--nunchaku-checkpoint',help='Custom packed Qwen 2.1 transformer checkpoint') + from nunchaku_backend.hybrid_v3 import ROLES + parser.add_argument('--bf16-source','--restore-source',dest='bf16_source',default=os.getenv('QWEN_BF16_SOURCE') or None, + help='Original BF16 snapshot for optional Nunchaku projection restoration') + parser.add_argument('--restore-role',dest='restore_roles',choices=ROLES,action='append',default=None, + help='Restore this projection role to BF16 in all 32 blocks; repeat for multiple roles') + parser.add_argument('--quant',choices=['nf4'],default='nf4') + parser.add_argument('--offload',action=argparse.BooleanOptionalAction,default=True) + parser.add_argument('--compile',action='store_true') + parser.add_argument('--flex',action='store_true') + parser.add_argument('--cache',action=argparse.BooleanOptionalAction,default=True) + parser.add_argument('--tiling',action=argparse.BooleanOptionalAction,default=False) + parser.add_argument('--release-kv',action=argparse.BooleanOptionalAction,default=True) + parser.add_argument('--stage-offload',action=argparse.BooleanOptionalAction,default=True) + parser.add_argument('--selective-nf4',action='store_true') + parser.add_argument('--bf16-vision',action='store_true') + parser.add_argument('--bf16-transformer',action='store_true') + parser.add_argument('--lean-encoder',action=argparse.BooleanOptionalAction,default=True) + parser.add_argument('--verify-lean',action='store_true') + parser.add_argument('--prequant') + parser.add_argument('--save-prequant') + args=parser.parse_args(argv) + if args.restore_roles is None: + args.restore_roles = _env_bf16_roles() + return args + + +def main(): + args=_parse_cli_args() + engine=Engine(args) + if args.jobs: + for job in json.loads(Path(args.jobs).read_text()): + try: + engine.generate(job) + except Exception: + traceback.print_exc() + raise + +if __name__=='__main__': + main() diff --git a/reproduction/samples/archive/2026-09-20-initial-poc/reference-clean-rgb.png b/reproduction/samples/archive/2026-09-20-initial-poc/reference-clean-rgb.png new file mode 100644 index 0000000000000000000000000000000000000000..6b9d24db631cc6e8c342fab61ee5e8cff10f2807 --- /dev/null +++ b/reproduction/samples/archive/2026-09-20-initial-poc/reference-clean-rgb.png @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:a191a4896478811560e3cf679c23a33b9530252c6d6a72e6136e10418d51c803 +size 1331584 diff --git a/reproduction/server.py b/reproduction/server.py new file mode 100644 index 0000000000000000000000000000000000000000..d60d07e25f38e1fb6f27d728f24680df451de986 --- /dev/null +++ b/reproduction/server.py @@ -0,0 +1,290 @@ +"""Local single-GPU RunPod-shaped POC API. Run: uvicorn server:app --host 0.0.0.0 --port 8091.""" +from __future__ import annotations + +import base64 +import binascii +import io +import ipaddress +import logging +import math +import socket +import tempfile +import threading +import time +import uuid +from contextlib import asynccontextmanager +from pathlib import Path +from typing import Literal +from urllib.parse import urlsplit, urlunsplit + +import httpx +from fastapi import FastAPI, Request +from fastapi.responses import JSONResponse +from PIL import Image, ImageOps +from pydantic import BaseModel, ConfigDict, Field, ValidationError, field_validator +from starlette.concurrency import run_in_threadpool + +MAX_REF_BYTES = 20 * 1024 * 1024 +MAX_BODY_BYTES = 60 * 1024 * 1024 +MAX_REF_PIXELS = 16_777_216 +LOGGER = logging.getLogger("qwen_image21_poc.server") +PUBLIC_NUMERIC_METRICS = frozenset({ + "seconds", "encode_prompt_seconds", "encode_reference_seconds", "decode_seconds", + "transformer_seconds", "load_seconds", "startup_seconds", "started_unix", "finished_unix", + "peak_allocated_mib", "peak_reserved_mib", "prompt_cache_hits", "vae_cache_hits", +}) + + +def public_metrics(metrics): + """Expose measured scalar values, never raw jobs, configuration or paths.""" + result = { + key: value for key, value in metrics.items() + if key in PUBLIC_NUMERIC_METRICS + and isinstance(value, (int, float)) and not isinstance(value, bool) + and math.isfinite(value) + } + size = metrics.get("output_size") + if isinstance(size, (list, tuple)) and len(size) == 2: + if all(isinstance(value, int) and not isinstance(value, bool) and value > 0 for value in size): + result.update(output_width=size[0], output_height=size[1]) + if metrics.get("output_mode") in ("RGB", "RGBA", "L", "LA"): + result["output_mode"] = metrics["output_mode"] + return result + + +class Inputs(BaseModel): + model_config = ConfigDict(extra="forbid", allow_inf_nan=False) + prompt: str = Field(min_length=1, max_length=16000) + width: int = Field(default=1024, ge=256, le=1024, strict=True) + height: int = Field(default=1024, ge=256, le=1024, strict=True) + steps: int = Field(default=40, ge=1, le=60, strict=True) + CFGScale: Literal[1.0] = 1.0 + referenceImages: list[str] = Field(default_factory=list, max_length=2) + outputFormat: Literal["PNG", "JPEG", "WEBP"] = "PNG" + outputQuality: int = Field(default=90, ge=1, le=100, strict=True) + uploadUrl: str | None = Field(default=None, max_length=16000) + includeCost: bool = False + taskUUID: str | None = Field(default=None, max_length=128) + seed: int = Field(default=42, ge=0, le=2**63 - 1, strict=True) + useKVCache: bool = True + + @field_validator("width", "height") + @classmethod + def dimensions(cls, value): + if value % 32: + raise ValueError("dimensions must be multiples of 32") + return value + + @field_validator("prompt") + @classmethod + def meaningful_prompt(cls, value): + if not value.strip(): + raise ValueError("prompt must contain text") + return value + + @field_validator("outputFormat", mode="before") + @classmethod + def normalize_format(cls, value): + if isinstance(value, str): + return {"JPG": "JPEG"}.get(value.upper(), value.upper()) + return value + + +class Envelope(BaseModel): + model_config = ConfigDict(extra="forbid") + input: Inputs + + +class InputError(Exception): + pass + + +def failure(message, status=400, task_id=None): + body = {"status": "FAILED", "error": message} + if task_id: + body["id"] = task_id + return JSONResponse(body, status_code=status) + + +def external_url(url): + """Public HTTP(S) only; redirects are deliberately not followed.""" + try: + parsed = urlsplit(url) + if parsed.scheme not in ("http", "https") or not parsed.hostname: + raise ValueError() + if parsed.username or parsed.password or parsed.fragment: + raise ValueError() + port = parsed.port or (443 if parsed.scheme == "https" else 80) + addresses = socket.getaddrinfo(parsed.hostname, port, type=socket.SOCK_STREAM) + if not addresses or any(not ipaddress.ip_address(a[4][0]).is_global for a in addresses): + raise ValueError() + except (ValueError, OSError): + raise InputError("URL must identify a public HTTP(S) host") from None + return url + + +def reference_bytes(source): + if source.startswith(("https://", "http://")): + external_url(source) + started = time.monotonic() + chunks, size = [], 0 + try: + with httpx.Client(timeout=httpx.Timeout(30, connect=5), follow_redirects=False, trust_env=False) as client: + with client.stream("GET", source) as response: + response.raise_for_status() + length = response.headers.get("content-length") + if length and int(length) > MAX_REF_BYTES: + raise InputError("Reference exceeds 20 MiB") + for chunk in response.iter_bytes(): + size += len(chunk) + if size > MAX_REF_BYTES: + raise InputError("Reference exceeds 20 MiB") + if time.monotonic() - started > 60: + raise InputError("Reference download exceeded time limit") + chunks.append(chunk) + return b"".join(chunks) + except (httpx.HTTPError, ValueError): + raise InputError("Reference download failed") from None + if source.startswith("data:"): + header, separator, source = source.partition(",") + if not separator or not header.startswith("data:image/") or not header.endswith(";base64"): + raise InputError("Reference data URI must contain a base64 image") + if len(source) > 4 * math.ceil(MAX_REF_BYTES / 3): + raise InputError("Reference exceeds 20 MiB") + try: + data = base64.b64decode(source, validate=True) + except (ValueError, binascii.Error): + raise InputError("Reference must be a public URL or base64 image") from None + if len(data) > MAX_REF_BYTES: + raise InputError("Reference exceeds 20 MiB") + return data + + +def save_reference(source, path): + data = reference_bytes(source) + try: + with Image.open(io.BytesIO(data)) as image: + if image.width * image.height > MAX_REF_PIXELS or image.width < 1 or image.height < 1: + raise InputError("Reference exceeds 16 megapixels") + image.verify() + with Image.open(io.BytesIO(data)) as image: + oriented = ImageOps.exif_transpose(image) + oriented.convert("RGBA" if "A" in oriented.getbands() else "RGB").save(path, format="PNG") + except InputError: + raise + except Exception: + raise InputError("Reference is not a valid supported raster image") from None + + +def encoded_result(path, fmt, quality): + with Image.open(path) as image: + if fmt == "JPEG": + rgba = image.convert("RGBA") + image = Image.new("RGB", rgba.size, "white") + image.paste(rgba, mask=rgba.getchannel("A")) + stream = io.BytesIO() + image.save(stream, format=fmt, **({"quality": quality} if fmt != "PNG" else {})) + return stream.getvalue(), image.width, image.height + + +def create_app(engine_factory=None): + @asynccontextmanager + async def lifespan(application): + application.state.ready = False + if engine_factory is None: + from runner import Engine, default_args + factory = lambda: Engine(default_args()) + else: + factory = engine_factory + application.state.engine = await run_in_threadpool(factory) + application.state.ready = True + try: + yield + finally: + application.state.ready = False + + application = FastAPI(title="Qwen Image 2.1 POC", lifespan=lifespan) + application.state.ready = False + application.state.lock = threading.Lock() + + @application.get("/healthz") + def health(): + return {"status": "ok"} + + @application.get("/readyz") + def ready(): + if not application.state.ready: + return JSONResponse({"ready": False}, status_code=503) + return {"ready": True, "busy": application.state.lock.locked()} + + def execute(inp): + task_id = inp.taskUUID or str(uuid.uuid4()) + if not application.state.ready: + return failure("Model is not ready", 503, task_id) + # The lock lives entirely in this synchronous function: client disconnects + # cannot release it while a CUDA operation is still executing. + if not application.state.lock.acquire(blocking=False): + response = failure("GPU is busy; retry later", 503, task_id) + response.headers["Retry-After"] = "5" + return response + started = time.monotonic() + try: + if inp.uploadUrl: + external_url(inp.uploadUrl) + with tempfile.TemporaryDirectory(prefix="qwen-request-") as directory: + paths = [] + for i, source in enumerate(inp.referenceImages): + path = str(Path(directory) / f"reference-{i}.png") + save_reference(source, path) + paths.append(path) + job = dict(prompt=inp.prompt, width=inp.width, height=inp.height, + steps=inp.steps, cfg=inp.CFGScale, images=paths, seed=inp.seed, + kv_cache=inp.useKVCache, label=f"api-{uuid.uuid4().hex}") + metrics = application.state.engine.generate(job) + data, width, height = encoded_result(metrics["path"], inp.outputFormat, inp.outputQuality) + row = {"taskUUID": task_id, "imageUUID": str(uuid.uuid4()), + "imageWidth": width, "imageHeight": height, "seed": inp.seed, + "outputFormat": inp.outputFormat} + if inp.uploadUrl: + mime = {"JPEG": "image/jpeg", "PNG": "image/png", "WEBP": "image/webp"}[inp.outputFormat] + try: + with httpx.Client(timeout=httpx.Timeout(60, connect=5), follow_redirects=False, trust_env=False) as client: + response = client.put(inp.uploadUrl, content=data, headers={"Content-Type": mime}) + response.raise_for_status() + except httpx.HTTPError: + return failure("Generated image upload failed", 502, task_id) + parts = urlsplit(inp.uploadUrl) + row["imageURL"] = urlunsplit((parts.scheme, parts.netloc, parts.path, "", "")) + else: + row["imageBase64Data"] = base64.b64encode(data).decode("ascii") + output = {"images": [row], "metrics": public_metrics(metrics)} + if inp.includeCost: + output["cost"] = None # No billing estimate is available for this POC. + return {"id": task_id, "status": "COMPLETED", "executionTime": round((time.monotonic() - started) * 1000), "output": output} + except InputError as exc: + return failure(str(exc), 400, task_id) + except Exception: + LOGGER.exception("Generation request failed") + return failure("Generation failed; inspect local experiment logs", 500, task_id) + finally: + application.state.lock.release() + + @application.post("/runsync") + async def runsync(request: Request): + body = bytearray() + async for chunk in request.stream(): + if len(body) + len(chunk) > MAX_BODY_BYTES: + return failure("Request exceeds 60 MiB", 413) + body.extend(chunk) + try: + envelope = Envelope.model_validate_json(bytes(body)) + except ValidationError as exc: + # No input values in diagnostics: references and signed URLs can be large/private. + issues = [".".join(map(str, e["loc"])) + ": " + e["msg"] for e in exc.errors(include_input=False)] + return failure("Invalid request: " + "; ".join(issues), 422) + return await run_in_threadpool(execute, envelope.input) + + return application + + +app = create_app() diff --git a/runtime/API.md b/runtime/API.md new file mode 100644 index 0000000000000000000000000000000000000000..9fce486503505ae8da6d6060023e44dd18213317 --- /dev/null +++ b/runtime/API.md @@ -0,0 +1,46 @@ +# Local POC API + +Inside the GPU-restricted POC container, run `uvicorn server:app --host 0.0.0.0 --port 8091`. Compose publishes this as `127.0.0.1:8091` on PC 2; use an SSH local forward for remote access. This POC has no authentication and must remain bound to loopback at the host. It does not change the production media router. + +`GET /healthz` is process liveness. `GET /readyz` reports `ready` and `busy`, returning 503 before readiness. Lifespan loads the engine before serving. Readiness means model loading completed; it does not imply every shape has been warmed. + +`POST /runsync` accepts: + +```json +{ + "input": { + "prompt": "A handmade red ceramic teapot on a wooden table", + "width": 1024, + "height": 1024, + "steps": 40, + "CFGScale": 1, + "seed": 42, + "referenceImages": [], + "outputFormat": "PNG", + "outputQuality": 90, + "includeCost": true, + "taskUUID": "example-job", + "useKVCache": true + } +} +``` + +Dimensions must be multiples of 32 from 256 through 1024. The API caps dimensions at the largest tested output size; native 2K is not validated by this POC. Steps are integers 1–60. At most two references are accepted, each a public HTTP(S) URL, base64 image data URI, or raw base64 image. Reference files are limited to 20 MiB and 16 megapixels each; JSON bodies to 60 MiB. Downloads use connection/read limits and a 60-second checked elapsed limit. Redirects and private/loopback targets are rejected. URL validation resolves addresses before the HTTP request; this local POC is not hardened against DNS rebinding and must not be exposed to untrusted internet traffic. + +Only `CFGScale=1` is supported: this POC supplies no negative prompt, so higher CFG values would be silently ineffective in the underlying pipeline. The model's recommended default remains 40 steps. Set `steps: 25` for the faster setting tested in the [expanded quality comparison](experiments/fidelity-v3/REPORT.md); instruction misses and visible changes occur at both settings. PNG preserves alpha; JPEG composites alpha over white; WEBP supports alpha. JPG is accepted as a JPEG alias. Output quality is 1–100 and applies to JPEG/WEBP. + +Successful responses contain `status: "COMPLETED"`, `id`, total `executionTime` in milliseconds, and `output.images[0]` with image dimensions, seed, IDs and `imageBase64Data`. `output.metrics` contains an explicit allowlist of finite scalar timing, peak-memory, cache-hit and image-size measurements. Raw jobs, settings, filesystem paths and per-step lists stay in local experiment records. Total API time also includes reference transfer, image conversion and upload. `includeCost` returns `output.cost: null` because no cost model is established. + +Optional `uploadUrl` performs an outbound presigned PUT with the matching content type and returns `imageURL` with the signed query removed. This presumes the object is readable at that unsigned URL, as with the existing router's presigned upload contract; uploading does not make a private object public. No local filesystem path is accepted as a reference or returned as the primary output. + +Failures contain `status: "FAILED"` and a concise top-level `error`. HTTP codes are 400 for unusable references/URLs, 413 for oversized JSON, 422 for schema errors, 503 with `Retry-After: 5` when busy, 500 for engine errors, and 502 for upload errors. Check the JSON status as well as HTTP status. A single lock covers reference handling, GPU work and result delivery. It remains held until the synchronous worker completes even if the HTTP caller disconnects; there is no claim that an HTTP timeout cancels a CUDA kernel. + +CPU contract tests use a fake engine and never import the GPU runner: + +```sh +python -m unittest -v test_server.py +``` + +The tests exercise schema rejection before inference, JPEG alpha handling, validated reference handoff and cleanup, request body bounds, error sanitization with local traceback logging, public-metric filtering, and concurrent busy rejection. + +The active POC backend is selected at startup with `QWEN_BACKEND=nunchaku` and `QWEN_NUNCHAKU_CHECKPOINT=/cache/qwen-nunchaku-v3-r128`. Use `compose.nunchaku-v3.yaml` with the base Compose file for the tested custom backend; use the base file alone to return to NF4. Both preserve the same request/response contract. The [expanded comparison](experiments/fidelity-v3/REPORT.md) records measured speed and quality tradeoffs. Same-seed repeats can differ visibly even with caching disabled. diff --git a/runtime/LICENSE b/runtime/LICENSE new file mode 100644 index 0000000000000000000000000000000000000000..13ae08d5a5828f508cf0e852b1b316c0c0bbb9b3 --- /dev/null +++ b/runtime/LICENSE @@ -0,0 +1,55 @@ +Qwen RESEARCH LICENSE AGREEMENT + +Qwen RESEARCH LICENSE AGREEMENT Release Date: September 20, 2026 + +By clicking to agree or by using or distributing any portion or element of the Qwen Materials, you will be deemed to have recognized and accepted the content of this Agreement, which is effective immediately. + +1. Definitions + a. This Qwen RESEARCH LICENSE AGREEMENT (this "Agreement") shall mean the terms and conditions for use, reproduction, distribution and modification of the Materials as defined by this Agreement. + b. "We" (or "Us") shall mean Hangzhou Tongyi Laboratory Technology Co., Ltd. + c. "You" (or "Your") shall mean a natural person or legal entity exercising the rights granted by this Agreement and/or using the Materials for any purpose and in any field of use. + d. "Third Parties" shall mean individuals or legal entities that are not under common control with us or you. + e. "Qwen" shall mean the large language models, diffusion models, and software and algorithms, consisting of trained model weights, parameters (including optimizer states), machine-learning model code, inference-enabling code, training-enabling code, fine-tuning enabling code and other elements of the foregoing distributed by us. + f. "Materials" shall mean, collectively, our proprietary Qwen and Documentation (and any portion thereof) made available under this Agreement. + g. "Source" form shall mean the preferred form for making modifications, including but not limited to model source code, documentation source, and configuration files. + h. "Object" form shall mean any form resulting from mechanical transformation or translation of a Source form, including but not limited to compiled object code, generated documentation, and conversions to other media types. + i. "Non-Commercial" shall mean for research or evaluation purposes only. + +2. Grant of Rights + a. You are granted a non-exclusive, worldwide, non-transferable and royalty-free limited license under our intellectual property or other rights owned by us embodied in the Materials to use, reproduce, distribute, copy, create derivative works of, and make modifications to the Materials FOR NON-COMMERCIAL PURPOSES ONLY. + b. You shall not use the Materials for any commercial purpose without obtaining a separate commercial license from us. If you wish to use the Materials commercially, you shall request a license from us at model-business@notice.qwencloud.com. + +3. Redistribution +Subject to Section 2 (Grant of Rights), you may distribute copies or make the Materials, or derivative works thereof, available as part of a product or service that contains any of them, with or without modifications, and in Source or Object form, provided that you meet the following conditions: + a. You shall give any other recipients of the Materials or derivative works a copy of this Agreement; + b. You shall cause any modified files to carry prominent notices stating that you changed the files; + c. You shall retain in all copies of the Materials that you distribute the following attribution notices within a "Notice" text file distributed as a part of such copies: "Qwen is licensed under the Qwen RESEARCH LICENSE AGREEMENT, Copyright (c) 2026 Hangzhou Tongyi Laboratory Technology Co., Ltd. All Rights Reserved."; and + d. You may add your own copyright statement to your modifications and may provide additional or different license terms and conditions for use, reproduction, or distribution of your modifications, or for any such derivative works as a whole, provided your use, reproduction, and distribution of the work otherwise complies with the terms and conditions of this Agreement. + +4. Rules of use + a. The Materials may be subject to export controls or restrictions in China, the United States or other countries or regions. You shall comply with applicable laws and regulations in your use of the Materials. + b. If you use the Materials or any outputs or results therefrom to create, train, fine-tune, or improve an AI model that is distributed or made available, you shall prominently display “Built with Qwen” or “Improved using Qwen” in the related product documentation. + c. You shall not use "Qwen" as the primary name or identifier of any derivative works or products; reasonable descriptive use (e.g., "fine-tuned from Qwen Image") is permitted. + +5. Intellectual Property + a. We retain ownership of all intellectual property rights in and to the Materials and derivatives made by or for us. Conditioned upon compliance with the terms and conditions of this Agreement, with respect to any derivative works and modifications of the Materials that are made by you, you are and will be the owner of such derivative works and modifications. + b. No trademark license is granted to use the trade names, trademarks, service marks, or product names of us, except as required to fulfill notice requirements under this Agreement or as required for reasonable and customary use in describing and redistributing the Materials. + c. If you commence a lawsuit or other proceedings (including a cross-claim or counterclaim in a lawsuit) against us or any entity alleging that the Materials or any output therefrom, or any part of the foregoing, infringe any intellectual property or other right owned or licensable by you, then all licenses granted to you under this Agreement shall terminate as of the date such lawsuit or other proceeding is commenced or brought. + +6. Disclaimer of Warranty and Limitation of Liability + a. We are not obligated to support, update, provide training for, or develop any further version of the Qwen Materials or to grant any license thereto. + b. THE MATERIALS ARE PROVIDED "AS IS" WITHOUT ANY EXPRESS OR IMPLIED WARRANTY OF ANY KIND INCLUDING WARRANTIES OF MERCHANTABILITY, NONINFRINGEMENT, OR FITNESS FOR A PARTICULAR PURPOSE. WE MAKE NO WARRANTY AND ASSUME NO RESPONSIBILITY FOR THE SAFETY OR STABILITY OF THE MATERIALS AND ANY OUTPUT THEREFROM. + c. IN NO EVENT SHALL WE BE LIABLE TO YOU FOR ANY DAMAGES, INCLUDING, BUT NOT LIMITED TO ANY DIRECT, OR INDIRECT, SPECIAL OR CONSEQUENTIAL DAMAGES ARISING FROM YOUR USE OR INABILITY TO USE THE MATERIALS OR ANY OUTPUT OF IT, NO MATTER HOW IT’S CAUSED. + d. You will defend, indemnify and hold harmless us from and against any claim by any third party arising out of or related to your use or distribution of the Materials. + +7. Survival and Termination. + a. The term of this Agreement shall commence upon your acceptance of this Agreement or access to the Materials and will continue in full force and effect until terminated in accordance with the terms and conditions herein. + b. We may terminate this Agreement if you breach any of the terms or conditions of this Agreement. Upon termination of this Agreement, you must delete and cease use of the Materials. Sections 6 and 8 shall survive the termination of this Agreement. + +8. Governing Law and Jurisdiction. + a. This Agreement and any dispute arising out of or relating to it will be governed by the laws of China, without regard to conflict of law principles, and the UN Convention on Contracts for the International Sale of Goods does not apply to this Agreement. + b. The People's Courts in Hangzhou City shall have exclusive jurisdiction over any dispute arising out of this Agreement. + +9. Other Terms and Conditions. + a. Any arrangements, understandings, or agreements regarding the Material not stated herein are separate from and independent of the terms and conditions of this Agreement. You shall request a separate license from us, if you use the Materials in ways not expressly agreed to in this Agreement. + b. We shall not be bound by any additional or different terms or conditions communicated by you unless expressly agreed. diff --git a/runtime/NOTICE b/runtime/NOTICE new file mode 100644 index 0000000000000000000000000000000000000000..8bd7a1155eaca7e16f36a644a6a7d7cb3eec7141 --- /dev/null +++ b/runtime/NOTICE @@ -0,0 +1,9 @@ +Built with Qwen + +Qwen is licensed under the Qwen RESEARCH LICENSE AGREEMENT, Copyright (c) 2026 Hangzhou Tongyi Laboratory Technology Co., Ltd. All Rights Reserved. + +Mesmer Image 21 Nunchaku is a modified, independently calibrated quantization of Qwen Image 2.1. The transformer uses Nunchaku signed INT4 weights and activations with BF16 rank128 residual branches. The text encoder is serialized in bitsandbytes NF4; processor, scheduler and VAE originate from the pinned upstream release. These modifications are by MesmerTech, September 2026, and are not an official Qwen or Nunchaku release. + +The model and derivatives are for non-commercial research and evaluation under the accompanying LICENSE. Commercial use requires a separate license from Qwen. + +Runtime dependencies retain their respective licenses. The custom linear runtime uses MIT HAN Lab Nunchaku; the conversion packing adapter uses DeepCompressor. See THIRD_PARTY_NOTICES.md for pinned source and license links. diff --git a/runtime/THIRD_PARTY_NOTICES.md b/runtime/THIRD_PARTY_NOTICES.md new file mode 100644 index 0000000000000000000000000000000000000000..f9a44a036fb0343f436b6a5db1b7bb059178cbe6 --- /dev/null +++ b/runtime/THIRD_PARTY_NOTICES.md @@ -0,0 +1,10 @@ +# Third-party components + +- Qwen Image 2.1, revision `b3179ad355be050328e483a9dfdd9e60cd62adfa`: [Qwen Research License](https://huggingface.co/Qwen/Qwen-Image-2.1/blob/b3179ad355be050328e483a9dfdd9e60cd62adfa/LICENSE). Full local copy: LICENSE. +- Nunchaku 1.2.1: [upstream source and license](https://github.com/nunchaku-tech/nunchaku/tree/v1.2.1). Generic SVDQW4A4Linear runtime; no claim of official Qwen Image 2.1 support. +- DeepCompressor `69f3473f5e1c1504bae35cc50c7858ef900a9b17`: [source and license](https://github.com/mit-han-lab/deepcompressor/tree/69f3473f5e1c1504bae35cc50c7858ef900a9b17). Conversion only. +- Diffusers `80c7ed262aeffbeb43ef13ae04baeb9b84515a69`: [Apache-2.0 source](https://github.com/huggingface/diffusers/tree/80c7ed262aeffbeb43ef13ae04baeb9b84515a69). +- Transformers 5.17.0: [Apache-2.0 source](https://github.com/huggingface/transformers). +- bitsandbytes 0.50.2: [MIT source](https://github.com/bitsandbytes-foundation/bitsandbytes). + +Dependency packages in the container include their own license metadata. These software licenses do not replace the model's research-only terms. diff --git a/runtime/lean_encoder.py b/runtime/lean_encoder.py new file mode 100644 index 0000000000000000000000000000000000000000..ce2f040f55f7dd1cb770b8f0e28ef30b09cc6b55 --- /dev/null +++ b/runtime/lean_encoder.py @@ -0,0 +1,164 @@ +# MesmerTech modified runtime for the September 2026 Qwen Image 2.1 research quantization. See NOTICE. +"""Optional Qwen3-VL encoder-only adapter for the pinned Qwen Image 2.1 POC. + +Save a normal (including quantized) checkpoint BEFORE calling enable_lean_encoder. +The adapted instance no longer has the architecture's lm_head; do not pass it to +save_pretrained or pipeline.save_pretrained. Reload the normal checkpoint and +reapply this adapter at runtime. No decoder layers or normalization are changed. + +Typical use, after loading and before serving concurrent requests:: + + report = verify_encoder_parity(te, processor, prompt="A red teapot") + assert report["passed"], report + enable_lean_encoder(te) + +The parity helper is optional, runs two encoder forwards, and temporarily changes +forward dispatch. Run it only while the engine is idle. The application must +serialize enablement and parity checks with generation requests. +""" +from types import MethodType + + +def _base_model_forward(self, *args, **kwargs): + # Image generation reads features, never autoregressive vocabulary logits. + kwargs["use_cache"] = False + return self.model(*args, **kwargs) + + +def _dispatch_attribute(te): + # Accelerate wraps forward and calls _old_forward. Replace its inner function + # to preserve onload/offload hooks; future remove/reinstall cycles retain it. + return "_old_forward" if hasattr(te, "_hf_hook") and hasattr(te, "_old_forward") else "forward" + + +def enable_lean_encoder(te): + """Remove only lm_head and forward through the existing multimodal base model. + + Returns small metadata; mutates te in place and is idempotent. This preserves + te.model.language_model so Diffusers' final-RMSNorm hook keeps working. + Existing Accelerate CPU-offload hooks are retained. Save checkpoints first. + """ + if getattr(te, "_qwen21_lean_encoder", False): + return {"enabled": True, "already_enabled": True} + if not hasattr(te, "model") or not hasattr(te.model, "language_model"): + raise TypeError("Expected Qwen3VLForConditionalGeneration with model.language_model") + if not hasattr(te, "lm_head"): + raise ValueError("lm_head is missing; load an unmodified checkpoint before enabling this adapter") + removed_parameters = sum(p.numel() for p in te.lm_head.parameters()) + dispatch = _dispatch_attribute(te) + setattr(te, dispatch, MethodType(_base_model_forward, te)) + te.config.use_cache = False + te.config.text_config.use_cache = False + del te.lm_head + te._qwen21_lean_encoder = True + return { + "enabled": True, + "already_enabled": False, + "removed_parameters": removed_parameters, + "dispatch_attribute": dispatch, + "checkpoint_save_supported": False, + } + + +def _prepare_inputs(processor, prompt, images, device): + from PIL import Image + system = "Comprehend and analyze the provided prompt." + vision_images = [] + for img in images or []: + if not isinstance(img, Image.Image): + raise TypeError("Parity references must be PIL images, already resized as for inference") + if img.mode == "RGBA": + white = Image.new("RGB", img.size, (255, 255, 255)) + white.paste(img, mask=img.getchannel("A")) + img = white + vision_images.append(img) + prefix = " ".join( + f"<|vision_start|><|image_pad|><|vision_end|>" + for i in range(1, len(vision_images) + 1) + ) + text = ( + f"<|im_start|>system\n{system}<|im_end|>\n" + f"<|im_start|>user\n{prefix}{prompt or ' '}<|im_end|>\n" + "<|im_start|>assistant\n" + ) + kwargs = {"text": [text], "padding": True, "padding_side": "left", "return_tensors": "pt"} + if vision_images: + kwargs["images"] = vision_images + encoded = processor(**kwargs).to(device) + forward = {"input_ids": encoded.input_ids, "attention_mask": encoded.attention_mask} + for key in ("pixel_values", "image_grid_thw", "mm_token_type_ids"): + if key in encoded: + forward[key] = encoded[key] + return forward + + +def verify_encoder_parity(te, processor=None, prompt="A red teapot", images=None, + model_inputs=None, device=None, atol=0.0, rtol=0.0): + """Compare full final pre-norm features with/without vocabulary projection. + + Does not remove lm_head or permanently change dispatch/configuration. Supply + either model_inputs (prepared forward kwargs, including reference pixels) or + processor/prompt/optional PIL images. Images should already have the same + resize used by the pipeline. CPU-offloaded encoders use their hook's execution + device by default. Returns a parity report; enable only when passed is true. + + Exact equality is the default because both paths execute the same base model. + Tolerances can be supplied explicitly if the runtime is nondeterministic. + The comparison mimics the pinned pipeline's norm hook and output_hidden_states + request. It compares full features rather than only the prompt-trimmed suffix. + """ + import torch + if getattr(te, "_qwen21_lean_encoder", False) or not hasattr(te, "lm_head"): + raise ValueError("Verify parity before enabling lean encoding") + if not hasattr(te, "model") or not hasattr(te.model, "language_model"): + raise TypeError("Expected Qwen3VLForConditionalGeneration with model.language_model") + if device is None: + device = getattr(getattr(te, "_hf_hook", None), "execution_device", None) + if device is None: + device = next(te.parameters()).device + if model_inputs is None: + if processor is None: + raise ValueError("Provide processor or prepared model_inputs") + model_inputs = _prepare_inputs(processor, prompt, images, device) + else: + model_inputs = { + k: v.to(device) if isinstance(v, torch.Tensor) else v + for k, v in dict(model_inputs).items() + } + model_inputs = {**model_inputs, "use_cache": False, "output_hidden_states": True, "return_dict": True} + # These conditional-generation-only options do not belong to base-model input. + for key in ("labels", "logits_to_keep"): + model_inputs.pop(key, None) + dispatch = _dispatch_attribute(te) + original_forward = getattr(te, dispatch) + was_training = te.training + norm = te.model.language_model.norm + handle = norm.register_forward_hook(lambda module, args, output: args[0]) + try: + te.eval() + with torch.inference_mode(): + baseline = te(**model_inputs).hidden_states[-1].detach().float().cpu().clone() + setattr(te, dispatch, MethodType(_base_model_forward, te)) + candidate = te(**model_inputs).hidden_states[-1].detach().float().cpu() + finally: + setattr(te, dispatch, original_forward) + handle.remove() + if was_training: + te.train() + if baseline.shape != candidate.shape: + return {"passed": False, "reason": "shape_mismatch", "baseline_shape": list(baseline.shape), + "candidate_shape": list(candidate.shape)} + delta = candidate - baseline + finite = bool(torch.isfinite(baseline).all() and torch.isfinite(candidate).all()) + return { + "passed": finite and bool(torch.allclose(baseline, candidate, atol=atol, rtol=rtol)), + "finite": finite, + "shape": list(baseline.shape), + "max_abs_difference": float(delta.abs().max()), + "rms_difference": float(delta.square().mean().sqrt()), + "baseline_rms": float(baseline.square().mean().sqrt()), + "atol": atol, + "rtol": rtol, + "reference_count": len(images or []) if images is not None else None, + "dispatch_attribute": dispatch, + } diff --git a/runtime/nunchaku_backend/__init__.py b/runtime/nunchaku_backend/__init__.py new file mode 100644 index 0000000000000000000000000000000000000000..f134a37e42a57426d4ef7e028855e356f71a6b4f --- /dev/null +++ b/runtime/nunchaku_backend/__init__.py @@ -0,0 +1,12 @@ +# MesmerTech modified runtime for the September 2026 Qwen Image 2.1 research quantization. See NOTICE. +"""Experimental Qwen Image2.1 SVDQuant/Nunchaku integration. + +This package does not alias NF4: its runtime uses packed signed INT4 weights, +INT4 activations, and a BF16 low-rank branch through Nunchaku CUDA kernels. +""" +FORMAT = "qwen21-nunchaku-svdq-int4-v1" + + +def load_transformer(*args, **kwargs): + from .runtime import load_transformer as implementation + return implementation(*args, **kwargs) diff --git a/runtime/nunchaku_backend/layout.py b/runtime/nunchaku_backend/layout.py new file mode 100644 index 0000000000000000000000000000000000000000..0b69faf332ad5054d7e57ec54119118edba6a84d --- /dev/null +++ b/runtime/nunchaku_backend/layout.py @@ -0,0 +1,8 @@ +# MesmerTech modified runtime for the September 2026 Qwen Image 2.1 research quantization. See NOTICE. +"""Shared Qwen2.1 block-linear selection; safe to import without Torch.""" +import re + +BLOCK_LINEAR_PATTERN = re.compile( + r"^transformer_blocks\.\d+\.(?:attn\.(?:to_q|to_k|to_v|to_out\.0)|img_mlp\.(?:proj|out|gate_layer))$" +) + diff --git a/runtime/nunchaku_backend/runtime.py b/runtime/nunchaku_backend/runtime.py new file mode 100644 index 0000000000000000000000000000000000000000..39659d413319d2ff016e33c4b9599d03e4b2c380 --- /dev/null +++ b/runtime/nunchaku_backend/runtime.py @@ -0,0 +1,131 @@ +# MesmerTech modified runtime for the September 2026 Qwen Image 2.1 research quantization. See NOTICE. +"""Load custom Qwen2.1 checkpoints using upstream generic Nunchaku linears. + +The original Diffusers model still owns positional encoding, shared modulation, +attention, prefix-KV caching, timestep handling and output projection. Only the +224 block projection matrices use SVDQW4A4Linear. This is not a port of the older +Qwen-Image architecture or an NF4 wrapper. +""" +import json +from pathlib import Path + +from . import FORMAT +from .layout import BLOCK_LINEAR_PATTERN + + +def probe_backend(): + """Import-only probe, safe in a container with NVIDIA_VISIBLE_DEVICES=void.""" + import inspect + import torch + import nunchaku + from importlib.metadata import version, PackageNotFoundError + from nunchaku.models.linear import SVDQW4A4Linear + parameters = inspect.signature(SVDQW4A4Linear.__init__).parameters + required = {"in_features", "out_features", "rank", "precision", "torch_dtype", "device"} + if not required.issubset(parameters): + raise RuntimeError(f"Unsupported Nunchaku SVDQW4A4Linear signature: {list(parameters)}") + try: + nunchaku_version = version("nunchaku") + except PackageNotFoundError: + nunchaku_version = getattr(nunchaku, "__version__", "unknown") + return {"torch": torch.__version__, "nunchaku": nunchaku_version, + "linear_class": SVDQW4A4Linear.__module__ + "." + SVDQW4A4Linear.__name__, + "precision": "int4", "weight_bits": 4, "activation_bits": 4} + + +def _read_manifest(directory): + manifest = json.loads((directory / "manifest.json").read_text()) + if manifest.get("backend_format") != FORMAT or manifest.get("complete") is not True: + raise ValueError("Not a compatible Qwen2.1 Nunchaku INT4 checkpoint") + layers = manifest.get("layers", {}) + if len(layers) != 224 or any(not BLOCK_LINEAR_PATTERN.fullmatch(n) for n in layers): + raise ValueError("Checkpoint must describe exactly224 Qwen2.1 block linears") + for name, info in layers.items(): + if info["in_features"] % 128 or info["out_features"] % 128 or info["rank"] % 16: + raise ValueError(f"Unsupported kernel dimensions: {name}: {info}") + if info.get("precision", "int4") != "int4": + raise ValueError(f"Unsupported precision for4070 Ti SUPER: {name}") + return manifest + + +def _state_files(directory): + index = directory / "model.safetensors.index.json" + if index.exists(): + weight_map = json.loads(index.read_text())["weight_map"] + filenames = sorted(set(weight_map.values())) + else: + filenames = ["model.safetensors"] + for filename in filenames: + path = directory / filename + if path.parent.resolve() != directory.resolve() or not path.is_file(): + raise ValueError(f"Invalid or missing checkpoint shard: {filename}") + yield path + + +def load_transformer(checkpoint, *, device="cpu", torch_dtype=None): + """Return a QwenImage21Transformer2DModel with real Nunchaku W4A4 blocks. + + Construction and checkpoint load happen on CPU; moving to CUDA is explicit. + `device='cuda:0'` means the GPU visible as0 in the isolated POC container. + Use normal Diffusers model CPU offload if desired. Load this checkpoint with + this function, not stock from_pretrained; its linears have a custom schema. + """ + import torch + from accelerate import init_empty_weights + from diffusers import QwenImage21Transformer2DModel + from nunchaku.models.linear import SVDQW4A4Linear + from safetensors.torch import load_file + + probe = probe_backend() + directory = Path(checkpoint) + manifest = _read_manifest(directory) + torch_dtype = torch.bfloat16 if torch_dtype is None else torch_dtype + if torch_dtype != torch.bfloat16: + raise ValueError("This initial checkpoint/runtime contract requires BF16 compute") + config = json.loads((directory / "config.json").read_text()) + with init_empty_weights(include_buffers=False): + model = QwenImage21Transformer2DModel.from_config(config) + for name, info in manifest["layers"].items(): + parent_name, child_name = name.rsplit(".", 1) + parent = model.get_submodule(parent_name) + original = parent.get_submodule(child_name) + if (original.in_features, original.out_features) != (info["in_features"], info["out_features"]): + raise ValueError(f"Model/checkpoint dimension mismatch: {name}") + if (original.bias is not None) != info["bias"]: + raise ValueError(f"Model/checkpoint bias mismatch: {name}") + parent._modules[child_name] = SVDQW4A4Linear( + info["in_features"], info["out_features"], rank=info["rank"], + bias=info["bias"], precision="int4", act_unsigned=False, + torch_dtype=torch_dtype, device="meta", + ) + expected = set(model.state_dict()) + seen = set() + for path in _state_files(directory): + state = load_file(str(path), device="cpu") + duplicate = seen.intersection(state) + unexpected = set(state).difference(expected) + if duplicate or unexpected: + raise ValueError(f"Invalid shard {path.name}: duplicate={sorted(duplicate)}, unexpected={sorted(unexpected)}") + model.load_state_dict(state, strict=False, assign=True) + seen.update(state) + del state + missing = expected - seen + if missing: + raise ValueError(f"Incomplete checkpoint; missing {sorted(missing)}") + meta = [n for n, p in model.named_parameters() if p.is_meta] + if meta: + raise ValueError(f"Uninitialized model parameters: {meta}") + for name, module in model.named_modules(): + if BLOCK_LINEAR_PATTERN.fullmatch(name): + if module.qweight.dtype != torch.int8 or module.wscales.dtype != torch_dtype: + raise ValueError(f"Wrong packed tensor dtypes: {name}") + model.eval().requires_grad_(False) + model._qwen21_backend = {"kind": "nunchaku-svdq-w4a4-int4", "checkpoint": str(directory), + "manifest": manifest, "runtime": probe} + # No dtype cast here: preserve integer kernel storage and BF16 scales. + model.to(device=device) + return model + + +if __name__ == "__main__": + print(json.dumps(probe_backend(), indent=2)) diff --git a/runtime/runner.py b/runtime/runner.py new file mode 100644 index 0000000000000000000000000000000000000000..b87ae8fa246c8810d35ca1c3462a6c55b564b721 --- /dev/null +++ b/runtime/runner.py @@ -0,0 +1,373 @@ +# MesmerTech modified runtime for the September 2026 Qwen Image 2.1 research quantization. See NOTICE. +"""Pinned Qwen Image 2.1 experiments. One process, one GPU, serialized jobs.""" +import argparse +from collections import OrderedDict +import functools +import hashlib +import json +import os +from pathlib import Path +import time +import traceback + +MODEL = 'Qwen/Qwen-Image-2.1' +REVISION = 'b3179ad355be050328e483a9dfdd9e60cd62adfa' +ROOT = Path(__file__).resolve().parent + + +def _env_bf16_roles(): + """Comma-separated projection roles; empty environment means no override.""" + return [role.strip() for role in os.getenv('QWEN_BF16_ROLES', '').split(',') if role.strip()] + + +def default_args(): + return argparse.Namespace(quant='nf4', sample_dir=os.getenv('QWEN_SAMPLE_DIR'), backend=os.getenv('QWEN_BACKEND','nf4'), nunchaku_checkpoint=os.getenv('QWEN_NUNCHAKU_CHECKPOINT'), offload=True, compile=os.getenv('QWEN_COMPILE','0')=='1', flex=False, + cache=True, prequant=os.getenv('QWEN_PREQUANT'), save_prequant=None, tiling=False, lean_encoder=True, verify_lean=False, bf16_transformer=False, selective_nf4=False, bf16_vision=False, stage_offload=True, release_kv=True, + bf16_source=os.getenv('QWEN_BF16_SOURCE') or None, restore_roles=_env_bf16_roles()) + + +def _validate_hybrid_args(args): + """Validate opt-in configuration before importing Torch or loading weights.""" + roles = getattr(args, 'restore_roles', None) or [] + source = getattr(args, 'bf16_source', None) + if not roles and not source: + return False + if not isinstance(roles, (list, tuple)) or any(not isinstance(role, str) for role in roles): + raise ValueError('restore_roles must be a list of exact projection-role names') + if bool(roles) != bool(source): + raise ValueError('BF16 restoration requires both --bf16-source/QWEN_BF16_SOURCE and --restore-role/QWEN_BF16_ROLES') + if getattr(args, 'backend', 'nf4') != 'nunchaku': + raise ValueError('BF16 projection restoration requires the Nunchaku backend') + from nunchaku_backend.hybrid_v3 import _selected_names + _selected_names(roles, []) + if getattr(args, 'save_prequant', None): + raise ValueError('Hybrid restoration is a runtime override; do not combine it with --save-prequant') + args.restore_roles = sorted(set(roles)) + args.bf16_source = str(source) + return True + + +def _apply_hybrid_override(transformer, args): + """Restore CPU modules before the pipeline installs device/offload hooks.""" + if not getattr(args, 'restore_roles', None): + return None + from nunchaku_backend.hybrid_v3 import restore_bf16_projections + report = restore_bf16_projections(transformer, args.bf16_source, roles=args.restore_roles) + # Measurements serialize args, so record what was actually restored rather + # than only the requested selectors or the original checkpoint precision. + args.hybrid_roles = report['roles'] + args.hybrid_names = report['restored_names'] + args.hybrid_restored_count = report['restored_count'] + args.hybrid_source = report['source_transformer'] + args.hybrid_source_config_sha256 = report['source_config_sha256'] + args.hybrid_bf16_state_bytes = report['bf16_state_bytes'] + args.hybrid_replaced_packed_state_bytes = report['replaced_packed_state_bytes'] + args.hybrid_state_bytes_delta = report['net_state_bytes_delta'] + args.hybrid_precision = 'Original BF16 restored projections; Nunchaku W4A4 remaining projections' + return report + + +class Engine: + def __init__(self, args=None): + engine_init_started=time.perf_counter() + self.args = args or default_args() + args = self.args + _validate_hybrid_args(args) + import torch + from diffusers import QwenImage21Pipeline, QwenImage21Transformer2DModel, BitsAndBytesConfig + from transformers import Qwen3VLForConditionalGeneration, BitsAndBytesConfig as TBits + self.torch = torch + torch.set_num_threads(8) + self.metrics = {} + self.prompt_cache, self.vae_cache = OrderedDict(), OrderedDict() + started = time.perf_counter() + source = args.prequant or MODEL + kwargs = {} if args.prequant else {'revision': REVISION} + q = dict(load_in_4bit=True, bnb_4bit_quant_type='nf4', bnb_4bit_use_double_quant=True, + bnb_4bit_compute_dtype=torch.bfloat16) + te = Qwen3VLForConditionalGeneration.from_pretrained( + MODEL if args.bf16_vision else source, subfolder='text_encoder', dtype=torch.bfloat16, + quantization_config=TBits(**q, **({'llm_int8_skip_modules':['lm_head','model.visual']} if args.bf16_vision else {})) if args.quant == 'nf4' and (not args.prequant or args.bf16_vision) else None, + device_map={'': 0} if args.quant == 'nf4' else {'': 'cpu'}, + **({'revision':REVISION} if args.bf16_vision else kwargs)) + te.config.use_cache = False + te.config.text_config.use_cache = False + print(json.dumps({'event': 'text_encoder_loaded', 'seconds': time.perf_counter()-started, + 'footprint_bytes': te.get_memory_footprint()}), flush=True) + if args.offload: + te.to('cpu') + torch.cuda.empty_cache() + if getattr(args,'backend','nf4') == 'nunchaku': + if not getattr(args,'nunchaku_checkpoint',None): + raise ValueError('--nunchaku-checkpoint is required for the Nunchaku backend') + from nunchaku_backend.runtime import load_transformer + transformer = load_transformer(args.nunchaku_checkpoint,device='cpu',torch_dtype=torch.bfloat16) + else: + transformer = QwenImage21Transformer2DModel.from_pretrained( + MODEL if args.bf16_transformer or args.selective_nf4 else source, subfolder='transformer', torch_dtype=torch.bfloat16, + quantization_config=BitsAndBytesConfig(**q, **({'llm_int8_skip_modules':['time_text_embed','txt_in','img_in','modulation','norm_out','proj_out']} if args.selective_nf4 else {})) if args.quant == 'nf4' and (not args.prequant or args.selective_nf4) and not args.bf16_transformer else None, + **({'revision':REVISION} if args.bf16_transformer or args.selective_nf4 else kwargs)) + self.hybrid_metadata = _apply_hybrid_override(transformer, args) + print(json.dumps({'event':'quantization_scope', + 'encoder_linear4bit':sum(type(m).__name__=='Linear4bit' for m in te.modules()), + 'vision_linear4bit':sum(type(m).__name__=='Linear4bit' for m in te.model.visual.modules()), + 'transformer_linear4bit':sum(type(m).__name__=='Linear4bit' for m in transformer.modules()), + 'transformer_nunchaku_w4a4':sum(type(m).__name__=='SVDQW4A4Linear' for m in transformer.modules()), + **({'hybrid_roles':args.hybrid_roles, 'hybrid_restored_count':args.hybrid_restored_count, + 'hybrid_bf16_state_bytes':args.hybrid_bf16_state_bytes, + 'hybrid_state_bytes_delta':args.hybrid_state_bytes_delta} if self.hybrid_metadata else {})}),flush=True) + self.pipe = QwenImage21Pipeline.from_pretrained( + source, text_encoder=te, transformer=transformer, torch_dtype=torch.bfloat16, **kwargs) + if args.save_prequant: + if args.prequant: + # Transformers 5.17 cannot reserialize its already-loaded bnb + # conversion graph. Copy unchanged components; save only fresh ones. + import shutil + src,dst=Path(args.prequant).resolve(),Path(args.save_prequant).resolve() + if dst==src or dst.is_relative_to(src): + raise ValueError('Save checkpoint to a separate sibling directory') + shutil.copytree(src,dst,dirs_exist_ok=True) + if args.selective_nf4 or args.bf16_transformer: + shutil.rmtree(dst/'transformer') + transformer.save_pretrained(dst/'transformer',safe_serialization=True) + if args.bf16_vision: + shutil.rmtree(dst/'text_encoder') + te.save_pretrained(dst/'text_encoder',safe_serialization=True) + else: + self.pipe.save_pretrained(args.save_prequant, safe_serialization=True) + if args.tiling: + self.pipe.vae.enable_tiling() + if args.offload: + self.pipe.enable_model_cpu_offload(gpu_id=0) + else: + self.pipe.to('cuda') + if args.verify_lean: + from lean_encoder import verify_encoder_parity + from PIL import Image + reports=[verify_encoder_parity(te,self.pipe.processor,prompt='A red ceramic coffee mug'), + verify_encoder_parity(te,self.pipe.processor,prompt='Change the mug to blue', + images=[Image.new('RGB',(512,512),(150,50,40))])] + print(json.dumps({'event':'encoder_parity','reports':reports}),flush=True) + if not all(r['passed'] for r in reports): + raise RuntimeError('Lean encoder parity failed') + self.pipe.maybe_free_model_hooks() + if args.lean_encoder: + from lean_encoder import enable_lean_encoder + print(json.dumps({'event':'lean_encoder',**enable_lean_encoder(te)}),flush=True) + if args.flex: + if not args.compile: + raise ValueError('Flex attention requires compilation') + from diffusers.models.transformers.transformer_qwenimage21 import QwenImage21FlexAttnProcessor + self.pipe.transformer.set_attn_processor(QwenImage21FlexAttnProcessor()) + if args.compile: + self.pipe.transformer.compile(fullgraph=False, mode='default') + self._install_instrumentation() + torch.cuda.synchronize() + self.load_seconds = time.perf_counter() - started + self.startup_seconds=time.perf_counter()-engine_init_started + (ROOT/'results').mkdir(exist_ok=True) + (ROOT/'samples').mkdir(exist_ok=True) + print(json.dumps({'event': 'ready', 'load_seconds': self.load_seconds, 'startup_seconds':self.startup_seconds, 'settings':vars(args)}),flush=True) + + def sync(self): + self.torch.cuda.synchronize() + + def timed(self, name, function): + @functools.wraps(function) + def wrapper(*args, **kwargs): + self.sync() + started = time.perf_counter() + try: + return function(*args, **kwargs) + finally: + self.sync() + self.metrics[name] = self.metrics.get(name, 0) + time.perf_counter()-started + return wrapper + + @staticmethod + def _pil_key(image): + return (image.mode, image.size, hashlib.sha256(image.tobytes()).hexdigest()) + + def _install_instrumentation(self): + pipe = self.pipe + original = pipe._get_qwen_prompt_embeds + def compute_prompt(prompt=None,image=None,device=None): + result=original(prompt=prompt,image=image,device=device) + if self.args.offload and self.args.stage_offload: + pipe.text_encoder.to('cpu') + self.torch.cuda.empty_cache() + return result + def cached_prompt(prompt=None, image=None, device=None): + if not self.args.cache: + return compute_prompt(prompt=prompt, image=image, device=device) + key = (json.dumps(prompt), tuple(self._pil_key(i) for i in (image or []))) + target = device or pipe._execution_device + if self.args.cache and key in self.prompt_cache: + self.metrics['prompt_cache_hits'] = self.metrics.get('prompt_cache_hits', 0)+1 + self.prompt_cache.move_to_end(key) + return tuple(t.to(target) for t in self.prompt_cache[key]) + result = compute_prompt(prompt=prompt, image=image, device=device) + if self.args.cache: + self.prompt_cache[key] = tuple(t.detach().cpu() for t in result) + while len(self.prompt_cache)>8: + self.prompt_cache.popitem(last=False) + return result + pipe._get_qwen_prompt_embeds = self.timed('encode_prompt_seconds', cached_prompt) + original_vae = pipe._encode_vae_image + def compute_vae(image,generator): + result=original_vae(image,generator) + if self.args.offload and self.args.stage_offload: + pipe.vae.to('cpu') + self.torch.cuda.empty_cache() + return result + def cached_vae(image, generator): + if not self.args.cache: + return compute_vae(image,generator) + # Upstream uses argmax, so cached reference latents do not depend on RNG seed. + raw=image.detach().cpu().contiguous() + key=(tuple(raw.shape),str(raw.dtype),hashlib.sha256(raw.view(self.torch.uint8).numpy().tobytes()).hexdigest()) + if self.args.cache and key in self.vae_cache: + self.metrics['vae_cache_hits'] = self.metrics.get('vae_cache_hits', 0)+1 + self.vae_cache.move_to_end(key) + return self.vae_cache[key].to(image.device) + result=compute_vae(image,generator) + if self.args.cache: + self.vae_cache[key]=result.detach().cpu() + while len(self.vae_cache)>8: + self.vae_cache.popitem(last=False) + return result + pipe._encode_vae_image=self.timed('encode_reference_seconds',cached_vae) + pipe.vae.decode=self.timed('decode_seconds',pipe.vae.decode) + # Module hooks survive Accelerate rebuilding forward/offload hooks after each job. + # Keep timing calls outside Dynamo graphs. + @self.torch.compiler.disable + def before_transformer(module, inputs, kwargs): + cache=kwargs.get('kv_cache') + if cache is not None: + self._kv_caches[id(cache)]=cache + self.sync() + self._transformer_start=time.perf_counter() + @self.torch.compiler.disable + def after_transformer(module, inputs, output): + self.sync() + self.metrics['transformer_seconds']=self.metrics.get('transformer_seconds',0)+time.perf_counter()-self._transformer_start + pipe.transformer.register_forward_pre_hook(before_transformer,with_kwargs=True) + pipe.transformer.register_forward_hook(after_transformer) + + def generate(self, job): + try: + return self._generate(job) + except Exception: + # Restore stage hooks after an interrupted/OOM request before a later retry. + self.pipe.maybe_free_model_hooks() + self.torch.cuda.empty_cache() + raise + + def _generate(self, job): + job={**job,'kv_cache':job.get('kv_cache',True)} + from PIL import Image + torch=self.torch + refs=[Image.open(p).copy() for p in job.get('images', [])] + self.metrics={} + self._kv_caches={} + torch.cuda.reset_peak_memory_stats() + self.sync() + started=time.perf_counter() + started_unix=time.time() + last=started + steps=[] + def callback(pipe,index,timestep,kwargs): + nonlocal last + self.sync() + now=time.perf_counter() + steps.append(now-last) + last=now + # Prefix KV is dead after the final sampler step. Upstream otherwise + # retains it while the large VAE decode runs (about 2 GiB for one 1K ref). + if self.args.release_kv and index+1==pipe.num_timesteps: + released=0 + for cache in self._kv_caches.values(): + for layer in cache.layer_caches: + for name in ('k','v'): + value=getattr(layer,name) + if value is not None: + released+=value.numel()*value.element_size() + setattr(layer,name,None) + self._kv_caches.clear() + self.metrics['released_kv_mib']=released/2**20 + torch.cuda.empty_cache() + if index == 0 or (index+1)%10 == 0: + print(json.dumps({'event':'step','label':job.get('label'),'step':index+1,'elapsed':now-started}),flush=True) + return kwargs + result=self.pipe( + prompt=job['prompt'], image=refs or None, + width=job.get('width',1024), height=job.get('height',1024), + output_resolution=max(job.get('width',1024),job.get('height',1024)), + num_inference_steps=job.get('steps',40), true_cfg_scale=job.get('cfg',1.0), + generator=torch.Generator(device='cuda').manual_seed(job.get('seed',42)), + use_kv_cache=job.get('kv_cache',True), callback_on_step_end=callback).images[0] + self.sync() + elapsed=time.perf_counter()-started + label=job.get('label',str(time.time_ns())) + if not label or any(c not in 'abcdefghijklmnopqrstuvwxyzABCDEFGHIJKLMNOPQRSTUVWXYZ0123456789_-' for c in label): + raise ValueError('label must contain only letters, numbers, underscores or hyphens') + sample_dir=Path(getattr(self.args,'sample_dir',None) or ROOT/'samples') + sample_dir.mkdir(parents=True,exist_ok=True) + path=sample_dir/f'{label}.png' + result.save(path) + metrics={**self.metrics, 'path':str(path),'seconds':elapsed,'started_unix':started_unix,'finished_unix':time.time(), + 'peak_allocated_mib':torch.cuda.max_memory_allocated()/2**20, + 'peak_reserved_mib':torch.cuda.max_memory_reserved()/2**20, + 'step_seconds':steps,'job':job,'settings':vars(self.args),'load_seconds':self.load_seconds,'startup_seconds':self.startup_seconds, + 'output_size':result.size,'output_mode':result.mode} + with (ROOT/'results'/'measurements.jsonl').open('a') as f: + f.write(json.dumps(metrics)+'\n') + print(json.dumps({'event':'result',**metrics}),flush=True) + return metrics + + +def _parse_cli_args(argv=None): + parser=argparse.ArgumentParser() + parser.add_argument('--jobs') + parser.add_argument('--sample-dir',help='Separate output directory for a quality experiment') + parser.add_argument('--backend',choices=['nf4','nunchaku'],default='nf4') + parser.add_argument('--nunchaku-checkpoint',help='Custom packed Qwen 2.1 transformer checkpoint') + from nunchaku_backend.hybrid_v3 import ROLES + parser.add_argument('--bf16-source','--restore-source',dest='bf16_source',default=os.getenv('QWEN_BF16_SOURCE') or None, + help='Original BF16 snapshot for optional Nunchaku projection restoration') + parser.add_argument('--restore-role',dest='restore_roles',choices=ROLES,action='append',default=None, + help='Restore this projection role to BF16 in all 32 blocks; repeat for multiple roles') + parser.add_argument('--quant',choices=['nf4'],default='nf4') + parser.add_argument('--offload',action=argparse.BooleanOptionalAction,default=True) + parser.add_argument('--compile',action='store_true') + parser.add_argument('--flex',action='store_true') + parser.add_argument('--cache',action=argparse.BooleanOptionalAction,default=True) + parser.add_argument('--tiling',action=argparse.BooleanOptionalAction,default=False) + parser.add_argument('--release-kv',action=argparse.BooleanOptionalAction,default=True) + parser.add_argument('--stage-offload',action=argparse.BooleanOptionalAction,default=True) + parser.add_argument('--selective-nf4',action='store_true') + parser.add_argument('--bf16-vision',action='store_true') + parser.add_argument('--bf16-transformer',action='store_true') + parser.add_argument('--lean-encoder',action=argparse.BooleanOptionalAction,default=True) + parser.add_argument('--verify-lean',action='store_true') + parser.add_argument('--prequant') + parser.add_argument('--save-prequant') + args=parser.parse_args(argv) + if args.restore_roles is None: + args.restore_roles = _env_bf16_roles() + return args + + +def main(): + args=_parse_cli_args() + engine=Engine(args) + if args.jobs: + for job in json.loads(Path(args.jobs).read_text()): + try: + engine.generate(job) + except Exception: + traceback.print_exc() + raise + +if __name__=='__main__': + main() diff --git a/runtime/runpod/Dockerfile b/runtime/runpod/Dockerfile new file mode 100644 index 0000000000000000000000000000000000000000..a40e8b3f90910562bced76a752f771ecc11dcb61 --- /dev/null +++ b/runtime/runpod/Dockerfile @@ -0,0 +1,46 @@ +# syntax=docker/dockerfile:1.7 +# Build context: qwen-image-2.1-poc/ (not runpod/). +# Clean PyTorch 2.8.0 / CUDA 12.8 / Python 3.11 runtime, pinned by digest. +FROM pytorch/pytorch@sha256:417bd75df6365104c283ea4c1651fb3530d9eb5a4c2fafa51943cff2a94e6385 +USER root +WORKDIR /app +ENV PYTHONUNBUFFERED=1 PIP_NO_CACHE_DIR=1 PYTHONPATH=/app + +COPY runpod/requirements.txt /tmp/requirements.txt +RUN python -c "import sys, torch; assert sys.version_info[:2] == (3,11); assert torch.__version__ == '2.8.0+cu128'" \ + && printf 'torch==2.8.0+cu128\ntriton==3.4.0\n' > /tmp/base-constraints.txt \ + && python -m pip install --no-cache-dir -c /tmp/base-constraints.txt -r /tmp/requirements.txt \ + && python -m pip install --no-cache-dir --no-deps torchvision==0.23.0 --index-url https://download.pytorch.org/whl/cu128 \ + && python -c "import torch, torchvision, nunchaku; from nunchaku.models.linear import SVDQW4A4Linear; from diffusers import QwenImage21Pipeline; from transformers import Qwen3VLForConditionalGeneration; assert torch.__version__ == '2.8.0+cu128'; assert torchvision.__version__ == '0.23.0+cu128'" \ + && python -m pip freeze > /opt/serving-environment.freeze.txt + +# Stable large model layer precedes application code. An empty/moving revision fails. +ARG MODEL_REPO=mesmertech/Mesmer-Image-21-Nunchaku +ARG MODEL_REVISION= +COPY runpod/download_models.py /tmp/download_models.py +RUN --mount=type=secret,id=hf_token \ + python /tmp/download_models.py --repo "$MODEL_REPO" --revision "$MODEL_REVISION" --destination /models \ + && rm -rf /root/.cache/huggingface /tmp/download_models.py + +# Explicit serving-only copies: no DeepCompressor, calibration code or NF4 DiT. +COPY runner.py server.py lean_encoder.py /app/ +COPY nunchaku_backend/__init__.py nunchaku_backend/runtime.py nunchaku_backend/layout.py /app/nunchaku_backend/ +COPY runpod/handler.py /app/runpod/handler.py +COPY LICENSE NOTICE THIRD_PARTY_NOTICES.md /app/ +RUN python -m compileall -q /app + +ARG IMAGE_TAG=mesmerlord/mesmer-image21-runpod:v1-int4 +ENV IMAGE_TAG=${IMAGE_TAG} MODEL_REPO=${MODEL_REPO} MODEL_REVISION=${MODEL_REVISION} \ + QWEN_BACKEND=nunchaku QWEN_PREQUANT=/models QWEN_NUNCHAKU_CHECKPOINT=/models/transformer \ + QWEN_COMPILE=0 QWEN_SAMPLE_DIR=/tmp/qwen-output \ + QWEN_APP_DIR=/app QWEN_PORT=8091 QWEN_STARTUP_TIMEOUT=900 QWEN_REQUEST_TIMEOUT=600 \ + HF_HUB_OFFLINE=1 TRANSFORMERS_OFFLINE=1 HF_HUB_DISABLE_TELEMETRY=1 USE_HUB_KERNELS=NO \ + HF_HOME=/tmp/huggingface TORCHINDUCTOR_CACHE_DIR=/tmp/torchinductor TRITON_CACHE_DIR=/tmp/triton \ + CONCURRENCY=1 RUNPOD_SKIP_AUTO_SYSTEM_CHECKS=true RUNPOD_SKIP_GPU_CHECK=true +LABEL org.opencontainers.image.title="Mesmer Image 21 Nunchaku evaluation worker" \ + org.opencontainers.image.description="Built with Qwen; Qwen Research License; Ada INT4 evaluation POC" \ + ai.mesmer.model.repository=${MODEL_REPO} ai.mesmer.model.revision=${MODEL_REVISION} +HEALTHCHECK --interval=30s --timeout=5s --start-period=900s --retries=3 \ + CMD python -c "import json,urllib.request; assert json.load(urllib.request.urlopen('http://127.0.0.1:8091/readyz',timeout=3))['ready']" +ENTRYPOINT [] +CMD ["python", "-u", "/app/runpod/handler.py"] diff --git a/runtime/runpod/download_models.py b/runtime/runpod/download_models.py new file mode 100644 index 0000000000000000000000000000000000000000..aec45d6ec86db459c9911c0620f49d92ceb2bc60 --- /dev/null +++ b/runtime/runpod/download_models.py @@ -0,0 +1,57 @@ +# MesmerTech modified runtime for the September 2026 Qwen Image 2.1 research quantization. See NOTICE. +"""Bake a revision-pinned complete pipeline without calibration dependencies.""" +from __future__ import annotations + +import argparse +import hashlib +import json +from pathlib import Path +import re + + +def validate_revision(revision): + if not re.fullmatch(r"[0-9a-f]{40}", revision or ""): + raise ValueError("MODEL_REVISION must be an explicit 40-character HF commit SHA") + return revision + + +def validate_snapshot(root): + root = Path(root) + for name in ("model_index.json", "processor", "scheduler", "text_encoder", "vae", "transformer/manifest.json"): + if not (root / name).exists(): + raise ValueError(f"Incomplete pipeline snapshot: missing {name}") + manifest_path = root / "transformer/manifest.json" + manifest = json.loads(manifest_path.read_text()) + if manifest.get("backend_format") != "qwen21-nunchaku-svdq-int4-v1" or manifest.get("complete") is not True: + raise ValueError("Expected a complete Nunchaku INT4 transformer manifest") + layers = manifest.get("layers", {}) + if len(layers) != 224 or any(item.get("rank") != 128 for item in layers.values()): + raise ValueError("This image requires the calibrated 224-layer rank128 transformer") + return hashlib.sha256(manifest_path.read_bytes()).hexdigest() + + +def main(): + parser = argparse.ArgumentParser() + parser.add_argument("--repo", required=True) + parser.add_argument("--revision", required=True) + parser.add_argument("--destination", required=True) + args = parser.parse_args() + validate_revision(args.revision) + from huggingface_hub import snapshot_download + secret = Path("/run/secrets/hf_token") + token = secret.read_text().strip() if secret.exists() else None + snapshot_download(repo_id=args.repo, revision=args.revision, local_dir=args.destination, + token=token, max_workers=8, + allow_patterns=["*.json", "processor/**", "scheduler/**", "text_encoder/**", + "vae/**", "transformer/**", "tokenizer/**", + "LICENSE", "NOTICE", "THIRD_PARTY_NOTICES.md"]) + manifest_hash = validate_snapshot(args.destination) + identity = {"repo_id": args.repo, "revision": args.revision, + "transformer_manifest_sha256": manifest_hash, + "backend": "nunchaku", "precision": "W4A4 + BF16 low-rank", "rank": 128} + (Path(args.destination) / "BUILD_IDENTITY.json").write_text(json.dumps(identity, indent=2) + "\n") + print(json.dumps({"event": "model_snapshot_baked", **identity}), flush=True) + + +if __name__ == "__main__": + main() diff --git a/runtime/runpod/handler.py b/runtime/runpod/handler.py new file mode 100644 index 0000000000000000000000000000000000000000..00e83a5f6f2aee568c9fc55b2daa144a5de61ccd --- /dev/null +++ b/runtime/runpod/handler.py @@ -0,0 +1,149 @@ +# MesmerTech modified runtime for the September 2026 Qwen Image 2.1 research quantization. See NOTICE. +"""Thin RunPod worker: one loopback shared server, one job at a time. + +Importing this module never imports Torch, RunPod or starts a subprocess. +""" +from __future__ import annotations + +import atexit +import json +import logging +import os +from pathlib import Path +import subprocess +import sys +import threading +import time + +import httpx + +LOGGER = logging.getLogger("mesmer_image21.runpod") + + +class Worker: + def __init__(self, client=None, popen=None, clock=None, sleep=None): + self.base_url = "http://127.0.0.1:" + os.getenv("QWEN_PORT", "8091") + self.startup_timeout = float(os.getenv("QWEN_STARTUP_TIMEOUT", "900")) + self.request_timeout = float(os.getenv("QWEN_REQUEST_TIMEOUT", "600")) + self.client = client or httpx.Client(trust_env=False, follow_redirects=False) + self.popen = popen or subprocess.Popen + self.clock = clock or time.monotonic + self.sleep = sleep or time.sleep + self.process = None + self.ready = False + self.lock = threading.Lock() + + def identity(self): + model = {"repo_id": os.getenv("MODEL_REPO", "mesmertech/Mesmer-Image-21-Nunchaku"), + "revision": os.getenv("MODEL_REVISION", "unbaked"), + "backend": "nunchaku", "precision": "W4A4 + BF16 low-rank", "rank": 128} + identity_path = Path(os.getenv("QWEN_PREQUANT", "/models")) / "BUILD_IDENTITY.json" + if identity_path.is_file(): + baked = json.loads(identity_path.read_text()) + model = {key: baked[key] for key in (*model.keys(), "transformer_manifest_sha256") if key in baked} + return {"handler": "mesmer-image21-shared-server-v1", + "image": os.getenv("IMAGE_TAG", "unbuilt"), "checkpoint": model} + + def fail(self, message, stage, http_status=None): + # A top-level error is the RunPod SDK failure contract, not a nested COMPLETED result. + result = {"error": message, "stage": stage, **self.identity()} + if http_status is not None: + result["inner_http_status"] = http_status + return result + + def ensure_ready(self): + if self.process is None or self.process.poll() is not None: + self.ready = False + self.process = self.popen([sys.executable, "-m", "uvicorn", "server:app", "--host", + "127.0.0.1", "--port", self.base_url.rsplit(":", 1)[1], + "--workers", "1", "--no-access-log"], + cwd=os.getenv("QWEN_APP_DIR", "/app"), env=dict(os.environ)) + if self.ready: + return + deadline = self.clock() + self.startup_timeout + while self.clock() < deadline: + if self.process.poll() is not None: + raise RuntimeError("Shared model server exited before readiness") + try: + response = self.client.get(self.base_url + "/readyz", timeout=3) + if response.status_code == 200 and response.json().get("ready") is True: + self.ready = True + return + except (httpx.HTTPError, ValueError, AttributeError): + pass + self.sleep(0.25) + raise TimeoutError("Shared model server readiness timed out") + + def handle(self, job): + if not isinstance(job, dict) or not isinstance(job.get("input"), dict): + return self.fail("Job input must be an object", "validation") + if not self.lock.acquire(blocking=False): + return self.fail("Worker is busy; concurrency must remain 1", "admission", 503) + try: + try: + self.ensure_ready() + except Exception: + LOGGER.exception("Shared server startup failed") + return self.fail("Shared model server failed readiness; inspect worker logs", "startup") + started = self.clock() + try: + response = self.client.post(self.base_url + "/runsync", json={"input": job["input"]}, + timeout=httpx.Timeout(self.request_timeout, connect=5)) + except httpx.HTTPError: + # No retry: the first request may still be running on the GPU. + LOGGER.exception("Shared server request failed") + return self.fail("Shared model server request failed; no automatic retry", "transport") + try: + body = response.json() + except ValueError: + return self.fail("Shared model server returned invalid JSON", "response", response.status_code) + if not isinstance(body, dict): + return self.fail("Shared model server returned an invalid envelope", "response", response.status_code) + if response.status_code != 200 or body.get("status") != "COMPLETED": + error = body.get("error") + error = error[:1000] if isinstance(error, str) else "Shared model server rejected generation" + return self.fail(error, "inference", response.status_code) + output = body.get("output") + if not isinstance(output, dict) or not isinstance(output.get("images"), list) or not output["images"]: + return self.fail("Shared model server returned no images", "response", response.status_code) + if any(not isinstance(row, dict) or not (row.get("imageURL") or row.get("imageBase64Data")) for row in output["images"]): + return self.fail("Shared model server returned an image without result data", "response", response.status_code) + return {**output, **self.identity(), "handler_seconds": round(self.clock() - started, 6)} + finally: + self.lock.release() + + def close(self): + if self.process is not None and self.process.poll() is None: + self.process.terminate() + try: + self.process.wait(timeout=10) + except subprocess.TimeoutExpired: + self.process.kill() + self.process.wait(timeout=5) + self.client.close() + + +_worker = None + + +def handler(job): + global _worker + if _worker is None: + _worker = Worker() + atexit.register(_worker.close) + return _worker.handle(job) + + +def main(): + global _worker + logging.basicConfig(level=logging.INFO) + import runpod + _worker = Worker() + atexit.register(_worker.close) + # Fail startup instead of accepting jobs before the actual model has loaded. + _worker.ensure_ready() + runpod.serverless.start({"handler": handler, "concurrency_modifier": lambda current: 1}) + + +if __name__ == "__main__": + main() diff --git a/runtime/runpod/requirements.txt b/runtime/runpod/requirements.txt new file mode 100644 index 0000000000000000000000000000000000000000..abc19195e64cd1c9939993e94dd3bf5c0728cb31 --- /dev/null +++ b/runtime/runpod/requirements.txt @@ -0,0 +1,24 @@ +# Serving dependencies taken from the measured POC environment.freeze.txt. +# Torch/triton come from the immutable PyTorch 2.8.0 CUDA 12.8 base. +accelerate==1.12.0 +bitsandbytes==0.50.2 +diffusers @ https://github.com/huggingface/diffusers/archive/80c7ed262aeffbeb43ef13ae04baeb9b84515a69.zip#sha256=7b0d00b5b44f5ca1745d477ae4a0c2da125d58b89cff8470c422d3181a797a09 +einops==0.8.2 +fastapi==0.116.1 +hf-xet==1.6.0 +httpx==0.28.1 +huggingface-hub==1.32.0 +nunchaku @ https://github.com/nunchaku-tech/nunchaku/releases/download/v1.2.1/nunchaku-1.2.1+cu12.8torch2.8-cp311-cp311-linux_x86_64.whl#sha256=77dab1a3abdff16d5cbff70e26a5700e6fccb77b4207c3d53127d5f0224bd82d +numpy==2.3.2 +peft==0.19.1 +pillow==12.2.0 +protobuf==7.35.1 +pydantic==2.13.4 +requests==2.34.2 +runpod==1.9.1 +safetensors==0.8.0 +sentencepiece==0.2.1 +starlette==0.47.3 +tokenizers==0.23.2 +transformers==5.17.0 +uvicorn==0.35.0 diff --git a/runtime/server.py b/runtime/server.py new file mode 100644 index 0000000000000000000000000000000000000000..e095e932e346e908728a5cbfa3c56772f4047d6f --- /dev/null +++ b/runtime/server.py @@ -0,0 +1,291 @@ +# MesmerTech modified runtime for the September 2026 Qwen Image 2.1 research quantization. See NOTICE. +"""Local single-GPU RunPod-shaped POC API. Run: uvicorn server:app --host 0.0.0.0 --port 8091.""" +from __future__ import annotations + +import base64 +import binascii +import io +import ipaddress +import logging +import math +import socket +import tempfile +import threading +import time +import uuid +from contextlib import asynccontextmanager +from pathlib import Path +from typing import Literal +from urllib.parse import urlsplit, urlunsplit + +import httpx +from fastapi import FastAPI, Request +from fastapi.responses import JSONResponse +from PIL import Image, ImageOps +from pydantic import BaseModel, ConfigDict, Field, ValidationError, field_validator +from starlette.concurrency import run_in_threadpool + +MAX_REF_BYTES = 20 * 1024 * 1024 +MAX_BODY_BYTES = 60 * 1024 * 1024 +MAX_REF_PIXELS = 16_777_216 +LOGGER = logging.getLogger("qwen_image21_poc.server") +PUBLIC_NUMERIC_METRICS = frozenset({ + "seconds", "encode_prompt_seconds", "encode_reference_seconds", "decode_seconds", + "transformer_seconds", "load_seconds", "startup_seconds", "started_unix", "finished_unix", + "peak_allocated_mib", "peak_reserved_mib", "prompt_cache_hits", "vae_cache_hits", +}) + + +def public_metrics(metrics): + """Expose measured scalar values, never raw jobs, configuration or paths.""" + result = { + key: value for key, value in metrics.items() + if key in PUBLIC_NUMERIC_METRICS + and isinstance(value, (int, float)) and not isinstance(value, bool) + and math.isfinite(value) + } + size = metrics.get("output_size") + if isinstance(size, (list, tuple)) and len(size) == 2: + if all(isinstance(value, int) and not isinstance(value, bool) and value > 0 for value in size): + result.update(output_width=size[0], output_height=size[1]) + if metrics.get("output_mode") in ("RGB", "RGBA", "L", "LA"): + result["output_mode"] = metrics["output_mode"] + return result + + +class Inputs(BaseModel): + model_config = ConfigDict(extra="forbid", allow_inf_nan=False) + prompt: str = Field(min_length=1, max_length=16000) + width: int = Field(default=1024, ge=256, le=1024, strict=True) + height: int = Field(default=1024, ge=256, le=1024, strict=True) + steps: int = Field(default=40, ge=1, le=60, strict=True) + CFGScale: Literal[1.0] = 1.0 + referenceImages: list[str] = Field(default_factory=list, max_length=2) + outputFormat: Literal["PNG", "JPEG", "WEBP"] = "PNG" + outputQuality: int = Field(default=90, ge=1, le=100, strict=True) + uploadUrl: str | None = Field(default=None, max_length=16000) + includeCost: bool = False + taskUUID: str | None = Field(default=None, max_length=128) + seed: int = Field(default=42, ge=0, le=2**63 - 1, strict=True) + useKVCache: bool = True + + @field_validator("width", "height") + @classmethod + def dimensions(cls, value): + if value % 32: + raise ValueError("dimensions must be multiples of 32") + return value + + @field_validator("prompt") + @classmethod + def meaningful_prompt(cls, value): + if not value.strip(): + raise ValueError("prompt must contain text") + return value + + @field_validator("outputFormat", mode="before") + @classmethod + def normalize_format(cls, value): + if isinstance(value, str): + return {"JPG": "JPEG"}.get(value.upper(), value.upper()) + return value + + +class Envelope(BaseModel): + model_config = ConfigDict(extra="forbid") + input: Inputs + + +class InputError(Exception): + pass + + +def failure(message, status=400, task_id=None): + body = {"status": "FAILED", "error": message} + if task_id: + body["id"] = task_id + return JSONResponse(body, status_code=status) + + +def external_url(url): + """Public HTTP(S) only; redirects are deliberately not followed.""" + try: + parsed = urlsplit(url) + if parsed.scheme not in ("http", "https") or not parsed.hostname: + raise ValueError() + if parsed.username or parsed.password or parsed.fragment: + raise ValueError() + port = parsed.port or (443 if parsed.scheme == "https" else 80) + addresses = socket.getaddrinfo(parsed.hostname, port, type=socket.SOCK_STREAM) + if not addresses or any(not ipaddress.ip_address(a[4][0]).is_global for a in addresses): + raise ValueError() + except (ValueError, OSError): + raise InputError("URL must identify a public HTTP(S) host") from None + return url + + +def reference_bytes(source): + if source.startswith(("https://", "http://")): + external_url(source) + started = time.monotonic() + chunks, size = [], 0 + try: + with httpx.Client(timeout=httpx.Timeout(30, connect=5), follow_redirects=False, trust_env=False) as client: + with client.stream("GET", source) as response: + response.raise_for_status() + length = response.headers.get("content-length") + if length and int(length) > MAX_REF_BYTES: + raise InputError("Reference exceeds 20 MiB") + for chunk in response.iter_bytes(): + size += len(chunk) + if size > MAX_REF_BYTES: + raise InputError("Reference exceeds 20 MiB") + if time.monotonic() - started > 60: + raise InputError("Reference download exceeded time limit") + chunks.append(chunk) + return b"".join(chunks) + except (httpx.HTTPError, ValueError): + raise InputError("Reference download failed") from None + if source.startswith("data:"): + header, separator, source = source.partition(",") + if not separator or not header.startswith("data:image/") or not header.endswith(";base64"): + raise InputError("Reference data URI must contain a base64 image") + if len(source) > 4 * math.ceil(MAX_REF_BYTES / 3): + raise InputError("Reference exceeds 20 MiB") + try: + data = base64.b64decode(source, validate=True) + except (ValueError, binascii.Error): + raise InputError("Reference must be a public URL or base64 image") from None + if len(data) > MAX_REF_BYTES: + raise InputError("Reference exceeds 20 MiB") + return data + + +def save_reference(source, path): + data = reference_bytes(source) + try: + with Image.open(io.BytesIO(data)) as image: + if image.width * image.height > MAX_REF_PIXELS or image.width < 1 or image.height < 1: + raise InputError("Reference exceeds 16 megapixels") + image.verify() + with Image.open(io.BytesIO(data)) as image: + oriented = ImageOps.exif_transpose(image) + oriented.convert("RGBA" if "A" in oriented.getbands() else "RGB").save(path, format="PNG") + except InputError: + raise + except Exception: + raise InputError("Reference is not a valid supported raster image") from None + + +def encoded_result(path, fmt, quality): + with Image.open(path) as image: + if fmt == "JPEG": + rgba = image.convert("RGBA") + image = Image.new("RGB", rgba.size, "white") + image.paste(rgba, mask=rgba.getchannel("A")) + stream = io.BytesIO() + image.save(stream, format=fmt, **({"quality": quality} if fmt != "PNG" else {})) + return stream.getvalue(), image.width, image.height + + +def create_app(engine_factory=None): + @asynccontextmanager + async def lifespan(application): + application.state.ready = False + if engine_factory is None: + from runner import Engine, default_args + factory = lambda: Engine(default_args()) + else: + factory = engine_factory + application.state.engine = await run_in_threadpool(factory) + application.state.ready = True + try: + yield + finally: + application.state.ready = False + + application = FastAPI(title="Qwen Image 2.1 POC", lifespan=lifespan) + application.state.ready = False + application.state.lock = threading.Lock() + + @application.get("/healthz") + def health(): + return {"status": "ok"} + + @application.get("/readyz") + def ready(): + if not application.state.ready: + return JSONResponse({"ready": False}, status_code=503) + return {"ready": True, "busy": application.state.lock.locked()} + + def execute(inp): + task_id = inp.taskUUID or str(uuid.uuid4()) + if not application.state.ready: + return failure("Model is not ready", 503, task_id) + # The lock lives entirely in this synchronous function: client disconnects + # cannot release it while a CUDA operation is still executing. + if not application.state.lock.acquire(blocking=False): + response = failure("GPU is busy; retry later", 503, task_id) + response.headers["Retry-After"] = "5" + return response + started = time.monotonic() + try: + if inp.uploadUrl: + external_url(inp.uploadUrl) + with tempfile.TemporaryDirectory(prefix="qwen-request-") as directory: + paths = [] + for i, source in enumerate(inp.referenceImages): + path = str(Path(directory) / f"reference-{i}.png") + save_reference(source, path) + paths.append(path) + job = dict(prompt=inp.prompt, width=inp.width, height=inp.height, + steps=inp.steps, cfg=inp.CFGScale, images=paths, seed=inp.seed, + kv_cache=inp.useKVCache, label=f"api-{uuid.uuid4().hex}") + metrics = application.state.engine.generate(job) + data, width, height = encoded_result(metrics["path"], inp.outputFormat, inp.outputQuality) + row = {"taskUUID": task_id, "imageUUID": str(uuid.uuid4()), + "imageWidth": width, "imageHeight": height, "seed": inp.seed, + "outputFormat": inp.outputFormat} + if inp.uploadUrl: + mime = {"JPEG": "image/jpeg", "PNG": "image/png", "WEBP": "image/webp"}[inp.outputFormat] + try: + with httpx.Client(timeout=httpx.Timeout(60, connect=5), follow_redirects=False, trust_env=False) as client: + response = client.put(inp.uploadUrl, content=data, headers={"Content-Type": mime}) + response.raise_for_status() + except httpx.HTTPError: + return failure("Generated image upload failed", 502, task_id) + parts = urlsplit(inp.uploadUrl) + row["imageURL"] = urlunsplit((parts.scheme, parts.netloc, parts.path, "", "")) + else: + row["imageBase64Data"] = base64.b64encode(data).decode("ascii") + output = {"images": [row], "metrics": public_metrics(metrics)} + if inp.includeCost: + output["cost"] = None # No billing estimate is available for this POC. + return {"id": task_id, "status": "COMPLETED", "executionTime": round((time.monotonic() - started) * 1000), "output": output} + except InputError as exc: + return failure(str(exc), 400, task_id) + except Exception: + LOGGER.exception("Generation request failed") + return failure("Generation failed; inspect local experiment logs", 500, task_id) + finally: + application.state.lock.release() + + @application.post("/runsync") + async def runsync(request: Request): + body = bytearray() + async for chunk in request.stream(): + if len(body) + len(chunk) > MAX_BODY_BYTES: + return failure("Request exceeds 60 MiB", 413) + body.extend(chunk) + try: + envelope = Envelope.model_validate_json(bytes(body)) + except ValidationError as exc: + # No input values in diagnostics: references and signed URLs can be large/private. + issues = [".".join(map(str, e["loc"])) + ": " + e["msg"] for e in exc.errors(include_input=False)] + return failure("Invalid request: " + "; ".join(issues), 422) + return await run_in_threadpool(execute, envelope.input) + + return application + + +app = create_app() diff --git a/scheduler/scheduler_config.json b/scheduler/scheduler_config.json new file mode 100644 index 0000000000000000000000000000000000000000..4802d62b6d80715af98fa13605ab6428d4764050 --- /dev/null +++ b/scheduler/scheduler_config.json @@ -0,0 +1,18 @@ +{ + "_class_name": "FlowMatchEulerDiscreteScheduler", + "_diffusers_version": "0.41.0.dev0", + "base_image_seq_len": 256, + "base_shift": 0.5, + "invert_sigmas": false, + "max_image_seq_len": 8192, + "max_shift": 0.9, + "num_train_timesteps": 1000, + "shift": 1.0, + "shift_terminal": 0.02, + "stochastic_sampling": false, + "time_shift_type": "exponential", + "use_beta_sigmas": false, + "use_dynamic_shifting": true, + "use_exponential_sigmas": false, + "use_karras_sigmas": false +} diff --git a/text_encoder/config.json b/text_encoder/config.json new file mode 100644 index 0000000000000000000000000000000000000000..4bffb6766b4003e8ddc4e331e0de5327d43867d8 --- /dev/null +++ b/text_encoder/config.json @@ -0,0 +1,85 @@ +{ + "architectures": [ + "Qwen3VLForConditionalGeneration" + ], + "dtype": "bfloat16", + "image_token_id": 151655, + "model_type": "qwen3_vl", + "quantization_config": { + "_load_in_4bit": true, + "_load_in_8bit": false, + "bnb_4bit_compute_dtype": "bfloat16", + "bnb_4bit_quant_storage": "uint8", + "bnb_4bit_quant_type": "nf4", + "bnb_4bit_use_double_quant": true, + "llm_int8_enable_fp32_cpu_offload": false, + "llm_int8_has_fp16_weight": false, + "llm_int8_skip_modules": null, + "llm_int8_threshold": 6.0, + "load_in_4bit": true, + "load_in_8bit": false, + "quant_method": "bitsandbytes" + }, + "text_config": { + "attention_bias": false, + "attention_dropout": 0.0, + "bos_token_id": 151643, + "dtype": "bfloat16", + "eos_token_id": 151645, + "head_dim": 128, + "hidden_act": "silu", + "hidden_size": 4096, + "initializer_range": 0.02, + "intermediate_size": 12288, + "max_position_embeddings": 262144, + "model_type": "qwen3_vl_text", + "num_attention_heads": 32, + "num_hidden_layers": 36, + "num_key_value_heads": 8, + "pad_token_id": null, + "rms_norm_eps": 1e-06, + "rope_parameters": { + "mrope_interleaved": true, + "mrope_section": [ + 24, + 20, + 20 + ], + "rope_theta": 5000000, + "rope_type": "default" + }, + "use_cache": false, + "vocab_size": 151936 + }, + "tie_word_embeddings": false, + "transformers_version": "5.17.0", + "use_cache": false, + "video_token_id": 151656, + "vision_config": { + "deepstack_visual_indexes": [ + 8, + 16, + 24 + ], + "depth": 27, + "dtype": "bfloat16", + "hidden_act": "gelu_pytorch_tanh", + "hidden_size": 1152, + "in_channels": 3, + "initializer_range": 0.02, + "intermediate_size": 4304, + "model_type": "qwen3_vl_vision", + "num_heads": 16, + "num_position_embeddings": 2304, + "out_hidden_size": 4096, + "patch_size": 16, + "rope_parameters": { + "rope_theta": 10000.0, + "rope_type": "axial" + }, + "spatial_merge_size": 2, + "temporal_patch_size": 2 + }, + "vision_end_token_id": 151653, + "vision_start_token_id": 151652 +} diff --git a/text_encoder/generation_config.json b/text_encoder/generation_config.json new file mode 100644 index 0000000000000000000000000000000000000000..c608c0ebfd3577d4d0c6968a81ad5c3b9743b85d --- /dev/null +++ b/text_encoder/generation_config.json @@ -0,0 +1,13 @@ +{ + "bos_token_id": 151643, + "do_sample": true, + "eos_token_id": [ + 151645, + 151643 + ], + "pad_token_id": 151643, + "temperature": 0.7, + "top_k": 20, + "top_p": 0.8, + "transformers_version": "5.17.0" +} diff --git a/text_encoder/model.safetensors b/text_encoder/model.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..a973e3ce8f6009d43c354c37db991a17f62f000b --- /dev/null +++ b/text_encoder/model.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:61c75987e552c49b9bb804b6665b70e6df10d808397d7dcfe2e6ce94e7292210 +size 6378433686 diff --git a/transformer/config.json b/transformer/config.json new file mode 100644 index 0000000000000000000000000000000000000000..984c717ead60a0c9d3366077aca86f5f85c3e373 --- /dev/null +++ b/transformer/config.json @@ -0,0 +1,19 @@ +{ + "_class_name": "QwenImage21Transformer2DModel", + "_diffusers_version": "0.37.0.dev0", + "attention_head_dim": 128, + "axes_dims_rope": [ + 16, + 56, + 56 + ], + "causal_condition": true, + "context_in_dim": 4096, + "eps": 1e-06, + "in_channels": 64, + "mlp_ratio": 3, + "num_attention_heads": 32, + "num_layers": 32, + "out_channels": 64, + "patch_size": 1 +} diff --git a/transformer/manifest.json b/transformer/manifest.json new file mode 100644 index 0000000000000000000000000000000000000000..e2b82b6f12f9fae0544d99331e40d74c6f2ef57d --- /dev/null +++ b/transformer/manifest.json @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:cf69e83a646981d22ee6df8d5239b46a50df25d8eb73c9f0478feae87323e6cb +size 54158724 diff --git a/transformer/model-boundary.safetensors b/transformer/model-boundary.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..52069ae4f11eda873e3143cf7540555194bc9e7f --- /dev/null +++ b/transformer/model-boundary.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:91cedb20bfd9f986b3dfd666508f7ef59bf2f4eba98784e15071715b9082aba6 +size 271613760 diff --git a/transformer/model-layer-000.safetensors b/transformer/model-layer-000.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..fe52e6f980034205930477042c3e2b210ef0ec43 --- /dev/null +++ b/transformer/model-layer-000.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:be954118c35254d2512ab5a469765dcc834b9cd9ce23c7f0206de1a63f0d04e5 +size 11027112 diff --git a/transformer/model-layer-001.safetensors b/transformer/model-layer-001.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..081b469fd6b2409fad1357c19c740c6bbd301564 --- /dev/null +++ b/transformer/model-layer-001.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:805910c5b758ae10aa3bd41bd35a0226a0a7ffda8dd5249ef5dc8642b17533f3 +size 11027136 diff --git a/transformer/model-layer-002.safetensors b/transformer/model-layer-002.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..c205097417e823cea242b4ae6294d35c2450e856 --- /dev/null +++ b/transformer/model-layer-002.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:dee0003b353df79eb67c4fe949e484999b194158c4cbbd144e8d4f7898e57569 +size 11027112 diff --git a/transformer/model-layer-003.safetensors b/transformer/model-layer-003.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..c0f1ebdb363a4d0c1d40e1e7f59a620096b6f9d5 --- /dev/null +++ b/transformer/model-layer-003.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:3b38108461d9afd2248223c28f208b175822cfdc03d6cac27bd9be8b58ce6b4a +size 11027112 diff --git a/transformer/model-layer-004.safetensors b/transformer/model-layer-004.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..a1b09f49290ef4fbcd758306a71715043282b17b --- /dev/null +++ b/transformer/model-layer-004.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:bf2dedbba5139a3ef770ed583bc6b5391598448fdf81e977f62210830e470493 +size 30950112 diff --git a/transformer/model-layer-005.safetensors b/transformer/model-layer-005.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..a06718ae65f6d4de713ae5e53eefcb70e62227ce --- /dev/null +++ b/transformer/model-layer-005.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:0d8b6b13096092574df713cfa59333c80d275a8eb983d041cb45b7ef510999c0 +size 30982840 diff --git a/transformer/model-layer-006.safetensors b/transformer/model-layer-006.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..8fbdfbcbbd7d5f0dc4624ccad3a0b55fd31e8692 --- /dev/null +++ b/transformer/model-layer-006.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:43e266de81565062f7e07e8b283ae1678deb0bf3124643e8a1746428a3816098 +size 30950072 diff --git a/transformer/model-layer-007.safetensors b/transformer/model-layer-007.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..ebacdccdbe4c829b59170bce8670936198e82ede --- /dev/null +++ b/transformer/model-layer-007.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:a385b1e5c2bf279906832f96fd028b35939ed4d5e8468ef8a7403695eafb010b +size 11027112 diff --git a/transformer/model-layer-008.safetensors b/transformer/model-layer-008.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..088536e5633a3dd676f07093ba1725d7ad48707c --- /dev/null +++ b/transformer/model-layer-008.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:438e4c30185fc9f196777196eeafeeca6fbcd0f738b9bbec1d7fc37dd1c51e21 +size 11027136 diff --git a/transformer/model-layer-009.safetensors b/transformer/model-layer-009.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..c538dba119595c99a148249c4968d0a4a4a9ee05 --- /dev/null +++ b/transformer/model-layer-009.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:f83cbe6924cec25cccd785e226af722ecc1e68a800e57bdfcedbac3d730cb651 +size 11027112 diff --git a/transformer/model-layer-010.safetensors b/transformer/model-layer-010.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..7ec1478358dd7eef835adac11422ecadb425921f --- /dev/null +++ b/transformer/model-layer-010.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:db6a92ce00cda731450c7b79d5d39fd7df6ce3ffdda6bf58297fd47431fc3bf7 +size 11027112 diff --git a/transformer/model-layer-011.safetensors b/transformer/model-layer-011.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..05fdf87796e8c2d2e35740f0334d449fbafe354a --- /dev/null +++ b/transformer/model-layer-011.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:4dc2ec89217357598e6829b46fc4f60c85c46567efbd0f0696cbef036ad11091 +size 30950112 diff --git a/transformer/model-layer-012.safetensors b/transformer/model-layer-012.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..7577ecdc2d50e1f5eda1521d4eebae3600bda4e1 --- /dev/null +++ b/transformer/model-layer-012.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:54d032380b5660ec2c2b9b1f73ef125fdd05f234164597eb2a6fcc13923945db +size 30982840 diff --git a/transformer/model-layer-013.safetensors b/transformer/model-layer-013.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..bd1e858965e69e519b671732537f4e61b0b1c63a --- /dev/null +++ b/transformer/model-layer-013.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:124ae09e4b97a1667808651d8f445a84b7d8da42bbd84f4d7379b41c6c9d86d2 +size 30950072 diff --git a/transformer/model-layer-014.safetensors b/transformer/model-layer-014.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..7225a8acafde224f767810027124169987d21689 --- /dev/null +++ b/transformer/model-layer-014.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:a166d614f6a6941fb7a7508d2a611610abc7bf7f6d1527f52f9c7d7db435f688 +size 11027112 diff --git a/transformer/model-layer-015.safetensors b/transformer/model-layer-015.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..193d82733bf477a60a66112c94b46be86459cee7 --- /dev/null +++ b/transformer/model-layer-015.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:4cf90f6c54659fbb2f259b1454da2b9d457f4b12bba9834568fd43bdf8e6acdb +size 11027136 diff --git a/transformer/model-layer-016.safetensors b/transformer/model-layer-016.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..983d7519627fd019ab2c342ffd12f5bb23f4dc77 --- /dev/null +++ b/transformer/model-layer-016.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:4fe7a02c345da061343ac51feea997dfbbee408588d1894ce52ca45804ee7554 +size 11027112 diff --git a/transformer/model-layer-017.safetensors b/transformer/model-layer-017.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..b044f0a1853a539ad4183266d859a0f21faee33f --- /dev/null +++ b/transformer/model-layer-017.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:403f07d5311ae7da8ed97ace3ece9d3d56284d0a764cee35d9419df1ce79a515 +size 11027112 diff --git a/transformer/model-layer-018.safetensors b/transformer/model-layer-018.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..80db23debb17bef83ed0ded1c336fa08e3347608 --- /dev/null +++ b/transformer/model-layer-018.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:fa11beb0cfdb4eaa571a6968872be1f75e25d06e00efef6e44179fc72e04a021 +size 30950112 diff --git a/transformer/model-layer-019.safetensors b/transformer/model-layer-019.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..8d9f46ad6c13ab43ccc3848c248e3f805ee58025 --- /dev/null +++ b/transformer/model-layer-019.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:4adcdb2c60bda8d4ec3c96c715cad036dd1c43da11d0ab225dfe72d09ca0dacb +size 30982840 diff --git a/transformer/model-layer-020.safetensors b/transformer/model-layer-020.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..f3a2563eb5cf5c4640ec390f9b672f94d8c7175a --- /dev/null +++ b/transformer/model-layer-020.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:95593972d08bd06bba59cc6a7ec9ee2a61a78d6b53ef5e0d01de34719b18cb9e +size 30950072 diff --git a/transformer/model-layer-021.safetensors b/transformer/model-layer-021.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..5dbce75dad82d8381b78ceb71128e22298197a17 --- /dev/null +++ b/transformer/model-layer-021.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:e68b95ecfc1df57b3b91dc813d57740acbf40302fa5e5f6989634ea44bd80d26 +size 11027112 diff --git a/transformer/model-layer-022.safetensors b/transformer/model-layer-022.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..a9881525cb3dd5f9ac58708182063cedbbc313c2 --- /dev/null +++ b/transformer/model-layer-022.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:6c13a089c4784a26862cfaa59312f09fc71f95b72c4d26b8065ee54b6bc7c421 +size 11027136 diff --git a/transformer/model-layer-023.safetensors b/transformer/model-layer-023.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..03669037c17b7d350fbdb0ff7f0dde9a3cd7d229 --- /dev/null +++ b/transformer/model-layer-023.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:121526e8517c376058d032eecbc2445df4e26faa26fa04361473d83fb94ec960 +size 11027112 diff --git a/transformer/model-layer-024.safetensors b/transformer/model-layer-024.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..8204a10665866a5f828bffd944a749b4414a8257 --- /dev/null +++ b/transformer/model-layer-024.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:238667fe2e13404a8adaf981a3283f2593ae859096d26bce454bf574d0a61434 +size 11027112 diff --git a/transformer/model-layer-025.safetensors b/transformer/model-layer-025.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..2867d0118cc9f2865e696de7ec15bd4c2be7026e --- /dev/null +++ b/transformer/model-layer-025.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:e3f3f259fb31dc98060e1de725cad827ca1b0eae0bf5bfa3ab4bc50ff0d11f55 +size 30950112 diff --git a/transformer/model-layer-026.safetensors b/transformer/model-layer-026.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..8c8ac68c81d2a88c13274ef75bc5e2efd53a3c1f --- /dev/null +++ b/transformer/model-layer-026.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:896efca07b1bcd0272af0cd3d580c64762591e738c9b9bff2eec81feda9b6cf2 +size 30982840 diff --git a/transformer/model-layer-027.safetensors b/transformer/model-layer-027.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..bc327687f11c8a84eefdfdc70de6794abfe2f300 --- /dev/null +++ b/transformer/model-layer-027.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:1feff8f4d8d07777e3c914c0152443d643257e9155a364829fe14fa52eaca563 +size 30950072 diff --git a/transformer/model-layer-028.safetensors b/transformer/model-layer-028.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..234b516785c2cc3a7e4f556d4ca2729b616c55fb --- /dev/null +++ b/transformer/model-layer-028.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:562ed669c40ca3bd447109beb66df73ad21c00aa82b7b91d50c5d34475fc88c8 +size 11027112 diff --git a/transformer/model-layer-029.safetensors b/transformer/model-layer-029.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..b061c387a1ddddea160ee2f0f0767d39f753fdf7 --- /dev/null +++ b/transformer/model-layer-029.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:eb27895adb5ff1402342f530c9244b02344edcaf5cdb7571b55170dc87aa8b66 +size 11027136 diff --git a/transformer/model-layer-030.safetensors b/transformer/model-layer-030.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..fa25338d538f550bdbddc6026d1b14a1f788ea1f --- /dev/null +++ b/transformer/model-layer-030.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:7ead55cbc062e1a32f48ca2e5bc17f311f0a80e6125e6bad1d4bd30c762434f3 +size 11027112 diff --git a/transformer/model-layer-031.safetensors b/transformer/model-layer-031.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..25f8684733b64aefd205e80475f759447d095787 --- /dev/null +++ b/transformer/model-layer-031.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:b01bbd72216709fb7888e201f1b25a70505740e34ab23d399c17465e324c75e9 +size 11027112 diff --git a/transformer/model-layer-032.safetensors b/transformer/model-layer-032.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..3d84374d76c26099bf5636714b8cd044abe9ee4f --- /dev/null +++ b/transformer/model-layer-032.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:49db7476493465a03fce9fc9085e27e596440246b7a0d182074e569114d52f06 +size 30950112 diff --git a/transformer/model-layer-033.safetensors b/transformer/model-layer-033.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..9e6f0b6e521d4af8ffeb3718923f14178a01bf2a --- /dev/null +++ b/transformer/model-layer-033.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:c67faa03d75949ce4a209210de02cdf58192877d445a0264b0030c55f7c374b9 +size 30982840 diff --git a/transformer/model-layer-034.safetensors b/transformer/model-layer-034.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..975d4453e281321ef0d5f9956d853b09922c0009 --- /dev/null +++ b/transformer/model-layer-034.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:38b273218108066350b01c2006a6ab667092868dbbc144b390dddf8550be2988 +size 30950072 diff --git a/transformer/model-layer-035.safetensors b/transformer/model-layer-035.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..f1ac7463a54af39271e6f21ee3353c3b78f7e36d --- /dev/null +++ b/transformer/model-layer-035.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:046ac87ebda321a370208bbc879d62f62d439463b297e05a11bbec12990edb7a +size 11027112 diff --git a/transformer/model-layer-036.safetensors b/transformer/model-layer-036.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..82038ffdf3adc96d6990339e6085446820c10c00 --- /dev/null +++ b/transformer/model-layer-036.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:e1e8afb5ccbb88266124e88a233e7545e27591e9e0ac07c8ca9fb45783569eaa +size 11027136 diff --git a/transformer/model-layer-037.safetensors b/transformer/model-layer-037.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..c8664e54094fa3f15ddc92296c13d54a476e3e4a --- /dev/null +++ b/transformer/model-layer-037.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:d3628d58f6e3394ebf0e790c8dd44685b76ee027154c6d5fff863308b5758ab6 +size 11027112 diff --git a/transformer/model-layer-038.safetensors b/transformer/model-layer-038.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..05231fd5c31575e5b4b7aba2120173471313a953 --- /dev/null +++ b/transformer/model-layer-038.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:e73d297573fb3fb87d909441a81d82ed56eed0facd8a38d42836ed8cc61fbb84 +size 11027112 diff --git a/transformer/model-layer-039.safetensors b/transformer/model-layer-039.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..9b5e5a47f3a28ab2c19f6ed8bf0463f5ec99e2a5 --- /dev/null +++ b/transformer/model-layer-039.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:6604928bf9b57f4ec5189fd534249e064758d2c3c319db2231b608cabd968a4e +size 30950112 diff --git a/transformer/model-layer-040.safetensors b/transformer/model-layer-040.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..0efec13e0d7b2ae1c9ed8ef343a754acf918d968 --- /dev/null +++ b/transformer/model-layer-040.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:2c07c7123bd68cb52e5f69abdee4e2d4cc945dc8c8436662b6ff004422f821dd +size 30982840 diff --git a/transformer/model-layer-041.safetensors b/transformer/model-layer-041.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..1d0ae6afdf4a0698b59b9e73b8bb1217152b08c2 --- /dev/null +++ b/transformer/model-layer-041.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:4177a378ff702ee8e079a980bd47cc759a8f8e25a57b63d6746110ee9ed5e257 +size 30950072 diff --git a/transformer/model-layer-042.safetensors b/transformer/model-layer-042.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..aad645728bb2bc6d35cf7ff7bbc50f039f96c238 --- /dev/null +++ b/transformer/model-layer-042.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:a200b9c9383eb3c969ea124941df9fb157821c327e4a2991a1073ca0b274943b +size 11027112 diff --git a/transformer/model-layer-043.safetensors b/transformer/model-layer-043.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..692718ef00dfcc87677033ac3bd5b4ab4a574274 --- /dev/null +++ b/transformer/model-layer-043.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:681f1584a3b9200fdea7ca0f55fc263ca01f1d8f5d80326fc8a00b09f73bbd9d +size 11027136 diff --git a/transformer/model-layer-044.safetensors b/transformer/model-layer-044.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..f8db1f3dc29fa9d99ba97db320dcf9e6280de2bd --- /dev/null +++ b/transformer/model-layer-044.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:241e6e0ba34c77390c9173c5b512d4544b74c27c3f71e1834d14d900af1f4cd2 +size 11027112 diff --git a/transformer/model-layer-045.safetensors b/transformer/model-layer-045.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..714c45a7db48797f91931bc665b23d934761722f --- /dev/null +++ b/transformer/model-layer-045.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:5276d95df07f1e59f18388a1f2d210e1c5cfe95f4136df5ea6aa27d2ce112a3f +size 11027112 diff --git a/transformer/model-layer-046.safetensors b/transformer/model-layer-046.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..e82dcd539417b85cff64b0d6bad098d92b26288a --- /dev/null +++ b/transformer/model-layer-046.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:18b52c52508cfa67e73a8c74db48dee8d9b2453a3234557d97e1f3b2ca1dfb9b +size 30950112 diff --git a/transformer/model-layer-047.safetensors b/transformer/model-layer-047.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..88f802f72e47f494037c8c8071f9c209fea09bcc --- /dev/null +++ b/transformer/model-layer-047.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:543bc08d4fc332aa6a290dc026b62ebfc22d05e181f16346ad36003ae1e8566e +size 30982840 diff --git a/transformer/model-layer-048.safetensors b/transformer/model-layer-048.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..5d635745d1c2024e50eba9ccb34ff1b8f29c3f18 --- /dev/null +++ b/transformer/model-layer-048.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:8717d187ac5dc06762b8c578b791d932221b2adffc85d499d4b8e3a9103af03c +size 30950072 diff --git a/transformer/model-layer-049.safetensors b/transformer/model-layer-049.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..9a1d606bbe64a80a16aec8d6fe1b3f9ffd37946d --- /dev/null +++ b/transformer/model-layer-049.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:56962839346c87c9dc42328b8af42f56a4d678c7bfbea57d36bc7b3dfc7ec9b1 +size 11027112 diff --git a/transformer/model-layer-050.safetensors b/transformer/model-layer-050.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..0f1bb555a387eb9bcf0e55675a607067d4a58b31 --- /dev/null +++ b/transformer/model-layer-050.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:ed43ccd63f133a9a191ac6e2037b43f4effeee1e44887f377ddc6459c9d1ef1f +size 11027136 diff --git a/transformer/model-layer-051.safetensors b/transformer/model-layer-051.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..59c79dc2b39ae95d0c55dd7eef4623de62093b43 --- /dev/null +++ b/transformer/model-layer-051.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:e3597b0ff8f3a2b798d0873873b9627f0c7e0ef5ceeac560251fdc66fe8f6195 +size 11027112 diff --git a/transformer/model-layer-052.safetensors b/transformer/model-layer-052.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..4641fb1b8ebe2dda5d7e61f449a78e0c4619f4e1 --- /dev/null +++ b/transformer/model-layer-052.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:fce9a7a76138781ec49ba8054f7d2f8b770aa7b9bcf78787d3db7d8cde159ff4 +size 11027112 diff --git a/transformer/model-layer-053.safetensors b/transformer/model-layer-053.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..07e42b3bd93d8971718609eef5371b60c9661d59 --- /dev/null +++ b/transformer/model-layer-053.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:22b532bffe7dd1ded582cb37b970fa483fe6eb197acdaa8778d6b700a740393a +size 30950112 diff --git a/transformer/model-layer-054.safetensors b/transformer/model-layer-054.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..15843140f948b5ff939b774f2fa004d8fb7f7fea --- /dev/null +++ b/transformer/model-layer-054.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:ae24df341390707a611a12cd08763386397a6c271e7b1e0dbf062f5116e56af5 +size 30982840 diff --git a/transformer/model-layer-055.safetensors b/transformer/model-layer-055.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..3041e631ffbd80881179736c2216c8e066a81e33 --- /dev/null +++ b/transformer/model-layer-055.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:089382d22679fe93818b79bd53a52b413c3984d366cdd1331bde0d4f959d4782 +size 30950072 diff --git a/transformer/model-layer-056.safetensors b/transformer/model-layer-056.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..431cf100321185548a25072f68e876f6a767338b --- /dev/null +++ b/transformer/model-layer-056.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:12a9d61e10e9623df25d571e42a1b1006aaf0e3eec573e89f53cf6c21a4bf120 +size 11027112 diff --git a/transformer/model-layer-057.safetensors b/transformer/model-layer-057.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..d1adc65265c1c7fe73069efd3814c72bd20c9b89 --- /dev/null +++ b/transformer/model-layer-057.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:6ca13e70f2ed1616e2c9c5604aff6daeebb226d21b04effeeb9b1de1de1ea99f +size 11027136 diff --git a/transformer/model-layer-058.safetensors b/transformer/model-layer-058.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..b381049c056893d58588836e44f456238014a626 --- /dev/null +++ b/transformer/model-layer-058.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:d01cfd242427e7bbe69d73245ec95e1d22e43f62bf45f53595019a703135968d +size 11027112 diff --git a/transformer/model-layer-059.safetensors b/transformer/model-layer-059.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..0606f5cb49f945de49d434d647881d5afc74fbcb --- /dev/null +++ b/transformer/model-layer-059.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:378f3c3fcdb57101f5760e988aff466923f5bf0332584860aedf24c3352b8617 +size 11027112 diff --git a/transformer/model-layer-060.safetensors b/transformer/model-layer-060.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..c9e90f8c0ba627cfa63f77eee0bb73844cd6fec4 --- /dev/null +++ b/transformer/model-layer-060.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:662886be033ee5f2ee3541ed7187d3b752dd8176c4688dbb72890d5aaa891532 +size 30950112 diff --git a/transformer/model-layer-061.safetensors b/transformer/model-layer-061.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..b14dcd88519283f14d717ba1802e93c8cec9fabf --- /dev/null +++ b/transformer/model-layer-061.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:64375e1fe523f359aca412649e1c30ab1496269f0737e230fade32adcfbf6347 +size 30982840 diff --git a/transformer/model-layer-062.safetensors b/transformer/model-layer-062.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..cc7b619c045a28e3fb0fe4c33f9ca39eb051da9e --- /dev/null +++ b/transformer/model-layer-062.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:37658199e6de2fb0e049ae0fe30cb16407db44922274cbc911245cec8fa3332b +size 30950072 diff --git a/transformer/model-layer-063.safetensors b/transformer/model-layer-063.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..efa049cd13894996b327f28d1225e1cf70be4c6a --- /dev/null +++ b/transformer/model-layer-063.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:702e1a78ebfd91fe2b6efbef78fefd1257d0509320636e21e557bd86f5e2153c +size 11027112 diff --git a/transformer/model-layer-064.safetensors b/transformer/model-layer-064.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..2fb0acbb9efa2b923a01761caa59f3024d14b17d --- /dev/null +++ b/transformer/model-layer-064.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:1c3f65152c3afba62069cda1c22df5798184a29c869de87d835e311798906530 +size 11027136 diff --git a/transformer/model-layer-065.safetensors b/transformer/model-layer-065.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..a940af965262f7279eec27b7e8d19b3238535efa --- /dev/null +++ b/transformer/model-layer-065.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:610ba7bd5f30ad75be63f2368b5e76320007ded4d3551873c0366d9602d681ca +size 11027112 diff --git a/transformer/model-layer-066.safetensors b/transformer/model-layer-066.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..586bd90651d470af85e64947722fe40123589f56 --- /dev/null +++ b/transformer/model-layer-066.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:f910131d9dbdc89949e03f087ac979b01fbfd1f43eed7ae1db583107d89de328 +size 11027112 diff --git a/transformer/model-layer-067.safetensors b/transformer/model-layer-067.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..fe154f4fdfd6d93a80d480ebaf6d1ce6e5139ce8 --- /dev/null +++ b/transformer/model-layer-067.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:bdeee3b356428a7e75a19982522273328c20708cc572f12fe442f1dde57a3d15 +size 30950112 diff --git a/transformer/model-layer-068.safetensors b/transformer/model-layer-068.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..e87c9e13111213d384e9eafd4272eb77b8cadccc --- /dev/null +++ b/transformer/model-layer-068.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:3fc30d0a256df124ca9d254f2c677e01c114b729239a589848b3039289de947c +size 30982840 diff --git a/transformer/model-layer-069.safetensors b/transformer/model-layer-069.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..1fb4e2ee2c043551969faf6edcdc99a4af99d735 --- /dev/null +++ b/transformer/model-layer-069.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:9b6ca9886ac5cf675000ad407d008857f46c0b50617f926d53af1a028a00e278 +size 30950072 diff --git a/transformer/model-layer-070.safetensors b/transformer/model-layer-070.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..8214c4e5dbe3bd353aba150e59a1087581d53109 --- /dev/null +++ b/transformer/model-layer-070.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:40ad6cdbd4a624bba006875b65055bfcb78dbf8a3dd22e719f2a5afafdde36d3 +size 11027112 diff --git a/transformer/model-layer-071.safetensors b/transformer/model-layer-071.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..71f882ec50ef504e93f25f4164ca4cabaa89af2c --- /dev/null +++ b/transformer/model-layer-071.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:333c2026b5f2c6f2746241824a1ef7560fac6a2bb9930c160cca9218b5c0e49d +size 11027136 diff --git a/transformer/model-layer-072.safetensors b/transformer/model-layer-072.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..7d3eb99cb99f1cb4e9884d1bcabb902c017fb231 --- /dev/null +++ b/transformer/model-layer-072.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:852177773dcff10bae0917265572f7658cd6cb06512a197b6cbba770f269c703 +size 11027112 diff --git a/transformer/model-layer-073.safetensors b/transformer/model-layer-073.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..5d1d3b227671c6a192f81772331041a43349ace1 --- /dev/null +++ b/transformer/model-layer-073.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:7ec3d0b4a46c5683f3dedd572298dab524cf3f6c22d2db4483786f2f4f1b7f4c +size 11027112 diff --git a/transformer/model-layer-074.safetensors b/transformer/model-layer-074.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..35c9c26d405cf30d65f5e6c9fcd15a9e0905c998 --- /dev/null +++ b/transformer/model-layer-074.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:e042abf43b93871fe6b0a0b3a6374d1d2435308643df8fe720d7e9899f25b8f5 +size 30950112 diff --git a/transformer/model-layer-075.safetensors b/transformer/model-layer-075.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..d32ebb76992f9ef8b9c159e27a19128c5e6c03db --- /dev/null +++ b/transformer/model-layer-075.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:ad29ca50a411ef8070919e4217520d1de0f41ca4666c6a2a3b475595a9e0e9ae +size 30982840 diff --git a/transformer/model-layer-076.safetensors b/transformer/model-layer-076.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..a6565b3b26fad4b9cde596a626f446e0b002217d --- /dev/null +++ b/transformer/model-layer-076.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:04471891428d103172959be5afc598594537ce52929fb8ef4e007cd869726a81 +size 30950080 diff --git a/transformer/model-layer-077.safetensors b/transformer/model-layer-077.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..7574529d480fd930363918a53a5983d5dd7da37b --- /dev/null +++ b/transformer/model-layer-077.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:2b94f5d716a82f220d3f846b70402346610671e48617084843ec161e558155ea +size 11027112 diff --git a/transformer/model-layer-078.safetensors b/transformer/model-layer-078.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..ae0f908f46559dcccb9ad9070aea402eee86ff42 --- /dev/null +++ b/transformer/model-layer-078.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:8f686e842a4abea031616d9be089dc679f2b953b2b3bf575a1e109df7eeedbaf +size 11027136 diff --git a/transformer/model-layer-079.safetensors b/transformer/model-layer-079.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..67f6ce11a59f934c724fe16bcb7907b3dbd74ee7 --- /dev/null +++ b/transformer/model-layer-079.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:f080bed05b77c27fa76a2e4f5b7394f59a3a7732c853e358788f95d564ab62f3 +size 11027112 diff --git a/transformer/model-layer-080.safetensors b/transformer/model-layer-080.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..5d618538d7572a00f84d4389e01450beba4f960b --- /dev/null +++ b/transformer/model-layer-080.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:60cdb41cdfecd42bcc2313e0c30c03109deaf88349472234a1f065ebe6165c96 +size 11027112 diff --git a/transformer/model-layer-081.safetensors b/transformer/model-layer-081.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..bdf3d1d957362213b9a89284281ab828f861af08 --- /dev/null +++ b/transformer/model-layer-081.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:c2fff5071e36336544a28ea5bab59f34a739b1887b201b1c0c77e4af4112d9ab +size 30950112 diff --git a/transformer/model-layer-082.safetensors b/transformer/model-layer-082.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..6de64bc5ca97e42d386c93ff378dc814aca7f9ac --- /dev/null +++ b/transformer/model-layer-082.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:622a1e8c1e6ebb3acff58b676d822cd3ab7cf849da5d692090f6dbf18e123fbd +size 30982840 diff --git a/transformer/model-layer-083.safetensors b/transformer/model-layer-083.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..aa297fca8d1b3e2ad8e6d51174504cae2e2201ca --- /dev/null +++ b/transformer/model-layer-083.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:b213327da54a43dd807d91785c035be6b5f1ff651b0e523e8a7d4215231517c8 +size 30950080 diff --git a/transformer/model-layer-084.safetensors b/transformer/model-layer-084.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..fb422e171db70f2d419b54b86743bd06fe521aa7 --- /dev/null +++ b/transformer/model-layer-084.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:ca594ed744f93b6c247582ce8b238f3f7741f92b11b8de175cd2f898c3fd1201 +size 11027112 diff --git a/transformer/model-layer-085.safetensors b/transformer/model-layer-085.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..3c9dd26017fe69e3170eee8e194ecd517fabc411 --- /dev/null +++ b/transformer/model-layer-085.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:6178d8df840e6cd804a85a93400ab524a723aa1779109278ed40c1915bf58176 +size 11027136 diff --git a/transformer/model-layer-086.safetensors b/transformer/model-layer-086.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..a2f93b2f429892bded7f8ea7eb31d068a20a22a1 --- /dev/null +++ b/transformer/model-layer-086.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:cc0fbfa403743d2b45586649e5842c25993c46e69e41842a4c0e95494791b4d5 +size 11027112 diff --git a/transformer/model-layer-087.safetensors b/transformer/model-layer-087.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..ee713f757d379b1780ce10fe5bee667b7c96a47c --- /dev/null +++ b/transformer/model-layer-087.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:5aaea51447d8f892bed089a2ed4bd30cffe8a017430423493a2e9fdeaa806fd8 +size 11027112 diff --git a/transformer/model-layer-088.safetensors b/transformer/model-layer-088.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..cb0d7d44c26c4748fbd4b4255ce93e491a86551e --- /dev/null +++ b/transformer/model-layer-088.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:1fb2952585109eb0e33e7fb280f950850b685207c694603f0e241ba19a0efc2c +size 30950112 diff --git a/transformer/model-layer-089.safetensors b/transformer/model-layer-089.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..c66124f4ef58a473b2d39edad2dddcd035f15faa --- /dev/null +++ b/transformer/model-layer-089.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:3595b8921f6308f9df9a8bc73c68d28b0628ee3346c2b755a9b0e152ad50dbff +size 30982840 diff --git a/transformer/model-layer-090.safetensors b/transformer/model-layer-090.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..7516e064f5557bbf1a038d43d06c19b2fed9d837 --- /dev/null +++ b/transformer/model-layer-090.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:779cba9de87994d34841be62fb9fb3098edfea5f8a32719a192820a95b04b570 +size 30950080 diff --git a/transformer/model-layer-091.safetensors b/transformer/model-layer-091.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..df2409620a0f6595e371346b828d59ae33b0592b --- /dev/null +++ b/transformer/model-layer-091.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:1d8d72051cf8d046e9d3804ad00c0feb887e8fa89adcd3c74190ebdb510b6241 +size 11027112 diff --git a/transformer/model-layer-092.safetensors b/transformer/model-layer-092.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..b865d928be7317ac4a4666506de802f97e7b2246 --- /dev/null +++ b/transformer/model-layer-092.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:c83ff768a390c18f08ff89120787f05661ba2646a88eb128a80cdefad310aa13 +size 11027136 diff --git a/transformer/model-layer-093.safetensors b/transformer/model-layer-093.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..9d1fd920dff7daaf26a4fe8dc99cf97d4f84f662 --- /dev/null +++ b/transformer/model-layer-093.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:b39f0766395452bcd08186a57659bdf8fde0e14302ed6d486913794281ff9cc5 +size 11027112 diff --git a/transformer/model-layer-094.safetensors b/transformer/model-layer-094.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..c428707b3a129be4ed7c94b694a0b8976ee665e5 --- /dev/null +++ b/transformer/model-layer-094.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:7ca843ab6f7c7c503f31370980a6509e092081518df1c8a5a190bafe761b46d9 +size 11027112 diff --git a/transformer/model-layer-095.safetensors b/transformer/model-layer-095.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..4f84a405da7fef629c41d18a3ff2f8c350acbef3 --- /dev/null +++ b/transformer/model-layer-095.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:4780e7617421b97b54d3fbd997a15410009298cd8ddf901a32dc3662fb6c0c45 +size 30950112 diff --git a/transformer/model-layer-096.safetensors b/transformer/model-layer-096.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..a25b1cad770363325c42444800480fc278f7faf4 --- /dev/null +++ b/transformer/model-layer-096.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:ef31e5488181a2e5a8efffaa15ab993ea1fbd698aa282a0aff12e787d702aa58 +size 30982840 diff --git a/transformer/model-layer-097.safetensors b/transformer/model-layer-097.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..bf1d5d328976c0b6114a84e478a63c6f2431f18f --- /dev/null +++ b/transformer/model-layer-097.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:fe65aa51d7ae66bd32c7647a32a9e6a6f8e205247da5ecedee6b820b42b457af +size 30950080 diff --git a/transformer/model-layer-098.safetensors b/transformer/model-layer-098.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..42ebf79d9a270f3d975ad85b8b1d0b14fa32b9ed --- /dev/null +++ b/transformer/model-layer-098.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:d9530418ef182c32baf1ebbf6597d4f395795e9597cc0ae3480f5a7bc3ae589f +size 11027112 diff --git a/transformer/model-layer-099.safetensors b/transformer/model-layer-099.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..76ac32c5992752ee38f7d98fccc09cd9296a352f --- /dev/null +++ b/transformer/model-layer-099.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:f061a9bd99376abf7896896f5e3fa39cc75b5dfee8dd3112dc30fd76cd6c99df +size 11027136 diff --git a/transformer/model-layer-100.safetensors b/transformer/model-layer-100.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..e35734b30dffc06895a49297db042999fcda328d --- /dev/null +++ b/transformer/model-layer-100.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:ef0d94bf7d96ddcf977127d461593fd1b1fe92287ebec8b1c039b72337ffbc49 +size 11027112 diff --git a/transformer/model-layer-101.safetensors b/transformer/model-layer-101.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..05e865d7017cb2473efef186b62c3a368614d240 --- /dev/null +++ b/transformer/model-layer-101.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:eb9e524affd0d6fff922656c540c695f09d5d0ef0dc3bda142c8155b2e03a33c +size 11027112 diff --git a/transformer/model-layer-102.safetensors b/transformer/model-layer-102.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..221e242a28670532a5cddf0ff86dca2f328ce4fd --- /dev/null +++ b/transformer/model-layer-102.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:17d356995ea24b6b3f78e421ef5dae1b3e4614f5688c761e0ac44782f97b3587 +size 30950112 diff --git a/transformer/model-layer-103.safetensors b/transformer/model-layer-103.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..c83e377e2120d3a2b42c6a95624a2290c168b236 --- /dev/null +++ b/transformer/model-layer-103.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:e943711c2e0f0c182336c85bbf3cea24a5c0e736498b54dd87e8f3b56fa7850f +size 30982840 diff --git a/transformer/model-layer-104.safetensors b/transformer/model-layer-104.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..124178cd028a247e3d29044f1b483351d916858c --- /dev/null +++ b/transformer/model-layer-104.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:7b89be9f61d558b436aa50cb76791fa27acee5ec2b4d5b39b9f81b52e40d2ba9 +size 30950080 diff --git a/transformer/model-layer-105.safetensors b/transformer/model-layer-105.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..e4d052eea7d681f2ad9f7ca0dcb7e7f5ed5a9033 --- /dev/null +++ b/transformer/model-layer-105.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:997415569ddeae044e68348e149237ce72997afe5807eeed9f153b8cb90ea9f6 +size 11027112 diff --git a/transformer/model-layer-106.safetensors b/transformer/model-layer-106.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..0fd9370464d86edfb6b9bde83743bfcbfb1b6917 --- /dev/null +++ b/transformer/model-layer-106.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:cb839042bf02639fb0f8d7f30245a644ab718865f8219b069eb11ef52a4b3700 +size 11027136 diff --git a/transformer/model-layer-107.safetensors b/transformer/model-layer-107.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..326163661e20034b4e5fa86340563fd6d2e97988 --- /dev/null +++ b/transformer/model-layer-107.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:37a4279e3632da1979e7e228bb56fa251f2af33b32423685139680ae37e3ed79 +size 11027112 diff --git a/transformer/model-layer-108.safetensors b/transformer/model-layer-108.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..cbc61e7ba92d4be803a20bd8a0a6895cad17b5f2 --- /dev/null +++ b/transformer/model-layer-108.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:a72a4ca1452fde96faeb4913cabeb3990afd41fc7a59881fadad9b449e707869 +size 11027112 diff --git a/transformer/model-layer-109.safetensors b/transformer/model-layer-109.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..7c6697a80d42b88ba69e9e948c32c58a481e622a --- /dev/null +++ b/transformer/model-layer-109.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:cdb26f30293cf565b39dcc48432cb34a1127680ddcd98d5a7843d4529026e583 +size 30950112 diff --git a/transformer/model-layer-110.safetensors b/transformer/model-layer-110.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..d831e884e9fa2ea10b505ef67937a61565aa1142 --- /dev/null +++ b/transformer/model-layer-110.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:1640282696d83ecd960d9e53e494a2a0c258c572c0baf40a5d7293e171bf9d70 +size 30982840 diff --git a/transformer/model-layer-111.safetensors b/transformer/model-layer-111.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..122775a61604cf914acde2fdd57575ca7bbeb2f2 --- /dev/null +++ b/transformer/model-layer-111.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:c9fdb01cbb69e977e15d19fd5c676a978bd8de8e77e4d87b128fe4e112f1443b +size 30950080 diff --git a/transformer/model-layer-112.safetensors b/transformer/model-layer-112.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..5246ac595562785f88e6df4da7e9965880293621 --- /dev/null +++ b/transformer/model-layer-112.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:bc8406984a60b35a95cc3d3df9fa4043598fd4c4a7369ba46701c1c82e295329 +size 11027112 diff --git a/transformer/model-layer-113.safetensors b/transformer/model-layer-113.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..480f9fcdae7bf6293c385d67ebc06b0305886f50 --- /dev/null +++ b/transformer/model-layer-113.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:153830f28286160f9c24d532bc060ff5a88f66377ee5599a88fa0b75d18206b8 +size 11027136 diff --git a/transformer/model-layer-114.safetensors b/transformer/model-layer-114.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..0f4dca0c3b116de8543c78012b45eaba83113f47 --- /dev/null +++ b/transformer/model-layer-114.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:7c716c2b9efde0599f0d6d5a9c8b3fa372a3ccd35caf1c2f7bae9d59895ce810 +size 11027112 diff --git a/transformer/model-layer-115.safetensors b/transformer/model-layer-115.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..62a3d28113762bd8d5a108bfecb9d0875ce5182b --- /dev/null +++ b/transformer/model-layer-115.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:1dc07c10fe6e9ae35c6fb80e6a1db2074a170a149129924908845e610c949414 +size 11027112 diff --git a/transformer/model-layer-116.safetensors b/transformer/model-layer-116.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..cd63668eca9146afe42433276f5d28bd9995b187 --- /dev/null +++ b/transformer/model-layer-116.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:a8a6a19494147a2b0b54d1b0c1a478ccb7124da5c85589bfa432f2dd2a79d705 +size 30950112 diff --git a/transformer/model-layer-117.safetensors b/transformer/model-layer-117.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..9c4dbddbc0770a5e2d4adaf0013044e484710cd4 --- /dev/null +++ b/transformer/model-layer-117.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:47073f8778699492cd7c3efd5f518cf43d594f42a5e2bb73163eb5562a0df56d +size 30982840 diff --git a/transformer/model-layer-118.safetensors b/transformer/model-layer-118.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..c40a8354f01c4ec4b0764168f3b9ae24c75463a9 --- /dev/null +++ b/transformer/model-layer-118.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:1e8a2eeee518d9b399cc6d462d86950fd82e2dcec2a1782e49b4f800365b10a5 +size 30950080 diff --git a/transformer/model-layer-119.safetensors b/transformer/model-layer-119.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..a9759291e90d0451571aa7daab92e83ece27a0ff --- /dev/null +++ b/transformer/model-layer-119.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:9ab2de6d358176842de9cb233f23eb72c979e3d4b930d6d1d1c810ed163b6b47 +size 11027112 diff --git a/transformer/model-layer-120.safetensors b/transformer/model-layer-120.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..371a87eaa491807bee93c53ab5389d89aeaeb939 --- /dev/null +++ b/transformer/model-layer-120.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:1b70e6dcd33f38c84f407d528e1244e392af653b9b20320b7e1b08426d7126d0 +size 11027136 diff --git a/transformer/model-layer-121.safetensors b/transformer/model-layer-121.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..d5ebacb3ee99eeab763242ef66226c4997e33889 --- /dev/null +++ b/transformer/model-layer-121.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:7c25bcdec9d6343adccc80d1d8736e41513043592df3c7d51b843ab9f4f42a56 +size 11027112 diff --git a/transformer/model-layer-122.safetensors b/transformer/model-layer-122.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..10d37d89537794927d9d8d48418e7540bcab3c05 --- /dev/null +++ b/transformer/model-layer-122.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:09e0aede5e349c62df8cb67a4fe4f384394cc632dabe683c6d60bd011b691b51 +size 11027112 diff --git a/transformer/model-layer-123.safetensors b/transformer/model-layer-123.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..6cece87ebb72ab792206e66610c2b85a705a67f6 --- /dev/null +++ b/transformer/model-layer-123.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:7406a84e2cb4384b5f91f7d096da029c90f05f61c27c433444cd6ab7eb29f0f5 +size 30950112 diff --git a/transformer/model-layer-124.safetensors b/transformer/model-layer-124.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..65430ec9b88cf073068d62891a79e26780420679 --- /dev/null +++ b/transformer/model-layer-124.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:7f554df5ac153938fe5f19c8da9b386e4e7d0c302cc41cb8c00c9bed86356809 +size 30982840 diff --git a/transformer/model-layer-125.safetensors b/transformer/model-layer-125.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..8ac6100ffad117e31af2175909d4fc9bbe69527d --- /dev/null +++ b/transformer/model-layer-125.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:0d379ea34ac4cb43b0395ccadc262c8d4f3c2467d8705358fbc2ee6ef80abc7e +size 30950080 diff --git a/transformer/model-layer-126.safetensors b/transformer/model-layer-126.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..23781be73d7ebbfd1ce8ac8dbe8c88d36594c884 --- /dev/null +++ b/transformer/model-layer-126.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:f6e6bdfd817ab781c7a897d7a1278baf418990d129c5e2aa9fa53da7c6b54445 +size 11027112 diff --git a/transformer/model-layer-127.safetensors b/transformer/model-layer-127.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..c28b557a7b7252115a0a034e8a086fc928147712 --- /dev/null +++ b/transformer/model-layer-127.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:6a053d3ce96af8761049b45ba3e9973260c589241075fb6c7aad97af85261d19 +size 11027136 diff --git a/transformer/model-layer-128.safetensors b/transformer/model-layer-128.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..f3c61f38067e18b2729b1cdc77b8abf07692a5ed --- /dev/null +++ b/transformer/model-layer-128.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:b3ee8d8f7c634f20269b03f2ed4f9b07aa66322b3a1778551ad4cc895eac2f64 +size 11027112 diff --git a/transformer/model-layer-129.safetensors b/transformer/model-layer-129.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..9699bdb1cacb9e22e9dcb5de7e211149837f1a9a --- /dev/null +++ b/transformer/model-layer-129.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:fc5d83671933ce3687969b89fa4d3cbe417d87791f6a8eff115df5b5d5d69507 +size 11027112 diff --git a/transformer/model-layer-130.safetensors b/transformer/model-layer-130.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..d92fc7726cabd9380cc391c74469626db2b5bc55 --- /dev/null +++ b/transformer/model-layer-130.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:bbd42cd7cf806884036cc47f3eb8ce63a39225074069dd937f358df36375f6f7 +size 30950112 diff --git a/transformer/model-layer-131.safetensors b/transformer/model-layer-131.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..bcf5360496c4176e65d0a18d578f63c6ee2e658d --- /dev/null +++ b/transformer/model-layer-131.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:35399d61dbfef25d553d2ad4a9ab9aaddaf09d34c7a8c1fb26dd81d0359a3b32 +size 30982840 diff --git a/transformer/model-layer-132.safetensors b/transformer/model-layer-132.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..66ba513e8ad313a709307279026c6aa4c7ce16f0 --- /dev/null +++ b/transformer/model-layer-132.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:76891b27ccbe864d5dc81287f1efcc62f24cf9bf2cad8a95873d836ab3df1362 +size 30950080 diff --git a/transformer/model-layer-133.safetensors b/transformer/model-layer-133.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..cd14ac1ca7ec06b07f5ac702d8df81fc51d5d4be --- /dev/null +++ b/transformer/model-layer-133.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:f652df6c51d36d2f84ac3fb2c64f31c90854b61fa8bf9f1316f541533e52c4b9 +size 11027112 diff --git a/transformer/model-layer-134.safetensors b/transformer/model-layer-134.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..03e6fcdf835169620fd44e16897dbc8bfd0c7475 --- /dev/null +++ b/transformer/model-layer-134.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:8092cc3d159db04987ec29ea2dc66f311e1dab723891a98be65c28564dd119b1 +size 11027136 diff --git a/transformer/model-layer-135.safetensors b/transformer/model-layer-135.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..cbe57bf83966248dface771dd1cba5adcc2dcca7 --- /dev/null +++ b/transformer/model-layer-135.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:86a53245dd56b67d3c6d0f82ead2b958499224c8cc4d4413856380ef5e7c59aa +size 11027112 diff --git a/transformer/model-layer-136.safetensors b/transformer/model-layer-136.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..7664bc590fec71df42432c6b016d6fd3b1c6b7dd --- /dev/null +++ b/transformer/model-layer-136.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:d7bd44c6263219fd40557e7ee70e0b6f32798933adb6c2a9f6a610905fb228c8 +size 11027112 diff --git a/transformer/model-layer-137.safetensors b/transformer/model-layer-137.safetensors new file mode 100644 index 0000000000000000000000000000000000000000..e8b7450ddb5abf7fc0af9b546483f7791ebf6f3c --- /dev/null +++ b/transformer/model-layer-137.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:8d1fa8656256b8c3de2e20012d8c417922d23c2dc2364d52c8e799bc61ef23c9 +size 30950112