[NIPS2026🔥] DC-SAE: Deep Compression Semantic Autoencoder for Faster Diffusion Convergence
Xu Huang1* · Ye Huang1* · Zijun Liao1* · Yuwei Niu1 · Xiaojie Li
Menghan Zhou2 · De Wen Soh2 · Xiaotong Li1 · Daquan Zhou1†
1 Peking University 2 Singapore University of Technology and Design
* Equal contribution † Corresponding author
English · 简体中文 · Installation & Usage
🎬 Demo
https://github.com/user-attachments/assets/b29a4cd4-d103-4f7c-a890-2cdde307ef36
▶ Watch the demo · MP4 · English subtitles
Overview
dc-sae brings image reconstruction and latent-space generation into one training and evaluation pipeline. A pretrained vision encoder provides semantic features, an HF branch complements them, and a decoder reconstructs the image. DiT learns the resulting latent distribution for class-conditional generation.
The release includes SAE and DiT training, latent-statistics estimation, reconstruction PSNR / rFID, and generation gFID evaluation at 256px and 512px, using configurations matched to each checkpoint.
| Component | Role |
|---|---|
| Semantic + HF encoding | Combine pretrained visual features with a complementary high-frequency branch. |
| Reconstruction decoder | Decode latents into images, with optional spatial demerger support. |
| Latent DiT | Train and sample a class-conditional generative model with the corresponding latent normalization. |
| Evaluation | Measure reconstruction PSNR/rFID and generation FID with explicit reference files and sampling settings. |
Installation & Usage
git clone -b main --single-branch https://github.com/DAGroup-PKU/DCSAE.git
cd DCSAE
See the Installation & Usage guide for dependencies, model downloads, weight placement, and complete commands. Run all commands from the repository root.
| Get started | Guide |
|---|---|
| Install dependencies | Environment setup |
| Download DINOv2 and prepare checkpoints | Data & pretrained models |
| Train the autoencoder | SAE training |
| Evaluate reconstruction quality | PSNR & rFID |
| Prepare latents and train DiT | Latent statistics · DiT training |
| Evaluate generation quality | 256px / 512px gFID |
Pretrained backbones, trained SAE/DiT checkpoints, datasets, and FID references are downloaded or supplied separately. The guide includes the official ImageNet reference links and explains how to match checkpoints, configs, and latent statistics.
Repository Structure
dc-sae/ SAE / DiT training, evaluation, model modules and example configs
models/ DiT / DDT models and shared utilities
data/ ImageNet WebDataset support
train_vae/ GAN components for SAE training
docs/ Installation and usage guides in English and Chinese
tests/ Offline CPU checks
Training uses standard torchrun for single-node or multi-node execution. The bundled DINOv2 + HF64 configs provide an end-to-end architecture example; other checkpoints require their matching configs.
License & Acknowledgements
Source code is distributed under the MIT License. See Third-party notices for retained dependencies and their licenses. Pretrained models and datasets remain subject to their own terms.
Implementation checks and their scope are recorded in VALIDATION.md.
Our work is based on the RAE code repository. We thank the authors for their great work and for making their code publicly available.
Citation
@misc{huang2026dcsaedeepcompressionsemantic,
title={DC-SAE: Deep Compression Semantic Autoencoder for Faster Diffusion Convergence},
author={Xu Huang and Ye Huang and Zijun Liao and Yuwei Niu and Xiaojie Li and Menghan Zhou and De Wen Soh and Xiaotong Li and Daquan Zhou},
year={2026},
eprint={2609.39222},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2609.39222},
}
- Downloads last month
- 15

