Abstract
Recently proposed Drifting Models shift iterative distribution refinement from inference to training, enabling effective one-step generation. However, their performance on complex image datasets depends strongly on the representation used to construct the drifting field: pixel-space drifting performs poorly, whereas pretrained feature spaces substantially improve sample quality for reasons that remain unclear. We trace this gap to the discriminative geometry of the representation, which determines sample weighting in kernel density estimation (KDE) and, consequently drift. We introduce persistent representation learning, which continuously learns a more discriminative representation geometry as the generator evolves across batches. We further establish a current-step gradient equivalence between the KDE ratio loss and drift regression loss under matched conditions, connecting density-ratio-based generator optimization to empirical drifting and motivating direct control of the drifting velocity. Across multiple datasets, our method learns effective discriminative representations directly from pixels and reduces FID by approximately 82-95% over the original pixel-space Drifting Models, without pretrained encoders. Adapting pretrained representations and applying velocity clipping provide further gains.
Community
We investigate why Drifting Models perform poorly in pixel space but improve substantially with pretrained features. We show that the key lies in the discriminative geometry of the representation, which determines sample weighting in KDE and the resulting drift. Based on this insight, we introduce persistent representation learning, enabling the model to learn effective representations directly from pixels as the generator evolves. Our approach reduces FID by approximately 82โ95% over the original pixel-space Drifting Models across multiple datasets, without relying on pretrained encoders.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Unifying Distributional Training for One-Step Visual Generation (2026)
- One-Step Generation via Riemannian Wasserstein Gradient Flows (2026)
- DriftSR: One-Step Real-World Image Super-Resolution via Distribution Drifting (2026)
- Enhancing Autoregressive Video Generation via Representation Adversarial Distillation (2026)
- Geometry-Aware Time Reparameterization for Flow-Map Distillation (2026)
- MaDeL: Manifold-Decomposed Feature Losses for Generative Modeling (2026)
- Discrete Wasserstein Flows for One-Step Generative Modeling (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2610.04703 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper