PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 27, 20240 citationsOpen Access

EM Distillation for One-step Diffusion Models

View Full Paper
SXSirui XieZXZhisheng XiaoDKDiederik P. Kingma

Key Points

  • One-step generator models enhance sampling efficiency in diffusion models without major quality loss.
  • EM Distillation achieves superior FID scores on ImageNet datasets, indicating better generative performance.
  • Application of a maximum likelihood-based approach enhances model training via a novel reparametrized sampling scheme and noise cancellation techniques to stabilize the process with minimal loss in quality gains from earlier models during distillation steps. Improved capabilities highlight EMD's effectiveness in generative tasks, showing promise in updating the diffusion model architecture.

Abstract

While diffusion models can learn complex distributions, sampling requires a computationally expensive iterative process. Existing distillation methods enable efficient sampling, but have notable limitations, such as performance degradation with very few sampling steps, reliance on training data access, or mode-seeking optimization that may fail to capture the full distribution. We propose EM Distillation (EMD), a maximum likelihood-based approach that distills a diffusion model to a one-step generator model with minimal loss of perceptual quality. Our approach is derived through the lens of Expectation-Maximization (EM), where the generator parameters are updated using samples from the joint distribution of the diffusion teacher prior and inferred generator latents. We develop a reparametrized sampling scheme and a noise cancellation technique that together stabilizes the distillation process. We further reveal an interesting connection of our method with existing methods that minimize mode-seeking KL. EMD outperforms existing one-step generative methods in terms of FID scores on ImageNet-64 and ImageNet-128, and compares favorably with prior work on distilling text-to-image diffusion models.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Xie et al. (2024) studied this question.

synapsesocial.com/papers/68e68593b6db64358760de3dhttps://doi.org/10.48550/arxiv.2405.16852
Ask AI
Helpful
Bookmark
Share
View Full Paper