PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 29, 20240 citationsOpen Access

Tuning-Free Alignment of Diffusion Models with Direct Noise Optimization

View Full Paper
ZTZhiwei TangJPJiangweizhi PengJTJiasheng Tang

Key Points

Key points are not available for this paper at this time.

Abstract

In this work, we focus on the alignment problem of diffusion models with a continuous reward function, which represents specific objectives for downstream tasks, such as improving human preference. The central goal of the alignment problem is to adjust the distribution learned by diffusion models such that the generated samples maximize the target reward function. We propose a novel alignment approach, named Direct Noise Optimization (DNO), that optimizes the injected noise during the sampling process of diffusion models. By design, DNO is tuning-free and prompt-agnostic, as the alignment occurs in an online fashion during generation. We rigorously study the theoretical properties of DNO and also propose variants to deal with non-differentiable reward functions. Furthermore, we identify that naive implementation of DNO occasionally suffers from the out-of-distribution reward hacking problem, where optimized samples have high rewards but are no longer in the support of the pretrained distribution. To remedy this issue, we leverage classical high-dimensional statistics theory and propose to augment the DNO loss with certain probability regularization. We conduct extensive experiments on several popular reward functions trained on human feedback data and demonstrate that the proposed DNO approach achieves state-of-the-art reward scores as well as high image quality, all within a reasonable time budget for generation.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Tang et al. (2024) studied this question.

synapsesocial.com/papers/68e67f72b6db643587608ff0https://doi.org/10.48550/arxiv.2405.18881
Ask AI
Helpful
Bookmark
Share
View Full Paper