PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 10, 20250 citationsOpen Access

Asynchronous Denoising Diffusion Models for Aligning Text-to-Image Generation

View Full Paper
ZHZijing HuYTYunze TongFZFengda Zhang

Key Points

  • Asynchronous diffusion models improve alignment between generated images and input prompts, enhancing output quality.
  • Experimental results show significant advancements in alignment, yielding clearer context and better pixel denoising.
  • The proposed framework allows individual pixel timesteps to adapt, leading to more effective image generation.
  • Optimizing pixel-wise denoising enhances prompt-related areas, demonstrating potential in diverse text-to-image scenarios.

Abstract

Diffusion models have achieved impressive results in generating high-quality images. Yet, they often struggle to faithfully align the generated images with the input prompts. This limitation arises from synchronous denoising, where all pixels simultaneously evolve from random noise to clear images. As a result, during generation, the prompt-related regions can only reference the unrelated regions at the same noise level, failing to obtain clear context and ultimately impairing text-to-image alignment. To address this issue, we propose asynchronous diffusion models -- a novel framework that allocates distinct timesteps to different pixels and reformulates the pixel-wise denoising process. By dynamically modulating the timestep schedules of individual pixels, prompt-related regions are denoised more gradually than unrelated regions, thereby allowing them to leverage clearer inter-pixel context. Consequently, these prompt-related regions achieve better alignment in the final images. Extensive experiments demonstrate that our asynchronous diffusion models can significantly improve text-to-image alignment across diverse prompts. The code repository for this work is available at https://github.com/hu-zijing/AsynDM.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Hu et al. (2025) studied this question.

synapsesocial.com/papers/68e997abe14057276da7f1cfhttps://doi.org/10.48550/arxiv.2510.04504
Ask AI
Helpful
Bookmark
Share
View Full Paper