PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 16, 20250 citationsOpen Access

Att-Adapter: A Robust and Precise Domain-Specific Multi-Attributes T2I Diffusion Adapter via Conditional Variational Autoencoder

View Full Paper
WCWonwoong ChoYCYanying ChenMKMatthew Klenk

Key Points

  • Att-Adapter enables fine control of multiple attributes simultaneously in image generation, enhancing T2I capabilities.
  • Evaluations show that Att-Adapter outperforms LoRA-based methods in controlling continuous attributes across various datasets.
  • By leveraging a conditional variational autoencoder, the Att-Adapter minimizes overfitting and adapts to diverse visual qualities.
  • Its plug-and-play design requires no paired data, making it more flexible and scalable for complex attributes.

Abstract

Text-to-Image (T2I) Diffusion Models have achieved remarkable performance in generating high quality images. However, enabling precise control of continuous attributes, especially multiple attributes simultaneously, in a new domain (e.g., numeric values like eye openness or car width) with text-only guidance remains a significant challenge. To address this, we introduce the Attribute (Att) Adapter, a novel plug-and-play module designed to enable fine-grained, multi-attributes control in pretrained diffusion models. Our approach learns a single control adapter from a set of sample images that can be unpaired and contain multiple visual attributes. The Att-Adapter leverages the decoupled cross attention module to naturally harmonize the multiple domain attributes with text conditioning. We further introduce Conditional Variational Autoencoder (CVAE) to the Att-Adapter to mitigate overfitting, matching the diverse nature of the visual world. Evaluations on two public datasets show that Att-Adapter outperforms all LoRA-based baselines in controlling continuous attributes. Additionally, our method enables a broader control range and also improves disentanglement across multiple attributes, surpassing StyleGAN-based techniques. Notably, Att-Adapter is flexible, requiring no paired synthetic data for training, and is easily scalable to multiple attributes within a single model.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Cho et al. (2025) studied this question.

synapsesocial.com/papers/68f163c79903599108abcd52https://doi.org/10.48550/arxiv.2503.11937
Ask AI
Helpful
Bookmark
Share
View Full Paper