PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 27, 20240 citationsOpen Access

TextCraftor: Your Text Encoder Can be Image Quality Controller

View Full Paper
YLYanyu LiXLXian LiuAKAnil Kag

Key Points

Key points are not available for this paper at this time.

Abstract

Diffusion-based text-to-image generative models, e.g., Stable Diffusion, have revolutionized the field of content generation, enabling significant advancements in areas like image editing and video synthesis. Despite their formidable capabilities, these models are not without their limitations. It is still challenging to synthesize an image that aligns well with the input text, and multiple runs with carefully crafted prompts are required to achieve satisfactory results. To mitigate these limitations, numerous studies have endeavored to fine-tune the pre-trained diffusion models, i.e., UNet, utilizing various technologies. Yet, amidst these efforts, a pivotal question of text-to-image diffusion model training has remained largely unexplored: Is it possible and feasible to fine-tune the text encoder to improve the performance of text-to-image diffusion models? Our findings reveal that, instead of replacing the CLIP text encoder used in Stable Diffusion with other large language models, we can enhance it through our proposed fine-tuning approach, TextCraftor, leading to substantial improvements in quantitative benchmarks and human assessments. Interestingly, our technique also empowers controllable image generation through the interpolation of different text encoders fine-tuned with various rewards. We also demonstrate that TextCraftor is orthogonal to UNet finetuning, and can be combined to further improve generative quality.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Li et al. (2024) studied this question.

synapsesocial.com/papers/68e72422b6db64358769d59ehttps://doi.org/10.48550/arxiv.2403.18978
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1CustomText: Customized Textual Image Generation using Diffusion Models2024
  2. 2Diffusion Dynamics Applied with Novel Methodologies2024 · 8 citations
  3. 3A Comprehensive Survey of Text Encoders for Text-to-Image Diffusion Models2024
  4. 4Text to Image Conversion using Stable Diffusion2024 · 4 citations
  5. 5ECNet: Effective Controllable Text-to-Image Diffusion Models2024 · 2 citations