PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 19, 2026IET Image Processing0 citationsOpen Access

End‐to‐End Multi‐Entity Customization

View Full Paper
WPWonhark ParkJLJaehyun LeeWSWonsik Shin

Key Points

  • The study aims to tackle the issue of concept-mixing in text-to-image synthesis by refining how text tokens interact.
  • Proposed a modifier-based approach to handle concept-mixing issues.
  • Implemented a loss-based finetuning technique for adaptability to various algorithms.
  • Conducted qualitative and quantitative evaluations to assess performance against baselines.
  • Successfully mitigated concept-mixing, allowing clearer object identities in generated images.
  • Demonstrated superior performance in both qualitative assessments and quantitative metrics compared to recent methods.

Abstract

ABSTRACT Recent advancements in text‐to‐image (T2I) models have enabled the synthesis of personalized images that align closely with user‐specified prompts, especially through the use of modifiers. However, generating multiple detailed objects with distinct modifiers in a single image remains challenging due to concept‐mixing, resulting from the difficulty of capturing interactions among text tokens. This paper proposes a modifier‐based approach to mitigate concept‐mixing by addressing the interaction among text tokens. Our method enables practical multi‐personalization while preserving the original T2I model's straightforward inference pipeline. Without structural guidance, it ensures seamless object interaction with enhanced consistency. Through a loss‐based finetuning approach, our method is adaptable to various concept‐learning algorithms, enabling plug‐and‐play functionality. Through both qualitative and quantitative evaluations, we demonstrate that our method effectively resolves concept‐mixing issues to better preserve concepts' identities and outperforms recent baselines in both quantitative and qualitative results. Our code will be publicly available.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Park et al. (2026) studied this question.

synapsesocial.com/papers/6996a7e3ecb39a600b3ee09ahttps://doi.org/10.1049/ipr2.70306
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Multi-SBoRA: regional and non-overlapping weight updates for multi-concept customization of diffusion models2025 · 1 citations
  2. 2Visual Concept-driven Image Generation with Text-to-Image Diffusion Model2025 · 4 citations
  3. 3The Unreasonable Effectiveness of Deep Features as a Perceptual Metric2018 · 13,585 citations
  4. 4Attention Calibration for Disentangled Text-to-Image Personalization2024 · 21 citations
  5. 5InstantBooth: Personalized Text-to-Image Generation without Test-Time Finetuning2024 · 128 citations