AdaEdit enhances image editing by improving consistency and targeting in text-conditioned scenarios, suggesting better editing outcomes.
Text-conditioned image editing aims to modify a source image into a target image according to a specified text description, tackling two core challenges: locating target editing regions and ensuring consistency in non-target editing areas. Existing approaches utilize manual selection or cross-modal attention to define editing regions and deploy diffusion models to generate edited images. Despite these recent advancements, two problems remain. First, current methods fail to locate editing areas described in the text but invisible in the image. Second, they struggle to ensure spatial consistency in non-targeted regions due to the global noise addition along with excessive denoising during the diffusion process. To overcome these limitations, we propose AdaEdit, which comprises an adaptive mask localization module and an adaptive denoising strategy for text-conditioned image editing. AdaEdit can accurately identify the editing area via the measurement of cross-modal semantic mismatch, even when the visual details are not explicitly described in the text inputs. The adaptive denoising strategy applies varying noise levels to differentiate between targeted and non-targeted regions, enhancing the stability and consistency of the non-edited areas. Extensive experiments demonstrate that our proposed method achieves excellent performance on MS-COCO, MagicBrush, and Laion. We also expand our application to iterative editing tasks, thereby extending its utility for generalized editing scenarios.
No takes yet. Share an insight, caveat, or question.
A 2025 study studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: