ABSTRACT Prevailing image editing methods heavily rely on user‐provided bounding boxes or pixel‐level masks to ensure visual consistency in nonedited regions. Although some attention‐based approaches eliminate the need for manually annotated input, they often unexpectedly alter nontarget areas due to semantic leakage between objects. Our goal is to address this semantic inconsistency challenge with minimal user input by leveraging the mutual exclusion of scene graph nodes, thereby enhancing both editability and background preservation without additional training costs. To address the challenge of semantic inconsistency, we propose a Scene graph‐based ImaGe editing method with Mutually exclusive Attention manipulation, namely SIGMA, to leverage the inherent semantic mutual exclusion properties between scene graph nodes for attention map distribution manipulation. Specifically, we propose a semantic decoupling module to disentangle the desired and nontarget editing areas. We also introduce a semantic injection module to facilitate both foreground editing and background preservation. We validated the effectiveness of SIGMA on the widely used image editing dataset PIE‐Bench. The experimental results demonstrate that SIGMA significantly outperforms existing approaches without any additional training cost.
He et al. (Thu,) studied this question.