PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 8, 2026International Journal of Computer Vision0 citationsOpen Access

Disentangling Local and Global Semantics in Diffusion Models for Image Editing

View Full Paper
MPManos PlitsisTKTheodoros KouzelisPKPanagiotis Koromilas

Key Points

  • This research aims to enhance localized image editing by disentangling local and global semantics in diffusion models.
  • Proposed an unsupervised method for localized image editing in diffusion models.
  • Used the Jacobian of the denoising network to map regions to latent subspaces.
  • Separated latent subspaces into shared and region-specific components to enable semantic control over local attributes.
  • Yielded more localized edits compared to existing methods.
  • Achieved high-fidelity edits without retraining.
  • Minimized manual supervision by inferring directions from a single reference image.

Abstract

Abstract Diffusion models have achieved state-of-the-art image synthesis, yet unlike GANs, they lack a well-structured latent space for intuitive image editing. Existing diffusion-based editing methods often rely on supervised fine-tuning or text-based guidance, while recent unsupervised techniques leveraging the model’s bottleneck layer suffer from one or more key limitations: (i) they focus only on global attributes, (ii) fail to disentangle local and global semantics, or (iii) require extensive human intervention. To fill this gap, we first propose an unsupervised method for localized image editing in pre-trained unconditional diffusion models that disentangles local and global semantics in the model’s latent space. Given an input image and a user-specified region of interest, our approach uses the denoising network’s Jacobian to map that region to a corresponding latent subspace. We then separate this subspace into shared (global) and region-specific components to uncover latent directions that control local attributes. These directions generalize across images, enabling semantically consistent edits without retraining. We go one step further by extending our method to minimize manual supervision by automatically inferring edit directions from a single reference image and generating region masks without human input. Experiments on multiple datasets show that our method yields more localized, high-fidelity edits than state-of-the-art approaches.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Plitsis et al. (2026) studied this question.

synapsesocial.com/papers/69acc58f32b0ef16a404ff78https://doi.org/10.1007/s11263-025-02694-y
Ask AI
Helpful
Bookmark
Share
View Full Paper