ABSTRACT Current co‐salient object detection (CoSOD) methods leverage common visual representations to identify recurring foreground objects across image groups. A significant challenge arises when distractors from different categories exhibit high visual similarity to the target objects—such as apples and bananas sharing comparable color and texture—making pure appearance‐based matching prone to failure. To address this issue, we introduce En‐ASCoD, an enhanced architecture that integrates both appearance and shape cues to form a more discriminative consensus representation. The model incorporates a Global Co‐appearance Module (GoAM) to extract group‐level appearance prototypes, a Local Co‐appearance Module (LoAM) that refines these features through contrastive learning within local contexts, and a Co‐shape Module (CoSM) that introduces structural constraints via cross‐attention operating on salient tokens selected through global average pooling. By jointly optimizing these modules in a cascaded manner, En‐ASCoD effectively suppresses visually similar distractors while maintaining robustness to intra‐group appearance variations. Extensive experiments on three challenging benchmarks including CoCA, CoSOD3k, and CoSal2015 show that our method consistently outperforms state‐of‐the‐art alternatives, achieving notable gains in both detection accuracy and generalization.
Guo et al. (Thu,) studied this question.