ABSTRACT Multi‐focus image fusion (MFIF) aims to synthesise an all‐in‐focus image from multiple source images captured with different focal depths. However, effectively distinguishing focused regions from defocused backgrounds remains challenging due to the smooth transitions and lack of ground‐truth supervision. MFIF synthesises an all‐in‐focus image by selecting sharp regions from images captured at different focal planes. However, many deep methods ignore the gradient discrepancy between focused and defocused areas during multi‐scale extraction, treat features without quantifying their reliability, and build decision maps using only horizontal and vertical gradients, leading to detail loss and structural misclassification. We propose MSFE‐Net, a gradient‐aware multi‐scale feature enhancement network comprising three core components. First, a gradient‑aware multi‑scale feature enhancement module (GradientAwareMFE) balances fine detail and contextual cues by fusing dual atrous spatial pyramid pooling branches with small and large dilation rates under a learned gradient gate aligned to the Sobel gradient magnitude. Second, focus‐reliability attention converts local‐variance statistics into a spatial reliability mask and applies squeeze‐and‐excitation to modulate channels. Third, enhanced spatial frequency fusion integrates four‐directional gradients with local context and morphological refinement to produce a robust decision map for pixel‐wise fusion. Trained in a self‐supervised manner on Microsoft Common Objects in Context (MS‐COCO) using mean squared error, structural similarity, and gradient‐consistency losses, and evaluated on Lytro, MFFW, and MFI‐WHU datasets, MSFE‐Net attains state‐of‐the‐art or competitive results on most metrics, delivering sharper edges and fewer artefacts.
Gong et al. (Thu,) studied this question.