Aiming at the challenges of high intra-class disparity and low inter-class disparity in fine-grained image classification, a multi-branch fine-grained image classification method based on ConvNeXt network as the backbone and using GradCAM heatmap for cropping and attention erasure is proposed. This method uses GradCAM to obtain the attention heatmap of the network through gradient reflow, locates the region with discriminative features, crops and enlarges the region, and makes the network focus on local deeper features. In addition, supervised contrastive learning is introduced to expand between-class differences and reduce intra-class differences. Finally, a heatmap attention erasure operation is performed to enable the network to focus on other regions useful for classification while focusing on the most discriminative features. The proposed method achieved 91.8%, 94.9%, 94.0% and 94.4% classification accuracy on CUB-200-2011, Stanford Cars, FGVC Aircraft, and Stanford Dogs datasets, respectively, which is better than many main-stream fine-grained image classification methods. And our method proposed in this paper achieves stop-3 and top-1 classification accuracy on the CUB-200-2011 and Stanford Dogs datasets, respectivel.
No takes yet. Share an insight, caveat, or question.
Liu et al. (2025) studied this question.
Synapse has enriched one closely related paper. Consider it for comparative context: