Remote sensing object detection within multi-scale frameworks remains challenging, largely due to structural degradation and semantic misalignment introduced during cross-scale semantic enhancement. As feature hierarchies deepen, high-frequency details for small-object localization decay, while nonlinear transformations and receptive field asymmetry cause cross-scale semantic and spatial offsets. While existing feature pyramid-based approaches improve detection performance through multi-scale fusion or semantic aggregation, they fail to fundamentally address the cumulative information degradation arising from hierarchical feature extraction. To this end, we propose CFBA-FPN, a unified shallow–deep cross-scale feature compensation framework that explicitly models both frequency discrepancies and semantic offsets across scales. Specifically, shallow features are exploited as structural and spatial anchors to inject lost high-frequency information into deeper layers, effectively mitigating structural degradation. Meanwhile, a cross-scale collaborative semantic alignment strategy is introduced to correct semantic inconsistencies and spatial misalignments among multi-scale features. Building upon these designs, a cascaded gated fusion mechanism is developed to adaptively balance shallow structural compensation and deep semantic representation, thereby suppressing background noise and enhancing small-object responses. Extensive experiments on the AI-TOD, VisDrone, and DIOR benchmarks demonstrate that CFBA-FPN consistently improves localization accuracy and recognition capability, validating its effectiveness and generalization ability in remote sensing object detection.
Yuan et al. (Wed,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: