High detection accuracy under computing restrictions is a major concern as remote sensing imagery grows exponentially. Light-YOLOv7+, a lightweight target detection system, improves multi-scale representation and inference performance by coordinated architectural modifications. The proposed design uses a Quantized Efficient Feature Fusion Module (QEFFM) and a Spatial–Temporal Dual Attention Module (STDAM) to boost discriminative features across pyramid levels by aggregating features at low cost and reducing redundancy. STDAM adaptively reweights spatial–contextual responses. In addition, an adjustable scale feature pyramid, dynamic anchor creation, and an ultra-lightweight decoding head reduce computing cost while maintaining fine-grained localization. Testing on NWPU VHR-10, DOTA, and RSOD shows continuous performance increases under model capacity constraints. On NWPU VHR-10, Light-YOLOv7 + has a mAP of 0.878 (+ 0.062 over YOLOv7) and a DOTA recall (0.847). While improving accuracy, the model also reduces complexity, FLOPs, and inference latency, making it suitable for resource-limited deployments. The framework is suitable for large-scale remote sensing applications like environmental monitoring and catastrophe assessment because lightweight feature fusion and attention modulation balance efficiency and detection accuracy.
Chengyu Yang (2026) studied this question.