End-to-end learning methods can model complex spatial and temporal relationships and achieve good rate-distortion performance. However, many existing methods still rely on single-scale or coarse features for motion estimation. These features cannot capture fine motion details well. As a result, high-frequency information is often lost during motion compensation, which reduces the quality of the reconstructed video. In this work, we propose an end-to-end video compression approach that operates in the deep feature space. The core component is a Multi-scale Progressive Fusion (MPF) module. It extracts features at different scales and fuses them in a progressive manner. This strategy improves motion estimation accuracy while alleviating high-frequency information loss during compression. We further introduce a Global Attention Prediction Enhancement (GAPE) module. By combining high-frequency features with channel attention, this module refines motion compensation and enhances the quality of reconstructed frames. Experimental results show that the proposed method performs better than the baseline models with respect to PSNR and MS-SSIM.
LIU et al. (Thu,) studied this question.