PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 17, 2026ISPRS Journal of Photogrammetry and Remote Sensing0 citationsOpen Access

TAIS-Net: Time adaptive implicit sampling diffusion model for arbitrary-scale UAV video super-resolution

View Full Paper
WLWenke LiCDChenguang DaiHFHuixin Fan

Key Points

  • This work aims to improve UAV video super-resolution by exploring diffusion models for arbitrary-scale reconstruction.
  • Introduced TAIS-Net, a time adaptive implicit sampling diffusion model for UAV video SR.
  • Utilized a reference frame alignment module with a pre-trained optical flow model for robust temporal alignment.
  • Implemented an implicit denoising U-Net to learn latent features, correcting prediction error propagation.
  • TAIS-Net achieved superior reconstruction quality and temporal consistency compared to existing methods.
  • Performed well across various motion intensities while handling additional Gaussian blur without retraining.
  • Experimental validation showed enhanced perceptual quality, effectively preserving high-frequency details.

Abstract

Emerging applications urgently require unmanned aerial vehicle (UAV) videos with high clarity and rich structural details. However, limitations in imaging sensors and non-ideal conditions often result in blurred videos lacking sufficient details. Recent studies indicate diffusion models have strong potential for natural scene video super-resolution (SR). However, the application of diffusion models to UAV video SR remains underexplored, particularly in the context of arbitrary-scale reconstruction. To tackle the challenges in UAV video SR at arbitrary scales, this work introduces a time adaptive implicit sampling diffusion model TAIS-Net. First, to achieve robust temporal alignment under large camera motions and texture scarcity commonly encountered in UAV videos, we use a reference frame alignment module to compensate previously reconstructed frames by incorporating motion cues estimated from a pre-trained optical flow model. Second, to enable scale-flexible reconstruction while preserving fine geometric details for UAV videos under arbitrary upscaling factors, we introduce an implicit denoising U-Net to learn latent features by leveraging implicit neural representations and a temporal conditioning module. Third, to reduce the prediction error propagation during sampling, we introduce a historical sample gain module that dynamically corrects and refines latent features at each sampling step, thereby suppressing temporal artifacts such as flickering or drifting. Finally, an arbitrary-scale implicit decoder is used to reconstruct high-resolution videos directly from these enhanced latent features, avoiding scale-dependent blurring while preserving high-frequency details in UAV videos. To validate the effectiveness of TAIS-Net, we construct two visible and one infrared UAV video SR datasets. The experimental results on these datasets demonstrate that TAIS-Net attains competitive performance relative to existing methods in terms of reconstruction quality, perceptual quality and temporal consistency. Moreover, experiments on various motion intensities and additional Gaussian blur beyond conventional bicubic downsampling degradation without requiring model retraining also demonstrate the superiority of our TAIS-Net in maintaining perceptual and temporal consistency under real-world UAV scenarios. The code and datasets are available at https://github.com/cyber-lwk/TAIS-Net .

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Li et al. (2026) studied this question.

synapsesocial.com/papers/6a095a877880e6d24efe0836https://doi.org/10.1016/j.isprsjprs.2026.04.060
Ask AI
Helpful
Bookmark
Share
View Full Paper