PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 10, 20240 citationsOpen Access

FastVideoEdit: Leveraging Consistency Models for Efficient Text-to-Video Editing

View Full Paper
YZYouyuan ZhangUniversity of Electronic Science and Technology of ChinaXJXuan JuHangzhou Seventh Peoples HospitalJCJames J. ClarkImperial College Healthcare NHS Trust

Key Points

Key points are not available for this paper at this time.

Abstract

Diffusion models have demonstrated remarkable capabilities in text-to-image and text-to-video generation, opening up possibilities for video editing based on textual input. However, the computational cost associated with sequential sampling in diffusion models poses challenges for efficient video editing. Existing approaches relying on image generation models for video editing suffer from time-consuming one-shot fine-tuning, additional condition extraction, or DDIM inversion, making real-time applications impractical. In this work, we propose FastVideoEdit, an efficient zero-shot video editing approach inspired by Consistency Models (CMs). By leveraging the self-consistency property of CMs, we eliminate the need for time-consuming inversion or additional condition extraction, reducing editing time. Our method enables direct mapping from source video to target video with strong preservation ability utilizing a special variance schedule. This results in improved speed advantages, as fewer sampling steps can be used while maintaining comparable generation quality. Experimental results validate the state-of-the-art performance and speed advantages of FastVideoEdit across evaluation metrics encompassing editing speed, temporal consistency, and text-video alignment.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Zhang et al. (2024) studied this question.

synapsesocial.com/papers/68e74cd0b6db6435876c527ehttps://doi.org/10.48550/arxiv.2403.06269
Ask AI
Helpful
Bookmark
Share
View Full Paper