PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 8, 20240 citationsOpen Access

VideoElevator: Elevating Video Generation Quality with Versatile Text-to-Image Diffusion Models

View Full Paper
YZYabo ZhangYWYuxiang WeiXLXianhui Lin

Key Points

Key points are not available for this paper at this time.

Abstract

Text-to-image diffusion models (T2I) have demonstrated unprecedented capabilities in creating realistic and aesthetic images. On the contrary, text-to-video diffusion models (T2V) still lag far behind in frame quality and text alignment, owing to insufficient quality and quantity of training videos. In this paper, we introduce VideoElevator, a training-free and plug-and-play method, which elevates the performance of T2V using superior capabilities of T2I. Different from conventional T2V sampling (i.e., temporal and spatial modeling), VideoElevator explicitly decomposes each sampling step into temporal motion refining and spatial quality elevating. Specifically, temporal motion refining uses encapsulated T2V to enhance temporal consistency, followed by inverting to the noise distribution required by T2I. Then, spatial quality elevating harnesses inflated T2I to directly predict less noisy latent, adding more photo-realistic details. We have conducted experiments in extensive prompts under the combination of various T2V and T2I. The results show that VideoElevator not only improves the performance of T2V baselines with foundational T2I, but also facilitates stylistic video synthesis with personalized T2I. Our code is available at https://github.com/YBYBZhang/VideoElevator.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Zhang et al. (2024) studied this question.

synapsesocial.com/papers/68e752dab6db6435876cb730https://doi.org/10.48550/arxiv.2403.05438
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1T2V-Turbo: Breaking the Quality Bottleneck of Video Consistency Model with Mixed Reward Feedback2024
  2. 2TI2V-Zero: Zero-Shot Image Conditioning for Text-to-Video Diffusion Models2024
  3. 3StreamingT2V: Consistent, Dynamic, and Extendable Long Video Generation from Text2024 · 4 citations
  4. 4I4VGen: Image as Stepping Stone for Text-to-Video Generation2024
  5. 5xGen-VideoSyn-1: High-fidelity Text-to-Video Synthesis with Compressed Representations2024