PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
December 23, 20250 citationsOpen Access

4Real-Video-V2: Fused View-Time Attention and Feedforward Reconstruction for 4D Scene Generation

View Full Paper
CWChaoyang WangHandan CollegeAMAshkan MirzaeiUniversity of TorontoVGVidit GoelSnap (United States)

Key Points

  • To develop a framework for efficient 4D scene generation combining video frames and 3D Gaussian particles.
  • Proposed a fused architecture for spatial and temporal attention in 4D video diffusion.
  • Introduced a Gaussian head and camera token replacement for improved 3D reconstruction.
  • Employed dynamic layers in the architecture for better performance.
  • Achieved state-of-the-art visual quality in 4D scene generation.
  • Enhanced reconstruction capability demonstrated through proposed methods.

Abstract

We propose the first framework capable of computing a 4D spatio-temporal grid of video frames and 3D Gaussian particles for each time step using a feed-forward architecture. Our architecture has two main components, a 4D video model and a 4D reconstruction model. In the first part, we analyze current 4D video diffusion architectures that perform spatial and temporal attention either sequentially or in parallel within a two-stream design. We highlight the limitations of existing approaches and introduce a novel fused architecture that performs spatial and temporal attention within a single layer. The key to our method is a sparse attention pattern, where tokens attend to others in the same frame, at the same timestamp, or from the same viewpoint. In the second part, we extend existing 3D reconstruction algorithms by introducing a Gaussian head, a camera token replacement algorithm, and additional dynamic layers and training. Overall, we establish a new state of the art for 4D generation, improving both visual quality and reconstruction capability.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Wang et al. (2025) studied this question.

synapsesocial.com/papers/6949ddb572f746a93d788e79https://doi.org/10.48550/arxiv.2506.18839
Ask AI
Helpful
Bookmark
Share
View Full Paper