PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
December 8, 2025ACM Transactions on Graphics1 citationsOpen Access

SS4D: Native 4D Generative Model via Structured Spacetime Latents

View Full Paper
DLDahua Lin

Key Points

  • Develop a 4D generative model for synthesizing dynamic 3D objects from video data.
  • Synthesize objects from monocular video
  • Use structured spacetime latents for high fidelity
  • Introduce temporal layers for consistency
  • Employ factorized 4D convolutions and temporal downsampling for efficiency
  • Model shows robustness against motion blur
  • Generates high-quality spatio-temporally consistent 4D objects
  • Surpasses existing models on synthetic datasets

Abstract

We present SS4D, a native 4D generative model that synthesizes dynamic 3D objects directly from monocular video. Unlike prior approaches that construct 4D representations by optimizing over 3D or video generative models, we train a generator directly on 4D data, achieving high fidelity, temporal coherence, and structural consistency. At the core of our method is a compressed set of structured spacetime latents. Specifically, (1) To address the scarcity of 4D training data, we build on a pre-trained single-image-to-3D model, preserving strong spatial consistency. (2) Temporal consistency is enforced by introducing dedicated temporal layers that reason across frames. (3) To support efficient training and inference over long video sequences, we compress the latent sequence along the temporal axis using factorized 4D convolutions and temporal downsampling blocks. In addition, we employ a carefully designed training strategy to enhance robustness against occlusion and motion blur, leading to high-quality generation. Extensive experiments show that SS4D produces spatio-temporally consistent 4D objects with superior quality and efficiency, significantly outperforming state-of-the-art methods on both synthetic and real-world datasets.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Dahua Lin (2025) studied this question.

synapsesocial.com/papers/693624ce4fa91c937236ceb6https://doi.org/10.1145/3763302
Ask AI
Helpful
Bookmark
Share
View Full Paper