PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 15, 2026IEEE Transactions on Image Processing0 citations

SCSV: Spatial-temporal Consistent Dynamic 3D Scene Generation from Sparse Views

View Full Paper
JLJie LiChinese Academy of SciencesJHJunjie HeHong Kong University of Science and TechnologyWLWenjie LiuHong Kong University of Science and Technology

Key Points

  • The aim is to develop a method for generating dynamic 3D scenes from sparse views while maintaining spatial-temporal consistency and improving rendering quality.
  • Developed a two-stage approach: scene reconstruction and scene expansion.
  • Interpolated input images using a video generation model for scene reconstruction.
  • Utilized uncertainty-aware Gaussian training for optimizing the reconstructed scene's consistency.
  • Employed a geometry-aware diffusion process for refining background views in scene expansion.
  • Generated human motion for the foreground, ensuring temporal coherence.
  • SCSV outperforms existing methods on multiple datasets including NeuMan, Bonn, and EMDB.
  • Significant improvements were noticed in both rendering quality and spatial-temporal consistency.

Abstract

Generating dynamic scenes from images has gained increasing attention. Existing methods have two major limitations: (1) They can hardly handle sparse images which exhibit limited geometry constraints and insufficient motion; (2) They struggle to maintain spatial-temporal consistency when rendering multi-view videos. To address these limitations, we propose SCSV, a spatial-temporal consistent dynamic scene generation method from sparse views. Our method consists of two stages: scene reconstruction and scene expansion, both of which decouple background and foreground. In the scene reconstruction stage, we first interpolate a set of images between the input images based on a video generation model, followed by the optimization of the scene Gaussian from the interpolated and input images. To improve the spatial-temporal consistency of the reconstructed scene, we propose an uncertainty-aware Gaussian training approach, which introduces adaptive weights of images and pixels. In the scene expansion stage, for background, we render novel views and refine them with a geometry-aware diffusion process. These refined images are then used to incrementally add the Gaussians. As to foreground, we generate human motion according to previous motion, enabling temporal coherent generation of motion. To further enhance the physical plausibility, we integrate the expanded foreground into the background using a gravity-aware alignment. Experiments on NeuMan, Bonn, and EMDB datasets demonstrate that our SCSV achieves superior performance compared to state-of-the-art methods. The code will be released upon acceptance.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Li et al. (2026) studied this question.

synapsesocial.com/papers/69b64ccdb42794e3e660df68https://doi.org/10.1109/tip.2026.3671692
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Sapiens: Foundation for Human Vision Models2024 · 98 citations
  2. 2Vivid-ZOO: Multi-View Video Generation with Diffusion Model2024 · 4 citations
  3. 33D Gaussian Splatting for Real-Time Radiance Field Rendering2023 · 5,257 citations
  4. 4Normalizing Flows for Human Pose Anomaly Detection2023 · 83 citations
  5. 5Magic3D: High-Resolution Text-to-3D Content Creation2023 · 782 citations