PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 2, 2026IEEE Transactions on Image Processing14 citations

ReconX: Reconstruct Any Scene from Sparse Views with Video Diffusion Model

View Full Paper
FLFangfu LiuWSWenqiang SunHWHanyang Wang

Key Points

  • This research aims to improve 3D scene reconstruction from sparse views using a new approach called ReconX.
  • Developing a novel 3D reconstruction paradigm called ReconX.
  • Constructing a global point cloud from limited views to create a 3D structure condition.
  • Utilizing a pre-trained video diffusion model for synthesizing video frames.
  • Implementing a confidence-aware 3D Gaussian Splatting for 3D scene recovery.
  • Achieved high-quality 3D scene reconstructions with improved detail and consistency.
  • Demonstrated superiority in reconstruction quality compared to existing state-of-the-art methods.
  • Ensured 3D consistency across multiple perspectives of the scene.

Abstract

Advancements in 3D scene reconstruction have transformed 2D images from the real world into 3D models, producing realistic 3D results from hundreds of input photos. Despite great success in dense-view reconstruction scenarios, rendering a detailed scene from sparse views is still an ill-posed optimization problem, often resulting in artifacts and distortions in unseen areas. In this paper, we propose ReconX, a novel 3D scene reconstruction paradigm that reframes the ambiguous reconstruction problem as a temporal generation task. The key insight is to unleash the strong generative prior of large pretrained video diffusion models for sparse-view reconstruction. Nevertheless, it is challenging to preserve 3D view consistency when directly generating video frames from pre-trained models. To address this issue, given limited input views, the proposed ReconX first constructs a global point cloud and encodes it into a contextual space as the 3D structure condition. Guided by the condition, the video diffusion model then synthesizes video frames that are detail-preserved and exhibit a high degree of 3D consistency, ensuring the coherence of the scene from various perspectives. Finally, we recover the 3D scene from the generated video through a confidence-aware 3D Gaussian Splatting optimization scheme. Extensive experiments on various real-world datasets show the superiority of ReconX over state-of-the-art methods in terms of quality and generalizability.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Liu et al. (2026) studied this question.

synapsesocial.com/papers/69a52920f1e85e5c73bf06c1https://doi.org/10.1109/tip.2026.3666733
Ask AI
Helpful
Bookmark
Share
View Full Paper