PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
September 29, 20250 citationsOpen Access

MVD-HuGaS: Human Gaussians from a Single Image via 3D Human Multi-view Diffusion Prior

View Full Paper
KXKaiqiang XiongYFYing FengQZQi Zhang

Key Points

  • MVD-HuGaS significantly enhances 3D human rendering fidelity from a single image, addressing previous artifacts.
  • The approach utilizes a multi-view diffusion model to generate high-quality human representations and optimize camera poses.
  • An alignment module facilitates the joint optimization of 3D Gaussians and camera poses, improving overall reconstruction accuracy.
  • Experiments on Thuman2.0 and 2K2K datasets demonstrate MVD-HuGaS's state-of-the-art performance in rendering complex human models.

Abstract

3D human reconstruction from a single image is a challenging problem and has been exclusively studied in the literature. Recently, some methods have resorted to diffusion models for guidance, optimizing a 3D representation via Score Distillation Sampling (SDS) or generating one back-view image for facilitating reconstruction. However, these methods tend to produce unsatisfactory artifacts (e. g. flattened human structure or over-smoothing results caused by inconsistent priors from multiple views) and struggle with real-world generalization in the wild. In this work, we present MVD-HuGaS, enabling free-view 3D human rendering from a single image via a multi-view human diffusion model. We first generate multi-view images from the single reference image with an enhanced multi-view diffusion model, which is well fine-tuned on high-quality 3D human datasets to incorporate 3D geometry priors and human structure priors. To infer accurate camera poses from the sparse generated multi-view images for reconstruction, an alignment module is introduced to facilitate joint optimization of 3D Gaussians and camera poses. Furthermore, we propose a depth-based Facial Distortion Mitigation module to refine the generated facial regions, thereby improving the overall fidelity of the reconstruction. Finally, leveraging the refined multi-view images, along with their accurate camera poses, MVD-HuGaS optimizes the 3D Gaussians of the target human for high-fidelity free-view renderings. Extensive experiments on Thuman2. 0 and 2K2K datasets show that the proposed MVD-HuGaS achieves state-of-the-art performance on single-view 3D human rendering.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Xiong et al. (2025) studied this question.

synapsesocial.com/papers/68da58c9c1728099cfd1092bhttps://doi.org/10.48550/arxiv.2503.08218
Ask AI
Helpful
Bookmark
Share
View Full Paper