PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 4, 20240 citationsOpen Access

ViewDiff: 3D-Consistent Image Generation with Text-to-Image Models

View Full Paper
LHLukas HölleinABAljaž BožičNMNorman Müller

Key Points

Key points are not available for this paper at this time.

Abstract

3D asset generation is getting massive amounts of attention, inspired by the recent success of text-guided 2D content creation. Existing text-to-3D methods use pretrained text-to-image diffusion models in an optimization problem or fine-tune them on synthetic data, which often results in non-photorealistic 3D objects without backgrounds. In this paper, we present a method that leverages pretrained text-to-image models as a prior, and learn to generate multi-view images in a single denoising process from real-world data. Concretely, we propose to integrate 3D volume-rendering and cross-frame-attention layers into each block of the existing U-Net network of the text-to-image model. Moreover, we design an autoregressive generation that renders more 3D-consistent images at any viewpoint. We train our model on real-world datasets of objects and showcase its capabilities to generate instances with a variety of high-quality shapes and textures in authentic surroundings. Compared to the existing methods, the results generated by our method are consistent, and have favorable visual quality (-30% FID, -37% KID).

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Höllein et al. (2024) studied this question.

synapsesocial.com/papers/68e75ddfb6db6435876d5084https://doi.org/10.48550/arxiv.2403.01807
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1IT3D: Improved Text-to-3D Generation with Explicit View Synthesis2024 · 51 citations
  2. 2IM-3D: Iterative Multiview Diffusion and Reconstruction for High-Quality 3D Generation2024 · 3 citations
  3. 3Dual3D: Efficient and Consistent Text-to-3D Generation with Dual-mode Multi-view Latent Diffusion2024 · 1 citations
  4. 4Grounded Compositional and Diverse Text-to-3D with Pretrained Multi-View Diffusion Model2024
  5. 5Meta 3D TextureGen: Fast and Consistent Texture Generation for 3D Objects2024 · 1 citations