PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 20, 20250 citationsOpen Access

WorldSplat: Gaussian-Centric Feed-Forward 4D Scene Generation for Autonomous Driving

View Full Paper
FZFanghua ZhangZWZhanqian WuZZZhenxin Zhu

Key Points

  • WorldSplat generates consistent multi-track videos, enhancing the realism of driving scenarios.
  • Extensive experiments show that the framework produces temporally and spatially consistent driving videos.
  • The approach utilizes a 4D-aware latent diffusion model to integrate multi-modal information effectively.
  • This novel framework addresses limitations of existing driving-scene generation methods, improving training data quality.

Abstract

Recent advances in driving-scene generation and reconstruction have demonstrated significant potential for enhancing autonomous driving systems by producing scalable and controllable training data. Existing generation methods primarily focus on synthesizing diverse and high-fidelity driving videos; however, due to limited 3D consistency and sparse viewpoint coverage, they struggle to support convenient and high-quality novel-view synthesis (NVS). Conversely, recent 3D/4D reconstruction approaches have significantly improved NVS for real-world driving scenes, yet inherently lack generative capabilities. To overcome this dilemma between scene generation and reconstruction, we propose WorldSplat, a novel feed-forward framework for 4D driving-scene generation. Our approach effectively generates consistent multi-track videos through two key steps: (i) We introduce a 4D-aware latent diffusion model integrating multi-modal information to produce pixel-aligned 4D Gaussians in a feed-forward manner. (ii) Subsequently, we refine the novel view videos rendered from these Gaussians using a enhanced video diffusion model. Extensive experiments conducted on benchmark datasets demonstrate that WorldSplat effectively generates high-fidelity, temporally and spatially consistent multi-track novel view driving videos. Project: https://wm-research.github.io/worldsplat/

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Zhang et al. (2025) studied this question.

synapsesocial.com/papers/68f6379bb481a140a36cf440https://doi.org/10.48550/arxiv.2509.23402
Ask AI
Helpful
Bookmark
Share
View Full Paper