PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 18, 20250 citationsOpen Access

IL3D: A Large-Scale Indoor Layout Dataset for LLM-Driven 3D Scene Generation

View Full Paper
WZWenhui ZhouNanjing UniversityKNKun NieGuangdong Academy of Medical SciencesHDHang DuBeijing University of Posts and Telecommunications

Key Points

  • IL3D features 27,816 indoor layouts across 18 room types, providing diverse training data for 3D scene generation.
  • Supervised fine-tuning (SFT) of LLMs on IL3D improves generalization compared to other datasets, boosting performance.
  • The dataset includes 29,215 high-fidelity 3D object assets and supports robust multimodal learning through natural language annotations.
  • IL3D offers flexible data export options for various visual tasks, advancing research in embodied intelligence and environment perception.

Abstract

In this study, we present IL3D, a large-scale dataset meticulously designed for large language model (LLM)-driven 3D scene generation, addressing the pressing demand for diverse, high-quality training data in indoor layout design. Comprising 27,816 indoor layouts across 18 prevalent room types and a library of 29,215 high-fidelity 3D object assets, IL3D is enriched with instance-level natural language annotations to support robust multimodal learning for vision-language tasks. We establish rigorous benchmarks to evaluate LLM-driven scene generation. Experimental results show that supervised fine-tuning (SFT) of LLMs on IL3D significantly improves generalization and surpasses the performance of SFT on other datasets. IL3D offers flexible multimodal data export capabilities, including point clouds, 3D bounding boxes, multiview images, depth maps, normal maps, and semantic masks, enabling seamless adaptation to various visual tasks. As a versatile and robust resource, IL3D significantly advances research in 3D scene generation and embodied intelligence, by providing high-fidelity scene data to support environment perception tasks of embodied agents.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Zhou et al. (2025) studied this question.

synapsesocial.com/papers/68f3b2fb3f213c1f8b4d3522https://doi.org/10.48550/arxiv.2510.12095
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1LLplace: The 3D Indoor Scene Layout Generation and Editing via Large Language Model2024 · 1 citations
  2. 2M3DLayout: A Multi-Source Dataset of 3D Indoor Layouts and Structured Descriptions for 3D Generation2025
  3. 3OptiScene: LLM-driven Indoor Scene Layout Generation via Scaled Human-aligned Data Synthesis and Multi-Stage Preference Optimization2025
  4. 4LLMI3D: MLLM-based 3D Perception from a Single 2D Image2024 · 2 citations
  5. 5Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning2024 · 9 citations