PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 26, 2026IET Image Processing0 citationsOpen Access

Efficient Multi‐Scale Feature Fusion Model for 3D Indoor Object Detection

View Full Paper
YLYing LiuGDGuoqing DengYLYing Liu

Key Points

  • The study aims to develop an efficient model for 3D object detection in indoor environments, addressing challenges like occlusion and small object detection.
  • Proposed an efficient multi-scale feature fusion model (EMF3D) for 3D indoor object detection.
  • Utilized sparse voxel convolution for a compact feature extraction network and integrated a squeeze-and-excitation block for feature channel weighting.
  • Conducted extensive experiments on three indoor point cloud datasets: ScanNet V2, SUN RGB-D, and S3DIS.
  • EMF3D outperformed major detection approaches across several evaluation metrics.
  • Achieved superior performance in small object detection compared to existing methods.
  • Offered a better balance between accuracy and efficiency for indoor point cloud scenarios.

Abstract

ABSTRACT Recently, 3D indoor object detection has become increasingly important in applications such as service robotics, augmented reality and smart homes. However, due to factors like densely arranged objects, severe occlusion and sparse point cloud data, existing methods often suffer from complex architectures, high computational costs or poor performance in small object detection, making it difficult to balance accuracy and efficiency. To address these challenges, we propose an efficient multi‐scale feature fusion model for 3D indoor object detection (EMF3D), offering an end‐to‐end, lightweight solution tailored for indoor point cloud scenarios. The method employs sparse voxel convolution to build a compact feature extraction network and introduces a sparse convolution‐based squeeze‐and‐excitation block at the feature fusion stage to adaptively learn feature channel weights. Furthermore, the enhanced fusion module (EFM) strengthens the perception of critical structures and improves feature discriminability by aggregating multi‐scale representations from earlier attention stages. Extensive experiments are conducted on three indoor point cloud datasets–ScanNet V2, SUN RGB‐D and S3DIS. Results show that the proposed method outperforms major detection approaches across multiple metrics. Compared with existing methods, EMF3D achieves superior performance in small object detection and offers a better trade‐off between accuracy and efficiency.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Liu et al. (2026) studied this question.

synapsesocial.com/papers/69edacdb4a46254e215b49e8https://doi.org/10.1049/ipr2.70339
Ask AI
Helpful
Bookmark
Share
View Full Paper