ABSTRACT Recently, 3D indoor object detection has become increasingly important in applications such as service robotics, augmented reality and smart homes. However, due to factors like densely arranged objects, severe occlusion and sparse point cloud data, existing methods often suffer from complex architectures, high computational costs or poor performance in small object detection, making it difficult to balance accuracy and efficiency. To address these challenges, we propose an efficient multi‐scale feature fusion model for 3D indoor object detection (EMF3D), offering an end‐to‐end, lightweight solution tailored for indoor point cloud scenarios. The method employs sparse voxel convolution to build a compact feature extraction network and introduces a sparse convolution‐based squeeze‐and‐excitation block at the feature fusion stage to adaptively learn feature channel weights. Furthermore, the enhanced fusion module (EFM) strengthens the perception of critical structures and improves feature discriminability by aggregating multi‐scale representations from earlier attention stages. Extensive experiments are conducted on three indoor point cloud datasets–ScanNet V2, SUN RGB‐D and S3DIS. Results show that the proposed method outperforms major detection approaches across multiple metrics. Compared with existing methods, EMF3D achieves superior performance in small object detection and offers a better trade‐off between accuracy and efficiency.
Liu et al. (Thu,) studied this question.