Experimental evaluation demonstrates enhanced 3D semantic segmentation via adaptive RGB-LiDAR feature integration, suggesting improved efficiency for real-time robotic perception.
Semantic segmentation of large-scale 3D point clouds is a fundamental task in robotic perception, semantic mapping, and urban scene understanding. Existing methods mainly rely on geometric information, which limits their ability to distinguish semantic categories with similar spatial structures. To address this issue, this paper proposes a lightweight cross-modal feature learning framework that adaptively integrates geometric coordinates and RGB color information. By exploiting the complementary characteristics of spatial structure and visual appearance during feature encoding, the proposed method enhances feature discriminability while maintaining a compact model scale. Experiments on the Semantic3D dataset show that the proposed method achieves an mIoU of 87.1%, outperforming the original RandLA-Net and several representative approaches. Additional runtime and LiDAR-only cross-dataset experiments indicate the potential of the proposed structure for online outdoor point cloud perception.
No takes yet. Share an insight, caveat, or question.
Zhai et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: