PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 3, 2026Advanced Engineering Informatics5 citationsOpen Access

Real-time multimodal fusion and semantic mapping for robotic tower crane perception

View Full Paper
YLYifan LuXDXiuzhi DENGPLPeter E.D. Love

Key Points

  • Higher semantic accuracy is achieved using an improved semantic segmentation method with RandLA-Net, enhancing scene understanding.
  • Field deployment on the tower crane provided the lowest global reconstruction errors, demonstrating effective perception in challenging environments.
  • The framework integrates LiDAR, camera, and IMU data for real-time perception, crucial for reliable crane operations during construction.
  • A Mahony-filter-based module optimizes stability under vibration, addressing a key limitation of conventional SLAM methods.

Abstract

Robotic tower crane operation requires real-time perception of complex and rapidly changing construction environments. Conventional Simultaneous Localization and Mapping (SLAM) methods assume smooth sensor motion and emphasize geometry over semantics, limiting their suitability for crane-mounted sensing affected by vibration, rotation, and intermittent movement. This research proposes a multimodal perception framework that integrates Light Detection and Ranging (LiDAR), camera, and Inertial Measurement Unit (IMU) data within a tightly coupled fusion and semantic reconstruction pipeline. A Mahony-filter-based attitude optimization module stabilizes high-frequency vibrations, while a Fast LiDAR-Inertial Odometry (FAST-LIVO2)-inspired LiDAR–visual–inertial fusion strategy achieves centimeter-level three-dimensional (3D) mapping. To enhance scene understanding, an improved Random Sampled and Lightweight Aggregated Network (RandLA-Net) jointly exploits geometric and visual cues for point-level semantic segmentation, with color-aware spatial encoding. Field deployment on an operational tower crane demonstrates superior performance, yielding the lowest global reconstruction errors and highest semantic accuracy. The framework provides a robust perception foundation for autonomous planning, safety monitoring, and intelligent lifting assistance.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Lu et al. (2026) studied this question.

synapsesocial.com/papers/69a7673ebadf0bb9e87e0279https://doi.org/10.1016/j.aei.2026.104373
Ask AI
Helpful
Bookmark
Share
View Full Paper