Key points are not available for this paper at this time.
The advent of beyond 5G/6G technologies is expected to revolutionize the development of digital twins. Recent development of sensors and deep learning techniques has significantly improved the precision of perception of surroundings, and multimodal object recognition approaches, which play an important role in the realization of the digital twin, have been studied. Previously, we developed a multimodal object recognition method using RGB and depth images. Although our method handled multimodal data sets, there are two problems with this method: Low accuracy of the depth image and dependence on the results of video recognition. The dependence is due to the fact that the recognition results of the RGB images are used to determine the distance between the object and the camera. In this paper, we propose advanced multimodal object recognition methods that expand on our previous method. The proposal includes an additional location-modal approach that utilizes PointNet for semantic segmentation and location estimation. We evaluate our method using our prepared data set consisting of RGB-D images and 3D point clouds, considering the situations where workers and robots collaborate in the warehouse. Our results show that multimodal recognition achieves a better precision score than only the video modality under uncertain measurements.
Ando et al. (Mon,) studied this question.