PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 21, 2026Artificial Intelligence in Agriculture2 citationsOpen Access

YOLO-GPP: End-to-end prediction of the grasp position and pose on tomato peduncle for robotic harvesting

View Full Paper
LGLi GaoYLYajun LiJWJianwei Wu

Key Points

  • The research aims to develop a robust model for predicting grasp positions and poses of tomato peduncles during robotic harvesting.
  • Developed the YOLOv8-GPP model for predicting grasp position and pose from RGB images.
  • Implemented DCNv4 and SDI-BiFPN modules for improved feature extraction.
  • Designed specialized grasp pose loss functions for enhanced model precision.
  • Conducted experiments to evaluate the model's accuracy and robustness in various conditions.
  • Achieved a vector detection accuracy of 95.3% mAP-V score.
  • Attained an average grasp position localization error of 7.07 pixels.
  • Reduced localization error by 47% compared to traditional keypoint detection methods.
  • Decreased pose estimation error by 30% compared to standard instance segmentation methods.

Abstract

A high fresh fruit harvesting success rate relies on the real-time and precise determination of the optimal grasping position and pose of the target fruit. This study proposes an end-to-end grasp pose prediction network, YOLOv8-GPP, which predicts the optimal cutting point and end-effector pose vector of target fruit clusters directly from RGB images by analyzing the positional relationships among the main stems, peduncles, and fruits. To enhance the model's adaptability to target deformation and scale variations in complex agricultural environments, this study introduces the latest DCNv4(Deformable Convolution v4) and SDI-BiFPN(Semantic and Detail Injection-Bidirectional Feature Pyramid Network) modules, significantly improving feature extraction robustness while maintaining lightweight characteristics. Furthermore, through the design of dedicated grasp pose loss functions and coarse-grained pose training strategies, the precision and stability of grasp vector prediction are further enhanced. Experimental results demonstrated that the YOLOv8-GPP model achieved a vector detection accuracy of 95.3% mAP-V score, with average grasp position localization error of 7.07 pixels and grasp pose estimation error of 2.9° in the test set. Compared to keypoint detection methods with post-processing, YOLOv8-GPP achieved a 47% reduction in localization error, and it reduced the pose error by 30% compared to prevalent instance segmentation methods with post-processing. The results indicated that the proposed network accurately predicted the grasp position and pose, contributing valuable visual guidance for servo control and end-effector motion planning in fresh fruit harvesting robots. • An end-to-end network (YOLOv8-GPP) is proposed for simultaneous prediction of grasp position and pose on tomato peduncles from RGB images. • A dedicated grasp vector branch enables direct estimation of “where to pick” and “how to pick” without post-processing. • DCNv4 and SDI-BiFPN improve robustness to slender structures, deformation, occlusion, and scale variation in close-range harvesting scenes. • A vector similarity (VS) metric and mAP-V are designed to jointly evaluate grasp position and pose accuracy.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Gao et al. (2026) studied this question.

synapsesocial.com/papers/69be35166e48c4981c6733f0https://doi.org/10.1016/j.aiia.2026.03.002
Ask AI
Helpful
Bookmark
Share
View Full Paper