Three-dimensional surface reconstruction is essential for accurately acquiring the external quality parameters of watermelons, such as size, volume, and defect area. Binocular stereo vision provides a low-cost and easily deployable solution for the single-view 3D reconstruction of watermelons. However, watermelons present highly similar surface textures, and as typical spheroid-like objects, the excessive angle between surface normals of edge regions and the camera optical axis leads to insufficient feature representation. Consequently, directly applying existing stereo matching algorithms often introduces matching ambiguities, and lightweight networks struggle to balance real-time performance with matching accuracy. This study focuses on the high-precision single-view point cloud generation of Kirin watermelons. To address these issues, we first construct a cross-modal, high-precision Kirin watermelon stereo matching dataset. Building upon the Fast-ACVNet+ architecture, we then propose MI-ACVNet, a lightweight stereo matching network tailored for high-precision watermelon point cloud acquisition. In the feature extraction stage, a Multi-Scale Stereo Feature Extraction (MSFE) module is adapted. By incorporating the re-parameterized network MobileOne and Epipolar-Enhanced Coordinate Attention (E2CA), MSFE improves the discriminative capability for weak and similar textures without compromising inference speed. For cost computation, a Coarse-to-Fine Cascaded Residual Correction (C2F-CRC) strategy is incorporated to construct a fine-grained cost volume via sub-pixel interpolation, enhancing the network’s ability to capture subtle surface fluctuations. Furthermore, a Semantics-Guided Region-Aware Loss (SGRA-Loss) is formulated, leveraging semantic masks to apply differentiated supervision weights across edge, center, and background regions to significantly improve edge matching accuracy. Ablation studies validate the effectiveness of the MSFE, C2F-CRC, and SGRA-Loss components. Compared to the baseline model, the full MI-ACVNet reduces the End-Point Error (EPE) by 19.5% and the Bad-0.5 error rate by 34.5% in the watermelon region. Furthermore, when compared against five mainstream algorithms (StereoNet, AANet, HSMNet, LightStereo-L, and NMRF-swint), MI-ACVNet achieves state-of-the-art performance: EPE and Bad-0.5 are reduced to 0.091 pixels and 1.159%, respectively, with a single-frame inference time of only 46 ms. The average depth error of the reconstructed point clouds is merely 0.26 mm. By ensuring both real-time efficiency and high-precision depth estimation, this method demonstrates promising potential for deployment in industrial Kirin watermelon sorting lines, driving sorting equipment toward higher precision and intelligence.
Li et al. (Mon,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: