State-of-the-art image matching methods have shown strong generalization across natural image datasets, but their effectiveness in complex surgical environments remains underexplored. Surgical scenes introduce unique challenges, including homogeneous tissue textures, variable lighting, and frequent occlusions, which can degrade the reliability of keypoint correspondences essential for downstream vision tasks such as camera pose estimation and structure-from-motion. In this study, we systematically evaluate leading image matching methods within laparoscopic surgical settings, emphasizing performance under resource-constrained conditions. We present an optimized evaluation pipeline that incorporates robust estimators to enhance correspondence filtering and assess their impact on pose estimation accuracy. Our approach also examines the influence of fine-tuning individual pipeline components, particularly robust estimators, on overall system performance. Mean Reprojection error is refined by thresholding the nearest ground truth projections, enabling a more precise characterization of matching accuracy. Across five robust estimators, FM₈PTS consistently demonstrates superior resilience to outliers. Our results establish RoMa as the leading model for balancing pose estimation accuracy, reprojection performance, and computational efficiency, making it suitable for real-time surgical applications. By providing the first systematic benchmark and actionable insights for optimizing image matching pipelines in surgical domains, this work sets a new standard and paves the way for more reliable, efficient, and clinically applicable image-guided tools in minimally invasive surgery.
Tan et al. (Mon,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: