In the context of minimally invasive spine surgery, accurately estimating the 3D coordinates of the vertebrae from intraoperative 2D X-ray images is crucial for aligning preoperative data with the patient’s real-time posture. However, existing methods are hindered by the ill-posed nature of 2D-to-3D localization and the distinctive anatomical features of the spinal column, leading to ambiguities and reduced accuracy. In this paper, we introduce X2P-net, a novel prompt-guided and semantic context-enhanced 2D/3D vertebra detection framework. To achieve this, we design a novel Transformer architecture, referred to as BrickFormer, which can automatically extract the refined vertebral foreground context at low computational cost using a dual-attention mechanism. Comprehensive experiments were conducted to validate the proposed approach on two datasets: a large-scale synthetic dataset (BiSpineX) and a sheep spine dataset (SheepSpineX). Results obtained from these experiments demonstrate superior landmark localization performance of the proposed method compared to other state-of-the-art methods. Specifically, on the BiSpineX dataset, X2P-Net achieves percentages of 96.9% and 98.8% at 10 mm and 20 mm thresholds, respectively, a mean position error of 2.99 mm, and an AUC of 0.9923. Similar superior performance was also observed when the proposed method was applied to the SheepSpineX dataset, with percentages of 98.4% and 100.0% at 10 mm and 20 mm thresholds, respectively, a mean position error of 1.08 mm, and an AUC of 0.9972.
Tao et al. (Tue,) studied this question.