Accurate and efficient segmentation of interventional instruments such as catheters, guidewires, and vessels is essential for the development of robot-assisted endovascular procedures. In this paper, we present a Multi-Task Learning (MTL) segmentation model that simultaneously extracts binary masks for all three structures from 512×512 grayscale intervention images. Our model combines a U-Net-style shared encoder with three task-specific decoders, enabling flexible handling of overlapping anatomical regions—an inherent limitation in traditional multi-class segmentation frameworks such as nnUNet. We evaluate our model on a private dataset and compare it with the nnU-Net baseline in terms of IoU, F1-score, and inference speed. Results show that our approach achieves comparable segmentation accuracy while significantly improving inference efficiency. We further introduce a lightweight variant inspired by YOLOP, which achieves over fivefold speed improvement with acceptable performance loss. These findings demonstrate the advantages of MTL for real-time segmentation in minimally invasive image-guided procedures.
zhijun bian (Fri,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: