PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 26, 2026Biomedicines0 citationsOpen Access

Multi-Patient Vision Transformer for Markerless Tumor Motion Forecasting

View Full Paper
GHGauthier Rotsart de HertaingDMDani ManjahBMBenoît Macq

Key Points

  • The aim is to improve the accuracy of lung tumor motion forecasting using a multi-patient model with vision transformers.
  • Used digitally reconstructed radiographs from 4D CT scans of lung cancer patients to train a multi-patient model.
  • Compared patient-specific models trained on planning data with the multi-patient model.
  • Fine-tuned the multi-patient model using limited patient-specific treatment images.
  • Evaluated model performance using root mean square error on sequences of 12 DRRs.
  • Low-resolution input images with larger patch sizes provided better performance by reducing noise.
  • Patient-specific models require more extensive data to achieve competitive performance compared to multi-patient models.
  • Fine-tuning with limited data yielded comparable or improved forecasting accuracy for the multi-patient model.

Abstract

Background: Accurate forecasting of lung tumor motion is crucial for precise radiotherapy. Deep-learning-based markerless tracking methods have been explored, but extending these approaches to predict future tumor trajectories remains largely unaddressed. We address this by framing markerless lung tumor motion forecasting as a spatio-temporal prediction task using a vision transformer to estimate three-dimensional tumor positions over short horizons. Methods: Digitally reconstructed radiographs (DRRs) generated from four-dimensional computed tomography scans of 12 lung cancer patients were used to train a multi-patient (MP) model. Patient-specific (PS) models trained solely on planning data were compared, and the MP model was further fine-tuned using a small number of patient-specific treatment images under realistic clinical constraints. Models processed sequences of 12 DRRs, with performance evaluated via root mean square error. Results: The results indicate that low-resolution inputs with larger patch sizes outperform higher-resolution configurations by reducing image noise. PS models require extensive data to match MP performance, whereas fine-tuning the MP model with limited patient-specific data achieves comparable or superior forecasting accuracy at a lower cost. Conclusions: These findings demonstrate that Vision Transformers can extend markerless tracking methods to accurate short-term forecasting and highlight fine-tuning as an efficient strategy for personalized prediction.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Hertaing et al. (2026) studied this question.

synapsesocial.com/papers/699fe3d995ddcd3a253e7ec8https://doi.org/10.3390/biomedicines14030496
Ask AI
Helpful
Bookmark
Share
View Full Paper