PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 19, 2026IEEE Transactions on Medical Imaging7 citations

M2Net: Multimodal Multitask Mutual Learning for Anti-VEGF Efficacy Prediction

View Full Paper
YWYang WenYZYing ZengLBLei Bi

Key Points

  • This research aims to enhance the prediction of treatment efficacy for age-related macular degeneration using M2Net, a multimodal approach.
  • Developed M2Net as a dual-branch network incorporating fundus photographs and OCT scans.
  • Implemented a Multimodal Collaborative Treatment Efficacy Prediction module to classify visual acuity changes before image generation.
  • Created a dataset (MMPD) with paired multimodal retinal images and corresponding visual acuity measurements.
  • Achieved classification accuracy of 96.03% for visual acuity prediction.
  • Obtained Structural Similarity Index (SSIM) scores of 0.6377 for OCT and 0.8347 for fundus images.
  • Demonstrated superior performance over existing prediction methods.

Abstract

Age-related macular degeneration with abnormal blood vessel growth (neovascular AMD) is the leading cause of vision loss in elderly populations. While anti-VEGF injections are the standard treatment, they present financial burdens for patients and vary in effectiveness. Predicting treatment efficacy is therefore crucial for patient care. Current prediction methods fail to fully integrate information from different imaging techniques, typically focusing on either forecasting vision improvements or generating post-treatment images-but not both simultaneously. This approach overlooks the important relationship between these tasks. We present M2Net, a novel joint generation and classification network based on Multimodal Multitask Mutual learning, to simultaneously predict changes in visual acuity and generate post-treatment retinal images. M2Net employs a dual-branch structure that processes both fundus photographs and Optical Coherence Tomography (OCT) scans to improve prediction accuracy. Our framework includes two key innovations: the Multimodal Collaborative Treatment Efficacy Prediction module, which interacts the features between the two modalities and provides initial visual acuity change classification to guide the generation of post-treatment images; and the Pre-Post Treatment Image Joint Analysis module, which identifies both common and changing features between pre-treatment and post-treatment images to enhance prediction accuracy. To validate our approach, we created the dataset (MMPD) containing paired multimodal retinal images with corresponding visual acuity measurements. Experiments on the dataset demonstrate that M2Net achieves superior performance compared to existing methods, with a classification accuracy of 96.03%, an SSIM of 0.6377 on the OCT modality, and an SSIM of 0.8347 on the fundus modality. Our code will be available at https://github.com/zengying123/M2Net.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Wen et al. (2026) studied this question.

synapsesocial.com/papers/69e47193010ef96374d8de14https://doi.org/10.1109/tmi.2026.3684331
Ask AI
Helpful
Bookmark
Share
View Full Paper