PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
June 4, 2026Electronics0 citationsOpen Access

Foundation Model-Based One-Shot Anatomical Landmark Detection with Mamba and Graph Refinement

View Full Paper
YTYinbing TianZWZiyang WangLGLi Guo

Key Points

  • This work aims to enhance anatomical landmark detection using a one-shot training approach with minimal expert annotations.
  • Utilize a frozen DINO Vision Transformer (ViT) as a backbone for landmark detection.
  • Implement Multi-Layer Multi-Facet module for feature fusion and reweighting.
  • Incorporate Topology-Constrained Graph Refinement to refine landmark configurations based on anatomical graphs.
  • Achieved strong performance on the Cephalometric dataset and Hand X-ray dataset.
  • Showed improved accuracy in landmark detection by integrating multi-source representations and efficient context aggregation.

Abstract

Accurate anatomical landmark detection is important for orthodontic analysis, surgical planning, and morphometric measurement, but fully supervised methods usually require large expert-annotated datasets. This work studies a one-shot setting, where only a single annotated template image is used for training. We propose a foundation-model-based landmark detection framework using a frozen DINO Vision Transformer (ViT) backbone. The proposed framework integrates three complementary components: a Multi-Layer Multi-Facet (MLMF) module that adaptively fuses key and value features from multiple ViT layers through global source-wise reweighting; a Mamba-Based Long-Range Context Aggregation (MLCA) module that injects global anatomical context into fused patch descriptors with linear complexity; and a Topology-Constrained Graph Refinement (TCGR) module that refines the predicted landmark configuration using anatomical graph constraints. Experiments on the Cephalometric dataset and the Hand X-ray dataset demonstrate that the proposed method achieves strong performance. Overall, the results show that jointly exploiting multi-source foundation-model representations, efficient long-range context aggregation, and topology-aware refinement improves annotation-efficient anatomical landmark detection.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Tian et al. (2026) studied this question.

synapsesocial.com/papers/6a211852d499ed480b170e97https://doi.org/10.3390/electronics15112414
Ask AI
Helpful
Bookmark
Share
View Full Paper