PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 18, 2026Industrial Lubrication and Tribology0 citations

MFT: Multi-scale fusion and text-guided segmentation for interactive wear debris analysis

View Full Paper
YHYongqiu HuangQZQinghua ZhangXSXinfa Shi

Key Points

  • The research aims to enhance interactive segmentation of ferrography images through a transformer-based model incorporating textual cues.
  • Developed the Multi-scale Fusion and Text-guided segmentation (MFT) model
  • Implemented a cross-modal fusion module to integrate text and visual features
  • Used two sequential decoders to improve visual-text alignment and segmentation accuracy
  • Employed a label enhancement strategy to reduce annotation costs using pretrained models and MLLMs
  • Evaluated MFT on a wear debris dataset with automated mask and description generation.
  • Achieved 72.20% mean Intersection over Union and 80.52% Prec@0.5
  • MFT showed more accurate segmentation than other transformer methods
  • Descriptions generated by MLLMs improved guidance over simple category names
  • Significantly reduced annotation costs while enhancing segmentation quality

Abstract

Purpose This study aims to apply a transformer-based network to achieve language-guided interactive segmentation for the automated analysis of ferrography images, while also alleviating the high cost of manual annotation. Design/methodology/approach To tackle the challenges of visual-linguistic alignment in referring image segmentation (RIS) and the complexity of ferrography images, a model named Multi-scale Fusion and Text-guided segmentation (MFT) is proposed. MFT injects textual cues into multi-scale visual features via a cross-modal fusion module. It then uses two sequential decoders to enhance cross-scale interaction and refine visual-text alignment for accurate segmentation. MFT is trained and evaluated on a wear debris data set with 11 categories. To reduce its annotation costs, a label enhancement strategy is introduced as a by-product. It leverages a pretrained segmentation model and multi-modal large language models (MLLMs) to automatically generate fine-grained masks and appearance-based descriptions from bounding box-annotated images, providing MFT mask-level supervision and rich textual guidance. Findings MFT achieves more accurate segmentation from referring expressions than other transformer-based methods. Moreover, MLLMs-generated descriptions guide the model more effectively than using category names alone. Originality/value MFT enables accurate interactive segmentation via multi-scale feature fusion and two sequential decoders, guided by either category names or appearance-based descriptions – especially effective with the latter. It achieves 72.20% mean Intersection over Union and 80.52% Prec@0.5 with only 124.6 M parameters, demonstrating competitive accuracy and efficiency. Combined with the enriched label, it also reduces annotation costs and improves segmentation quality. Peer review The peer review history for this article is available at: https://publons.com/publon/10.1108/ILT-08-2025-0356/

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Huang et al. (2026) studied this question.

synapsesocial.com/papers/69e3203440886becb653f491https://doi.org/10.1108/ilt-08-2025-0356
Ask AI
Helpful
Bookmark
Share
View Full Paper