PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 7, 2026Applied Sciences2 citationsOpen Access

Large Language and Foundation Models for Machinery Health Monitoring: A Systematic Review

View Full Paper
CTChristos TsallisHAH. AlbrechtRMRadu Munteanu

Key Points

  • This review aims to analyze the impact of language and foundation models on machinery health monitoring and their practical industrial applications.
  • Systematic review of 58 Scopus studies from 2022 to 2026
  • Focus on practical deployment rather than algorithmic performance
  • Analysis aligned with PRISMA 2020 guidelines
  • Mapping the evolution of text-centric models to autonomous agents
  • Addressing challenges like data scarcity and real-time determinism
  • Multimodal models increased out-of-distribution accuracy to 71.95%, compared to 18.25% with conventional models.
  • Achieved ~98% fault diagnosis accuracy with only 1.2% labeled samples.
  • RAG and KGs reduced hallucinations and improved reliability.
  • Digital twins decreased false positive alarms by up to 67%.

Abstract

The rapid adoption of large language models (LLMs) and foundation models is reshaping machinery health monitoring. This shift moves the field beyond task-specific deep learning (DL) toward more generalist and multimodal intelligence. Addressing this transition’s fragmented methodology, this systematic review uniquely shifts the focus from purely algorithmic performance to practical industrial deployment. It achieves this by mapping the evolution of text-centric LLMs into autonomous, cyber-physical industrial agents. Following the preferred reporting items for systematic reviews and meta-analyses (PRISMA) 2020 guidelines, an analysis of 58 Scopus studies published between 2022 and early 2026 was conducted to answer six core research questions (RQs). The synthesized literature demonstrates striking quantitative gains. Multimodal foundation models improve out-of-distribution accuracy to 71.95%, up from 18.25% in conventional models. Furthermore, they achieve approximately 98% fault diagnosis accuracy using merely 1.2% labeled samples. To ensure reliability, integrating retrieval-augmented generation (RAG) and knowledge graphs (KGs) mitigates hallucinations. Meanwhile, autonomous agentic architectures within digital twins (DTs) reduce false positive alarms by up to 67%. Despite generative artificial intelligence (GenAI) tackling data scarcity via synthetic data generation, challenges remain regarding real-time determinism, corpus poisoning, and edge deployment. Ultimately, real-world adoption demands targeted physics-AI hybridization and hardware-embedded DTs over generic compression.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Tsallis et al. (2026) studied this question.

synapsesocial.com/papers/69abc2555af8044f7a4ebde3https://doi.org/10.3390/app16052493
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Beyond the Sensor: A Systematic Review of AI’s Role in Next-Generation Machine Health Monitoring2025 · 12 citations
  2. 2Agentic AI in Smart Manufacturing: Enabling Human-Centric Predictive Maintenance Ecosystems2025 · 20 citations
  3. 3Hybrid Fine-Tuning in Large Language Model Learning for Machinery Fault Diagnosis2024 · 7 citations
  4. 4Leveraging Pre-Trained GPT Models for Equipment Remaining Useful Life Prognostics2025 · 12 citations
  5. 5Large Language Model and Digital Twins Empowered Asynchronous Federated Learning for Secure Data Sharing in Intelligent Labeling2024 · 5 citations