The rapid adoption of large language models (LLMs) and foundation models is reshaping machinery health monitoring. This shift moves the field beyond task-specific deep learning (DL) toward more generalist and multimodal intelligence. Addressing this transition’s fragmented methodology, this systematic review uniquely shifts the focus from purely algorithmic performance to practical industrial deployment. It achieves this by mapping the evolution of text-centric LLMs into autonomous, cyber-physical industrial agents. Following the preferred reporting items for systematic reviews and meta-analyses (PRISMA) 2020 guidelines, an analysis of 58 Scopus studies published between 2022 and early 2026 was conducted to answer six core research questions (RQs). The synthesized literature demonstrates striking quantitative gains. Multimodal foundation models improve out-of-distribution accuracy to 71.95%, up from 18.25% in conventional models. Furthermore, they achieve approximately 98% fault diagnosis accuracy using merely 1.2% labeled samples. To ensure reliability, integrating retrieval-augmented generation (RAG) and knowledge graphs (KGs) mitigates hallucinations. Meanwhile, autonomous agentic architectures within digital twins (DTs) reduce false positive alarms by up to 67%. Despite generative artificial intelligence (GenAI) tackling data scarcity via synthetic data generation, challenges remain regarding real-time determinism, corpus poisoning, and edge deployment. Ultimately, real-world adoption demands targeted physics-AI hybridization and hardware-embedded DTs over generic compression.
Tsallis et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: