Abstract Large language models (LLMs) can support predictive maintenance by reasoning over text-based expert knowledge, but their reliability depends on how such knowledge is structured and retrieved. This work presents an expert-knowledge-grounded LLM inference pipeline that combines Delphi- and FMEA-derived maintenance knowledge, retrieval-augmented prompting, and engineered telemetry summaries. In a multi-robot case study, we evaluate basic knowledge queries, complex diagnostic questions, and telemetry description tasks. Retrieved expert knowledge improves LLM-based diagnostic answers, while a deterministic knowledge-graph reasoner performs best on threshold-driven telemetry questions. The results indicate a complementary design in which graph inference provides traceable rule execution and expert-knowledge-grounded LLMs synthesize diagnostic explanations from selected evidence.
Ren et al. (Thu,) studied this question.