Benchmarking evaluation demonstrates high diagnostic accuracy across cloud-operation systems, highlighting the value of multi-agent collaboration with human oversight.
Key Points
To develop and evaluate ChatRCA, a multi-agent framework that structures root cause analysis tasks across specialized large language models while incorporating targeted human feedback.
Decomposed root cause analysis into specialized subtasks assigned to Manager, Observation, Architecture, Operation, and Expert LLM agents based on an empirical study of diagnostic workflows.
Integrated human-in-the-loop checkpoints at work-order verification and root-cause adjudication using a consensus-then-arbitration protocol with three blinded operations engineers.
Evaluated diagnostic performance across three datasets: TrainTicket, a private CMCC cloud-operation dataset, and GAIA.
Achieved root-cause category Top-1 accuracy of 91.11% on TrainTicket, 86.67% on the CMCC cloud dataset, and 87.80% on GAIA.
Generated high-quality root cause explanations on the CMCC dataset, attaining a BLEU-4 score of 63.45 and a BERTScore of 86.92.