Reliable cloud-based storage systems require accurate failure prediction for Solid-State Drives (SSDs) and Hard Disk Drives (HDDs) to reduce data loss, enable proactive maintenance, support service-level reliability, and lower operational costs. In this survey, we review over 150 prior studies on storage failure prediction and related tasks, and provide a structured overview and evaluation of currently available techniques for storage failure prediction, spanning traditional statistical methods, machine learning, and deep learning approaches. We focus on device-level predictions and compare the performance, constraints, and implementation overhead of prior works in real-world scenarios. Challenges such as data imbalance, fail-slow degradation, and evolving failure patterns are discussed to identify current research gaps, such as the limited interpretability of advanced models, and the need for standardized benchmarks. Our main contribution is the introduction of structured decision frameworks that guide practitioners to choose suitable evaluation metrics, predictive models, and data preparation methods based on certain operational scenarios. These frameworks are complemented by comparative analysis of models, evaluation metrics, interpretability methods and computational overhead across deployment contexts. Our survey discusses open challenges and research directions in the domain, and offers useful insights and a structured methodology for translating research into practical deployment strategies.
Chakraborttii et al. (Sat,) studied this question.