As sensor-rich systems become increasingly prevalent, data-driven applications face growing pressure to balance information richness against practical constraints, such as storage, communication, and computational cost. This tension raises a fundamental question: how much data is enough? The challenge is particularly critical in anomaly detection (AD), where early identification of subtle deviations can prevent failures and safety incidents, yet excessive data collection may be infeasible. In this paper we study this tradeoff in the context of battery storage systems, a representative and safety-critical industrial application. Battery packs are composed of large numbers of electrically connected cells and can be monitored either at fine-grained cell level or using aggregated pack-level measurements. While cell-level monitoring promises higher fidelity, most real-world systems rely on aggregated pack-level measurements due to scalability limitations. We systematically investigate whether access to fine-grained cell-level data consistently improves AD performance, or whether an intelligent selection of a small number of pack-level variables can achieve comparable results. We benchmark state-of-the-art deep learning-based AD methods against feature-engineering-based approaches under both data granularity assumptions. Furthermore, leveraging a highly flexible battery pack simulation framework, we analyze systematic differences between synthetic and real-world battery data, highlighting reality gaps that are often overlooked in AD research. Our results offer practical insights into data efficiency, model selection, and monitoring design, with implications that extend beyond batteries to a wide range of industrial condition monitoring applications.
Wüest et al. (Thu,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: