• An ensemble feature extraction is presented as a strategy to address the information security challenge commonly faced during designing ML-ADSs for IEC 61850 GOOSE communications. • An in-depth analysis of IEC 61850 GOOSE packet features is provided to identify the optimal input feature subset, examine the output, and use these to evaluate the effectiveness of our proposed strategy. • The proposed strategy is evaluated to show that it can enhance ML-ADSs performance in detecting abnormal GOOSE patterns. The evaluation process includes two main phases. First, a one-class classifier (OCC) baseline is evaluated by employing the LOF algorithm. Then, the DNN and Autoencoder are evaluated under different test scenarios. IEC 61850 is a powerful standard for digital substations due to its fast transmission of information, high efficiency, and robust interoperability. However, the security of IEC 61850 communication protocols poses a critical challenge, potentially leading to disruptions and actual blackouts. For example, three Ukrainian electric companies were compromised through phishing attacks, resulting in blackouts that affected at least 230000 people. However, detecting and preparing for such cyber-attacks against IEC 61850-based substations remains challenging, due to Machine learning-based anomaly detection systems (ML-ADSs) have been shown to enhance information security in a digital substation environment. An effective detection strategy relies on extracting optimal input feature sets and examining the ADS output; however, to the best of our knowledge, no existing study has considered their impacts in the realm of feature engineering and selection for anomaly GOOSE detection. To fill this gap, we propose a novel strategy based on ensemble-based feature extraction, a strategy that extracts a high-dimensional feature space to derive the most informative feature subset. Conducting an in-depth analysis of IEC 61850 GOOSE packet features, we show that an optimal feature list can significantly improves ADS performance in detecting abnormal GOOSE patterns by prioritizing protocol-aware feature engineering. Using different GOOSE feature configurations, we show that the performance of machine learning models is enhanced compared with the baseline test. The results reveal that the Autoencoder achieved higher accuracy compared with the initial configurations, from 98.95% to 100%. After analyzing the most selective attributes, we conclude that duration and inter-arrival times are solid evidence for distinguishing attacks from benign patterns, and there are persistent timing characteristics in spontaneous traffic. Collectively, this work supports the improvement of reliable ADS within substations, where data security is a problem in IEC 61850 substation cybersecurity.
Yasseen et al. (Wed,) studied this question.