This paper proposes a hybrid machine learning framework for detecting electricity fraud within the broader context of Non-Technical Losses (NTLs) in power-distribution systems. The framework combines unsupervised anomaly detection using Isolation Forest with supervised classification through XGBoost, exploiting the complementary strengths of both algorithms. Using real consumption data from a Peruvian utility, the approach integrates domain-informed feature engineering to capture behavioral, temporal, and contextual indicators of irregular usage. To address the extreme class imbalance inherent to fraud datasets, the SMOTETomek hybrid resampling technique was applied, enhancing minority-class representation and decision boundary clarity. Experimental results achieved high predictive performance on the test set (AUC-ROC = 0.999, F1-score = 0.77) using an optimized decision threshold of 0.6. Moreover, SHAP-based interpretability analysis identified extreme monthly variations, prolonged low-consumption periods, and tariff category as key behavioral predictors of fraudulent activity. The robustness of the proposed framework was further validated through a 5-fold cross-validation procedure during the training phase, ensuring consistent performance across different data partitions. Overall, the proposed framework demonstrates not only robust and explainable performance but also practical operational value, providing utilities with a scalable data-driven tool to optimize inspection strategies and maximize recovery of non-technical losses.
Helder Chávez (Fri,) studied this question.