This paper explores the effectiveness of various machine learning algorithms in predicting financial fraud and recidivism using a hand-collected dataset of cases prosecuted by UK financial regulators from 2010 to 2022. The study aims to identify key factors for predicting financial fraud and recidivism. Comparative analysis of machine learning algorithms, including Logistic Regression, Ridge Regression, Support Vector Machine, Decision Tree, Random Forest, Artificial Neural Network, and Adaptive Least Absolute Shrinkage and Selection Operator, reveals that the Random Forest model consistently outperforms others in AUC and precision rate for both financial fraud and recidivism. The study identifies crucial factors, including financial factors such as leverage, firm size, market value, and tangibility, and non-financial factors such as firm age, tone, gender ratio, and the number of directors, contributing significantly to financial fraud detection. The insights provide valuable guidance to accountants, independent directors and regulators for developing effective early warning systems for financial fraud.
Wang et al. (Mon,) studied this question.