Financial fraud in listed companies has attracted increasing academic attention due to its severe implications for investors and market stability. Traditional methods for detecting fraudulent financial statements have proven insufficient in addressing the growing complexity and volume of financial data. This study proposes a novel hybrid model combining recursive feature elimination with cross-validation and parallel random forest (RFECV-PRF) for feature selection and a genetic algorithm-optimized LightGBM (GA-LightGBM) for financial fraud detection. The RFECV-PRF method effectively evaluates feature importance and selects the optimal subset of financial indicators, while the GA-LightGBM model enhances prediction accuracy and efficiency through global optimization of hyperparameters. Empirical tests using data from five industries—energy, materials, industrials, information technology, and healthcare—demonstrated significant improvements in classification accuracy, precision, recall, and F1 scores. The proposed framework outperforms traditional machine learning approaches such as random forest and XGBoost, achieving superior results in detecting fraudulent financial activities. This research highlights the potential of integrating advanced feature selection methods with optimized machine learning models to address the challenges of financial fraud detection and improve the reliability of corporate financial reporting.
Cui Hu (Sun,) studied this question.