This article examines a model's fraud detection performance and regulatory compliance, suggesting improved transparency.
This article examines whether a high-performing fraud detection model can also meet the demands of auditability, documentation, and regulatory transparency. Using the publicly available European credit card fraud dataset of 284,807 transactions, including 492 fraudulent cases, this study compares weighted logistic regression with XGBoost under severe class imbalance. Model performance is assessed through precision, recall, F1 score, ROC AUC, and precision–recall AUC, with particular attention to alert burden and fraud capture. Results show that XGBoost materially outperforms logistic regression in operational terms. While logistic regression achieves slightly higher recall, XGBoost raises precision from 0.061 to 0.562, improves PR AUC from 0.719 to 0.863, and reduces false positives from 1386 to 67. The PR AUC of 0.863 refers to the cross-validated average reported in the model comparison, while the holdout test result reported later in this paper is 0.852. It cuts the review queue from 1476 alerts to 153 while still identifying 86 of 98 fraud cases in the test set. Explainability is then introduced through SHAP, which provides both global feature attribution and transaction-level reasoning. The findings show that SHAP makes the boosted model readable at the level of both overall model behaviour and individual fraud flags, thereby supporting audit review, model validation, and regulatory scrutiny. The article argues that the combination of XGBoost and SHAP offers a stronger fit for auditing than either a weaker but transparent linear model or a stronger opaque classifier. One limit remains, since the dataset contains anonymised principal components rather than original business variables, which restricts semantic interpretation. Even so, the workflow provides a practical bridge between predictive fraud analytics and the demands of explainable, reviewable, and accountable AI in auditing.
No takes yet. Share an insight, caveat, or question.
Alessio Faccia (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: