PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 30, 2026Quality & Quantity1 citationsOpen Access

Risk-based predictive modelling for audit verification: evidence from EU-funded programmes

View Full Paper
EVElisa VernaGGGianfranco GentaMGMaurizio Galetto

Key Points

  • This research aims to develop a framework using machine learning for risk-based verification in EU fund audits.
  • Utilized an imbalanced three-class classification task with ordered outcomes
  • Trained on over ninety thousand documents from an Italian Operational Programme
  • Employed the CatBoost gradient-boosting algorithm to address methodological challenges
  • The model provided interpretable probability estimates for validated, partially validated, and not validated cases
  • Achieved satisfactory predictive performance, aiding resource allocation in audits
  • Variable-importance analysis highlighted the significance of financial and administrative variables in predicting irregularities

Abstract

This study proposes a machine learning framework to support risk‑based verification of expenditure declarations in European Structural and Investment Funds, reflecting the current regulatory emphasis on proportional and data‑driven audit strategies. Quantitatively, the problem is formulated as an imbalanced three-class classification task with ordered outcomes on high-dimensional administrative data; the ordinal structure is exploited ex post in evaluation and error interpretation. The framework classifies expense documents as validated, partially validated, or not validated, and provides audit authorities with interpretable probability estimates for each case. A predictive model was trained and validated on more than ninety thousand expense documents from the Italian Regional Operational Programme co‑funded by the European Regional Development Fund (2014–2020). Methodological challenges—ordered outcomes, severe class imbalance, and mixed‑type features—were addressed through targeted preprocessing and the CatBoost gradient‑boosting algorithm. The model achieved satisfactory predictive performance, offering probabilistic outputs aligned with the ordered structure of audit outcomes. Variable‑importance analysis confirmed the relevance of both financial and administrative variables in predicting irregularities. The framework is designed with operational integration in mind and could underpin risk‑based sampling in expenditure verification, subject to further validation across time, programmes, and beneficiary structures. Departing from a literature that largely focuses on binary classification or fraud detection, the study addresses the understudied challenge of multi‑class prediction in public expenditure control and provides an interpretable prototype decision‑support tool. The model could support public authorities in prioritizing controls and allocating resources more efficiently, contributing to the modernization of European Union fund management and promoting data‑driven, proportionate oversight—conditional on governance arrangements and external validation.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Verna et al. (2026) studied this question.

synapsesocial.com/papers/69c9c5c5f8fdd13afe0bdb82https://doi.org/10.1007/s11135-026-02725-x
Ask AI
Helpful
Bookmark
Share
View Full Paper