Benchmarking study demonstrates that machine learning models match but do not surpass bookmaker odds in predicting football match outcomes, indicating high betting market efficiency.
Football match outcomes are influenced by a complex interplay of dynamic team strategies and stochastic match events, posing a significant challenge for predictive analytics. In this paper, we present a unified benchmarking study for football match outcome classification, using a dataset of 225,474 matches spanning 2018–2026 and a strictly chronological train–test split. Six classifiers spanning several modelling approaches, Logistic Regression, Generalized Additive Models, FastTree, Random Forest, Radial Basis Function networks and Multi-Layer Perceptron, are trained and evaluated under identical conditions. Bookmaker odds are converted into margin-free probabilities by the power method, with the exponent found by the Newton–Raphson method, and are included among the input features. In addition to the six classifiers, we also report a majority-class baseline and the de-vigged bookmaker prediction itself. Results are reported for both a binary one-vs-rest formulation and the original three-class (1, X, 2) formulation, with 95% bootstrap CIs for every metric and every model. Across the six classifiers, the differences are small: macro-averaged Precision ranges from 64.83% to 65.16% for Home Win and from 68.05% to 68.46% for Away Win, and macro-averaged Recall from 63.16% to 63.51% and from 59.10% to 59.72%, respectively, with substantially overlapping confidence intervals. In the three-class formulation, models trained without market data reach 48.0% accuracy against 43.4% for a trivial baseline and 51.9% for the bookmaker. Models trained with market data match the bookmaker but do not improve upon it on log-loss, Brier score or the ranked probability score. The only statistically distinguishable improvement of any kind is a 0.086 point accuracy advantage for the additive model, which is not accompanied by any improvement in the proper scoring rules and therefore does not indicate a practical advantage. Probability calibration and betting-signal generation are outside the scope of this study.
No takes yet. Share an insight, caveat, or question.
Sallis et al. (2026) studied this question.
Synapse has enriched 3 closely related papers on similar clinical questions. Consider them for comparative context: