Empirical evaluation reveals limited net abnormal returns from daily machine learning stock rankings in Nasdaq equities, indicating trading costs eliminate statistical predictive advantages.
This study examines whether daily machine learning stock rankings based on technical information contain out-of-sample ordering information and whether that information can be converted into economically implementable returns. Using a dynamically screened Nasdaq source universe from 2021 to 2026, four XGBoost objectives are evaluated in a chronological walk-forward design. Test NDCG converges to 0.495–0.504, but permutation analysis places the corresponding random-ranking mean near 0.45, indicating statistically detectable but modest cross-sectional ordering information. Economic performance is substantially weaker. Under the execution convention implied by the next-day open-to-close target, every invested portfolio is bought at the open and liquidated at the close, so round-trip turnover equals two. Pseudo-Huber Top-1, treated as an ex-post concentration diagnostic, produces a 34.9% gross annual geometric return with 89.5% volatility and an 86.5% maximum drawdown; the Newey–West mean-return test is not significant (p = 0.106). At five basis points per trading leg, its zero-cash net CAGR falls to 4.8%; crediting idle capital with the daily risk-free rate raises total-return CAGR to 9.3%, but the excess-return inference is unchanged (p = 0.297). A matched-horizon regression on SPY open-to-close returns yields a statistically insignificant net annualized alpha (p = 0.436). Hansen’s SPA test across the synchronized 12-strategy family gives p = 0.207, and the Deflated Sharpe Ratio probability for Pseudo-Huber Top-1 is 0.462. The result is also highly time- and tail-dependent. The evidence therefore supports a distinction between statistically detectable ranking information and robust implementable abnormal performance rather than a persistent trading anomaly.
No takes yet. Share an insight, caveat, or question.
Ferdinantos Kottas (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: