We test whether machine learning can improve setup selection for a discretionary intraday strategy on the Nasdaq-100 future (NQ). Three pre-registered milestones (bar features, order flow from 62 million real ticks on a stratified day sample, and a model boosted from the geometric prior) find that the model does not beat the prior p = 1/(1 + rr) — the gambler's-ruin probability — beyond its calibration. A later review identified a systematic artefact in the comparison metric: labelling time-stopped trades as full losses compares a finite-horizon frequency with an infinitehorizon probability. Those exits are worth +0.605 R on average, not −1; the bias shifts expected value by 0.13 R per trade and selectively penalises wide-stop geometries. The corrected form needs no volatility or horizon estimate: by the optional stopping theorem, under a martingale the expectation of any bounded exit rule is zero, so the gross mark-to-market return is simply tested against zero. Re-measured, the setup is not below the random walk (gross EV −0.017 R, CI containing zero) but is made unprofitable by costs (−0.093 R). In the near-target band a genuine drift appears, +0.038 R [+0.012, +0.063], separated from microstructure by a mirror test and increasing with excursion depth (−0.9 to +2.0 points per quintile). A screen of nine intraday setups with the corrected instrument (53 confirmatory tests, false-discovery-rate control) yields no discovery; the apparently best candidate — the overnight hold — vanishes once the null includes the measured market drift (0 of 4 geometries, p = 1.000), because a long-only strategy on a rising index beats a zero-drift null by construction. The main contribution is methodological: the artefact, its correction, the mirror test and the drift-adjusted null transfer to any strategy with a time-based exit. A final review conducted with a second model identified two false positives in subsequent exploratory measurements, both arising from sample construction rather than analysis. The first credited the setup with a +0.065 R timing advantage over a random entry: the control group copied the side decided at the signal and applied it to earlier instants, and against a control free of this anticipation the advantage falls to +0.0008 ATR. The second, +0.149 ATR, came from a sample that weighted days by the number of value-area re-entries, i.e. selected on the outcome being measured: on the declared population the value is −0.066 in training and +0.020 in the extension, with opposite signs. Both defects were identified from the description of the procedure alone, before any number was seen. The out-of-sample window 2023-10 → 2026-03 was never opened. Part II (version 4.0). The work continued as an open laboratory: seven rounds, 24 hypothesis families, 5,013 logged tests and 480,181 automatically generated rules, judged on three time blocks by a controller that opens validation and confirmation only once per candidate. No directional strategy survives: the best of half a million intraday rules is indistinguishable from the best of 700 placebo worlds (corrected p 0.31), and the only strategy that passed validation (+12.1 net points per trade) returns −2.6 in confirmation. Two volatility regularities are confirmed in all three blocks: a HAR model forecasts tomorrow's variance better than a moving average, and the overnight range anticipates the day range. A power analysis of the verdicts shows that many negative results mean "undecidable" rather than "absent": confirmation of the best strategy had 10% power against half the validation estimate. Machine learning applied where signal exists (range quantiles) does worse than a linear regression (−3.2%), and the comparison between the volatility forecast and option prices cannot be decided with 2010-2018 data. Two tools follow — a fixed-coefficient expected-range indicator and a backtest auditor — reported with their measured error rates.
No takes yet. Share an insight, caveat, or question.
Giovanni Febo (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: