Pre-registered benchmark study reveals failure of four machine learning paradigms to beat naive baselines across five market streams, highlighting the need for auditable forecasting protocols.
Directional forecasting claims about markets are, in common practice, unfalsifiable: the search behind a published strategy is unreported, the benchmark is chosen alongside the results, and failed attempts leave no record. We present a proving protocol that makes such claims checkable without trusting the authors. Every evaluative decision is committed to a cryptographically timestamped append-only record before it can be compared with outcomes. We report the protocol's first completed application: a baseline stage across five market target streams, in which four model paradigms -- from regularized linear models to a large language model -- all failed to exceed floors set by a pre-registered naive family operating on price alone, net of costs. The floors, the statistical machinery (a cumulative hypothesis family under Romano-Wolf stepdown control, validated on synthetic ground truth), and the enforcement patterns that make protocol violations machine-detectable or impossible by construction are published in full. On this record we announce, in registered-report form, the project's next examinations: an echo state network stage, and a quantum reservoir stage evaluated as the performance difference Δ, evaluated two-sided, against an exactly twinned classical arm. To our knowledge this is the first pre-registered forward test of quantum reservoir computing against a matched classical counterpart in financial forecasting. The design invariants are timestamped with this paper; the full designs undergo registered-report review before any run. Their outcomes, and the per-arm grid of the baseline stage, are published in a separate results report under the disclosure policy fixed here, whatever the sign of Δ. Version 1.1 — errata. Editorial corrections only; no change to methods, data, results, or conclusions. A substantive revision (v2) responding to external reviewer comments is in preparation. (1) Terminology: "symmetric difference" replaced by "performance difference Δ, evaluated two-sided" throughout (abstract, §1, §8.1, Fig. 1) [R-E1]. (2) Factual correction, §7.2: the smallest adjusted p-value (0.9388) was described as "three orders of magnitude" above the threshold; against the registry-level per-stream threshold of 0.01 the ratio is roughly 94, just under two orders of magnitude [R-E2]. (3) §7.2: the claim that numeric blinding makes corpus leakage "structurally inapplicable" is weakened to "substantially reduces the risk" [R-E3]. (4) References: four internal editorial notes removed and metadata completed [R-E4]. Concept DOI unchanged. SSRN mirror revised with the same PDF.
No takes yet. Share an insight, caveat, or question.
Timur Karatyhin (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: