We present Epsilon, an autonomous multi-agent research engine that conducts end-to-end scientific investigations, from hypothesis generation to statistical validation, while enforcing epistemic integrity by design. The system introduces a strict architectural separation between a Design Agent and an Evaluation Agent, preventing common sources of bias such as p-hacking, metric switching, and post-hoc hypothesis rewriting. Epsilon incorporates a self-correcting feedback loop for iterative experiment refinement and a three-tier memory architecture (Evidence Memory, Knowledge Memory, and Run Memory) that enables full auditability and cross-run learning. We evaluate Epsilon on MLAgentBench-derived tasks spanning tabular classification, natural language processing, and hyperparameter optimization, demonstrating that autonomous agents can achieve benchmark targets while adhering to rigorous statistical protocols and transparent reporting standards. All experimental artifacts, run logs, and source code are available in the accompanying GitHub repository. All results are fully reproducible using the provided code and logged experiment traces.
Ritvik Jhawar (Mon,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: