Autonomous AI research agents have so far been demonstrated in single runs lasting hours, too short to accumulate signal across a biomedical research programme. We present ARIA, a multi-agent system that conducts sustained autonomous scientific research across multi-week deployments, and report a 50-day continuous deployment in retinal imaging. ARIA combines a weighted-pool idea queue, an adversarial cross-model critique layer, and a hybrid local-to-cloud compute path. Running unsupervised on the NIH Bridge2AI AI-READI v3. 0 cohort, it designed, executed and analyzed experiments on retinal fundus images and paired clinical biomarkers, evaluating linear probing regression heads over a frozen convolutional feature encoder. The resulting models establish an exploratory cross-platform age-regression generalization map: a ConvNeXt-Tiny model trained on the union of three fundus cameras (iCare Eidon, Topcon Triton, Optomed Aurora) reached r = 0. 594 (paired-bootstrap 95% CI 0. 524, 0. 660) across 309 subjects imaged on all three devices, against a single-platform baseline of r = 0. 771 ± 0. 038 (Eidon, n = 2, 120, 20-shuffle Monte Carlo cross-validation). A paired Steiger Z-test detected no device-specific advantage within the cross-platform cohort (p > 0. 5). We report the cross-versus-single gap descriptively rather than as a confirmatory test, because the two estimators differ in cohort size, training distribution and resampling structure. Across a wider 18-week window spanning seven instances and four scientific domains, ARIA generated 19, 364 commits and approximately 17, 800 autonomous sessions, achieved a 97. 8% autonomous resumption rate (314 of 321 recovery events), and ran 5, 196 consecutive commits without human intervention. Total cloud compute spend was 86. 92 across 900 preemptible runs. The kernel framework is released under the Apache 2. 0 license.
Johnson et al. (Wed,) studied this question.