PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 12, 2025ACM SIGIR Forum3 citations

Evaluating Sakana's AI Scientist: Bold Claims, Mixed Results, and a Promising Future?

View Full Paper
JBJoeran BeelBerkeley CollegeMKMin‐Yen KanNational University of SingaporeMBMoritz BaumgartUniversity of Siegen

Key Points

  • AI Scientist struggles with novelty assessments, incorrectly classifying established concepts as new, which undermines its credibility.
  • Only 5 out of 34 citations were from 2020 or later, highlighting the AI's inability to leverage recent research effectively.
  • 42% of attempted experiments failed due to coding errors, indicating serious issues with the AI's robustness in experiment execution.
  • The AI's ability to produce manuscripts reflects an undergraduate's work quality, suggesting a significant gap in output and human researcher expectations.

Abstract

Recently, Sakana. ai introduced the AI Scientist, a system claiming to automate the entire research lifecycle and conduct research autonomously, a concept we term Artificial Research Intelligence (ARI). Achieving ARI would be a major milestone toward Artificial General Intelligence (AGI) and a prerequisite to achieving Super Intelligence. The AI Scientist received much attention in the academic and broader AI community. A thorough evaluation of the AI Scientist, however, had not yet been conducted. 1 We evaluated the AI Scientist and found several critical shortcomings. The system's literature review process is inadequate, relying on simplistic keyword searches rather than profound synthesis, which leads to poor novelty assessments. In our experiments, several generated research ideas were incorrectly classified as novel, including well-established concepts such as micro-batching for stochastic gradient descent (SGD). The AI Scientist also lacks robustness in experiment execution—five out of twelve proposed experiments (42%) failed due to coding errors, and those that did run often produced logically flawed or misleading results. In one case, an experiment designed to optimize energy efficiency reported improvements in accuracy while consuming more computational resources, contradicting its stated goal. Furthermore, the system modifies experimental code minimally, with each iteration adding only 8% more characters on average, suggesting limited adaptability. The generated manuscripts were poorly substantiated, with a median of just five citations per paper—most of which were outdated (only five out of 34 citations were from 2020 or later). Structural errors were frequent, including missing figures, repeated sections, and placeholder text such as "Conclusions Here". Hallucinated numerical results were contained in several manuscripts, undermining the reliability of its outputs. Despite its limitations, the AI Scientist represents a significant leap forward in research automation. It produces complete research manuscripts with minimal human intervention, challenging conventional expectations of AI-generated scientific work. Many reviewers or university instructors conducting only a superficial assessment may struggle to distinguish its output from that of human researchers, demonstrating how far AI has progressed in mimicking academic writing and structuring scientific arguments. While the quality of its manuscripts currently aligns with that of an unmotivated undergraduate student rushing to meet a deadline, this level of autonomy in research generation is remarkable. More strikingly, it achieves this at an unprecedented speed and cost efficiency—our analysis indicates that generating a full research paper costs only 6–15, with just 3. 5 hours of human involvement. This is significantly faster than traditional human researchers. Given that AI research automation was nearly nonexistent just a few years ago, the AI Scientist marks a substantial milestone toward Artificial Research Intelligence (ARI), signalling the acceleration of AI-driven scientific discovery. The AI Scientist also illustrates the urgent need for a discussion within the Information Retrieval (IR) and broader scientific communities. Whether and when ARI becomes a reality depends on how the academic and AI communities shape its development and governance. We propose concrete steps, including pilot projects and competitions, and standardized attribution frameworks such as research logs and markup languages.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Beel et al. (2025) studied this question.

synapsesocial.com/papers/68ebffcfdef9fcb308ff231fhttps://doi.org/10.1145/3769733.3769747
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Towards end-to-end automation of AI research2026 · 73 citations
  2. 2The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery2024 · 114 citations
  3. 3AI-Researcher: Autonomous Scientific Innovation2025 · 2 citations
  4. 4Virtuous Machines: Towards Artificial General Science2025
  5. 5What Research Is Suitable for AI for Science—and What Is Not2026