Researchers have developed the first quantitative benchmark for AI Scientist systems, using automated peer review to evaluate autonomous research generation frameworks.

The study, published on arXiv, assessed four leading AI Scientist platforms: Sakana AI (versions 1 and 2), CycleResearcher, and Data-to-Paper. Each system generated papers from 15 research proposals originally published by FARS, a commercial autonomous AI scientist company.

The evaluation used three frontier language models — GPT-5.4, Gemini, and Claude — as independent reviewers. Papers were scored across four dimensions: originality, scientific rigor, clarity, and significance on a 1-5 scale.

FARS outperforms all competitors

FARS benchmark papers significantly outscored all competing frameworks, achieving mean scores of 2.14-2.47 compared to 1.00-1.87 for other systems. FARS scored more than twice as high as the next-best systems on both Gemini and Claude evaluations.

The researchers found strong agreement between Gemini and Claude reviewers (correlation coefficient 0.907), with both correlating extremely strongly with synthesis scores (0.961). However, GPT-5.4 showed weaker agreement with other models (correlation approximately 0.32), suggesting it applies different evaluation criteria.

The study generated 60 papers total from the four competing frameworks, evaluating them alongside 15 FARS benchmark papers. This created the first systematic comparison of autonomous research generation quality.

"These results establish the first quantitative benchmark for AI Scientist systems and demonstrate that multi-model LLM evaluation provides a scalable, consistent framework for assessing autonomous research quality," the authors wrote.

The research addresses a critical challenge in AI-driven scientific discovery: how to evaluate and compare AI-generated papers at scale. The automated peer-review approach could enable more systematic development of autonomous research systems.

The benchmark will likely influence development priorities for AI Scientist platforms, providing concrete metrics for measuring progress in autonomous scientific paper generation.