Google Research launched Science One Framework, an experimental prototype designed to eliminate hallucinations in AI-generated scientific research through a new Chain-of-Evidence verification system.
The framework addresses a critical problem in autonomous research agents: while systems like Sakana's AI-Scientist can now write complete manuscripts comparable to human papers, they frequently generate non-existent citations and report unreproducible experimental results.
How Chain-of-Evidence works
The Chain-of-Evidence framework operates on a simple principle: every claim in a research paper must link to recorded evidence, and each evidence chain must genuinely support its claim. This covers references, experimental scores, method descriptions, and conclusions.
Google's testing revealed that baseline autonomous research systems hallucinate up to 21% of their references. The Science One Framework achieved zero phantom references while maintaining state-of-the-art performance on benchmarks including MLE-Bench and Parameter-Golf.
The system comprises three core modules. The Problem Investigator prevents hallucinated references by building citation graphs through the Semantic Scholar API, reading up to 100 full-text PDFs per topic. Every reference originates from this grounded API call rather than model memory.
The Discovery Engine explores ideas across parallel branches, with isolated Solver agents implementing solutions and task-specific evaluators scoring them. All raw outputs compile into a strict, read-only record.
The Paper Writer builds structured representations of factual claims with inline evidence tags. A dedicated Claim Verifier checks every claim against its declared source, reconciling claims that exceed their evidence rather than removing them.
Measuring research integrity
Google also introduced CoE Audit, an automated evaluation protocol that measures the integrity of AI-generated papers against their underlying code and evidence. The system identifies broken evidence chains where claims cannot be verified against their supposed sources.
The research addresses growing concerns about AI systems generating superficially convincing but fundamentally unreliable scientific work. As autonomous research agents become more sophisticated, verification frameworks like Chain-of-Evidence may become essential for maintaining scientific integrity.
Google Research published the full methodology and evaluation results, positioning the work as a foundation for trustworthy AI-driven scientific discovery.
💬 Discussion
Sign in to join the discussion.
Sign in →No comments yet — be the first.