A theoretical chemist at Zhejiang Lab in Hangzhou discovered that an AI model predicting molecular boiling points produced values conflicting with a 75-year-old reference database. Sebastian Pios initially suspected the model was faulty, but manual verification of original literature confirmed the database entries were incorrect.
The AI system identified two additional categories of errors: a typographical mistake in a published paper and inaccurate century-old boiling-point measurements that had been incorporated into the scientific canon. Pios noted these errors likely caused significant difficulties for researchers relying on the database.
This case reflects a broader trend of scientists deploying AI agents to audit scientific knowledge at scale. Researchers at SAI Labs used AI to assess 168 papers selected for oral presentation at the 2026 International Conference on Machine Learning, extracting claims and attempting to reproduce reported results.
Of 92 papers with at least five assessable claims, the AI agents could reproduce more than two claims from only 34 papers, and successfully repeated over 80% of claims from just eight papers. The analysis included rerunning experiments where possible and comparing outcomes with author-reported results.
Computer scientist Odd Erik Gundersen at the Norwegian University of Science and Technology cautions that AI fact-checking tools remain unreliable arbiters, making mistakes similar to humans and requiring manual oversight for quality control.
James Zou at Stanford University emphasizes the speed advantage of AI systems in scanning scientific literature compared to human reviewers. Zou and colleagues deployed an AI checker on papers from the NeurIPS conference, finding objective errors rose from 3.8 per paper in 2021 to 5.9 in 2025, a 55% increase.
The analysis focused on objective errors in formulae, calculations, and figures, excluding subjective judgments about data interpretation or novelty. Co-author Federico Bianchi at Together AI stated this was a deliberate design choice to keep significance assessments in human hands.
Zou warned that errors in published papers, which serve as foundational knowledge for subsequent research, can propagate and undermine follow-on studies. The growing movement aims to catch such mistakes before they compound through the scientific literature.
AI agents are checking the scientific literature — and spotting decades-old errors
This is an independent summary. The complete reporting, supporting context and any primary documents remain with Nature.
