Artificial intelligence (AI) is increasingly being used to detect errors in scientific research papers, an effort that has gained traction following a high-profile mistake involving black plastic cooking utensils. The error, a miscalculation that exaggerated the toxicity of a key chemical, demonstrated how AI could swiftly identify such issues. This revelation has led to the emergence of two AI-driven projects—The Black Spatula Project and YesNoError—that aim to scrutinize scientific literature for inaccuracies.
The Black Spatula Project, an open-source AI initiative, has analyzed about 500 papers so far. Instead of publicizing its findings, the project is approaching affected authors directly. “Already, it’s catching many errors,” said Joaquin Gulloso, an independent AI researcher helping to coordinate the project. Meanwhile, YesNoError, inspired by The Black Spatula Project, has taken a broader approach, analyzing over 37,000 papers in just two months. Founded by AI entrepreneur Matt Schlicht, YesNoError’s website flags potentially flawed papers, though many flagged errors have yet to undergo human verification.
These projects seek to integrate AI into the scientific review process, encouraging researchers and journals to use AI tools before submission and publication. While many experts in research integrity tentatively support the movement, concerns remain about accuracy and potential reputational damage. Michèle Nuijten, a metascience researcher at Tilburg University, warns that false positives could unjustly harm researchers if mistakes are wrongly attributed.
Despite these risks, some argue that AI-powered scrutiny is necessary. “It’s much easier to churn out shoddy papers than it is to retract them,” said James Heathers, a forensic metascientist at Linnaeus University, who consults for the Black Spatula Project. He believes AI could serve as a triage mechanism, identifying suspect papers for deeper scrutiny.
AI-based integrity checks are not entirely new. Various tools already exist for spotting specific issues in scientific papers, but proponents hope AI can conduct broader analyses across larger volumes of research. Both The Black Spatula Project and YesNoError use large language models (LLMs) to identify factual, methodological, and referencing errors. These systems extract information from papers, use complex instructions to analyze errors, and sometimes cross-check results. Depending on the complexity of analysis, each paper costs between 15 cents and a few dollars to evaluate.
A significant challenge for AI-driven paper analysis is minimizing false positives. The Black Spatula Project’s system is reportedly wrong 10% of the time, requiring subject-matter experts to manually verify flagged errors. Schlicht’s YesNoError team has tested their system against 100 detected mathematical errors, with 90% of authors who responded agreeing with the AI’s findings.
However, some experts remain skeptical. Nick Brown, a research integrity specialist at Linnaeus University, reviewed 40 papers flagged by YesNoError and found 14 false positives, such as incorrect claims that figures were missing from papers. Brown worries these initiatives may overwhelm the scientific community with minor errors, including typos, that should be caught during peer review.
A major controversy surrounding YesNoError is its use of cryptocurrency to influence which papers get scrutinized first. While Schlicht argues that this approach will prioritize widely discussed research, critics fear it could be exploited to target politically sensitive studies, such as those on climate science.
As these AI initiatives evolve, they face a crucial test: refining their accuracy while maintaining transparency. As reported in Nature, some experts believe that if these tools are properly developed, they could expose systemic flaws in research integrity. Brown notes, “If someone actually built a highly effective version of these tools, in some fields, I think it would be like turning on the light in a room full of cockroaches.”

