A new analysis has raised alarm over an explosion of low-quality biomedical research papers that may have been generated or heavily assisted by artificial intelligence (AI). The study, published this month in PLoS Biology, found hundreds of papers making simplistic and potentially misleading health claims based on publicly available health data.
The research team examined over 300 papers that relied on data from the U.S. National Health and Nutrition Examination Survey (NHANES), a publicly accessible dataset that includes information on health, diet, and lifestyle from thousands of Americans. Many of these studies, the authors say, followed a “cookie-cutter” template—linking a single variable, such as vitamin D levels or hours of sleep, to complex diseases like depression or heart conditions, while ignoring other known contributing factors.
“We’ve seen a sudden explosion in publication rates of papers that are extremely formulaic—so much so that they could easily have been generated by large language models,” said Matt Spick, co-author of the study and a biomedical scientist at the University of Surrey.
The researchers found significant statistical issues in many of the papers. For example, a subset of 28 studies that linked single variables to depression showed that after applying basic statistical correction techniques, only 13 of those findings remained valid.
Other concerns included cherry-picking of data and omitting parts of the NHANES dataset without explanation. Of 14 papers that looked at inflammation markers, only 4 used the full available data. The rest selectively analyzed specific age groups or time periods—often without justifying why.
“This is like taking an exam, checking which questions you got right, and then removing the ones you got wrong. That’s essentially what’s happening,” explained Charlie Harrison, a computational biologist at Aberystwyth University and co-author of the study.
The trend appears to have accelerated in recent years, especially after 2022, when large language models like ChatGPT became more widely available. In 2024 alone, over 2,200 studies using NHANES data were published—more than in any previous year.
Although the study did not confirm whether these papers were products of so-called “paper mills”—organizations that mass-produce low-quality or fraudulent academic papers—experts say the ease of using AI to analyze open datasets makes them a tempting target.
“These papers have fingerprints of being produced to a formula,” Spick added. “You can basically run through all possible combinations of variables until something looks statistically significant.”
The authors recommend stricter regulations for the use of public datasets. One suggestion is to require researchers to register their study designs before accessing NHANES or similar data—making it harder to manipulate findings for the sake of publication.
“If we don’t act, we risk drowning out meaningful scientific discoveries with waves of weak and repetitive studies,” Harrison warned.
Ioana Alina Cristea, a clinical psychologist and meta-researcher at the University of Padua who was not involved in the study, agreed. “It’s not useful anymore to show that a single factor is related to depression, because we already know it’s a multifactorial condition,” she said. “We need to focus on quality, not quantity.”

