AI ‘Scientist’ Sparks Debate Over Plagiarism and Originality in Research

The debate over AI-generated research is still in its early stages, but it is forcing scientists to grapple with fundamental questions.

4 mins read
Artificial Intelligence [Aerps.com/Unsplash]

In January this year, South Korean researcher Byeongjun Park opened an e-mail that would draw him into one of the most heated debates in modern science. Two computer scientists in India told him that an artificial intelligence tool had produced a manuscript borrowing methods from his work without giving him credit. The manuscript had not been formally published but appeared online, clearly labeled as the product of a system called The AI Scientist, developed by Tokyo-based company Sakana AI. The tool, announced in 2024, is designed to function as a fully automated research system: generating ideas with a large language model, running code, and then writing papers that present its findings. For Park, who works at the Korea Advanced Institute of Science and Technology, the experience was jarring. “I was surprised by how closely the core methodology resembled that of my paper,” he told Nature.

The AI-generated paper did not copy his text directly. Instead, it proposed a new architecture for diffusion models — the backbone of many image generators — while Park’s research focused on improving how those models are trained. The similarity, however, was enough to prompt questions. Indian researchers Tarun Gupta and Danish Pruthi, who had alerted Park, soon uncovered broader concerns. In a study published earlier this year, they reported multiple cases of AI-generated manuscripts that, according to external experts they consulted, re-used methods or ideas from existing research without attribution. They described the phenomenon as “idea plagiarism” — a subtle but serious erosion of academic credit.

Their findings, which won an outstanding paper award at a major computational linguistics conference in Vienna, suggest that nearly one-quarter of AI-generated manuscripts they examined showed strong overlaps with prior work. Yet the claims are contested. The team behind The AI Scientist told Nature that they reject the plagiarism label, calling the accusations “unfounded, inaccurate, extreme, and should be ignored.” In Park’s case, one independent reviewer found the overlap too weak to qualify as plagiarism, while others disagreed. Park himself acknowledged what he saw as a strong methodological resemblance but hesitated to use the word plagiarism.

The dispute reflects a deeper question: what counts as originality in scientific research? In fields like computer science, where thousands of papers are published each year, it is already difficult for human researchers to track novelty. Large language models complicate the picture further, because they are designed to remix existing patterns of knowledge rather than produce ideas entirely from scratch. “A significant portion of LLM-generated research ideas appear novel on the surface but are actually skillfully plagiarized in ways that make their originality difficult to verify,” Gupta and Pruthi wrote.

Some experts argue that this is not a new issue. Debora Weber-Wulff, a plagiarism researcher at the University of Applied Sciences in Berlin, told Nature that idea plagiarism has long existed in human-authored papers, and AI tools will likely worsen the trend. But proving it is notoriously difficult. Unlike traditional plagiarism, which involves copying sentences, idea plagiarism involves borrowing concepts or methods, and assessing such overlaps often depends on subjective judgment. “There’s no one way to prove idea plagiarism,” Weber-Wulff said.

The controversy came into sharper focus after Gupta and Pruthi scrutinized a set of AI-generated papers released both by Sakana AI and by Stanford University researcher Chenglei Si, whose team had asked humans and AIs to propose novel research ideas. Independent experts reviewing these manuscripts for Gupta and Pruthi found multiple cases where AI-generated ideas bore striking similarities to existing work. In some instances, original authors agreed, saying the AI outputs were “basically very similar” to their papers despite superficial differences. In one example, an AI manuscript that had even passed through a stage of peer review for a major machine-learning conference was accused of reusing contributions from a 2015 paper without citation.

For critics, such overlaps highlight the risks of allowing AI systems to produce research at scale without robust mechanisms for crediting sources. For defenders of The AI Scientist, however, the situation is not so different from the common lapses of human authors who also fail to cite all relevant work. “It would have been good for The AI Scientist to cite them,” the team admitted in response to questions about specific omissions, but they insisted that missing references happen “every day” in conventional academia.

The debate also reveals differing views about the meaning of plagiarism itself. The AI Scientist team argues that plagiarism implies intentional fraud, which cannot apply to a machine with no intent. But Weber-Wulff counters that intent should not matter: if a manuscript presents ideas attributable to someone else without citing them, it meets the definition of plagiarism. She points to a standard definition developed by the International Center for Academic Integrity, which emphasizes the absence of proper attribution regardless of the author’s intention.

As Nature reports, the clash reflects not only questions of ethics but also the limitations of current technology. Tools like Turnitin, widely used to detect text plagiarism, fail to catch idea-level borrowing in AI manuscripts. More specialized approaches, such as systems that search academic databases for semantic similarities, remain far from reliable. Even among human reviewers, judgments about novelty often vary, underscoring how subjective the concept can be.

Despite the controversy, the team behind The AI Scientist insists that their tool marks an important milestone — showing that AI can already draft research papers, even if imperfectly. They argue that the system is best seen as a source of inspiration, with human researchers responsible for validating and refining its outputs. “Ultimately, The AI Scientist and systems like it will soon be making obviously new, major discoveries,” they said. Critics, meanwhile, warn that without stronger safeguards, such tools risk accelerating a culture of uncredited borrowing and diluting the meaning of originality in science.

The debate over AI-generated research is still in its early stages, but it is forcing scientists to grapple with fundamental questions. How should novelty be defined in an age when machines remix existing knowledge? How much overlap is acceptable before a paper ceases to be original? And can automated systems ever be trusted to uphold academic standards? As Nature highlights, there are no clear answers yet, but what is certain is that AI is changing the way research is produced, reviewed and credited — and the scientific community must now decide how to adapt.

Sri Lanka Guardian

The Sri Lanka Guardian is an online web portal founded in August 2007 by a group of concerned Sri Lankan citizens including journalists, activists, academics and retired civil servants. We are independent and non-profit. Email: editor@slguardian.org

Leave a Reply

Your email address will not be published.

Latest from Blog