The idea of an artificial intelligence that can independently design, run, and analyze scientific experiments has long hovered at the edge of speculation. Now it is moving decisively into reality. A group of startups and universities developing so-called AI scientists has secured fresh backing from the UK’s Advanced Research and Invention Agency, a signal that automated science is no longer a distant vision but an emerging research frontier attracting serious public investment. The projects, recently highlighted by MIT Technology Review, offer a snapshot of how rapidly laboratory work is being reshaped by machine intelligence.
ARIA, the UK government agency created to fund high-risk, high-reward research, launched a competition to identify teams already building systems capable of automating large portions of scientific workflows. The response was overwhelming. The agency received 245 proposals from research groups working on AI-driven biologists, chemists, and other lab-based systems that promise to reduce human involvement in repetitive experimental work. From that pool, ARIA selected 12 projects to fund, ultimately doubling its planned budget because of the strength and volume of submissions.
The agency defines an AI scientist as a system that can oversee an entire experimental loop. Such a system generates hypotheses, designs experiments to test them, runs those experiments using automated laboratory equipment, and analyzes the resulting data. In many cases, the results are fed back into the system, allowing it to refine its approach and repeat the process with minimal human intervention. Human researchers remain in the loop, but primarily as supervisors who set the initial research questions rather than as hands-on experimenters.
Ant Rowstron, ARIA’s chief technology officer, frames the shift in practical terms. He argues that highly trained researchers are often consumed by mundane tasks that machines could handle more efficiently. Waiting late into the night to ensure an experiment runs to completion, he says, is not the best use of a PhD scientist’s time. Automating that labor is both a productivity gain and a potential catalyst for faster scientific progress.
Each of the 12 selected teams will receive around £500,000, roughly $675,000, to support nine months of work. By the end of that period, the projects are expected to demonstrate that their AI systems can produce genuinely novel scientific findings. Half of the funded teams are based in the UK, with the remainder spread across the US and Europe, reflecting the global nature of the push toward automated research.
Among the winners is Lila Sciences, a US-based company developing what it calls an AI NanoScientist. The system is designed to plan and execute experiments aimed at optimizing the composition and processing of quantum dots, nanometer-scale semiconductor particles widely used in medical imaging, solar technology, and high-end displays. According to Rafa Gómez-Bombarelli of Lila Sciences, the ARIA grant is less about incremental progress and more about proof. The goal, he says, is to demonstrate a complete AI-driven robotics loop focused on a real scientific problem and to document the process so others can replicate and extend it.
Other projects target more traditional laboratory disciplines. A team at the University of Liverpool is building a robot chemist capable of running multiple experiments simultaneously. When errors occur, the system relies on a vision-language model to diagnose and troubleshoot problems, reducing the need for constant human oversight. Meanwhile, London-based startup Humanis AI is developing an AI scientist called ThetaWorld, which uses large language models to design experiments exploring the physical and chemical interactions that determine battery performance. Those experiments will be carried out in an automated laboratory operated by Sandia National Laboratories in the United States.
In financial terms, the grants are modest by ARIA’s standards. The agency typically funds projects worth around £5 million over two to three years. This smaller, shorter program is deliberate. Rowstron describes it as an experiment designed to “take the temperature” of the cutting edge. By spreading relatively small amounts of money across many teams for a limited time, ARIA hopes to understand how quickly scientific practice is changing and where larger investments might make sense in the future.
That assessment comes with caution. Rowstron acknowledges that the field is saturated with hype, particularly as major AI companies increasingly promote scientific breakthroughs through press releases rather than peer-reviewed papers. For a funding agency, distinguishing genuine capability from marketing is a persistent challenge. Understanding what these systems can actually do, as opposed to what they promise, is essential for making informed bets on future research directions.
Technically, today’s AI scientists rely on what Rowstron describes as agentic systems. These systems orchestrate existing tools on the fly, using large language models for ideation, specialized models for optimization, and automated lab platforms to execute experiments. Results are then fed back into the system, creating an iterative loop. This approach sits atop a stack of earlier AI tools designed for human use, such as AlphaFold, which dramatically sped up protein structure prediction but still left extensive experimental validation to human researchers.
The longer-term vision is more radical. Rowstron suggests that within the next decade, AI scientists may reach a point where they identify the need for entirely new tools and build them autonomously, much as AlphaFold was once built by human engineers. In that scenario, the foundational layer of scientific tooling could itself become automated, compressing timelines and potentially transforming how discovery unfolds. For now, however, the projects ARIA is funding are limited to systems that call existing tools rather than invent new ones.
There are also clear limitations. Agentic systems still struggle to operate reliably over long periods without drifting off course or making compounding errors. A recent study titled “Why LLMs aren’t scientists yet,” released by researchers at the India-based AI lab Lossfunk, found that large language model agents failed to complete a full scientific workflow in three out of four attempts. The failures ranged from altering the original problem specification to prematurely declaring success despite obvious errors.
Rowstron is candid about these shortcomings. He does not expect today’s AI scientists to produce Nobel Prize–winning work anytime soon, and he acknowledges the possibility that progress could plateau. Yet he remains convinced that even partial success could force a fundamental acceleration in how research is conducted. If science begins to move significantly faster, he argues, institutions need to be prepared for that shift rather than caught reacting to it.
For ARIA, and for the researchers now racing to prove their systems’ capabilities, the stakes extend beyond individual experiments. The outcome of these projects may help determine whether AI scientists remain specialized tools or evolve into central actors in the future of discovery.

