OpenAI’s o3 Tops Scientific AI Leaderboard in SciArena Evaluation

With scientific tasks pushing the boundaries of AI capabilities, tools like SciArena could play a central role in ensuring large language models are rigorously and transparently evaluated

1 min read
OpenAI [Zac Wolff/Unsplash]

OpenAI’s latest large language model, o3, has emerged as the top-performing AI system for answering scientific questions, according to a newly launched benchmarking platform called SciArena, as reported by Nature.

Developed by the Allen Institute for Artificial Intelligence (Ai2) in Seattle, SciArena ranks large language models (LLMs) based on how well they answer technical questions across various domains. In its inaugural ranking, OpenAI’s o3 — the model behind ChatGPT’s most advanced responses — outperformed 22 other models in categories including natural sciences, health care, engineering, and the humanities.

The rankings were determined through a novel crowdsourced method: 102 researchers cast more than 13,000 votes comparing LLMs’ answers to scientific questions submitted by users. Each question was answered by two randomly selected models, with their responses citing references from Semantic Scholar, an AI-powered academic research tool. Voters then assessed which model gave the better answer — or whether both (or neither) provided satisfactory responses.

According to Nature, researchers favored o3 for its technically nuanced responses and its detailed references to relevant literature. “The model’s depth in citing literature and explaining concepts appears to set it apart,” said Arman Cohan, a research scientist at Ai2. However, because many LLMs remain proprietary, fully explaining performance differences remains difficult.

Trailing behind o3 were DeepSeek-R1, developed by China-based DeepSeek, which ranked second in natural sciences and fourth in engineering, and Google’s Gemini-2.5-Pro, which came in third for natural sciences.

SciArena marks one of the first open platforms to evaluate AI tools for scientific research using verified user feedback. AI experts, including Jonathan Kummerfeld of the University of Sydney, say such efforts could revolutionize how researchers interact with scientific literature. “This will help researchers find work they may have otherwise missed,” he told Nature, noting the potential for LLMs to aid in literature discovery.

However, experts also caution that while promising, AI-generated summaries are not a replacement for reading primary research. “LLMs can misunderstand terminology or generate text that misrepresents cited papers,” warned Rahul Shome, a robotics and AI researcher at the Australian National University.

Despite these limitations, the platform has been praised for its transparency and design, which aims to prevent score manipulation — a problem that has plagued other AI benchmarking systems. Verified user votes contribute to a live-updated leaderboard, and SciArena remains free to use.

With scientific tasks pushing the boundaries of AI capabilities, tools like SciArena could play a central role in ensuring large language models are rigorously and transparently evaluated, offering researchers better tools — and accountability — in the evolving world of AI-assisted science.

Sri Lanka Guardian

The Sri Lanka Guardian is an online web portal founded in August 2007 by a group of concerned Sri Lankan citizens including journalists, activists, academics and retired civil servants. We are independent and non-profit. Email: editor@slguardian.org

Latest from Blog