Artificial intelligence researchers are shifting their focus from models that generate responses to AI agents that can carry out complex tasks in the real world, with the goal of creating systems that can help humans perform economically valuable work across industries.
Dawn Song, Meta Platforms’ newly appointed vice-president of AI research, said the next frontier of artificial intelligence is developing agents that can operate more effectively in important real-world domains while supporting human workers rather than replacing them.
“The goal is not to replace humans,” Song told the South China Morning Post during the World Economic Forum in Dalian, also known as Summer Davos, shortly before joining Meta. “But we want these AI agents to be more effective in these important real-world domains and help humans do this work better and provide more economic value,” she said.
Song, a professor of computer science at the University of California, Berkeley, and co-director of the university’s Centre for Responsible, Decentralised Intelligence (RDI), has focused much of her career on artificial intelligence security. She is also a co-founder of enterprise AI safety start-up Virtue AI.
On Friday, Song announced on X that she and many members of the Virtue AI team were joining Meta’s Superintelligence Labs, where she will work on AI safety and security efforts.
Her move comes as researchers race to measure how close AI systems are to performing complex professional tasks. Earlier this month, UC Berkeley’s RDI centre introduced Agents’ Last Exam (ALE), a benchmark designed to test AI agents across more than 1,500 real-world tasks spanning 55 industries.
The benchmark was designed to examine whether AI agents can handle practical assignments with economic value. Examples include using video editing software such as DaVinci Resolve to create a finished video from raw footage, or performing the work of a neuroimaging analyst by reviewing the quality of brain MRI scans.
Song said the tests were intentionally difficult, resulting in low success rates even among the most advanced AI systems currently available.
According to ALE results, OpenAI’s GPT-5.5 paired with the Codex harness recorded the highest performance, with a pass rate of 24.3%. Anthropic’s Claude Fable 5 model using the Claude Code harness ranked third with a 22% pass rate, although ALE noted that its performance could have been affected by Anthropic’s policy for powerful models.
Among Chinese AI models tested, ByteDance’s Seed2.1 Pro using a Claude Code harness achieved the highest ranking, recording a 19.5% pass rate and placing eighth globally. DeepSeek V4 Pro, Alibaba Group Holding’s Qwen3.7-Max and Zhipu AI’s GLM 5.1 ranked lower in the global comparison.
Alibaba owns the South China Morning Post.
The hardest tier of ALE testing showed the continuing limitations of current AI agents. Song said almost every frontier model tested achieved a zero success rate on the most challenging tasks, while GPT-5.5 was the only system to record any success, with a score of 2.6%.
The results, she said, showed that today’s AI agents remain far from human-level performance. Song said future versions of the benchmark would become increasingly difficult as AI capabilities continue to develop.
“We hope the Agents’ Last Exam can be a guidepost for the community to actually build solutions to the models and agents that can really help solve these economically valuable tasks,” Song said.
Song and her team have also developed other benchmarks focused on AI cybersecurity capabilities, including CyberGym and ExploitGym. These tests examine how AI models perform in areas such as analysing cyber vulnerabilities and conducting cyberattacks.
Cybersecurity has become a key area of attention as AI systems become more capable. Anthropic highlighted the performance of its Claude Mythos Preview model on the CyberGym benchmark, while OpenAI CEO Sam Altman cited GPT-5.5-Cyber’s performance on the same benchmark as evidence of its cybersecurity capabilities.
During her appearance at Summer Davos, Song, who was born in Dalian, met with business leaders and policymakers and urged organisations to strengthen their cybersecurity capabilities through AI. She said open-source Chinese AI models with significant cybersecurity implications could emerge soon.
“AI is a dual-use technology, where, for example, coding and cybersecurity are two sides of the same coin,” Song said. “We’re not far away from a world where we have open-weight models of very powerful capabilities.”
As researchers continue developing more advanced AI agents, new testing systems such as Agents’ Last Exam are being used to measure whether these technologies can move beyond demonstrations and begin handling complex tasks across real-world industries.

