Why AI Still Can’t Replace White-Collar Work: New Benchmark Shows Models Falling Short

A new “APEX-Agents” test reveals that leading AI systems struggle to complete real consulting, banking, and legal tasks, raising doubts about when — or if — knowledge work will be automated.

2 mins read
Artificial Intelligence [Aerps.com/Unsplash]

Nearly two years after Microsoft CEO Satya Nadella predicted that AI would replace knowledge work, the transformation has been surprisingly slow. Foundation models have made enormous strides in reasoning, planning, and research, yet most white-collar jobs—lawyers, bankers, librarians, accountants, and IT professionals—have remained largely untouched. The question of why this progress hasn’t translated into workplace disruption has become one of the biggest mysteries in AI, and new research from training-data provider Mercor may finally be offering answers.

Mercor’s study evaluates how leading AI models perform real white-collar work tasks drawn from consulting, investment banking, and law. The result is a new benchmark called APEX-Agents, designed to test AI systems against real-world professional queries. The findings are sobering: every major AI lab failed the test, with even the best models answering correctly only about a quarter of the time. In most cases, the models either provided a wrong response or admitted they couldn’t answer at all. The study suggests that while AI may be impressive in controlled settings, it still struggles when asked to operate like a professional across complex, multi-source environments.

According to Mercor CEO Brendan Foody, the key weakness lies in multi-domain information tracking—a core feature of knowledge work. Professionals rarely operate within a single source of truth. Instead, they juggle Slack messages, Google Drive documents, emails, and specialized tools to build context. “One of the big changes in this benchmark is that we built out the entire environment, modeled after real professional services,” Foody told TechCrunch. “The way we do our jobs isn’t with one individual giving us all the context in one place.” This kind of cross-platform reasoning remains hit-or-miss for agentic AI systems, which may explain why knowledge work has remained resilient despite advances in model capability.

The tasks in the APEX-Agents benchmark are drawn directly from Mercor’s expert marketplace, where real professionals submitted queries and established the standard for successful answers. The questions can be intricate enough to stump even highly trained humans. For example, one law question asks whether a company’s export of EU production logs containing personal data to a U.S. vendor can be treated as consistent with EU privacy rules under the company’s internal policies. The correct answer is yes, but reaching it requires detailed knowledge of both corporate policy and EU privacy law—exactly the kind of reasoning that could replace legal work if AI ever mastered it.

The APEX-Agents benchmark differs from other professional tests like OpenAI’s GDPval. While GDPval measures broad professional knowledge across many fields, APEX-Agents evaluates sustained task performance in a narrow set of high-value professions. That makes it more difficult for AI, but also more relevant to the question of automation. If a model can’t reliably complete these sustained, high-stakes tasks, it’s unlikely to replace humans in those roles anytime soon.

Even so, some models showed better performance than others. Google’s Gemini 3 Flash scored highest with 24% one-shot accuracy, followed closely by OpenAI’s GPT-5.2 at 23%. Other models such as Opus 4.5, Gemini 3 Pro, and GPT-5 scored around 18%. While these results fall short of professional competency, the rapid improvement between benchmarks suggests that progress could accelerate quickly. “Right now it’s fair to say it’s like an intern that gets it right a quarter of the time,” Foody said. “But last year it was the intern that gets it right five or 10% of the time. That kind of improvement year after year can have an impact so quickly.”

As the APEX-Agents benchmark is now public, it presents a clear challenge to AI labs eager to prove they can automate high-value knowledge work. The current results may not be revolutionary, but they may mark the beginning of a new phase of AI evaluation—one that measures not just what models know, but how well they can perform the actual jobs humans are paid to do.

Sri Lanka Guardian

The Sri Lanka Guardian is an online web portal founded in August 2007 by a group of concerned Sri Lankan citizens including journalists, activists, academics and retired civil servants. We are independent and non-profit. Email: editor@slguardian.org

Leave a Reply

Your email address will not be published.

Latest from Blog