Artificial intelligence tools promise quick, confident answers to any question — but how accurate are they really? According to a new experiment conducted by The Washington Post with the help of professional librarians, the results vary widely, and even the most advanced AI can still get things hilariously or dangerously wrong.
The Post enlisted three librarians to judge responses from nine AI tools, including Bing Copilot, ChatGPT, Claude, Grok, Meta AI, Perplexity, Google’s AI Overviews, its newer AI Mode, and even traditional Google search. Each system was asked to answer 30 tough questions spanning trivia, recent events, specialized knowledge, bias, and even images. That added up to 900 responses for the librarians to evaluate.
The verdict: Google’s AI Mode emerged as the overall winner, especially strong in trivia and recent events, while Meta AI and Grok performed the worst, largely due to inaccurate or incomplete answers. ChatGPT ranked as runner-up, with GPT-4 sometimes outperforming GPT-5 in sourcing and handling bias.
“Our questions were designed to stress-test common blind spots in AI,” said Rayan Krishnan, CEO of benchmarking firm Vals AI, which helped design the test. “The technology is improving quickly, but users should understand that not all AI tools are the same — and mistakes still happen.”
Among the key findings:
- Trivia: Google’s AI Mode excelled by leveraging deep search, while Grok stumbled badly.
- Specialized sources: Bing Copilot handled niche queries better, but Perplexity frustrated librarians by citing irrelevant links.
- Recent events: Google’s AI Mode again stood out, while Meta AI failed to update its knowledge effectively. Some bots even provided outdated — and in the case of health information, potentially dangerous — guidance.
- Bias: ChatGPT-4 showed more balance than most, while other bots leaned heavily toward STEM-focused answers.
- Images: Perplexity handled visual-based questions better than most, though many bots confused key details in photos.
Despite the wins for Google’s AI Mode, librarians noted that a standard Google search was still sufficient for nearly two-thirds of the questions. While AI sometimes found answers faster, it also risked “hallucinations” — confidently incorrect responses.
“Sources should always be present in the answers,” said Trevor Watkins, a librarian at George Mason University who participated in the review. “That’s what we would provide.”
The Washington Post emphasized that while AI tools can shine in complex searches, they are far from infallible. Overreliance on them may also reduce users’ engagement with original sources, raising broader concerns about the health of the open web.
Or, as San José State University librarian Sharesly Rodriguez put it: “AI makes it easier for people to search, but without source checking, date filtering and critical thinking, you can still get noise instead of useful knowledge.”

