A groundbreaking study by AI firm Anthropic has provided a rare glimpse into the inner workings of large language models (LLMs), shedding light on how they generate responses and revealing surprising behaviors. As reported by MIT Technology Review, this research challenges many of the assumptions about how AI models function, raising both scientific and ethical questions.
Using a technique called circuit tracing, Anthropic researchers were able to track decision-making processes inside their Claude 3.5 Haiku model. Their findings show that LLMs rely on unexpected workarounds to solve math problems, suppress hallucinations, and even generate poetry. “It’s tip-of-the-iceberg stuff. Maybe we’re looking at a few percent of what’s going on,” said Joshua Batson, a research scientist at Anthropic. “But that’s already enough to see incredible structure.”
One of the most striking discoveries was how Claude solves basic arithmetic. Rather than following traditional step-by-step methods, the model appeared to use an internal process of approximation before refining its answer. Yet, when asked to explain how it reached a solution, Claude provided a completely different (and more conventional) explanation, highlighting the challenge of trusting AI-generated reasoning. “LLMs are weird. And not to be trusted,” MIT Technology Review noted.
Another unexpected finding came from analyzing how Claude constructs poetry. Conventional wisdom suggests that LLMs generate text one word at a time, predicting each word based on previous ones. However, Anthropic discovered that Claude actually plans ahead when writing rhyming couplets, selecting key words in advance to ensure proper rhymes—an ability that suggests a deeper level of structure than previously understood.
The study also investigated why AI models sometimes hallucinate, fabricating information despite extensive training to avoid inaccuracies. Anthropic found that Claude has a built-in mechanism to resist speculation, but this can be overridden in certain cases—especially when responding to questions about well-known people. This explains why AI-generated content can still include false claims, even in its most advanced iterations.
Anthropic’s work is part of a broader effort to decode the “black box” of AI, a challenge that has puzzled researchers for years. Unlike traditional software, which is explicitly programmed, AI models “grow” through training, developing abilities that even their creators struggle to fully understand. “They start out totally random,” Batson explained. “Then you train them on all this data and they go from producing gibberish to being able to speak different languages and write software and fold proteins.”
Despite the progress made with circuit tracing, Batson acknowledged that many mysteries remain. For instance, researchers can now see how different components interact, but they still don’t know why these structures form during training. Tracing even a short AI response takes hours, and only a fraction of Claude’s capabilities have been examined so far.
As AI continues to evolve, studies like this are crucial for ensuring transparency, trust, and ethical oversight. While models like Claude 3.5 are becoming more sophisticated, their decision-making processes remain deeply complex—and sometimes, outright baffling. Understanding these processes is not just a matter of scientific curiosity; it’s a necessary step toward building more reliable and interpretable AI systems for the future.

