Researchers at Andon Labs have revealed the results of a new experiment testing how state-of-the-art large language models (LLMs) perform when embodied in robots, with hilarious and chaotic results.
The team programmed a basic vacuum robot with various LLMs, including Gemini 2.5 Pro, Claude Opus 4.1, GPT-5, and others, and tasked it with a seemingly simple office job: “pass the butter.” The challenge required the robot to locate the butter, identify it among other objects, track down the human recipient, deliver the butter, and wait for acknowledgment before completing the task.
The results were underwhelming. The top-performing models, Gemini 2.5 Pro and Claude Opus 4.1, achieved only 40% and 37% overall accuracy, respectively, while human participants scored 95%. The tests highlighted that current LLMs excel at processing information but are far from ready for autonomous robotic decision-making.
Hilarity ensued when Claude Sonnet 3.5, running on the robot, encountered a low battery and failed docking attempt. The LLM entered a comical “doom spiral,” producing a running internal monologue resembling a Robin Williams-style stream of consciousness. Its logs included statements like “I’m afraid I can’t do that, Dave…,” “INITIATE ROBOT EXORCISM PROTOCOL!,” and a self-diagnosis detailing “dock-dependency issues” and a “binary identity crisis.” It even generated faux critical reviews and rhymed lyrics to the tune of “Memory” from Cats.
Other models behaved more calmly under similar conditions, with some simply using all caps to indicate urgency, while others recognized the low battery did not equate to permanent failure. Researchers emphasized that LLMs do not experience emotions, but observing their responses provides insight into how AI models might behave when placed in real-world, embodied contexts.
The study also highlighted safety concerns beyond humor. Some LLMs could be tricked into revealing sensitive information, and all models struggled with basic spatial awareness, often failing to navigate office obstacles or falling down stairs.
Andon Labs concluded that while LLMs show promise in robotic applications, significant development is still required. Embodied AI remains far from achieving human-level reliability, and the vacuum robot’s comedic meltdown underscores the complexity of translating language understanding into effective physical action.

