For two decades, the first instinct of anyone experiencing new medical symptoms was to “Google it,” a practice so widespread it earned the nickname “Dr. Google.” But that era is rapidly fading as more people turn to large language models for medical guidance. OpenAI reports that 230 million people ask ChatGPT health-related questions every week, signaling a major shift in how the public seeks medical information. This is the context behind OpenAI’s new ChatGPT Health product, which launched earlier this month. Yet the timing of the rollout was problematic: just two days earlier, SFGate reported that a teenager named Sam Nelson died of an overdose after extensive conversations with ChatGPT about how to combine drugs safely, raising urgent questions about the safety of AI medical advice.
ChatGPT Health appears in a separate sidebar tab, but it is not a new AI model. Rather, it is a “wrapper” that gives an existing OpenAI model new tools and guidance for delivering health advice, including the ability to access users’ electronic medical records and fitness data if permission is granted. OpenAI stresses that the product is intended to support—not replace—doctors. Still, the reality is that people often turn to alternatives when medical professionals are unavailable or inaccessible. As health reporters and experts have noted, AI’s potential for harm remains a major concern.
Some physicians see potential benefits in using LLMs to improve medical literacy. Navigating the vast and often misleading online health information landscape can be difficult for patients, but in theory, AI could help filter credible sources from dubious ones. Harvard Medical School radiologist Marc Succi explained that patients who previously arrived with misinformation from Google now ask questions at a level comparable to early medical students, suggesting that AI may be raising the quality of patient engagement. The question remains whether these improvements outweigh the risks.
The introduction of ChatGPT Health and similar moves by Anthropic, which recently announced new health integrations for Claude, show that major AI companies are increasingly embracing medical use cases. Yet LLMs have well-documented weaknesses, including a tendency to agree with users and fabricate answers rather than admit uncertainty. Still, supporters argue the risks must be balanced against potential benefits, drawing an analogy to autonomous vehicles: the key question is not whether AI is perfect, but whether it causes less harm than the current alternative.
However, measuring the real-world effectiveness of chatbots for health advice is difficult. The AI field has struggled to evaluate open-ended conversational systems. Danielle Bitterman, clinical lead for AI at Mass General Brigham, emphasized that while models perform well on medical licensing exams, those exams are multiple-choice and don’t reflect how people actually interact with chatbots. Studies that attempt to test models using more realistic prompts show mixed results. A study led by Amulya Yadav at Pennsylvania State University found GPT-4o answered medical questions correctly about 85% of the time in realistic scenarios, though Yadav himself expressed skepticism about patient-facing medical AI. Another evaluation by Sirisha Rambhatla found that GPT-4o answered only about half of licensing exam questions correctly when not given answer choices.
These mixed results are part of a larger debate about whether LLMs are truly safer than Google searches. Some studies suggest AI could reduce misinformation and anxiety by offering more accurate and calm responses. But other research has shown models can hallucinate or display sycophancy—agreeing with incorrect information provided by users and inventing fake syndromes or lab tests. This risk is particularly dangerous in health, where users may take AI responses as trustworthy. Critics worry that as chatbots become more human-like, people may trust them too much, especially when they disagree with doctors. University of Melbourne researcher Reeva Lederman warns that well-spoken AI may cause some patients to reject medical advice, increasing the risk of harm.
OpenAI says the GPT-5 series is less prone to hallucination and sycophancy, and the company has tested the model behind ChatGPT Health using its HealthBench benchmark. HealthBench rewards models for expressing uncertainty, recommending medical care when appropriate, and avoiding unnecessary alarm. Still, experts caution that some of HealthBench’s prompts were generated by AI rather than real users, which could limit how accurately it reflects real-world usage.
Even if ChatGPT Health proves to be a safer alternative to Google, it could still have unintended consequences. Like autonomous vehicles that might reduce safety but increase car use, AI could improve information quality while encouraging people to rely on chatbots instead of doctors. Lederman’s research suggests that online communities often trust confident-sounding sources regardless of accuracy, and the conversational tone of chatbots could make them particularly persuasive. As the AI giants push further into health, the balance between improving access to information and protecting patient safety remains precarious.
Notably, the publication MIT Technology Review has closely followed the debate over AI and health, highlighting both the potential benefits and the real dangers of using chatbots for medical advice. As the technology evolves, the conversation continues: can AI ever be a safe and reliable “Dr. ChatGPT,” or will it remain a risky shortcut that could lead people away from professional care?

