The rapid rise of artificial intelligence in health care is ushering in a new era of digital medicine—one that is as promising as it is uncertain. In recent months, major technology companies have accelerated the rollout of AI-powered health assistants, positioning them as tools that could transform access to care. Yet, as highlighted in reporting by MIT Technology Review, a critical question remains unresolved: do these systems actually work safely and effectively?
Earlier this month, Microsoft introduced Copilot Health, a feature embedded within its Copilot platform that allows users to connect medical records and ask personalized health questions. Around the same time, Amazon expanded access to its Health AI assistant, previously limited to subscribers of its One Medical service. These offerings join a growing ecosystem that includes OpenAI’s ChatGPT Health and Anthropic’s Claude, both capable of engaging with sensitive medical data when granted permission. Together, they signal a clear shift: AI-driven health advice is no longer experimental—it is becoming mainstream.
The momentum behind these tools is fueled by two powerful forces: technological progress and unmet demand. Developers argue that large language models have reached a level of sophistication where they can provide meaningful health guidance. At the same time, millions of users are already turning to AI for answers. Microsoft reports that its Copilot system fields tens of millions of health-related queries daily, making health the most common topic discussed on its platform.
This surge in usage reflects more than curiosity. For many people, especially those in underserved or resource-constrained environments, access to traditional health care remains limited. Long wait times, high costs, and geographic barriers often leave patients searching for alternatives. AI chatbots, available around the clock and free from judgment, offer an appealing substitute. Experts suggest that this demand is not accidental—it is a direct response to systemic gaps in health care delivery.
In theory, AI health assistants could help bridge those gaps. One of their most promising applications is triage, the process of determining whether a patient needs urgent medical attention. If effective, chatbots could guide users toward appropriate care, encouraging those with serious symptoms to seek help sooner while reassuring others that their conditions can be managed at home. Such a system could reduce strain on overcrowded hospitals and clinics.
However, early evidence raises concerns about whether these tools are ready for such responsibility. A recent study conducted by researchers at Mount Sinai found that AI systems can misjudge medical situations, sometimes recommending unnecessary care for minor issues while failing to recognize emergencies. Although some developers dispute aspects of the study’s methodology, the findings underscore a broader issue: the lack of independent, transparent evaluation.
This concern is echoed across the academic community. While companies like OpenAI have developed internal benchmarks such as HealthBench to measure performance, these evaluations often rely on simulated scenarios rather than real-world interactions. Critics argue that such testing may not fully capture how users actually engage with AI systems, particularly those without medical expertise. A chatbot may perform well when given clear, structured information, but real patients often struggle to describe symptoms accurately or interpret complex responses.
Research suggests that this gap between model capability and user understanding can significantly affect outcomes. In some cases, individuals using AI assistance are far less likely to arrive at correct conclusions than the models themselves. This discrepancy highlights a fundamental challenge: health care is not just about providing answers, but about asking the right questions—a task that requires context, experience, and nuance.
Despite these limitations, companies continue to push forward with public releases. Disclaimers stating that AI tools are not intended for diagnosis or treatment are standard, but experts widely agree that such warnings are unlikely to deter users. In practice, many people will rely on these systems for precisely those purposes, particularly when other options are unavailable.
The debate over how to evaluate AI health tools is ongoing. Some researchers advocate for rigorous clinical-style trials involving human participants, while others argue that the fast pace of AI development makes traditional approaches impractical. Even when high-quality studies are conducted, as in Google’s recent research on its experimental AMIE chatbot, companies may choose to delay public deployment until further testing addresses concerns around safety, fairness, and equity.
What most experts agree on is the need for independent oversight. Allowing companies to assess their own products creates potential conflicts of interest and increases the risk of blind spots. Third-party evaluations, conducted by diverse research groups, could provide a more balanced and comprehensive understanding of how these systems perform in real-world settings.
At the same time, expectations must remain realistic. No medical system, human or artificial, is error-free. Physicians make mistakes, and in many parts of the world, access to qualified professionals is sporadic at best. In such contexts, even an imperfect AI assistant could represent a meaningful improvement—provided its errors are not severe or systematically biased.
For now, the evidence is insufficient to draw firm conclusions. The current generation of AI health tools sits at a crossroads between innovation and uncertainty. They hold the potential to democratize access to medical knowledge, empower patients, and alleviate pressure on strained health systems. Yet they also carry risks that are not fully understood, particularly when deployed at scale without comprehensive validation.
As the technology continues to evolve, the stakes will only grow higher. Health care is an area where mistakes can have life-altering consequences, and the margin for error is thin. Whether AI chatbots ultimately become trusted allies in medicine or cautionary tales of premature deployment will depend on how seriously developers, researchers, and regulators address the need for robust, transparent evaluation.
Until then, the promise of AI-driven health care remains just that—a promise, waiting to be proven.

