AI Medical Tools May Risk Biased Care for Women and Minorities, Studies Warn

Despite the challenges, experts emphasize the potential benefits of AI in healthcare.

2 mins read
Marzyeh Ghassemi, associate professor at MIT’s Jameel Clinic

Artificial intelligence tools increasingly used in healthcare could worsen outcomes for women and ethnic minorities, according to research highlighted by the Financial Times. Studies suggest that large language models (LLMs) employed to assist doctors often downplay symptoms in these groups, raising concerns over biased medical decision-making.

Researchers in the US and UK have found that AI models—including OpenAI’s GPT-4, Meta’s Llama 3, and healthcare-focused systems such as Palmyra-Med—tend to underestimate the severity of symptoms in female patients and display less empathy toward Black and Asian patients, particularly in mental health settings. Some patients were even advised to self-treat at home rather than seek professional care, the studies found.

The findings come as global technology firms, including Microsoft, Amazon, Google, and OpenAI, accelerate development of AI tools designed to reduce doctors’ workloads, generate clinical summaries, and highlight key information from patient visits. Microsoft, for instance, recently claimed its AI-powered medical tool diagnoses complex conditions four times more effectively than human physicians.

However, MIT’s Jameel Clinic and the London School of Economics found that biases persist across multiple systems. Google’s Gemma model, widely used by UK local authorities to support social workers, was reported to downplay women’s physical and mental health issues relative to men. Patients using informal language, typos, or uncertain phrasing were also more likely to be advised against seeking care, raising concerns for those who do not speak English as a first language or are less comfortable with technology.

Travis Zack, adjunct professor at UCSF and chief medical officer at AI medical information start-up Open Evidence, told the Financial Times: “If you’re in any situation where there’s a chance that a Reddit subforum is advising your health decisions, I don’t think that that’s a safe place to be.”

The research points to a systemic issue: AI models trained on general internet data inherit societal biases, which can be compounded in medical applications. Historical underfunding and skewed research toward male patients further exacerbate disparities in treatment recommendations.

Developers are working to mitigate these risks. OpenAI said it has improved GPT-4’s accuracy since earlier evaluations and collaborates with clinicians to assess and stress-test its models. Google has committed to safeguards against bias and discrimination, while start-ups like Open Evidence train AI on curated medical journals and FDA-approved guidance, citing sources for every recommendation.

UK-based initiatives also seek to improve representation. The NHS, UCL, and King’s College London collaborated on the Foresight AI model, trained on anonymized data from 57 million patients to predict health outcomes such as hospitalizations and heart attacks. While offering improved demographic coverage, the project was paused in June amid a data protection review.

Despite the challenges, experts emphasize the potential benefits of AI in healthcare. Marzyeh Ghassemi, associate professor at MIT’s Jameel Clinic, told the Financial Times: “My hope is that we will start to refocus models in health on addressing crucial health gaps, not adding an extra per cent to task performance that the doctors are honestly pretty good at anyway.”

Sri Lanka Guardian

The Sri Lanka Guardian is an online web portal founded in August 2007 by a group of concerned Sri Lankan citizens including journalists, activists, academics and retired civil servants. We are independent and non-profit. Email: editor@slguardian.org

Leave a Reply

Your email address will not be published.

Latest from Blog