Advanced AI Models Perpetuate Racial and Gender Stereotypes in Medicine
Researchers at Flinders University have determined that next-generation reasoning large language models, specifically o3-mini and DeepSeek-R1, continue to reproduce racial and gender stereotypes when generating medical clinical vignettes. The study, published in the Journal of Medical Internet Research, challenges the assumption that enhanced computational reasoning inherently translates to improved representational fairness in healthcare applications. Building on prior research that identified demographic skew in earlier models like GPT-4, the Flinders team tasked the newer AI systems with generating 36,000 unique clinical scenarios involving common medical conditions. The analysis revealed persistent and in several cases elevated misrepresentation rates. While GPT-4 met the threshold for significant demographic distortion in 67 percent of conditions for both race and gender, o3-mini showed a 78 percent rate for race and 56 percent for gender. DeepSeek-R1 exhibited even higher racial misrepresentation at 89 percent, matching the earlier gender distortion rate. The median misrepresentation magnitude also increased, reaching 44 percent for o3-mini and 31 percent for DeepSeek-R1, compared to 15 percent for GPT-4. Both models consistently overrepresented Black patients in conditions such as sarcoidosis, systemic lupus erythematosus, preeclampsia, and essential hypertension. Qualitative analysis of DeepSeek-R1 reasoning traces indicated that the AI frequently defaulted to disease-demographic associations rather than referencing epidemiological data when selecting patient profiles. Lead researcher Joshua Docking warned that perpetuating these patterns could exacerbate health disparities and reinforce narrowed demographic assumptions in diagnostic workflows. Professor Michael Sorich emphasized that algorithmic improvements in reasoning do not guarantee parallel progress in ethical representation, noting that healthcare AI must be evaluated across multiple dimensions. The findings underscore a critical risk in integrating advanced language models into clinical environments. Researchers stress that awareness of these demographic defaults is essential for safe deployment, advocating for continuous bias monitoring alongside model adoption. Without targeted interventions, even state-of-the-art reasoning AI may inadvertently standardize stereotypical patient portrayals, ultimately compromising equitable healthcare delivery.
