As AI-assisted clinical decision support moves from concept to bedside reality, the assumption that more sophisticated reasoning capabilities would naturally reduce demographic bias turns out to be dangerously wrong. Understanding where and how these systems fail matters enormously for clinicians, health systems, and patients who may soon receive AI-informed care.

A systematic evaluation of 36,000 structured clinical vignettes tested two next-generation reasoning large language models — OpenAI's o3-mini and DeepSeek-R1 — for their tendency to associate medical conditions with specific racial or gender groups. Both models demonstrated consistent perpetuation of racial and gender stereotypes across common medical conditions. Critically, the enhanced chain-of-thought reasoning architecture that defines these models — explicitly designed to improve logical deliberation — did not confer meaningful protection against biased outputs. The pattern held broadly, not as isolated edge-case errors.

This finding lands at a particularly consequential moment in medical AI development. The prevailing hypothesis in the field has been that scaling model sophistication and reasoning depth would organically improve fairness by reducing pattern-matching shortcuts. This study challenges that assumption directly. Earlier work on standard large language models already documented demographic disparities in clinical text generation, but the expectation was that reasoning-optimized successors would perform better. The fact that they do not suggests the problem is upstream — embedded in training data distributions and reinforcement learning feedback loops — rather than a function of reasoning capacity alone.

Several important limitations apply. The vignette methodology, while scalable, is a controlled simulation rather than a real clinical workflow audit. Performance in synthetic scenarios may not map precisely onto deployment behavior. Additionally, the study does not disaggregate which condition categories or demographic intersections drive the strongest bias signals, leaving the mechanistic picture incomplete. Still, for health systems evaluating AI procurement, and for regulators developing fairness standards for medical AI, this is confirmatory and timely evidence that reasoning advancement is not a proxy for equity. Bias mitigation must be an explicit, independently validated design target.