The persistent gap between what clinicians know about a patient's biology and what they know about a patient's life circumstances — housing instability, food insecurity, transportation barriers — may be one of the most underappreciated drivers of poor health outcomes. A structured effort to close that gap using conversational AI represents a meaningful shift in how health systems could routinely capture social context at scale.

Published in JMIR Formative Research, this mixed-methods study describes the development and pre-deployment evaluation of a large language model-powered chatbot designed to collect social determinants of health (SDoH) data directly from patients. Researchers built a 10-criterion evaluation rubric derived from established healthcare AI frameworks and stress-tested it across 27 synthetic clinical scenarios representing diverse SDoH profiles. A licensed clinical social worker role-played patient interactions, and three multidisciplinary raters — a social worker, nurse practitioner, and physician — independently scored chatbot performance. Across simulated cases, the chatbot achieved high proportions of positive ratings, though ceiling effects in several domains complicated statistical interpretation via Fleiss kappa, prompting the team to rely on percent agreement metrics alongside iterative qualitative refinement.

The significance here is methodological as much as technological. The use of synthetic data simulation before clinical exposure addresses a genuine ethical bottleneck in patient-facing AI development — how to rigorously evaluate sensitive conversational tools without exposing vulnerable populations to undertested systems. This iterative, pre-deployment validation approach could serve as a replicable template for AI tools in other sensitive health domains. That said, key limitations are substantial: the 27-scenario sample is small, all interactions were simulated rather than drawn from real patients, and performance in synthetic conditions frequently overestimates real-world reliability. The chatbot has not yet been deployed clinically, so outcomes data — the ultimate test — remain absent. This work is best characterized as a promising methodological proof-of-concept, not yet evidence of clinical efficacy.