The gap between what AI promises and what it reliably delivers in the operating room has never been more consequential — or more honestly assessed. As large language models penetrate clinical workflows, a rigorous appraisal of their surgical applications reveals a technology that is genuinely more capable than its predecessors yet still constrained by foundational problems that have persisted for decades.
The historical arc traced in this analysis is instructive. Early clinical decision support in surgery, exemplified by Bayesian systems like AAPHelp for acute abdominal pain diagnosis in the 1970s, collapsed under the weight of complexity and opacity in the 1980s. The current revival, powered by machine learning advances and large clinical datasets from the 2010s onward, has produced tools capable of individualized postoperative complication risk prediction and automated clinical documentation. Modern large language models extend these capabilities further, yet the same structural obstacles — algorithmic bias, limited transparency, reliability concerns, and difficult workflow integration — continue to limit real-world deployment.
This historical pattern deserves serious weight. The surgical AI field has now cycled through at least two waves of optimism separated by stagnation, suggesting that technical capability alone is insufficient without parallel advances in validation methodology, regulatory frameworks, and clinician trust. From a health-systems perspective, postoperative complication prediction represents perhaps the highest-value near-term application, given that complications drive substantial morbidity and cost. However, most current models are trained on narrow institutional datasets, limiting generalizability. The bias problem is particularly acute in surgery, where patient populations, operative techniques, and outcomes vary enormously across demographics and institutions. This review reads as confirmatory rather than paradigm-shifting — it synthesizes a well-recognized tension — but its value lies in calibrating expectations precisely when commercial enthusiasm risks outpacing clinical evidence.