Every minute matters in a crowded emergency department, and the decision about who gets seen first can determine survival. The question of whether artificial intelligence can reliably perform or enhance that judgment — across diverse patient populations, under real-world conditions — has enormous implications for healthcare systems already straining under capacity pressure. This scoping review offers the most systematic mapping of that question to date.

Drawing from five major academic databases and spanning publications from 2015 to 2026, the review screened 1,865 records and identified 27 studies meeting rigorous inclusion criteria. AI approaches evaluated included conventional machine learning, deep learning, natural language processing (NLP), artificial neural networks, and large language models (LLMs). Applications ranged from acuity classification and sepsis detection to ICU admission prediction, mortality forecasting, and patient-flow optimization. Models incorporating structured triage variables — vital signs, chief complaint codes, and nursing assessments — alongside NLP-derived text from clinical notes generally demonstrated the strongest predictive performance. LLMs, while emergent in the landscape, remained largely under-validated in live ED settings.

What this review reveals is not a technology problem but a translation problem. Despite technically promising predictive metrics in controlled evaluations, the evidence base is fragmented by heterogeneous study designs, inconsistent outcome definitions, and — critically — near-universal underreporting of algorithmic performance across demographic subgroups. Equity concerns are not a footnote here; they are a central unresolved challenge. If AI triage tools inherit historical biases embedded in training data, they risk systematically deprioritizing already-underserved populations at their most vulnerable moments. Additionally, most studies evaluated AI in retrospective or siloed settings, leaving real-world integration — workflow friction, clinician trust calibration, liability frameworks — largely unaddressed. This review is confirmatory in signaling AI's potential but sobering in mapping how far the field remains from deployment-ready, equitable clinical tools.