The integration of artificial intelligence into clinical care has moved well beyond theoretical debate — nursing, the largest healthcare workforce globally, is now actively grappling with where AI tools help and where they may quietly harm. Understanding these boundaries matters because nurses make thousands of judgment calls daily, and AI errors embedded in those workflows could cascade through patient outcomes in ways that aggregate data may not immediately surface.
This narrative review, drawing on literature from January 2019 through March 2026 across Medline, Scopus, and arXiv, examined large language model (LLM) performance across three nursing domains: education, clinical practice, and workflow management. In educational settings, tools like ChatGPT functioned as adaptive cognitive scaffolds — supporting simulation training and virtual patient encounters — but unregulated use was linked to risks around academic integrity and the development of independent clinical reasoning. In clinical practice, LLMs performed adequately for preliminary symptom triage and generating patient education materials, yet accuracy degraded substantially in complex or data-sparse scenarios. Critically, hallucination rates — instances where the model generates plausible but factually incorrect output — remained at clinically significant levels, meaning errors were not rare edge cases but a systematic concern. Workflow management applications showed efficiency gains, though the review flagged governance gaps as a persistent structural issue.
This synthesis arrives at a moment when healthcare institutions are making binding deployment decisions without mature regulatory frameworks. The review's findings align with a growing body of evidence suggesting LLMs perform best as decision-support tools under human supervision, not as autonomous clinical agents. A key limitation here is the narrative review design itself: without meta-analytic pooling, effect sizes and error rates cannot be precisely quantified across studies. Additionally, nursing's extraordinary contextual diversity — from ICU to community health — means generalizations carry risk. The finding on hallucinations is arguably the most important: any deployment strategy that does not build systematic verification steps around LLM outputs is accepting patient safety risk that current evidence does not justify.