The assumption that artificial intelligence in medicine would remain a background assistant — surfacing information but leaving every clinical decision to humans — may be due for revision. A system capable of navigating the full arc of patient care, from history-taking through diagnosis to treatment planning, challenges not just the technology but the governance frameworks surrounding clinical AI adoption.
MIRA (Medical Intelligence for Reasoning and Action) is a large-language-model-based autonomous agent designed to operate directly inside a sandboxed electronic health record environment. Unlike previous clinical AI tools that addressed isolated subtasks or produced free-text suggestions, MIRA executes structured EHR actions: ordering and interpreting labs, imaging, and microbiology; generating differential diagnoses; prescribing medications; scheduling procedures; and planning admissions. Evaluated on real patient cases spanning multiple diagnostic categories, MIRA outperformed physicians on diagnostic accuracy and demonstrated guideline-concordant, medication-safe, and clinically appropriate admission decisions — a multi-dimensional performance benchmark that prior narrow systems have not approached.
This work sits at an inflection point in clinical AI research. Most published LLM evaluations test models on multiple-choice licensing exams or isolated question-answering benchmarks, which bear little resemblance to the iterative, action-heavy reality of clinical workflow. MIRA's sandboxed EHR architecture is a meaningful methodological step forward, though it also introduces the central limitation: simulation is not clinical deployment. Real-world EHR environments carry noise, incomplete records, time pressure, and liability structures that controlled sandboxes cannot replicate. The cohort of patient cases, while described as spanning multiple diagnoses, is not yet characterized by scale or demographic breadth sufficient to generalize. Physician performance benchmarks also vary substantially by specialty, experience, and case complexity. That said, if these findings replicate under prospective clinical trial conditions, they would represent a genuinely paradigm-shifting shift — moving AI from passive advisor to active clinical co-operator, with profound implications for diagnostic equity, workload reduction, and patient safety.