As artificial intelligence systems increasingly influence clinical decision-making, the absence of a rigorous standard for what constitutes genuine medical superintelligence represents a significant gap — one with real consequences for how hospitals, regulators, and patients trust and deploy these tools. The argument emerging from Nature Medicine is that the field is flying partly blind, relying on benchmarks that were never designed to capture the full complexity of real-world medical reasoning.

The core proposition advanced here is that existing evaluation frameworks for medical AI — predominantly multiple-choice question datasets, imaging classification tasks, and static knowledge tests — are structurally inadequate for assessing whether a system has crossed into superintelligent territory. The authors call for a task-based framework that would define medical AI superintelligence through measurable, clinically grounded performance criteria rather than proxy metrics. The argument is not that current AI is superintelligent, but that the field lacks the conceptual and empirical scaffolding to know when and if it ever gets there.

This perspective arrives at a moment when the gap between benchmark performance and clinical utility has become increasingly difficult to ignore. Systems that ace USMLE-style questions can still fail on ambiguous, longitudinal, or multi-modal patient cases — precisely the situations where superintelligence would matter most. The call for task-based measurement echoes longstanding critiques in AI evaluation more broadly, where Goodhart's Law — that a measure ceases to be a good measure once it becomes a target — has repeatedly undermined progress claims. From a health and longevity standpoint, the stakes are high: misplaced confidence in AI diagnostic or treatment-recommendation systems could delay appropriate care or introduce systematic errors at scale. This is a conceptually important contribution, but as a perspective piece rather than an empirical study, its practical impact will depend entirely on whether the research community and regulatory bodies adopt the proposed framework.