The assumption that regulatory clearance guarantees superior clinical performance is one of medicine's quietly held beliefs — and a new benchmarking study challenges it directly. When board-certified breast cancer specialists evaluated AI-generated treatment plans without knowing their source, a commercially available general-purpose model consistently outperformed two medically specialized, regulatory-cleared systems. The implications touch every clinician, hospital system, and patient who may soon interact with AI-assisted oncology care.

The blinded, multicenter study enrolled evaluators from seven university breast cancer centers across a structured set of 20 standardized patient cases. Three AI systems were tested: Prof. Valmed and OpenEvidence — both holding regulatory clearance as medical devices — and ChatGPT-5 Thinking, a general-purpose large language model. Specialist raters assessed outputs across six dimensions: safety, guideline adherence, medical adequacy, completeness, overall quality, and logical coherence. ChatGPT-5 Thinking ranked significantly higher across all six categories. The two specialized systems showed no meaningful performance differences from each other. The tradeoff was processing time: ChatGPT-5 Thinking required roughly 159 seconds per case versus 35 seconds for Prof. Valmed and just 9 seconds for OpenEvidence — a fourfold to eighteenfold difference.

This finding sits at an uncomfortable intersection of regulatory science and clinical AI deployment. Regulatory clearance frameworks, including those from the FDA and EU MDR, currently emphasize safety validation and intended-use documentation over head-to-head performance benchmarking against frontier general-purpose models. The pace of general-purpose LLM development may be outrunning specialized clinical AI, at least on benchmark tasks. Key limitations warrant caution: 20 cases is a modest evaluation set, the study tests standardized rather than real-world cases, and specialist rating of AI outputs does not capture downstream patient outcomes. Whether higher specialist ratings translate to fewer medical errors or better survival in practice remains undemonstrated. Still, for health systems evaluating AI procurement, this study introduces a meaningful evidence-based argument for regularly stress-testing specialized clinical tools against general-purpose alternatives — not merely accepting regulatory clearance as a performance proxy.