As oncology care grows more complex and multidisciplinary tumor boards (MTBs) face mounting resource pressures, the question of whether AI can meaningfully assist clinical decision-making has moved from theoretical to urgent. A new single-center retrospective study probes that question directly — and the answer is more nuanced than either AI optimists or skeptics might expect.
Analyzing 100 cases of primary liver cancer — 50 cholangiocarcinoma (CCA) and 50 hepatocellular carcinoma (HCC) — from 2022–2023 MTB protocols, investigators submitted identical structured clinical summaries to ChatGPT and Claude and compared their outputs to actual board recommendations. For CCA, ChatGPT achieved 80% exact concordance with MTB decisions, with substantial Cohen's kappa agreement (κ = 0.688) and a strong Spearman correlation (r = 0.725). Performance dropped for HCC — 66% concordance, κ = 0.604, r = 0.484 — but remained statistically meaningful. Claude's performance diverged sharply: 56% concordance for CCA and a striking 38% for HCC, with only fair agreement (κ = 0.314) and a correlation indistinguishable from chance (r = 0.086, p = 0.551).
These findings surface several layers worth unpacking. First, the dramatic performance gap between two leading LLMs — both trained on vast biomedical literature — underscores that model architecture, training data curation, and fine-tuning choices produce clinically non-interchangeable outputs. ChatGPT's stronger alignment likely reflects superior integration of structured oncologic guidelines, but without model transparency this remains speculative. Second, the study captures cases only at the point of MTB presentation, meaning LLMs were not exposed to dynamic deliberation, imaging nuance, or the patient-specific context that boards routinely weigh — which may explain residual discordance. Third, as a single-center retrospective study with 100 cases, generalizability is limited; different institutional protocols and patient populations could shift concordance substantially. This work is best read as confirmatory that LLMs can approximate guideline-concordant reasoning in structured oncologic scenarios — incremental rather than paradigm-shifting — but the inter-model variability alone is a meaningful finding that should inform any clinical AI deployment decision.