The quality of cancer outcomes research depends almost entirely on the accuracy and completeness of clinical registries — and those registries are riddled with gaps that no one has had the bandwidth to fix. A new AI framework challenges the assumption that high-fidelity endpoint extraction from messy, real-world clinical narratives requires armies of manual abstractors.
SCRIBE, a training-free, open-source system built on multi-stage large language model reasoning, was evaluated on a multi-center cohort of 2,065 lung cancer patients. Without requiring any task-specific model fine-tuning, SCRIBE extracted temporally precise recurrence events from unstructured clinical notes, halving the temporal localization error compared to simpler note-level inference approaches. The system compressed multi-year patient documentation by nearly two orders of magnitude in token volume, a critical efficiency gain for scalable deployment. Crucially, SCRIBE retained verbatim evidence linked to source records, enabling auditable expert verification rather than opaque black-box classification. Perhaps most striking: expert adjudication found that 47.8% of SCRIBE's apparent false positives were actually genuine recurrence events absent from official registries — meaning the AI was more complete than the gold standard it was being measured against.
This finding carries significant methodological weight for oncology research broadly. Clinical trials and observational studies that rely on registry-based recurrence endpoints may be systematically undercounting events, which could bias survival analyses and time-to-recurrence estimates in ways researchers haven't fully appreciated. SCRIBE's architecture — reconciling longitudinal evidence across time rather than processing isolated notes — reflects a maturing understanding of how LLMs can be applied to temporally complex medical records. Key limitations to acknowledge: this is a preprint, the cohort is lung cancer-specific, and real-world generalizability across tumor types and EHR systems remains untested. Still, the registry-auditing implication alone makes this an incrementally important contribution with potential to reshape how oncology endpoints are validated at scale.