A large language model — Llama 3.3 70B — achieved 100% sensitivity and 93.8% specificity in identifying percutaneous coronary intervention (PCI) procedures within 1,412 cardiac catheterization reports drawn from three Yale New Haven Health hospitals. For classifying procedures as "complex PCI" — defined by criteria including chronic total occlusion, bifurcation stenting, ≥3 stents, or total stent length ≥60 mm — the model reached 97.7% sensitivity and 99.2% negative predictive value across 590 evaluable notes, though positive predictive value lagged at 57.6%, reflecting meaningful false-positive complex classifications.
The finding matters because manual chart abstraction for cardiovascular quality registries and research is notoriously expensive and slow, creating bottlenecks that limit population-scale insights. Automating this pipeline could dramatically accelerate real-world evidence generation for interventional cardiology — a field where procedural complexity directly predicts outcomes and reimbursement. Notably, Llama 3.3 70B outperformed both domain-specific medical LLMs (Meditron-7B and BioMistral-7B), reinforcing a growing pattern: raw model scale often beats narrow fine-tuning for clinical NLP tasks. The lower performance on contextually demanding variables like lesion counting and bifurcation classification highlights that LLMs still struggle with implicit clinical reasoning embedded in unstructured prose. The relatively modest positive predictive value for complex PCI (57.6%) warrants caution before deploying this as a standalone classifier without human oversight. As a preprint not yet peer-reviewed, these results require independent validation before clinical or administrative adoption.