Applying GPT-4o within a secure clinical enclave to 45,794 echocardiography reports from MIMIC-III, LLM-based extraction identified right ventricular pressure/volume overload in 31.6% of patients versus only 7.8% captured by conventional rule-based dictionary schema — a roughly four-fold gap. More strikingly, the co-occurrence of RV overload with pulmonary hypertension was detected in 21.3% of cases by the LLM, compared to just 1.6% using structured fields alone. Pulmonary hypertension prevalence remained consistent between methods (33.6%), suggesting the LLM's advantage lies specifically in recovering multidimensional, narrative-embedded phenotypes rather than simple categorical variables.

Right ventricular dysfunction is notoriously underdiagnosed in clinical practice because its hallmarks — tricuspid annular planar systolic excursion, fractional area change, S' velocity — are often documented narratively rather than in discrete database fields. This study positions LLMs as a potentially transformative phenotyping tool for cardiovascular epidemiology and retrospective cohort assembly, areas where structured EHR data chronically underperforms. The implications extend to heart failure registries, pulmonary arterial hypertension trials, and population-level cardiovascular risk stratification.

Critical caveats apply: this is a preprint not yet peer-reviewed, so findings should be interpreted cautiously. The study is retrospective and non-causal, GPT-4o outputs were not systematically validated against manual cardiologist adjudication at scale, and MIMIC-III's ICU-heavy population limits generalizability. Still, the magnitude of the detection gap is striking enough to warrant prospective validation. Editorially, this reads as paradigm-shifting for clinical NLP — if confirmed.