Lipoprotein(a) remains one of cardiovascular medicine's most underdiagnosed risk factors — genetically determined, largely unresponsive to statins, and present at dangerous levels in roughly one in five adults. The bottleneck isn't treatment; it's identification. A machine learning approach that can prioritize who most warrants testing could meaningfully shift that equation for the tens of millions living with atherosclerotic cardiovascular disease.
The FIND Lp(a) model, developed within the Family Heart Foundation's quality improvement framework, applied a Light Gradient-Boosting Machine algorithm to deidentified electronic health records from 344,987 adults with confirmed atherosclerotic cardiovascular disease. Trained on 90% of the cohort and validated on the remaining 10% (n = 34,499), the model flagged 1,553 individuals as high-probability for elevated Lp(a). Among those flagged, 55.1% had confirmed Lp(a) at or above 125 nmol/L — the threshold associated with significantly elevated cardiovascular risk — compared to just 24.8% in the unfiltered test population. That represents a 2.2-fold screening enrichment. Submodels targeting higher thresholds (≥150 and ≥200 nmol/L) achieved enrichment ratios of 2.3 and 2.7, respectively. Medication use was the algorithm's strongest predictive feature category, outweighing diagnoses and lipid laboratory values.
This work sits at the intersection of two accelerating trends: the emergence of targeted Lp(a)-lowering therapies in late-stage trials, and the growing pressure on health systems to operationalize precision screening at scale. Enriching screening populations more than twofold could substantially reduce the number needed to test while concentrating resources on highest-risk individuals — a meaningful efficiency gain in resource-constrained clinical environments. Key caveats deserve attention: the cohort is drawn from a single database (Family Heart Database), the model was not prospectively validated across diverse health systems, and algorithmic performance may vary with different EHR data completeness. Still, as Lp(a)-specific therapies near regulatory consideration, this type of infrastructure-level screening tool could prove clinically consequential — making this an incrementally important, practically oriented contribution rather than a paradigm shift.