A preliminary head-to-head comparison tested whether paper-derived ECGs, digitized via ECGScan software, could substitute for native digital XML files in AI-driven cardiovascular risk modeling. Fifteen ECGs processed through an ensemble of five convolutional neural networks predicting 10-year heart failure risk showed strong rank-order correlation between native and digitized versions (Pearson r=0.804; Spearman ρ=0.893), though digitized ECGs produced slightly lower mean predicted probabilities (0.171 vs. 0.195).
The finding matters because decades of large epidemiologic cohorts — Framingham, ARIC, and similar studies — house millions of ECGs only as paper tracings or scanned PDFs, effectively locking them out of modern ECG-AI pipelines. If digitization tools can faithfully preserve the signal features that deep learning models rely on, those archives become actionable for retrospective cardiovascular risk research without costly re-collection.
However, critical caution is warranted. With only 15 ECGs, this is a proof-of-concept pilot, not a validation study — statistical conclusions are essentially anecdotal at this scale. The consistent downward shift in predicted probabilities from digitized files raises unresolved questions about systematic bias that could affect clinical thresholds. Scan quality, paper degradation, and printer resolution variability across real-world archives could widen divergence substantially. As an unreviewed preprint on medRxiv, these results have not undergone peer scrutiny and should be treated as exploratory. Overall, the finding is directionally encouraging but firmly incremental, requiring orders-of-magnitude larger validation before influencing research practice.