A Transformer-based deep learning model trained on 100,056 Japanese adults without baseline cardiovascular disease achieved a 10-year ROC-AUC of 0.821 — outperforming Cox regression, XGBoost, multilayer perceptron, the Framingham Risk Score, and the Hisayama Risk Score. Trained on 2010–2024 longitudinal health checkup records and externally validated in 79,756 Kanazawa City participants (where AUC held at 0.762 despite 26.6% event rate), the model used anthropometric, laboratory, and self-reported lifestyle inputs. SHAP analysis flagged age, ECG abnormality, antihypertensive medication use, and sex as top predictors, while an attention network revealed that daily exercise and weight gain dynamically modulated age-related risk.

The result matters because traditional linear risk scores like Framingham were built on relatively small, demographically homogeneous cohorts and struggle with nonlinear feature interactions. A 0.821 AUC in a real-world population of 100,000+ represents a meaningful leap — roughly 4–8 AUC points above typical Framingham performance in Asian populations. The external validation is a genuine strength; many AI health models fail here. However, CVD events were self-reported physician diagnoses, introducing recall and ascertainment bias. The cohort is Japanese, limiting direct generalizability to other ethnicities. Causal inference is not possible from this observational design. Critically, this is a preprint posted on medRxiv and has not yet undergone peer review — findings, particularly the performance metrics and clinical interpretations, should be treated as preliminary until independently validated. If confirmed, this architecture could meaningfully reshape population-level cardiac screening programs.