An XGBoost gradient-boosted model trained on 9,050 hospitalized heart failure patients across eight Singapore hospitals achieved external-validation AUROCs of 0.862 for 30-day mortality and 0.749 for 1-year mortality — outperforming three recalibrated established scores (MAGGIC, Singapore HF Score, OPTIMIZE-HF-Asia) that ranged from 0.711–0.754 and 0.682–0.699, respectively. Critically, just ten bedside variables reproduced the full model's 30-day discrimination, and among high-risk patients with reduced ejection fraction, 36.3% received none of the four guideline-directed medical therapy pillars versus only 2.0% of low-risk patients.

Heart failure risk stratification has long leaned on scores built in predominantly white Western cohorts — a significant limitation given documented ethnic differences in HF phenotype, comorbidity burden, and drug response. This multi-ethnic Asian dataset spanning Chinese, Malay, and Indian populations addresses a genuine evidence gap. The prediction-to-action framing — flagging GDMT gaps in the highest-risk patients — moves beyond prognosis toward clinical utility, which is where most AI health tools stall. The ten-variable shortlist is especially practical for resource-constrained settings. However, important caveats apply: this is a retrospective observational dataset, so causal inference about GDMT gaps is limited. The finding that the 30-day advantage over MAGGIC disappeared without medication variables warrants scrutiny. As a preprint not yet peer-reviewed, results may change substantially upon independent review. If validated prospectively, this represents a genuinely actionable advance over current standard-of-care risk tools.