Low calcium in critically ill patients is a deceptively common and underappreciated threat — one that traditional ICU risk tools were never designed to handle well. For clinicians and patients alike, better mortality prediction in this population could mean earlier, more targeted interventions and fewer preventable deaths in intensive care settings.
This multicenter study assembled 13,979 adult ICU admissions with confirmed hypocalcemia (total serum calcium below 2.12 mmol/L) drawn from the widely used MIMIC-III and MIMIC-IV databases alongside two Grade III Level A Chinese hospitals — providing both internal and genuinely external validation cohorts. Eight machine learning algorithms were benchmarked against each other, with variable selection handled via LASSO regression to reduce overfitting. XGBoost and LightGBM gradient-boosting frameworks emerged as top performers, substantially outperforming conventional severity metrics such as APACHE II and SOFA on discrimination and calibration. Critically, SHAP (SHapley Additive exPlanations) interpretability analysis was incorporated to identify which physiological variables drove individual predictions — a step toward clinical trust and actionability.
This work sits at the intersection of two growing research streams: the recognition that ionized and total calcium dysregulation carries independent prognostic weight in sepsis and multi-organ failure, and the push to replace fixed-weight scoring systems with adaptive, data-driven risk stratification. The key methodological strength is genuine external validation across geographically distinct populations — a bar many single-center ML studies fail to clear. Limitations worth noting include the retrospective design, reliance on de-identified database records that may miss granular treatment context, a relatively small Chinese external cohort (n=209), and the absence of ionized calcium measurements, which are mechanistically more relevant than total calcium. This study is best characterized as confirmatory-progressive: it validates ML's superiority over legacy scores in a specific high-risk subgroup without yet establishing whether model deployment changes clinical outcomes — the essential next step.