In the Bio-Hermes-001 cohort — 945 participants including a notably diverse sample (roughly one-third from populations underrepresented in dementia research) — only the BDSI algorithm exceeded conventional discrimination thresholds (AUC ≥ 0.8) for healthy cognition vs. probable Alzheimer's disease classification. Crucially, when the BDSI's functional assessment item was removed, DeLong tests confirmed statistically significant drops in discriminatory performance across all comparisons, including amyloid PET and phosphorylated tau-217 (pTau-217) binary outcomes. The two oldest published algorithms performed at near-chance accuracy.
This finding matters for several reasons. Biomarker-defined Alzheimer's staging — using amyloid PET and pTau-217 blood tests — is rapidly reshaping both clinical diagnosis and trial recruitment, yet most legacy risk algorithms were calibrated in older, less diverse cohorts before these biomarkers were standard. The demonstration that functional status captures variance that demographic and cognitive variables alone miss aligns with emerging evidence that instrumental activities of daily living deteriorate subtly well before formal dementia thresholds, potentially serving as a low-cost, scalable screening signal. Practically, clinicians and trialists should reconsider relying on older scoring tools without functional components. Limitations include the cross-sectional design, which cannot establish whether functional decline predicts future conversion, and the relatively modest pTau-217 negative sample (n=166). As this is a preprint posted on medRxiv and not yet peer-reviewed, findings require independent replication before influencing clinical protocols.