When clinical guidelines steer physicians away from imaging and toward physical examination alone to diagnose childhood pneumonia, the accuracy of that guidance hinges on a critical assumption: that two trained clinicians examining the same child will reach the same conclusions. New prospective data challenge how confidently that assumption can be held.
Drawn from the PedCAPS cohort — an ongoing multicenter study spanning seven academic pediatric emergency departments across the United States — this planned analysis enrolled 252 children aged three months to seventeen years presenting with community-acquired pneumonia between August 2023 and May 2025. Each child was independently assessed by two examiners within a 60-minute window, generating paired observations of standard physical findings. Interrater reliability (IRR) was quantified using Fleiss κ, with a lower 95% confidence interval bound of 0.4 set as the threshold for acceptable agreement. The study design is notable for its real-world ED setting, prospective structure, and explicit exclusion of higher-complexity patients — factors that would be expected to make agreement easier, not harder, to achieve.
This finding matters most in context: pediatric community-acquired pneumonia drives roughly two million outpatient encounters and 375,000 emergency department visits annually in the US, and current guidelines explicitly discourage routine chest radiography in children eligible for outpatient management. If the physical signs anchoring that decision — auscultatory findings, respiratory pattern abnormalities, work-of-breathing assessments — carry meaningful examiner-to-examiner variability, the diagnostic chain linking examination to treatment becomes less reliable than guidelines may imply. IRR research in adult pneumonia has similarly revealed that lung auscultation findings like crackles and decreased breath sounds show only fair-to-moderate agreement between observers, suggesting this is not a pediatric-specific limitation but a broader challenge in clinical respiratory assessment. The study's strength is its prospective multicenter design and rigorous paired-observer methodology; a key constraint is that academic pediatric EDs may not reflect community practice, where examiner experience varies more widely. Overall, this qualifies as confirmatory but clinically important work that should prompt scrutiny of how much diagnostic weight guidelines place on unvalidated examination findings.