Spinal curvature assessment remains one of the most measurement-dependent judgments in musculoskeletal radiology — two experienced clinicians examining the same radiograph can reach meaningfully different Cobb angle estimates, which directly influences whether a patient is monitored, braced, or referred for surgery. A deep learning pipeline that can reliably classify scoliosis severity from standard imaging would reduce that variability and potentially compress diagnostic timelines.

This retrospective cohort study applied RadImageNet-based transfer learning to 816 anteroposterior full-body biplanar radiographs drawn from adult patients with either degenerative or idiopathic scoliosis, plus scoliosis-negative controls. The classification task was framed across three clinically meaningful severity tiers anchored to established Cobb angle thresholds: no scoliosis (0°–10°), mild (10°–30°), and severe (>30°). The model was trained on an 80/20 split, with Cobb angle labels derived from resident-generated vertebral segmentations subsequently reviewed for quality. Critically, no external validation cohort was included, limiting generalizability claims.

RadImageNet as a pretraining foundation is a meaningful methodological choice worth contextualizing. Unlike ImageNet weights trained on natural photographs, RadImageNet embeddings are built from annotated medical images spanning CT, MRI, and radiograph modalities — theoretically providing feature representations more attuned to radiological texture and anatomy. Prior scoliosis AI literature has largely focused on adolescent idiopathic scoliosis or on direct angle regression rather than severity classification in adults, making this application genuinely distinct. However, a cohort of 816 cases remains modest for a three-class deep learning task, and the absence of external validation is a material limitation that prevents confident clinical translation. The resident-generated ground truth also introduces labeling noise that could affect ceiling performance. Overall, this is a methodologically reasonable proof-of-concept with incremental significance — useful to the field as a benchmark, but requiring prospective multicenter validation before influencing clinical workflow.