Detecting vanishingly small populations of residual cancer cells after treatment is one of oncology's most consequential measurement challenges. In acute leukemia, even a fraction of a percent of surviving leukemic cells can predict relapse, making the accuracy and reproducibility of minimal residual disease (MRD) assessment directly tied to patient survival decisions. A new feasibility study suggests that a class of unsupervised machine learning techniques could reduce the analyst-dependence that currently undermines this critical measurement.

Researchers retrospectively applied three dimensionality reduction algorithms — t-SNE, UMAP, and PaCMAP — to 58 bone marrow specimens from 32 patients with either acute myeloid leukemia or B-lineage acute lymphoblastic leukemia, all collected at low leukemic burden (under 5% malignant cells). Using a FlowJo plugin after lineage-based pre-gating with markers including CD19, CD34, and CD117, the approach identified MRD in 46 specimens with leukemic cell percentages closely concordant with expert manual analysis. Notably, two specimens previously classified as MRD-negative by routine gating were reclassified as positive — cases where abundant normal cellular counterparts had obscured the malignant population during conventional analysis.

This finding carries genuine clinical weight. Standard multiparametric flow cytometry MRD analysis is notoriously operator-dependent: results can diverge meaningfully between laboratories or even between analysts within the same lab, creating real ambiguity in therapy stratification. The appeal of this approach is its reference-free design — unlike supervised machine learning models, it requires no large annotated training datasets or standardized panel harmonization across institutions, both persistent barriers to clinical AI adoption in hematology. However, the cohort of 32 patients is small and retrospective, and the two MRD reclassifications, while intriguing, cannot yet be validated against downstream clinical outcomes. Whether those reclassifications represent true positives or algorithmic artifacts remains an open question. This is incremental but technically promising work — a credible proof-of-concept that warrants prospective multicenter validation before clinical deployment.