The ability to rapidly pinpoint which antibodies in a complex immune response will actually bind a target pathogen has long been a bottleneck in vaccine design, therapeutic development, and pandemic response. A new preprint from Research Square proposes a computational shortcut that could meaningfully compress that timeline by eliminating the need for labor-intensive experimental selection steps.
The approach applies parameter-efficient fine-tuning to ESM-2, Meta's large protein language model, training it on mRNA-derived heavy-chain V(D)J sequences — the genetic segments that encode the core antigen-binding regions of antibodies. The resulting model, called Antigen Specificity Predictor (ASP), achieved high recognition accuracy for three clinically important antigens: SARS-CoV-2 spike protein, influenza hemagglutinin, and HIV gp120. When deployed on single-cell deep-sequencing data from immunized mice, ASP identified antigen-specific B-cell receptors at high frequency, and those predictions were experimentally validated. Critically, the model also showed meaningful performance on previously unseen human B-cell receptor sequences obtained through affinity-selection barcoding, and demonstrated biology-consistent patterns in bulk human repertoire data — suggesting the model learned generalizable immunological features rather than memorizing training sequences.
This work sits at the intersection of two fast-moving fields: large language model applications in structural biology, and next-generation immune repertoire analysis. Prior efforts to predict antibody-antigen binding computationally have generally required either high-resolution structural data or extensive experimental screening; ASP's reliance on raw sequencing data from unselected peripheral blood is a meaningful step toward practical deployment. That said, several important caveats apply. The study is a preprint, not yet peer-reviewed. Performance was demonstrated on three antigens, and generalizability to structurally diverse or novel targets remains to be established. The mouse-to-human translation step, while promising, involves additional uncertainty. If the approach holds up to peer review and broader antigen testing, it could accelerate both vaccine candidate screening and therapeutic antibody discovery pipelines substantially — an incremental but potentially high-leverage advance in applied immunology.