Designing an effective antibody drug is like searching for the right key in a warehouse of locks. Scientists may begin with millions-or even billions-of antibody candidates, but only a tiny fraction will recognize and bind tightly to the disease target. Identifying those rare candidates has long been one of the biggest challenges in developing antibody medicines. 

Boston University researchers have now developed an antibody-specific AI framework that dramatically narrows that search. Rather than building a larger AI model, the team redesigned how AI learns, focusing it on the small regions of antibodies that recognize disease targets. 

Instead of treating antibodies like generic proteins, we designed an antibody-specific language model that learns the fundamental patterns in the regions responsible for antigen binding. That focused approach helps researchers identify the most promising therapeutic candidates before they ever enter the laboratory." 

Diane Joseph-McCarthy, PhD, study's principal investigator and executive director of Boston University's Bioengineering Technology & Entrepreneurship Center

The study was published today in the Nature Portfolio journal Communications AI & Computing. Researchers found that the approach improved predictions of antibody binding strength, known as binding affinity, by as much as 27 percent while requiring far fewer computational resources than many existing antibody AI models.

Why antibodies challenge AI

Artificial intelligence has transformed biology by identifying patterns across millions of protein sequences. Like ChatGPT predicts missing words, protein language models predict masked amino acids to learn the "language" of proteins. 

For most proteins, randomly hiding amino acids throughout a sequence is an effective training strategy because biologically important information is distributed across the molecule. Reconstructing the missing pieces helps the model discover the patterns that determine protein structure and function. 

Antibodies, however, are different. 

The regions responsible for recognizing viruses, bacteria, and cancer cells continually evolve so the immune system can adapt to new threats. That diversity makes antibodies extraordinarily powerful-and more difficult for AI to model. 

Most of an antibody serves as a structural scaffold. The information that determines what an antibody recognizes and how tightly it binds is concentrated within six tiny loops called complementarity-determining regions, or CDRs. 

"Think of an antibody like a screwdriver," says John Misasi, MD, a study co-author and assistant professor of virology, immunology, and microbiology at Boston University's Chobanian and Avedisian School of Medicine. "It doesn't matter whether it's long or short-the shape of the tip determines what kind of screw it fits. The CDRs are like that tip: they determine which target the antibody recognizes." 

Teaching AI the biology that matters

Recognizing that antibodies break many of the assumptions behind general protein language models, the researchers redesigned the training process around antibody biology instead of treating every amino acid as equally important. 

The model focused its learning on the CDRs-the regions directly responsible for recognizing disease targets-and was trained using more than 1.6 million naturally paired antibody heavy and light chains that together form the binding site. During training, the researchers deliberately masked up to half of the amino acids within the CDRs while leaving most of the surrounding antibody structure intact, repeatedly challenging the AI to reconstruct the regions most critical for binding. 

"A lot of AI research has focused on building larger models," says Ioannis (Yannis) Paschalidis, PhD, a co-author of the study and director of Boston University's Hariri Institute for Computing. "We asked a different question: How can we teach the model the biology that matters most? That turned out to be a much more effective strategy." 

The result was a smaller, more focused model containing about 600 million parameters that matched or outperformed much larger antibody language models on multiple benchmark tests. It improved binding affinity prediction by as much as 27 percent across datasets containing more than 90,000 engineered antibody variants targeting six different antigens. By training on millions rather than billions of antibody sequences, it required substantially less computational effort. 

"One of the exciting findings is that we didn't need a larger model or vastly more data," says Paschalidis. "That's similar to what's been observed with human-language AI models, where smaller domain-specific models trained on high-quality data can often outperform much larger, more general ones." 

The approach emerged from a convergent research effort that brought together expertise in artificial intelligence, immunology, structural biology, and experimental science-not simply to apply AI to biology, but to redesign how AI learns using biological knowledge. 

From prediction to prioritization

The study addresses one of the biggest bottlenecks in antibody discovery: deciding which candidates to test. Even small changes to an antibody's sequence can create trillions of possible variants-far more than laboratories can realistically evaluate experimentally. The model helps narrow those possibilities before laboratory testing, reducing unnecessary experiments and accelerating antibody optimization. 

"Knowing not just whether an antibody binds, but how strongly it binds, gives researchers a much better starting point for deciding which candidates to move forward," says Misasi, core faculty at BU's National Emerging Infectious Diseases Laboratories (NEIDL). "If a computer can narrow millions of possibilities down to the few hundred most promising candidates, that saves an enormous amount of time, labor, and cost in the laboratory." 

Beyond selecting antibody candidates for testing, the approach could improve antibody engineering. Researchers could use the model to predict which sequence changes are most likely to strengthen existing antibodies, helping optimize therapies against evolving viruses or other disease targets before moving those designs into experimental testing. 

"If another infectious disease outbreak occurs, we'd like to identify promising antibody candidates as quickly as possible," says Misasi. "Computational tools like this could help us find those candidates sooner and even suggest how existing antibodies might be adapted as viruses change over time." 

Looking ahead

The findings suggest that biologically informed AI may offer a more effective path for antibody discovery. Beyond this study, the researchers believe biologically informed AI could ultimately transform therapeutic antibody development by helping scientists better understand antibody-antigen recognition, prioritize candidates for laboratory testing, and design more effective therapies. 

"Understanding how antibodies recognize their targets is fundamental to developing better antibody therapies, diagnostic tests, and vaccines," says Joseph-McCarthy. "By teaching AI the biology that matters most, we hope to give researchers better tools to discover, optimize, and ultimately design the next generation of antibody therapies." 

Source:

Journal reference:

Talaei, M., et al. (2026). Preferential CDR masking in paired antibody language models improves binding affinity prediction. Communications AI & Computing. DOI: 10.1038/s44488-026-00010-2. https://www.nature.com/articles/s44488-026-00010-2