Boston University researchers have created an antibody-specific artificial intelligence framework that improves predictions of how tightly antibodies bind to disease targets by up to 27% compared with existing methods. The study, published in Communications AI & Computing, describes a model containing about 600 million parameters that matches or outperforms much larger antibody language models on benchmark tests.

Unlike general protein language models that treat all amino acids as equally important, the new approach focuses learning on the six complementarity-determining regions (CDRs) — tiny loops that determine which target an antibody recognizes and how strongly it binds. The researchers trained the model on more than 1.6 million naturally paired antibody heavy and light chains, deliberately masking up to half of the amino acids within the CDRs during training while leaving the surrounding scaffold largely intact.

Principal investigator Diane Joseph-McCarthy, executive director of BU's Bioengineering Technology & Entrepreneurship Center, said the team designed an antibody-specific language model that learns fundamental patterns in the regions responsible for antigen binding rather than treating antibodies like generic proteins. Co-author Ioannis Paschalidis, director of BU's Hariri Institute for Computing, noted that smaller domain-specific models trained on high-quality data can often outperform larger, more general ones.

The model was evaluated on datasets containing more than 90,000 engineered antibody variants targeting six different antigens. By training on millions rather than billions of sequences, it required substantially less computational effort than many existing antibody AI models while delivering improved binding affinity predictions.

A major bottleneck in antibody discovery is deciding which candidates to test experimentally, since even small sequence changes can create trillions of possible variants. Co-author John Misasi, assistant professor at BU's Chobanian and Avedisian School of Medicine and core faculty at the National Emerging Infectious Diseases Laboratories, said the model helps narrow millions of possibilities down to the few hundred most promising candidates, saving time, labor and cost in the laboratory.

Beyond candidate selection, the approach could improve antibody engineering by predicting which sequence changes are most likely to strengthen existing antibodies. This could help optimize therapies against evolving viruses or other disease targets before moving designs into experimental testing, potentially enabling faster responses to infectious disease outbreaks.

The research emerged from a convergent effort combining expertise in artificial intelligence, immunology, structural biology and experimental science. The researchers believe biologically informed AI may ultimately transform therapeutic antibody development by helping scientists better understand antibody-antigen recognition, prioritize candidates for testing and design more effective therapies.

Sources and further reading

Teaching AI the biology of antibodies speeds drug discovery

This is an independent summary. The complete reporting, supporting context and any primary documents remain with Phys.org.