AI Predicts High-Performance Proteins in One Round of Testing
- Designing proteins for specific functions – whether for new medicines, biofuels, or even laundry detergents – is traditionally a lengthy and iterative process.
- Published in the journal Science on February 19, 2026, MULTI-evolve predicts how proteins will behave when multiple amino acids are altered simultaneously.
- “It’s this very high-dimensional search problem where we effectively do guess and check,” explains Patrick Hsu, a bioengineer at the University of California, Berkeley, and the Arc Institute...
Designing proteins for specific functions – whether for new medicines, biofuels, or even laundry detergents – is traditionally a lengthy and iterative process. Researchers must systematically tweak the building blocks of proteins, known as amino acids, and then test the resulting changes in the lab. Now, a new machine learning framework called MULTI-evolve is poised to dramatically accelerate this process, potentially unlocking a new era of protein engineering.
Published in the journal Science on , MULTI-evolve predicts how proteins will behave when multiple amino acids are altered simultaneously. This is a significant advancement, as changing one amino acid can influence how subsequent changes affect the protein’s overall function. Finding optimal combinations of mutations often requires numerous rounds of modification and laboratory testing – a process that can be both time-consuming and resource-intensive.
“It’s this very high-dimensional search problem where we effectively do guess and check,” explains Patrick Hsu, a bioengineer at the University of California, Berkeley, and the Arc Institute in Palo Alto, Calif. MULTI-evolve aims to eliminate much of that guesswork.
The workflow behind MULTI-evolve involves a three-step process. First, researchers leverage existing data or machine learning techniques to predict the impact of single amino acid substitutions on protein function. Next, to understand how these mutations interact, they create and test proteins with pairs of those substitutions in the laboratory. This experimental data is then used to train a machine learning model, which is then tasked with predicting the performance of the target protein with five or more mutations.
The team successfully tested MULTI-evolve on three different proteins, including an antibody relevant to autoimmune diseases and a protein used in CRISPR gene editing. In each instance, the model identified several combinations of mutations that, when tested in the lab, outperformed the original, unmodified proteins. This suggests that the model is capable of identifying synergistic combinations of changes that would be difficult to discover through traditional methods.
The potential applications of this technology are broad. Hsu highlights two particularly promising areas: developing proteins that can track the movement of other molecules within cells and creating improved gene therapies for individuals with enzyme deficiencies. “We’re excited about this work,” Hsu says. “I think there’s tremendous interest in how this actually changes the practice of science.”
The development of MULTI-evolve builds upon recent advances in artificial intelligence and protein structure prediction. A tool called AlphaFold2, which can accurately predict protein structures, has generated a vast database of over 214 million predicted structures. This wealth of data provides a foundation for machine learning approaches like MULTI-evolve, allowing researchers to quickly prioritize which proteins to study further.
Researchers at the Innovative Genomics Institute (IGI) have developed a computational approach that allows scientists to rapidly search these massive datasets to identify proteins with useful functions, such as new gene editors or therapeutic candidates. This new method, detailed in a paper published in Nature Communications, provides a practical way to narrow down large datasets and focus laboratory efforts on the most promising candidates.
The integration of laboratory experiments with machine learning is a key feature of MULTI-evolve. By combining computational predictions with empirical data, the framework can refine its accuracy and identify mutations that might be overlooked by purely computational approaches. This iterative process of prediction, experimentation, and refinement is likely to become increasingly common in the field of protein engineering.
While the initial results are promising, further research is needed to fully explore the capabilities of MULTI-evolve and to optimize its performance for a wider range of proteins. However, this new framework represents a significant step forward in our ability to design and engineer proteins with tailored functions, potentially leading to breakthroughs in medicine, biotechnology, and beyond.
