Breakthroughs of AI Enzyme Molecular Design! Insufficient Data Fine-tuning of Large Models Enables Rapid Evolution of Enzyme Stability
Source: Hzymes Market Center
Date: 2025-01-09
Views: 812

AI-assisted enzyme modification technology has achieved another breakthrough. The founder and CTO of Hzymes Biotech, Professor Dr. Guangyu Yang from the School of Life Sciences and Technology at Shanghai Jiao Tong University, in collaboration with Professor Liang Hong’s team from the Institute of Natural Sciences at Shanghai Jiao Tong University, has applied the Pro-PRIME protein language large model and efficient model fine-tuning to achieve a significant increase in protein stability in just two rounds, with a 100% gain success rate for composite mutations.


Professor Guangyu Yang CTO, Hzymes Biotech


This technology is based on protein large models, using only a limited dataset of target enzyme proteins’ (mutation-activity-stability) annotated data (<100), to accurately predict the impact of composite mutations on the stability of target enzyme proteins, assisting in the sequence design process of enzyme stability modification. Compared to the traditional high-throughput screening method, it has the advantages of high efficiency, low cost, less workload, and easy screening, providing a new approach and direction for enzyme modification based on AI technology.




Overview


Optimizing enzyme thermostability is crucial for advancements in protein science and industrial applications. Currently, (semi-)rational design and random mutagenesis methods can accurately design multiple single-point mutations that enhance enzyme thermostability. However, when combining multiple mutations, complex epistatic interactions often arise, leading to the complete inactivation of combinatorial mutants. As a result, constructing an optimized enzyme often requires repeated rounds of design to incrementally incorporate single mutation sites, which is highly time-consuming.

Recently, an article titled “Optimizing enzyme thermostability by combining multiple mutations using protein language model” by the team of Professor Guangyu Yang, the founder and CTO of our company, from the School of Life Sciences and Technology at Shanghai Jiao Tong University, was officially published in mLife, with Professor Liang Hong from the Institute of Natural Sciences at Shanghai Jiao Tong University as a co-corresponding author. The research team proposed an AI-assisted enzyme thermostability engineering strategy that can efficiently combine multiple beneficial single-point mutations. In the evolution example of creatinase, only two rounds of design were needed to obtain 50 combinatorial mutants with superior thermostability, achieving a success rate of 100%. The model, after fine-tuning with a small amount of experimental data, can effectively capture epistatic effects in combinatorial mutants from the dataset.



Original Abstract


“Optimizing enzyme thermostability is essential for advancements in protein science and industrial applications. Currently, (semi-)rational design and random mutagenesis methods can accurately identify single-point mutations that enhance enzyme thermostability. However, complex epistatic interactions often arise when multiple mutation sites are combined, leading to the complete inactivation of combinatorial mutants. As a result, constructing an optimized enzyme often requires repeated rounds of design to incrementally incorporate single mutation sites, which is highly time-consuming. In this study, we developed an AI-aided strategy for enzyme thermostability engineering that efficiently facilitates the recombination of beneficial single-point mutations. We utilized thermostability data from creatinase, including 18 single-point mutants, 22 double-point mutants, 21 triple-point mutants, and 12 quadruple-point mutants. Using these data as inputs, we used a temperature-guided protein language model, Pro-PRIME, to learn epistatic features and design combinatorial mutants. After two rounds of design, we obtained 50 combinatorial mutants with superior thermostability, achieving a success rate of 100%. The best mutant, 13M4, contained 13 mutation sites and maintained nearly full catalytic activity compared to the wild-type. It showed a 10.19°C increase in the melting temperature and an ~655-fold increase in the half-life at 58°C. Additionally, the model successfully captured epistasis in high-order combinatorial mutants, including sign epistasis (K351E) and synergistic epistasis (D17V/I149V). We elucidated the mechanism of long-range epistasis in detail using a dynamics cross-correlation matrix method. Our work provides an efficient framework for designing enzyme thermostability and studying high-order epistatic effects in protein-directed evolution.”



Main Content


In this study, the authors used an AI-assisted enzyme thermostability engineering strategy to predict the stability and activity of combinatorial mutants by fine-tuning the Pro-PRIME model with a small amount of experimental data. The Pro-PRIME model is a protein language model trained on the optimal growth temperature data of 96 million host bacterial strains, which performs excellently in designing and optimizing high-temperature enzymes. The initial dataset used for fine-tuning includes sequence-thermostability and activity data from 73 low-order mutants of creatinase. Then, the fine-tuned model was used to predict the thermostability and activity of all possible mutants from 18 single-point mutants. The main goal was to identify mutants that maintain at least 60% relative activity (compared to the wild-type) while enhancing thermostability (Figure 1).



Figure 1. Strategy for combining mutations based on protein language models.

The entire process includes four steps:

1.Data collection
2.Fine-tuning of the protein language model
3.Predicting all mutants in the combinatorial sequence space
4.Verifying the selected mutants. The red dashed line indicates the second round of model fine-tuning.



To further improve prediction accuracy, researchers integrated the experimental characterization results of the first round of predictions into the dataset and conducted a second round of fine-tuning, prediction, and selection. The two rounds of fine-tuning and prediction took only two weeks, resulting in the design of 50 combinatorial mutants with a 100% success rate in thermostability design (Figure 2).


Figure 2. Thermostability and relative activity data of combinatorial mutants.

Yellow circles indicate relative activity data. Bar charts indicate thermostability data of mutants, with blue, cyan, and orange representing the initial dataset, the first round, and the second round of predicted datasets, respectively.

Among them, the best mutant 13M4 contains 13 mutation sites. Compared to the wild-type, its activity remains almost unchanged, with a Tm increase of 10.19°C and a half-life increase of about 655 times at 58°C.

Upon reviewing the data, it was found that even some mutations that are spatially distant from each other exhibit complex high-order epistatic effects. For example, the single-point mutation K351E is a negative mutation, but it behaves as a positive mutation in high-order mutants. In addition, single-point mutations D17V and I149V show significant synergistic effects. The results indicate that fine-tuning the model parameters with high-quality experimental data can help the model accurately capture existing epistatic effects in the dataset and predict the fitness of subsequent high-order combinatorial mutants.

The results of the dynamic correlation matrix analysis show that mutations affecting stability not only affect the dynamics of their local environment but also affect the dynamics of distant structural regions in some cases (Figure 3). This technology can serve as an effective tool for future research or design of epistatic effects.

Figure 3. Epistatic effect analysis between mutations.

Epistatic effects of K351E (A) and D17V/I149V (B) on Tm values. Blue indicates negative effects, and orange indicates positive effects. (C) Dynamic cross-correlation matrix map of wild-type creatinase and corresponding mutants. Correlation coefficients (Cij) are represented by different colors. Mutation sites are indicated by red arrows, and significant dynamic correlation regions around mutations are highlighted by red boxes. (D) Standardized RMSF changes comparing the structure of mutants with the wild-type structure.



Main Highlights


PART.01


The AI-assisted enzyme thermostability engineering strategy proposed in this study can efficiently combine multiple beneficial single-point mutations. Only two rounds of design were needed to characterize 50 combinatorial mutants, with a 100% success rate in stability design. Compared to the wild-type, the best mutant 13M4 increased Tm by 10.19°C, increased the half-life at 58°C by 655 times, and maintained unchanged catalytic activity.

PART.02


By fine-tuning the parameters of the protein language model with a small amount of high-quality experimental data, the fine-tuned model can accurately capture epistatic effects in the initial dataset, including sign and synergistic epistatic effects. This indicates that experimental data is crucial for improving the model’s prediction performance for high-order combinatorial mutants.

PART.03


Using dynamic correlation matrix analysis, this study revealed the mechanism of long-range epistasis, showing the dynamic correlation between distant mutations, which together affect the stability of the mutants.

PART.04


By adopting this strategy, the research team comprehensively explored over 260,000 possible mutants in the combinatorial sequence space through only two rounds of design. The best mutant contained 13 mutations, greatly reducing the number of evolutionary rounds required by traditional methods and improving the efficiency of protein engineering.

PART.05


The study emphasizes that combining data from protein engineering with advanced AI models can further enhance the model’s predictive performance, thereby improving the efficiency of protein engineering. This strategy can be widely applied to the evolutionary tasks of various key enzyme molecules.

Scan the QR code to browse the original article immediately



Message
Leave Your Message
Name *
Company *
Tel/WhatsApp *
Mail *
Nation *
Descriptions

Please contact on WhatsApp

Service Hotline: +86 400-808-5320

Large-scale production base: Building 6, Precision Medical Industry Base, Wuhan, China

Logistics & Supply Chain Center:417 Main St, Little Rock, AR 72201. United States.

Global Marketing Center: Hzymes Building, Fengxian District, Shanghai, China.

  • iso_copy_copy
  • iso_copy
  • iso
  • iso_copy_copy
  • iso_copy
  • iso
Contact Us

Service Hotline: +86 400-808-5320

Large-scale production base: Building 6, Precision Medical Industry Base, Wuhan, China.

Logistics & Supply Chain Center:417 Main St, Little Rock, AR 72201. United States.

Global Marketing Center: Hzymes Building, Fengxian District, Shanghai, China.

Copyright © Hzymes Biotechnology Co., Ltd. All Rights Reserved Web design

Site Map | Legal Notice | Privacy Policy |