Home · News · Articles ·
AI-Powered Biocatalysis: How CATNIP Enables Global Enzyme–Substrate Prediction
Source: Hzymes Market Center
Date: 2026-05-28
Views: 335


Introduction: Transforming Biocatalysis with AI


Biocatalysis has become an increasingly important technology in pharmaceutical synthesis, natural product modification, and green chemistry. Compared with traditional chemical catalysis, enzyme-catalyzed reactions offer significant advantages, including high selectivity, mild reaction conditions, fewer synthetic steps, and improved sustainability.


Despite these benefits, one major challenge continues to limit broader industrial application: identifying the right enzyme for a specific substrate.


In most cases, there are no reliable prior indicators to determine which enzyme can catalyze a target molecule. Researchers often rely on time-consuming trial-and-error screening, which increases development costs and reduces discovery efficiency.


To overcome this bottleneck, researchers from the University of Michigan and Carnegie Mellon University developed CATNIP, an AI-driven biocatalysis platform that combines high-throughput experimentation with machine learning to predict enzyme–substrate compatibility at scale.




By connecting chemical space with enzyme sequence space, CATNIP enables both:


  • Substrate-to-enzyme prediction — identifying enzymes capable of catalyzing a target compound
  • Enzyme-to-substrate prediction — predicting potential substrates for a given enzyme


This breakthrough represents a major step toward data-driven enzyme discovery and intelligent biocatalyst design.



Figure 1. Current landscape of biocatalytic reaction discovery



Mapping Chemical Space and Enzyme Sequence Space


One of the biggest limitations in biocatalysis research is the lack of known connections between enzymes and substrates. Today, less than 0.3% of possible enzyme–substrate relationships have been experimentally characterized.


As a result, researchers are often restricted to exploring only small, local regions of sequence space and chemical space, making it difficult to discover entirely new catalytic functions.


To address this issue, the CATNIP study systematically mapped the catalytic landscape of α-ketoglutarate-dependent non-heme iron enzymes (α-KG NHIs).


The researchers began with more than 265,000 protein sequences and used bioinformatics filtering to identify 27,005 representative candidates. Through sequence similarity network (SSN) analysis, they narrowed the dataset to 314 wild-type enzymes and constructed the highly diverse aKGLib1 enzyme library.


Notably, the average sequence similarity among these enzymes was only 13.7%, highlighting the exceptional diversity of the library.




Figure 2. Construction of the diverse α-KG NHI enzyme library aKGLib1


To maximize chemical diversity, the researchers also selected 111 structurally diverse substrates, including:


  • Amino acids
  • Natural products
  • Pharmaceuticals
  • Heterocyclic compounds


Using a high-throughput workflow based on 96-well microreaction systems and LC-MS analysis, the team screened all enzyme–substrate combinations for multiple reaction types, including:


  • Hydroxylation
  • Desaturation
  • Rearrangement reactions


The results demonstrated the enormous untapped potential of biocatalysis:


  • 32% of substrates were successfully transformed
  • 38% of enzymes showed detectable catalytic activity
  • 215 previously undiscovered biocatalytic reactions were identified



Figure 3. High-throughput discovery of diverse biocatalytic reactions



This large-scale experimental matrix effectively created a “biocatalytic navigation map,” filling critical gaps between chemical space and enzyme sequence space while generating high-quality training data for machine learning models.



AI-Based Bidirectional Prediction of Enzyme–Substrate Pairs


Using the experimentally validated dataset, the researchers developed a machine learning framework capable of bidirectional enzyme–substrate prediction.


To build the model, they established the BioCatSet1 dataset, containing 354 experimentally validated enzyme–substrate pairs collected from both experimental screening and published literature.


The model converts:

  • Substrates into chemical representations using SMILES encoding and molecular descriptors
  • Enzymes into sequence-space representations using sequence alignment information


Based on these encoded features, the system can accurately predict:


Substrate → Enzyme

Given a target compound, the model identifies enzymes most likely to catalyze it.


Enzyme → Substrate

Given an enzyme sequence, the model predicts compatible substrates and potential catalytic activity.




Figure 4. Construction of the CATNIP prediction model


To improve prediction performance, the team optimized the system using Gradient Boosting Machine (GBM) algorithms.


Compared with random screening approaches, the top-ranked predictions achieved more than a sevenfold improvement in accuracy, significantly increasing the efficiency of enzyme discovery and reaction development.




Figure 5. Optimization of AI prediction performance


The researchers also launched the CATNIP online platform, allowing users to directly input either:

  • A chemical structure
  • An enzyme sequence


The platform then returns predicted catalytic matches, enabling researchers to evaluate reaction feasibility before conducting laboratory experiments.


This AI-assisted workflow transforms enzyme discovery from blind experimental screening into intelligent navigation through high-dimensional biological and chemical space.



Experimental Validation: From AI Prediction to Real Catalytic Products


To validate the practical performance of CATNIP, the researchers experimentally tested multiple predicted enzyme–substrate combinations.


For the plant alkaloid (-)-Sparteine, seven of the top ten AI-predicted enzymes successfully catalyzed hydroxylation reactions, achieving product yields up to 35%.


Similarly, when testing the steroid substrate 6-Methyleneandrost-4-ene-3,17-dione, several predicted enzymes demonstrated catalytic activity, including enzymes with no previously reported activity toward this substrate class.


These findings demonstrate the platform’s ability to identify hidden functional relationships beyond existing biochemical knowledge.


The researchers further evaluated CATNIP using entirely unknown enzymes. In one example, the uncharacterized enzyme TqaL from Streptomyces violaceusniger successfully catalyzed several AI-predicted substrates, confirming the model’s strong generalization capability.




Figure 6. Experimental validation of CATNIP predictions


These results establish a complete workflow from:


AI prediction → Experimental validation → Product generation

This closed-loop strategy highlights the real-world value of AI-guided biocatalysis for enzyme engineering, synthetic biology, and pharmaceutical development.



The Future of AI-Driven Enzyme Discovery


CATNIP demonstrates how artificial intelligence and high-throughput experimentation can fundamentally reshape biocatalysis research.


By connecting chemical space with enzyme sequence space, the platform enables researchers to move beyond traditional trial-and-error screening toward predictive, data-driven enzyme discovery.


Although current studies focus on a specific enzyme family, the same strategy could eventually expand across broader protein families and more complex reaction systems.


In the future, scientists may be able to design enzymatic synthesis pathways as easily as using navigation software — rapidly identifying the most efficient, selective, and sustainable catalytic routes for molecular production.


AI-powered biocatalysis is no longer a concept for the future. Platforms like CATNIP are already laying the foundation for the next generation of intelligent enzyme engineering.



Open-Source Resources and Tools


CATNIP Web Platform

Direct online access without registration:

https://catnip.cheme.cmu.edu/


Open-Source Code Repository

Complete model training and data processing workflows:

https://doi.org/10.5281/zenodo.16779318


BioCatSet1 Dataset

Includes substrate SMILES, enzyme sequences, and LC-MS reaction data:
https://github.com/gomesgroup/catnip

Message
Leave Your Message
Name *
Company *
Tel/WhatsApp *
Mail *
Nation *
Descriptions

Please contact on WhatsApp

Service Hotline: +86 400-808-5320

Large-scale production base: Building 6, Precision Medical Industry Base, Wuhan, China

Logistics & Supply Chain Center:417 Main St, Little Rock, AR 72201. United States.

Global Marketing Center: Hzymes Building, Fengxian District, Shanghai, China.

  • iso_copy_copy
  • iso_copy
  • iso
  • iso_copy_copy
  • iso_copy
  • iso
Contact Us

Service Hotline: +86 400-808-5320

Large-scale production base: Building 6, Precision Medical Industry Base, Wuhan, China.

Logistics & Supply Chain Center:417 Main St, Little Rock, AR 72201. United States.

Global Marketing Center: Hzymes Building, Fengxian District, Shanghai, China.

Copyright © Hzymes Biotechnology Co., Ltd. All Rights Reserved Web design

Site Map | Legal Notice | Privacy Policy |