Key points are not available for this paper at this time.
Reductive amination with arylamines is a valuable but challenging transformation in pharmaceutical biocatalysis. To address this gap area in the biocatalytic toolkit, supervised machine learning was used to identify untested enzymes capable of catalyzing these reactions. A dataset comprising 269 enzymes tested on 9 different reductive amination reactions was constructed and used to train a K‐Nearest‐Neighbors regression model. The model was applied to generate predictions for 6588 untested wild‐type enzymes, from which 42 top candidates were identified. It was observed that these top 42 predicted enzymes were substantially enriched in active variants compared to the training dataset, and improved variants for 7 of the 9 reactions were identified. It was also found that these top‐predicted enzymes were significantly enriched in sequences with lower ESM1b pseudoperplexity (a measure of evolutionary plausibility) despite the model not being trained on such information. The results demonstrate that supervised machine learning is an effective tool for identifying highly active wild‐type biocatalysts for difficult transformations and suggest that protein language models may also be useful in identifying such enzymes.
Brennan et al. (Tue,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: