Prediction of cell penetrating peptides and their uptake efficiency using random forest‐based feature selections

Peng Liu,Yijie Ding,Ying Rong,Dong Chen
DOI: https://doi.org/10.1002/aic.17781
IF: 4.167
2022-05-27
AIChE Journal
Abstract:Cell penetrating peptides (CPPs) are short peptides that can carry biomolecules of varying sizes across the cell membrane into the cytoplasm. Correctly identifying CPPs is the basis for studying their functions and mechanisms. Here, we propose a novel CPP predictor that is able to predict CPPs and their uptake efficiency. In our method, five feature descriptors are applied to encode the sequence and compose a hybrid feature vector. Afterward, the wrapper + random forest algorithm is employed, which combines feature selection with the prediction process to find features that are crucial for identifying CPPs. The jackknife cross validation result shows that our predictor is comparable to state‐of‐the‐art CPP predictors, and our method reduces the feature dimension, which improves computational efficiency and avoids overfitting, allowing our predictor to be adopted to identify large‐scale CPP data. This article is protected by copyright. All rights reserved.
engineering, chemical
What problem does this paper attempt to address?