Abstract:Abstract Background It is well known that most of the binding free energy of protein interaction is contributed by a few key hot spot residues. These residues are crucial for understanding the function of proteins and studying their interactions. Experimental hot spots detection methods such as alanine scanning mutagenesis are not applicable on a large scale since they are time consuming and expensive. Therefore, reliable and efficient computational methods for identifying hot spots are greatly desired and urgently required. Results In this work, we introduce an efficient approach that uses support vector machine (SVM) to predict hot spot residues in protein interfaces. We systematically investigate a wide variety of 62 features from a combination of protein sequence and structure information. Then, to remove redundant and irrelevant features and improve the prediction performance, feature selection is employed using the F-score method. Based on the selected features, nine individual-feature based predictors are developed to identify hot spots using SVMs. Furthermore, a new ensemble classifier, namely APIS (A combined model based on Protrusion Index and Solvent accessibility), is developed to further improve the prediction accuracy. The results on two benchmark datasets, ASEdb and BID, show that this proposed method yields significantly better prediction accuracy than those previously published in the literature. In addition, we also demonstrate the predictive power of our proposed method by modelling two protein complexes: the calmodulin/myosin light chain kinase complex and the heat shock locus gene products U and V complex, which indicate that our method can identify more hot spots in these two complexes compared with other state-of-the-art methods. Conclusion We have developed an accurate prediction model for hot spot residues, given the structure of a protein complex. A major contribution of this study is to propose several new features based on the protrusion index of amino acid residues, which has been shown to significantly improve the prediction performance of hot spots. Moreover, we identify a compact and useful feature subset that has an important implication for identifying hot spot residues. Our results indicate that these features are more effective than the conventional evolutionary conservation, pairwise residue potentials and other traditional features considered previously, and that the combination of our and traditional features may support the creation of a discriminative feature set for efficient prediction of hot spot residues. The data and source code are available on web site http://home.ustc.edu.cn/~jfxia/hotspot.html .

Abstract 473: Identification of Apolipoproteins Using Feature Selection Technique

Identification of Apolipoprotein Using Feature Selection Technique

ApoPred: Identification of Apolipoproteins and Their Subfamilies With Multifarious Features

Prediction of Functional Class of Proteins and Peptides Irrespective of Sequence Homology by Support Vector Machines.

Learning the Drug Target-Likeness of A Protein

Identifying Antioxidant Proteins by Using Optimal Dipeptide Compositions

Support Vector Machines for Predicting Protein Homo- Oligomers by Incorporating Pseudo-Amino Acid Composition #

Analyzing HAb18G/CD147 antagonistic peptides using bioinformatics

AOPM: Application of Antioxidant Protein Classification Model in Predicting the Composition of Antioxidant Drugs

APIS: accurate prediction of hot spots in protein interfaces by combining protrusion index with solvent accessibility

Identification of properties important to protein aggregation using feature selection

Abstract 107: Protein Comigration Pattern Analysis to Identify High-Density Lipoprotein Subspecies

DP-AOP: A novel SVM-based antioxidant proteins identifier

PREDAIP: Computational Prediction and Analysis for Anti-inflammatory Peptide Via a Hybrid Feature Selection Technique

Effective identification and differential analysis of anticancer peptides

Identification of Immunoglobulins Using Chou'S Pseudo Amino Acid Composition with Feature Selection Technique

Prediction of Anti-apoptosis Proteins and Pro-apoptosis Proteins Based on Chemical Shift and Pseudo-amino Acid Composition Information

HBPred: a Tool to Identify Growth Hormone-Binding Proteins

Prediction of cell-penetrating peptides with feature selection techniques

Functional Classification of Secreted Proteins by Position Specific Scoring Matrix and Auto Covariance

Identification of Antioxidant Proteins with Deep Learning from Sequence Information