A Data Mining Approach to Predict Protein Secondary Structure

Bingru Yang,Haifeng Sui,Qu Wu,Lijun Wang
DOI: https://doi.org/10.1109/iccasm.2010.5620659
2010-01-01
Abstract:In bioinformatics,proteins are coded by strings, called “primary structures”. Biologists have long enough gathered these primary structures in large databases. Numerous experiments and analyses of primary structures have revealed that the protein primary structure closely correlates with the protein second structure. In this paper, we present a data mining approach based on machine learning techniques to predict protein second structure. Based on majority voting mechanism, the approach combine the predictions of homology analysis classifier, Support vector machine(SVM) classifier and modified Knowledge Discovery in Databases (KDD*) process. They are validated with 2 different datasets. Their predictive accuracy results outperform the best secondary structure predictors by 2.00% on average.
What problem does this paper attempt to address?