Reliability and Effectiveness of Cross-validation in Feature Selection

ghadah aldehim,wenjia wang
DOI: https://doi.org/10.1007/978-3-319-12069-0_13
2014-01-01
Abstract:Feature selection is increasingly important in data analysis and machine learning in the big data era. However, how to use the data in feature selection has become a serious issue as the conventional practice of using ALL the data in feature selection may lead to selection bias and some suggest to use PART of the data instead. This paper investigates the reliability and effectiveness of a PART approach implemented by cross validation mechanism in feature selection filters and compares it with the ALL approach. The reliability is measured by an Inter-system Average Tanimoto Index and the effectiveness of the selected features is measured by the mean generalisation accuracy of classification. The experiments are carried out by using synthetic datasets generated with a fixed number of relevant features and varied numbers of irrelevant features and instances, and different level of noise, to mimic some possible real world environments. The results indicate that the PART approach is more effective in reducing the bias when the dataset is small but starts to lose its advantage as the dataset size increases.
What problem does this paper attempt to address?