Optimal Feature Selection for Sparse Linear Discriminant Analysis and Its Applications in Gene Expression Data

Cheng Wang,Longbing Cao,Baiqi Miao
DOI: https://doi.org/10.1016/j.csda.2013.04.003
IF: 2.035
2013-01-01
Computational Statistics & Data Analysis
Abstract:This work studies the theoretical rules of feature selection in linear discriminant analysis (LDA), and a new feature selection method is proposed for sparse linear discriminant analysis. An l1 minimization method is used to select the important features from which the LDA will be constructed. The asymptotic results of this proposed two-stage LDA (TLDA) are studied, demonstrating that TLDA is an optimal classification rule whose convergence rate is the best compared to existing methods. The experiments on simulated and real datasets are consistent with the theoretical results and show that TLDA performs favorably in comparison with current methods. Overall, TLDA uses a lower minimum number of features or genes than other approaches to achieve a better result with a reduced misclassification rate.
What problem does this paper attempt to address?