Mistakes in validating the accuracy of a prediction classifier in high-dimensional but small-sample microarray data

Sunho Lee
DOI: https://doi.org/10.1177/0962280207084839
IF: 2.494
2008-03-28
Statistical Methods in Medical Research
Abstract:A major interest in gene expression microarray studies is to develop an accurate classifier which can be adopted in clinical practice. The usage of large numbers of genes with small data samples may lead to overfitting in classification, and generate promising, but often nonreproducible results. Therefore, assessing the reproducibility of a classifier is necessary. Appropriate methods for validating a developed classifier and estimating its predicting accuracy are discussed. In addition, some mistakes that can arise in the cross validation process are reviewed using published articles in prominent medical journals, to prevent the indefinite results of a classifier development from leading to inappropriate treatment.
health care sciences & services,medical informatics,mathematical & computational biology,statistics & probability
What problem does this paper attempt to address?