A Study on the Effiect of Class Distribution Using Cost-Sensitive Learning

Kai Ming Ting
DOI: https://doi.org/10.1007/3-540-36182-0_11
2002-01-01
Abstract:This paper investigates the effect of class distribution on the predictive performance of classification models using cost-sensitive learning, rather than the sampling approach employed previously by a similar study. The predictive performance is measured using the cost space representation, which is a dual to the ROC representation. This study shows that distributions which range between the natural distribution and the balanced distribution can also produce the best models, contrary to the finding of the previous study. In addition, we find that the best models are larger in size than those trained using the natural distribution. We also show two different ways to achieve the same effect of the corrected probability estimates proposed by the previous study.
What problem does this paper attempt to address?