Abstract:This paper considers the challenge of evaluating a set of classifiers, as done in shared task evaluations like the KDD Cup or NIST TREC, without expert labels. While expert labels provide the traditional cornerstone for evaluating statistical learners, limited or expensive access to experts represents a practical bottleneck. Instead, we seek methodology for estimating performance of the classifiers which is more scalable than expert labeling yet preserves high correlation with evaluation based on expert labels. We consider both: 1) using only labels automatically generated by the classifiers (blind evaluation); and 2) using labels obtained via crowdsourcing. While crowdsourcing methods are lauded for scalability, using such data for evaluation raises serious concerns given the prevalence of label noise. In regard to blind evaluation, two broad strategies are investigated: combine & score and score & combine methods infer a single pseudo-gold label set by aggregating classifier labels; classifiers are then evaluated based on this single pseudo-gold label set. On the other hand, score & combine methods: 1) sample multiple label sets from classifier outputs, 2) evaluate classifiers on each label set, and 3) average classifier performance across label sets. When additional crowd labels are also collected, we investigate two alternative avenues for exploiting them: 1) direct evaluation of classifiers; or 2) supervision of combine & score methods. To assess generality of our techniques, classifier performance is measured using four common classification metrics, with statistical significance tests. Finally, we measure both score and rank correlations between estimated classifier performance vs. actual performance according to expert judgments. Rigorous evaluation of classifiers from the TREC 2011 Crowdsourcing Track shows reliable evaluation can be achieved without reliance on expert labels.

Semisupervised Classifier Evaluation and Recalibration

Semi-supervised Contrastive Outlier removal for Pseudo Expectation Maximization (SCOPE)

Evaluating Classifiers Without Expert Labels

An interpretable semi-supervised classifier using two different strategies for amended self-labeling

Semi-Supervised Approaches to Efficient Evaluation of Model Prediction Performance

When less is more: on the value of "co-training" for semi-supervised software defect predictors

Semi-supervised Learning For Robust Speech Evaluation

Progressive Semisupervised Learning of Multiple Classifiers.

Learning with Inadequate and Incorrect Supervision

Minimum-Risk Recalibration of Classifiers

Penalizing Proposals using Classifiers for Semi-Supervised Object Detection

Weak Supervision Performance Evaluation via Partial Identification

Semi-supervised Transfer Learning for Evaluation of Model Classification Performance

Efficient Lifelong Model Evaluation in an Era of Rapid Progress

On the Importance of Calibration in Semi-supervised Learning

Classifier Calibration: A survey on how to assess and improve predicted class probabilities

SemiReward: A General Reward Model for Semi-supervised Learning

Enabling Precision/Recall Preferences for Semi-supervised SVM Training.

Learning with Confidence: Training Better Classifiers from Soft Labels

Are Labels Always Necessary for Classifier Accuracy Evaluation?

Efficient Evaluation of Prediction Rules in Semi-Supervised Settings under Stratified Sampling