Abstract:The effectiveness of automatic evaluation of generative models is typically measured by comparing it to human evaluation using correlation metrics. However, metrics like Krippendorff's $\alpha$ and Randolph's $\kappa$, originally designed to measure the reliability of human labeling, make assumptions about human behavior and the labeling process. In this paper, we show how *relying on a single aggregate correlation score* can obscure fundamental differences between human behavior and automatic evaluation methods, including LLM-as-a-Judge. Specifically, we demonstrate that when the proportion of samples with variation or uncertainty in human labels (gathered during human evaluation) is relatively high, machine labels (generated by automatic evaluation methods) may superficially appear to have similar or better correlation with the human majority label compared to human-to-human (HH) correlation. This can create the misleading impression that automatic evaluation is accurate enough to approximate the human majority label. However, as the proportion of samples with consistent human labels increases, the correlation between machine labels and human majority labels declines, falling below HH correlation. Based on these findings, we first propose stratifying results by human label uncertainty to provide a more robust analysis of automatic evaluation performance. Second, recognizing that uncertainty and variation are inherent in perception-based human evaluations, such as those involving attitudes or preferences, we introduce a new metric - *binned Jensen-Shannon Divergence for perception* for such scenarios to better measure the effectiveness of automatic evaluations. Third, we present visualization techniques -- *perception charts*, to compare the strengths and limitations of automatic evaluation and to contextualize correlation measures appropriately

TrueLabel + Confusions: A Spectrum of Probabilistic Models in Analyzing Multiple Ratings

Towards more accurate multi-label software behavior learning

Who Said What: Modeling Individual Labelers Improves Classification

Label-free Evaluation for Performance of Fault Diagnosis Model on Unknown Distribution Dataset

Accurate Label Refinement from Multiannotator of Remote Sensing Data

Are LLMs Better than Reported? Detecting Label Errors and Mitigating Their Effect on Model Performance

Crowdsourcing with Difficulty: A Bayesian Rating Model for Heterogeneous Items

Beyond confusion matrix: learning from multiple annotators with awareness of instance features

Beyond correlation: The impact of human uncertainty in measuring the effectiveness of automatic evaluation and LLM-as-a-judge

How many labelers do you have? A closer look at gold-standard labels

Two Wrongs Don't Make a Right: Combating Confirmation Bias in Learning with Label Noise.

Uncovering the Latent Structures of Crowd Labeling.

"Seeing the Big through the Small": Can LLMs Approximate Human Judgment Distributions on NLI from a Few Explanations?

Capturing Correlations of Multiple Labels: A Generative Probabilistic Model for Multi-Label Learning

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd

CalibraEval: Calibrating Prediction Distribution to Mitigate Selection Bias in LLMs-as-Judges

An Online Learning Approach to Improving the Quality of Crowd-Sourcing

Label-Driven Learning Framework: Towards More Accurate Bayesian Network Classifiers Through Discrimination of High-Confidence Labels

Predicting Label Distribution From Tie-Allowed Multi-Label Ranking

Bayesian modeling of multi-species labeling errors in ecological studies

Don't Blame the Data, Blame the Model: Understanding Noise and Bias When Learning from Subjective Annotations