Abstract:Traditional machine learning paradigms are based on the assumption that both training and test data follow the same statistical pattern, which is mathematically referred to as Independent and Identically Distributed ($i.i.d.$). However, in real-world applications, this $i.i.d.$ assumption often fails to hold due to unforeseen distributional shifts, leading to considerable degradation in model performance upon deployment. This observed discrepancy indicates the significance of investigating the Out-of-Distribution (OOD) generalization problem. OOD generalization is an emerging topic of machine learning research that focuses on complex scenarios wherein the distributions of the test data differ from those of the training data. This paper represents the first comprehensive, systematic review of OOD generalization, encompassing a spectrum of aspects from problem definition, methodological development, and evaluation procedures, to the implications and future directions of the field. Our discussion begins with a precise, formal characterization of the OOD generalization problem. Following that, we categorize existing methodologies into three segments: unsupervised representation learning, supervised model learning, and optimization, according to their positions within the overarching learning process. We provide an in-depth discussion on representative methodologies for each category, further elucidating the theoretical links between them. Subsequently, we outline the prevailing benchmark datasets employed in OOD generalization studies. To conclude, we overview the existing body of work in this domain and suggest potential avenues for future research on OOD generalization. A summary of the OOD generalization methodologies surveyed in this paper can be accessed at <a class="link-external link-http" href="http://out-of-distribution-generalization.com" rel="external noopener nofollow">this http URL</a>.

Our Evaluation Metric Needs an Update to Encourage Generalization

A Survey on Evaluation of Out-of-Distribution Generalization

Model-Agnostic Random Weighting for Out-of-Distribution Generalization

Generalizing to any diverse distribution: uniformity, gentle finetuning and rebalancing

The Value of Out-of-Distribution Data

Towards Out-Of-Distribution Generalization: A Survey

Is your benchmark truly adversarial? AdvScore: Evaluating Human-Grounded Adversarialness

Out-of-Distribution Generalization Analysis via Influence Function

Quantifying Variance in Evaluation Benchmarks

Efficient Lifelong Model Evaluation in an Era of Rapid Progress

Towards a Better Evaluation of Out-of-Domain Generalization

Unsupervised Evaluation of Out-of-distribution Detection: A Data-centric Perspective

Developing a Dataset-Adaptive, Normalized Metric for Machine Learning Model Assessment: Integrating Size, Complexity, and Class Imbalance

Towards a Theoretical Framework of Out-of-Distribution Generalization

Beyond correlation: The impact of human uncertainty in measuring the effectiveness of automatic evaluation and LLM-as-a-judge

Rethinking the Evaluation Protocol of Domain Generalization.

Mitigating Graph Covariate Shift via Score-based Out-of-distribution Augmentation

Are Some Words Worth More than Others?

Attribute Based Interpretable Evaluation Metrics for Generative Models

DIVE: Subgraph Disagreement for Graph Out-of-Distribution Generalization

Rethinking Out-of-Distribution Detection From a Human-Centric Perspective