Abstract:Traditional machine learning paradigms are based on the assumption that both training and test data follow the same statistical pattern, which is mathematically referred to as Independent and Identically Distributed ($i.i.d.$). However, in real-world applications, this $i.i.d.$ assumption often fails to hold due to unforeseen distributional shifts, leading to considerable degradation in model performance upon deployment. This observed discrepancy indicates the significance of investigating the Out-of-Distribution (OOD) generalization problem. OOD generalization is an emerging topic of machine learning research that focuses on complex scenarios wherein the distributions of the test data differ from those of the training data. This paper represents the first comprehensive, systematic review of OOD generalization, encompassing a spectrum of aspects from problem definition, methodological development, and evaluation procedures, to the implications and future directions of the field. Our discussion begins with a precise, formal characterization of the OOD generalization problem. Following that, we categorize existing methodologies into three segments: unsupervised representation learning, supervised model learning, and optimization, according to their positions within the overarching learning process. We provide an in-depth discussion on representative methodologies for each category, further elucidating the theoretical links between them. Subsequently, we outline the prevailing benchmark datasets employed in OOD generalization studies. To conclude, we overview the existing body of work in this domain and suggest potential avenues for future research on OOD generalization. A summary of the OOD generalization methodologies surveyed in this paper can be accessed at <a class="link-external link-http" href="http://out-of-distribution-generalization.com" rel="external noopener nofollow">this http URL</a>.

The Importance of Generalizability in Machine Learning for Systems

A Survey on Machine Learning for Geo-Distributed Cloud Data Center Management

Probing out-of-distribution generalization in machine learning for materials

Towards Out-Of-Distribution Generalization: A Survey

Fairness and Accuracy Under Domain Generalization

Out-of-Distribution Generalization in Natural Language Processing: Past, Present, and Future

Machine Learning vs Deep Learning: The Generalization Problem

Model-Based Domain Generalization

Out-of-Distribution Generalization in Text Classification: Past, Present, and Future

Robustness, Evaluation and Adaptation of Machine Learning Models in the Wild

A Holistic Assessment of the Reliability of Machine Learning Systems

Towards Reliable Learning in the Wild: Generalization and Adaptation

Modeling Generalization in Machine Learning: A Methodological and Computational Study

Verifying the Generalization of Deep Learning to Out-of-Distribution Domains

A Survey on Evaluation of Out-of-Distribution Generalization

Certifiable Out-of-Distribution Generalization.

Generalization in medical AI: a perspective on developing scalable models

Cloudy with high chance of DBMS: A 10-year prediction for Enterprise-Grade ML

Generalizing in the Real World with Representation Learning

Accuracy on the Line: On the Strong Correlation Between Out-of-Distribution and In-Distribution Generalization

Generalization Analysis for Game-Theoretic Machine Learning