Abstract:Deep learning models struggle with compositional generalization, i.e. the ability to recognize or generate novel combinations of observed elementary concepts. In hopes of enabling compositional generalization, various unsupervised learning algorithms have been proposed with inductive biases that aim to induce compositional structure in learned representations (e.g. disentangled representation and emergent language learning). In this work, we evaluate these unsupervised learning algorithms in terms of how well they enable compositional generalization. Specifically, our evaluation protocol focuses on whether or not it is easy to train a simple model on top of the learned representation that generalizes to new combinations of compositional factors. We systematically study three unsupervised representation learning algorithms - $\beta$-VAE, $\beta$-TCVAE, and emergent language (EL) autoencoders - on two datasets that allow directly testing compositional generalization. We find that directly using the bottleneck representation with simple models and few labels may lead to worse generalization than using representations from layers before or after the learned representation itself. In addition, we find that the previously proposed metrics for evaluating the levels of compositionality are not correlated with actual compositional generalization in our framework. Surprisingly, we find that increasing pressure to produce a disentangled representation produces representations with worse generalization, while representations from EL models show strong compositional generalization. Taken together, our results shed new light on the compositional generalization behavior of different unsupervised learning algorithms with a new setting to rigorously test this behavior, and suggest the potential benefits of delevoping EL learning algorithms for more generalizable representations.

Generalization in Multimodal Language Learning from Simulation

In-Context Compositional Generalization for Large Vision-Language Models

Improving Compositional Generalization Using Iterated Learning and Simplicial Embeddings

On the generalization capacity of neural networks during generic multimodal reasoning

A Study of Compositional Generalization in Neural Models

Learning to generalize to new compositions in image understanding

Sequential Compositional Generalization in Multimodal Models

Multimodal Generative Models for Compositional Representation Learning

Dynamics of Concept Learning and Compositional Generalization

Semi-supervised Multimodal Representation Learning through a Global Workspace

Compositional generalization through abstract representations in human and artificial neural networks

A Theory of Multimodal Learning

Compositional Generalization in Unsupervised Compositional Representation Learning: A Study on Disentanglement and Emergent Language

Compositional Generalization by Learning Analytical Expressions.

Towards Understanding the Relationship between In-context Learning and Compositional Generalization

Model Composition for Multimodal Large Language Models

Development of Compositionality and Generalization through Interactive Learning of Language and Action of Robots

SimMMDG: A Simple and Effective Framework for Multi-modal Domain Generalization

Learning Unseen Modality Interaction

Compositional diversity in visual concept learning

Compositional learning of functions in humans and machines