Abstract:In this paper, we contend that a natural objective of representation learning is to compress and transform the distribution of the data, say sets of tokens, towards a low-dimensional Gaussian mixture supported on incoherent subspaces. The goodness of such a representation can be evaluated by a principled measure, called sparse rate reduction, that simultaneously maximizes the intrinsic information gain and extrinsic sparsity of the learned representation. From this perspective, popular deep network architectures, including transformers, can be viewed as realizing iterative schemes to optimize this measure. Particularly, we derive a transformer block from alternating optimization on parts of this objective: the multi-head self-attention operator compresses the representation by implementing an approximate gradient descent step on the coding rate of the features, and the subsequent multi-layer perceptron sparsifies the features. This leads to a family of white-box transformer-like deep network architectures, named CRATE, which are mathematically fully interpretable. We show, by way of a novel connection between denoising and compression, that the inverse to the aforementioned compressive encoding can be realized by the same class of CRATE architectures. Thus, the so-derived white-box architectures are universal to both encoders and decoders. Experiments show that these networks, despite their simplicity, indeed learn to compress and sparsify representations of large-scale real-world image and text datasets, and achieve performance very close to highly engineered transformer-based models: ViT, MAE, DINO, BERT, and GPT2. We believe the proposed computational framework demonstrates great potential in bridging the gap between theory and practice of deep learning, from a unified perspective of data compression. Code is available at: <a class="link-external link-https" href="https://ma-lab-berkeley.github.io/CRATE" rel="external noopener nofollow">this https URL</a> .

Temporal Latent Bottleneck: Synthesis of Fast and Slow Processing Mechanisms in Sequence Learning

Spectral Transform Forms Scalable Transformer

Funnel-Transformer: Filtering out Sequential Redundancy for Efficient Language Processing

TransformerG2G: Adaptive time-stepping for learning temporal graph embeddings using transformers

TSLANet: Rethinking Transformers for Time Series Representation Learning

Generating Long Sequences with Sparse Transformers

Discrete Key-Value Bottleneck

Chunk, Align, Select: A Simple Long-sequence Processing Method for Transformers

Sparse Binary Transformers for Multivariate Time Series Modeling

Latte: Latent Attention for Linear Time Transformers

Fast Decoding in Sequence Models using Discrete Latent Variables

Sequential Recommendation with Bidirectional Chronological Augmentation of Transformer

Sequence Compression Speeds Up Credit Assignment in Reinforcement Learning

Contrastive Bidirectional Transformer for Temporal Representation Learning

Scalable and Efficient Temporal Graph Representation Learning via Forward Recent Sampling

White-Box Transformers via Sparse Rate Reduction: Compression Is All There Is?

Online Transformers with Spiking Neurons for Fast Prosthetic Hand Control

Multi-Layer Transformers Gradient Can be Approximated in Almost Linear Time

Can Looped Transformers Learn to Implement Multi-step Gradient Descent for In-context Learning?

Transformers are Multi-State RNNs

Long-term Leap Attention, Short-term Periodic Shift for Video Classification