Abstract:The versatility of self-attention mechanism earned transformers great success in almost all data modalities, with limitations on the quadratic complexity and difficulty of training. Efficient transformers, on the other hand, often rely on clever data-modality-dependent construction to get over the quadratic complexity of transformers. This greatly hinders their applications on different data modalities, which is one of the pillars of contemporary foundational modeling. In this paper, we lay the groundwork for efficient foundational modeling by proposing SAMSA - SAMpling-Self-Attention, a context-aware linear complexity self-attention mechanism that works well on multiple data modalities. Our mechanism is based on a differentiable sampling without replacement method we discovered. This enables the self-attention module to attend to the most important token set, where the importance is defined by data. Moreover, as differentiability is not needed in inference, the sparse formulation of our method costs little time overhead, further lowering computational costs. In short, SAMSA achieved competitive or even SOTA results on many benchmarks, while being faster in inference, compared to other very specialized models. Against full self-attention, real inference time significantly decreases while performance ranges from negligible degradation to outperformance. We release our source code in the repository: <a class="link-external link-https" href="https://github.com/HySonLab/SAMSA" rel="external noopener nofollow">this https URL</a>

Understanding Self-Attention of Self-Supervised Audio Transformers

Hand-crafted Attention is All You Need? A Study of Attention on Self-supervised Audio Transformer.

SAC: Accelerating and Structuring Self-Attention Via Sparse Adaptive Connection.

Hand-crafted Attention is All You Need? A Study of Attention on Self-supervised Audio Transformer

Input-independent Attention Weights Are Expressive Enough: A Study of Attention in Self-supervised Audio Transformers

Probing self-attention in self-supervised speech models for cross-linguistic differences

Unveiling the Hidden Structure of Self-Attention via Kernel Principal Component Analysis

SparseBERT: Rethinking the Importance Analysis in Self-attention

Adaptive Sparse and Monotonic Attention for Transformer-based Automatic Speech Recognition

Structured Self-Attention Weights Encode Semantics in Sentiment Analysis

Attention Flows: Analyzing and Comparing Attention Mechanisms in Language Models

How Does Attention Work in Vision Transformers? A Visual Analytics Attempt

Theory, Analysis, and Best Practices for Sigmoid Self-Attention

When to Use Efficient Self Attention? Profiling Text, Speech and Image Transformer Variants

An Empirical Study of Spatial Attention Mechanisms in Deep Networks

DARTFormer: Finding The Best Type Of Attention

Audio Mamba: Selective State Spaces for Self-Supervised Audio Representations

SAMSA: Efficient Transformer for Many Data Modalities

Exploring Self-Attention Mechanisms for Speech Separation

AutoAttend: Automated Attention Representation Search

Rethinking Self-Attention: Towards Interpretability in Neural Parsing