Abstract:The synergy of long-range dependencies from transformers and local representations of image content from convolutional neural networks (CNNs) has led to advanced architectures and increased performance for various medical image analysis tasks due to their complementary benefits. However, compared with CNNs, transformers require considerably more training data, due to a larger number of parameters and an absence of inductive bias. The need for increasingly large datasets continues to be problematic, particularly in the context of medical imaging, where both annotation efforts and data protection result in limited data availability. In this work, inspired by the human decision-making process of correlating new "evidence" with previously memorized "experience", we propose a Memorizing Vision Transformer (MoViT) to alleviate the need for large-scale datasets to successfully train and deploy transformer-based architectures. MoViT leverages an external memory structure to cache history attention snapshots during the training stage. To prevent overfitting, we incorporate an innovative memory update scheme, attention temporal moving average, to update the stored external memories with the historical moving average. For inference speedup, we design a prototypical attention learning method to distill the external memory into smaller representative subsets. We evaluate our method on a public histology image dataset and an in-house MRI dataset, demonstrating that MoViT applied to varied medical image analysis tasks, can outperform vanilla transformer models across varied data regimes, especially in cases where only a small amount of annotated data is available. More importantly, MoViT can reach a competitive performance of ViT with only 3.0% of the training data. In conclusion, MoViT provides a simple plug-in for transformer architectures which may contribute to reducing the training data needed to achieve acceptable models for a broad range of medical image analysis tasks.

MEViT: Motion Enhanced Video Transformer for Video Classification

TransVOS: Video Object Segmentation with Transformers

LS-VIT: Vision Transformer for action recognition based on long and short-term temporal difference

Disentangling 3D/4D Facial Affect Recognition with Faster Multi-View Transformer

MFEViT: A Robust Lightweight Transformer-based Network for Multimodal 2D+3D Facial Expression Recognition

FMViT: A multiple-frequency mixing Vision Transformer

Separable Self-attention for Mobile Vision Transformers

Efficient Video Transformers via Spatial-Temporal Token Merging for Action Recognition

MoViT: Memorizing Vision Transformers for Medical Image Analysis

DctViT: Discrete Cosine Transform Meet Vision Transformers

EfficientViT: Memory Efficient Vision Transformer with Cascaded Group Attention

Wave-ViT: Unifying Wavelet and Transformers for Visual Representation Learning

SimViT: Exploring a Simple Vision Transformer with sliding windows

Multi-entity Video Transformers for Fine-Grained Video Representation Learning

Adapting Short-Term Transformers for Action Detection in Untrimmed Videos

Vision Transformer with Cross-attention by Temporal Shift for Efficient Action Recognition

ASAFormer: Visual tracking with convolutional vision transformer and asymmetric selective attention

ViTAE: Vision Transformer Advanced by Exploring Intrinsic Inductive Bias

MaxViT: Multi-Axis Vision Transformer

AdaViT: Adaptive Vision Transformers for Efficient Image Recognition