Abstract:Recently, Vision Transformers (ViTs) have been broadly explored in visual recognition. With low efficiency in encoding fine-level features, the performance of ViTs is still inferior to the state-of-the-art CNNs when trained from scratch on a midsize dataset like ImageNet. Through experimental analysis, we find it is because of two reasons: 1) the simple tokenization of input images fails to model the important local structure such as edges and lines, leading to low training sample efficiency; 2) the redundant attention backbone design of ViTs leads to limited feature richness for fixed computation budgets and limited training samples. To overcome such limitations, we present a new simple and generic architecture, termed Vision Outlooker (VOLO), which implements a novel outlook attention operation that dynamically conduct the local feature aggregation mechanism in a sliding window manner across the input image. Unlike self-attention that focuses on modeling global dependencies of local features at a coarse level, our outlook attention targets at encoding finer-level features, which is critical for recognition but ignored by self-attention. Outlook attention breaks the bottleneck of self-attention whose computation cost scales quadratically with the input spatial dimension, and thus is much more memory efficient. Compared to our Tokens-To-Token Vision Transformer (T2T-ViT), VOLO can more efficiently encode fine-level features that are essential for high-performance visual recognition. Experiments show that with only 26.6 M learnable parameters, VOLO achieves 84.2% top-1 accuracy on ImageNet-1 K without using extra training data, 2.7% better than T2T-ViT with a comparable number of parameters. When the model size is scaled up to 296 M parameters, its performance can be further improved to 87.1%, setting a new record for ImageNet-1 K classification. In addition, we also take the proposed VOLO as pretr- ined models and report superior performance on downstream tasks, such as semantic segmentation. Code is available at https://github.com/sail-sg/volo.

ZoomViT: an Observation Behavior-Based Fine-Grained Recognition Scheme

SLViT: Scale-Wise Language-Guided Vision Transformer for Referring Image Segmentation.

Recombining Vision Transformer Architecture for Fine-Grained Visual Categorization.

Ts-vit: feature-enhanced transformer via token selection for fine-grained image recognition

VOLO: Vision Outlooker for Visual Recognition

TransFG: A Transformer Architecture for Fine-Grained Recognition

A free lunch from ViT:Adaptive Attention Multi-scale Fusion Transformer for Fine-grained Visual Recognition

A FREE LUNCH FROM VIT: ADAPTIVE ATTENTION MULTI-SCALE FUSION TRANSFORMER FOR FINE-GRAINED VISUAL RECOGNITION

A Vision Transformer for Fine-Grained Classification by Reducing Noise and Enhancing Discriminative Information

ViT-FOD: A Vision Transformer based Fine-grained Object Discriminator

Scalable Vision Transformers with Hierarchical Pooling.

So-ViT: Mind Visual Tokens for Vision Transformer

Multi-level information fusion Transformer with background filter for fine-grained image recognition

DeepViT: Towards Deeper Vision Transformer

Plant and Animal Species Recognition Based on Dynamic Vision Transformer Architecture

CF-ViT: A General Coarse-to-Fine Method for Vision Transformer

A Sequence-selective Fine-grained Image Recognition Strategy Using Vision Transformer

Progressive Learning Vision Transformer for Open Set Recognition of Fine-Grained Objects in Remote Sensing Images.

ReViT: Enhancing Vision Transformers Feature Diversity with Attention Residual Connections

A Transformer Architecture with Adaptive Attention for Fine-Grained Visual Classification