Abstract:It is a challenging task to learn discriminative representation from images and videos, due to large local redundancy and complex global dependency in these visual data. Convolution neural networks (CNNs) and vision transformers (ViTs) have been two dominant frameworks in the past few years. Though CNNs can efficiently decrease local redundancy by convolution within a small neighborhood, the limited receptive field makes it hard to capture global dependency. Alternatively, ViTs can effectively capture long-range dependency via self-attention, while blind similarity comparisons among all the tokens lead to high redundancy. To resolve these problems, we propose a novel Unified transFormer (UniFormer), which can seamlessly integrate the merits of convolution and self-attention in a concise transformer format. Different from the typical transformer blocks, the relation aggregators in our UniFormer block are equipped with local and global token affinity respectively in shallow and deep layers, allowing tackling both redundancy and dependency for efficient and effective representation learning. Finally, we flexibly stack our blocks into a new powerful backbone, and adopt it for various vision tasks from image to video domain, from classification to dense prediction. Without any extra training data, our UniFormer achieves 86.3 top-1 accuracy on ImageNet-1 K classification task. With only ImageNet-1 K pre-training, it can simply achieve state-of-the-art performance in a broad range of downstream tasks. It obtains 82.9/84.8 top-1 accuracy on Kinetics-400/600, 60.9/71.2 top-1 accuracy on Something-Something V1/V2 video classification tasks, 53.8 box AP and 46.4 mask AP on COCO object detection task, 50.8 mIoU on ADE20 K semantic segmentation task, and 77.4 AP on COCO pose estimation task. Moreover, we build an efficient UniFormer with a concise hourglass design of token shrinking and recovering, which achieves 2-4[Formula: see text] higher throughput than the recent lightweight models. Code is available at https://github.com/Sense-X/UniFormer.

UniTR: A Unified TRansformer-based Framework for Co-object and Multi-modal Saliency Detection

Learning Spatiotemporal Relationships with a Unified Framework for Video Object Segmentation

Transformer Union Convolution Network for Visual Object Tracking

UniTR: A Unified and Efficient Multi-Modal Transformer for Bird's-Eye-View Representation

A Unified Framework for 3D Scene Understanding

Discriminative Co-Saliency and Background Mining Transformer for Co-Salient Object Detection

SOTR: Segmenting Objects with Transformers

UniST: Towards Unifying Saliency Transformer for Video Saliency Prediction and Detection

UniHead: Unifying Multi-Perception for Detection Heads

UniFormer: Unifying Convolution and Self-Attention for Visual Recognition

Unifying Visual Perception by Dispersible Points Learning

MuTrans: Multiple Transformers for Fusing Feature Pyramid on 2D and 3D Object Detection

Masked-attention Mask Transformer for Universal Image Segmentation

UniTranSeR: A Unified Transformer Semantic Representation Framework for Multimodal Task-Oriented Dialog System

DOCTR: Disentangled Object-Centric Transformer for Point Scene Understanding

Co-Training Transformer for Remote Sensing Image Classification, Segmentation, and Detection

UNetFormer: A Unified Vision Transformer Model and Pre-Training Framework for 3D Medical Image Segmentation

OneFormer3D: One Transformer for Unified Point Cloud Segmentation

Referred by Multi-Modality: A Unified Temporal Transformer for Video Object Segmentation

Uni3DETR: Unified 3D Detection Transformer

GSTran: Joint Geometric and Semantic Coherence for Point Cloud Segmentation