Abstract:Object detection technology plays a crucial role in people's everyday lives, as well as enterprise production and modern national defense. Most current object detection networks, such as YOLOX, employ convolutional neural networks instead of a Transformer as a backbone. However, these techniques lack a global understanding of the images and may lose meaningful information, such as the precise location of the most active feature detector. Recently, a Transformer with larger receptive fields showed superior performance to corresponding convolutional neural networks in computer vision tasks. The Transformer splits the image into patches and subsequently feeds them to the Transformer in a sequence structure similar to word embeddings. This makes it capable of global modeling of entire images and implies global understanding of images. However, simply using a Transformer with a larger receptive field raises several concerns. For example, self-attention in the Swin Transformer backbone will limit its ability to model long range relations, resulting in poor feature extraction results and low convergence speed during training. To address the above problems, first, we propose an important region-based Reconstructed Deformable Self-Attention that shifts attention to important regions for efficient global modeling. Second, based on the Reconstructed Deformable Self-Attention, we propose the Swin Deformable Transformer backbone, which improves the feature extraction ability and convergence speed. Finally, based on the Swin Deformable Transformer backbone, we propose a novel object detection network, namely, Swin Deformable Transformer-BiPAFPN-YOLOX. experimental results on the COCO dataset show that the training period is reduced by 55.4%, average precision is increased by 2.4%, average precision of small objects is increased by 3.7%, and inference speed is increased by 35%.

SwinNet: Swin Transformer drives edge-aware RGB-D and RGB-T salient object detection

Swin Transformer-Based Edge Guidance Network for RGB-D Salient Object Detection

SwinSOD: Salient object detection using swin-transformer

Swin Transformer: Hierarchical Vision Transformer using Shifted Windows

SwinHCST: a deep learning network architecture for scene classification of remote sensing images based on improved CNN and Transformer

SwinSUNet: Pure Transformer Network for Remote Sensing Image Change Detection

SwinFuse: A Residual Swin Transformer Fusion Network for Infrared and Visible Images

SwinTFNet: Dual-Stream Transformer With Cross Attention Fusion for Land Cover Classification

Swin Transformer coupling CNNs Makes Strong Contextual Encoders for VHR Image Road Extraction

Improved deep learning image classification algorithm based on Swin Transformer V2

Target detection based on improved swin transformer and cascade RCNN

Class-Guided Swin Transformer for Semantic Segmentation of Remote Sensing Imagery

P-Swin: Parallel Swin transformer multi-scale semantic segmentation network for land cover classification

Swin transformer and ResNet based deep networks for low-light image enhancement

CSWin Transformer: A General Vision Transformer Backbone with Cross-Shaped Windows

Object Detection Based on Swin Deformable Transformer-BiPAFPN-YOLOX

SSTrans-Net: Smart Swin Transformer Network for medical image segmentation

An Improved Swin Transformer-Based Model for Remote Sensing Object Detection and Instance Segmentation