Abstract:Recently, the scenes in large high-resolution remote sensing (HRRS) datasets have been classified using convolutional neural network (CNN)-based methods. Such methods are well-suited for spatial feature extraction and can classify images with relatively high accuracy. However, CNNs do not adequately learn the long-distance dependencies between images and features in image processing, despite this being necessary for HRRS image processing as the semantic content of the scenes in these images is closely related to their spatial relationship. CNNs also have limitations in solving problems related to large intra-class differences and high inter-class similarity. To overcome these challenges, in this study we combine the channel-spatial attention (CSA) mechanism with the Vision Transformer method to propose an effective HRRS image scene classification framework using Channel-Spatial Attention Transformers (CSAT). The proposed model extracts the channel and spatial features of HRRS images using CSA and the Multi-head Self-Attention (MSA) mechanism in the transformer module. First, the HRRS image is mapped into a series of multiple planar 2D patch vectors after passing to the CSA. Second, the ordered vector is obtained via the linear transformation of each vector, and the position and learnable embedding vectors are added to the sequence vector to capture the inter-feature dependencies at a distance from the generated image. Next, we use MSA to extract image features and the residual network structure to complete the encoder construction to solve the gradient disappearance problem and avoid overfitting. Finally, a multi-layer perceptron is used to classify the scenes in the HRRS images. The CSAT network is evaluated using three public remote sensing scene image datasets: UC-Merced, AID, and NWPU-RESISC45. The experimental results show that the proposed CSAT network outperforms a selection of state-of-the-art methods in terms of scene classification.

SwinHCST: a deep learning network architecture for scene classification of remote sensing images based on improved CNN and Transformer

A Novel Transformer Network with a CNN-Enhanced Cross-Attention Mechanism for Hyperspectral Image Classification

Spectral Swin Transformer Network for Hyperspectral Image Classification

Remote sensing image scene classification based on transfer learning and Swin transformer mode

SwinSUNet: Pure Transformer Network for Remote Sensing Image Change Detection

SwinNet: Swin Transformer drives edge-aware RGB-D and RGB-T salient object detection

Hyperspectral Remote-Sensing Classification Combining Transformer and Multiscale Residual Mechanisms

Improved deep learning image classification algorithm based on Swin Transformer V2

A Lightweight Dual-Branch Swin Transformer for Remote Sensing Scene Classification

Class-Guided Swin Transformer for Semantic Segmentation of Remote Sensing Imagery

Swin-CFNet: An Attempt at Fine-Grained Urban Green Space Classification Using Swin Transformer and Convolutional Neural Network

P-Swin: Parallel Swin transformer multi-scale semantic segmentation network for land cover classification

An Improved Swin Transformer-Based Model for Remote Sensing Object Detection and Instance Segmentation

SwinTFNet: Dual-Stream Transformer With Cross Attention Fusion for Land Cover Classification

An Efficient Hybrid CNN-Transformer Approach for Remote Sensing Super-Resolution

Spectral-Swin Transformer with Spatial Feature Extraction Enhancement for Hyperspectral Image Classification

A Siamese Swin-Unet for image change detection

Transformer based on channel-spatial attention for accurate classification of scenes in remote sensing image

A Swin Transformer-Based Fusion Approach for Hyperspectral Image Super-Resolution