Abstract:In recent years, computer vision has witnessed remarkable advancements in image classification, specifically in the domains of fully convolutional neural networks (FCNs) and self-attention mechanisms. Nevertheless, both approaches exhibit certain limitations. FCNs tend to prioritize local information, potentially overlooking crucial global contexts, whereas self-attention mechanisms are computationally intensive despite their adaptability. In order to surmount these challenges, this paper proposes cross-and-diagonal networks (CDNet), innovative network architecture that adeptly captures global information in images while preserving local details in a more computationally efficient manner. CDNet achieves this by establishing long-range relationships between pixels within an image, enabling the indirect acquisition of contextual information. This inventive indirect self-attention mechanism significantly enhances the network's capacity. In CDNet, a new attention mechanism named "cross and diagonal attention" is proposed. This mechanism adopts an indirect approach by integrating two distinct components, cross attention and diagonal attention. By computing attention in different directions, specifically vertical and diagonal, CDNet effectively establishes remote dependencies among pixels, resulting in improved performance in image classification tasks. Experimental results highlight several advantages of CDNet. Firstly, it introduces an indirect self-attention mechanism that can be effortlessly integrated as a module into any convolutional neural network (CNN). Additionally, the computational cost of the self-attention mechanism has been effectively reduced, resulting in improved overall computational efficiency. Lastly, CDNet attains state-of-the-art performance on three benchmark datasets for similar types of image classification networks. In essence, CDNet addresses the constraints of conventional approaches and provides an efficient and effective solution for capturing global context in image classification tasks.

Attending Category Disentangled Global Context for Image Classification

Graph Attention Mechanism with Global Contextual Information for Multi-Label Image Recognition

GEA-net - Global Embedded Attention Neural Network for Image Classification.

Spatial Global Context Attention for Convolutional Neural Networks: an Efficient Method

Modeling Local and Global Contexts for Image Captioning

Learning visual relationship and context-aware attention for image captioning

Contextuality Helps Representation Learning for Generalized Category Discovery

Spatial Context-Aware Object-Attentional Network for Multi-Label Image Classification

Cross-and-Diagonal Networks: An Indirect Self-Attention Mechanism for Image Classification

Attentive Contexts for Object Detection

Ground-Based Remote Sensing Cloud Classification Via Context Graph Attention Network

Global Context Dependencies Aware Network for Efficient Semantic Segmentation of Fine-Resolution Remoted Sensing Images

DeepGCNs-Att: Point Cloud Semantic Segmentation with Contextual Point Representations

One Shot Object Detection with Mutual Global Context.

Attention-based Dual Context Aggregation for Image Semantic Segmentation

Generalized Category Discovery in Aerial Image Classification via Slot Attention

Learning More Discriminative Clues with Gradual Attention for Fine-Grained Visual Categorization.

Hyperspectral Image Classification with Context-Aware Dynamic Graph Convolutional Network

Attention Graph: Learning Effective Visual Features for Large-Scale Image Classification

Cross-Modal Attentional Context Learning for RGB-D Object Detection

Leveraging Attention-Based Visual Clue Extraction for Image Classification