Abstract:Vision transformers (ViTs) have been trending in image classification tasks due to their promising performance when compared to convolutional neural networks (CNNs). As a result, many researchers have tried to incorporate ViTs in hyperspectral image (HSI) classification tasks. To achieve satisfactory performance, close to that of CNNs, transformers need fewer parameters. ViTs and other similar transformers use an external classification (CLS) token which is randomly initialized and often fails to generalize well, whereas other sources of multimodal datasets, such as light detection and ranging (LiDAR) offer the potential to improve these models by means of a CLS. In this paper, we introduce a new multimodal fusion transformer (MFT) network which comprises a multihead cross patch attention (mCrossPA) for HSI land-cover classification. Our mCrossPA utilizes other sources of complementary information in addition to the HSI in the transformer encoder to achieve better generalization. The concept of tokenization is used to generate CLS and HSI patch tokens, helping to learn a {distinctive representation} in a reduced and hierarchical feature space. Extensive experiments are carried out on {widely used benchmark} datasets {i.e.,} the University of Houston, Trento, University of Southern Mississippi Gulfpark (MUUFL), and Augsburg. We compare the results of the proposed MFT model with other state-of-the-art transformers, classical CNNs, and conventional classifiers models. The superior performance achieved by the proposed model is due to the use of multihead cross patch attention. The source code will be made available publicly at \url{<a class="link-external link-https" href="https://github.com/AnkurDeria/MFT" rel="external noopener nofollow">this https URL</a>}.}

MCT-VHD: Multi-modal contrastive transformer for video highlight detection

MH-DETR: Video Moment and Highlight Detection with Cross-modal Transformer

HiT: Hierarchical Transformer with Momentum Contrast for Video-Text Retrieval

UMT: Unified Multi-modal Transformers for Joint Video Moment Retrieval and Highlight Detection

VideoLights: Feature Refinement and Cross-Task Alignment Transformer for Joint Video Highlight Detection and Moment Retrieval

Efficient Multiscale Multimodal Bottleneck Transformer for Audio-Video Classification

HMTV: hierarchical multimodal transformer for video highlight query on baseball

Multi-Scale Temporal Difference Transformer for Video-Text Retrieval

Bridging the Gap: A Unified Video Comprehension Framework for Moment Retrieval and Highlight Detection

Unsupervised Modality-Transferable Video Highlight Detection With Representation Activation Sequence Learning

Efficient Selective Audio Masked Multimodal Bottleneck Transformer for Audio-Video Classification

TR-DETR: Task-Reciprocal Transformer for Joint Moment Retrieval and Highlight Detection

MTCAM: A Novel Weakly-Supervised Audio-Visual Saliency Prediction Model with Multi-Modal Transformer

Multimodal Fusion Transformer for Remote Sensing Image Classification

A Multimodal Transformer for Live Streaming Highlight Prediction

Transformer-Based Interactive Multi-Modal Attention Network for Video Sentiment Detection

Temporal Cue Guided Video Highlight Detection with Low-Rank Audio-Visual Fusion

Audiovisual Highlight Detection in Videos

Hybrid multi-attention transformer for robust video object detection

Multi-Modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection