Abstract:For video recognition task, a global representation summarizing the whole contents of the video snippets plays an important role for the final performance. However, existing video architectures usually generate it by using a simple, global average pooling (GAP) method, which has limited ability to capture complex dynamics of videos. For image recognition task, there exist evidences showing that covariance pooling has stronger representation ability than GAP. Unfortunately, such plain covariance pooling used in image recognition is an orderless representative, which cannot model spatio-temporal structure inherent in videos. Therefore, this paper proposes a Temporal-attentive Covariance Pooling(TCP), inserted at the end of deep architectures, to produce powerful video representations. Specifically, our TCP first develops a temporal attention module to adaptively calibrate spatio-temporal features for the succeeding covariance pooling, approximatively producing attentive covariance representations. Then, a temporal covariance pooling performs temporal pooling of the attentive covariance representations to characterize both intra-frame correlations and inter-frame cross-correlations of the calibrated features. As such, the proposed TCP can capture complex temporal dynamics. Finally, a fast matrix power normalization is introduced to exploit geometry of covariance representations. Note that our TCP is model-agnostic and can be flexibly integrated into any video architectures, resulting in TCPNet for effective video recognition. The extensive experiments on six benchmarks (e.g., Kinetics, Something-Something V1 and Charades) using various video architectures show our TCPNet is clearly superior to its counterparts, while having strong generalization ability. The source code is publicly available.

Multi-Dimensional Attentive Hierarchical Graph Pooling Network for Video-Text Retrieval.

Multilevel Spatial-Temporal Feature Aggregation for Video Object Detection

Local-Global Graph Pooling Via Mutual Information Maximization for Video-Paragraph Retrieval

Stacked Convolutional Deep Encoding Network for Video-Text Retrieval.

Text-Video Retrieval via Variational Multi-Modal Hypergraph Networks

A Multi-interaction Model with Cross-Branch Feature Fusion for Video-Text Retrieval.

Fine-grained Video-Text Retrieval with Hierarchical Graph Reasoning

Tencent Text-Video Retrieval: Hierarchical Cross-Modal Interactions with Multi-Level Representations

MGSGA: Multi-grained and Semantic-Guided Alignment for Text-Video Retrieval

Hierarchical Cross-Modal Graph Consistency Learning for Video-Text Retrieval.

Video–text retrieval via multi-modal masked transformer and adaptive attribute-aware graph convolutional network

Multi-granularity graph pooling for video-based person re-identification

Temporal Multimodal Graph Transformer With Global-Local Alignment for Video-Text Retrieval

Multi-Granularity and Multi-modal Feature Interaction Approach for Text Video Retrieval

Video Captioning with Aggregated Features Based on Dual Graphs and Gated Fusion

Hierarchical Multi-View Graph Pooling with Structure Learning

Multiple Hypergraph Ranking for Video Concept Detection

Adversarial Multi-Grained Embedding Network for Cross-Modal Text-Video Retrieval

HANet: Hierarchical Alignment Networks for Video-Text Retrieval

Temporal-attentive Covariance Pooling Networks for Video Recognition

Dig into Multi-modal Cues for Video Retrieval with Hierarchical Alignment.