Abstract:Compared with spatial counterparts, temporal relationships between frames and their influences on video quality assessment (VQA) are still relatively under-studied in existing works. These relationships lead to two important types of effects for video quality. Firstly, some meaningless temporal variations (such as shaking, flicker, and unsmooth scene transitions) cause temporal distortions that degrade quality of videos. Secondly, the human visual system often has different attention to frames with different contents, resulting in their different importance to the overall video quality. Based on prominent time-series modeling ability of transformers, we propose a novel and effective transformer-based VQA method to tackle these two issues. To better differentiate temporal variations and thus capture the temporal distortions, we design the Spatial-Temporal Distortion Extraction (STDE) module that extracts multi-level spatial-temporal features with a video swin transformer tiny (Swin-T) backbone and uses temporal difference layer to further capture these distortions. To tackle with temporal quality attention, we propose the encoder-decoder-like temporal content transformer (TCT). We also introduce the temporal sampling on features to reduce the input length for the TCT, so as to improve the learning effectiveness and efficiency of this module. Consisting of the STDE and the TCT, the proposed Temporal Distortion-Content Transformers for Video Quality Assessment (DisCoVQA) reaches state-of-the-art performance on several VQA benchmarks without any extra pre-training datasets and up to 10% better generalization ability than existing methods. We also conduct extensive ablation experiments to prove the effectiveness of each part in our proposed model, and provide visualizations to prove that the proposed modules achieve our intention on modeling these temporal issues. Our code is published at https://github.com/QualityAssessment/DisCoVQA.

Learning Spatiotemporal Interactions for User-Generated Video Quality Assessment

Spatio-Temporal Deformable Convolution for Compressed Video Quality Enhancement

Capturing Co-existing Distortions in User-Generated Content for No-reference Video Quality Assessment

Blind Video Quality Assessment Via Space-Time Slice Statistics

DisCoVQA: Temporal Distortion-Content Transformers for Video Quality Assessment

Video Transformer based Video Quality Assessment with Spatiotemporally adaptive Token Selection and Assembly

Video Quality Assessment for Spatio-Temporal Resolution Adaptive Coding

A Spatial-Temporal Video Quality Assessment Method via Comprehensive HVS Simulation

Novel Spatio-Temporal Structural Information Based Video Quality Metric

Spatio-temporal Ssim Index for Video Quality Assessment

STARVQA: SPACE-TIME ATTENTION FOR VIDEO QUALITY ASSESSMENT

Using Spatial‐Temporal Attention for Video Quality Evaluation

STC: Spatio-Temporal Contrastive Learning for Video Instance Segmentation.

Blind Video Quality Prediction by Uncovering Human Video Perceptual Representation

Deep Neural Networks for End-to-End Spatiotemporal Video Quality Prediction and Aggregation

Video Quality Assessment Based on Swin Transformer with Spatio-Temporal Feature Fusion and Data Augmentation

StarVQA+: Co-training Space-Time Attention for Video Quality Assessment

Space-time video super-resolution using long-term temporal feature aggregation

Learning Generalized Spatial-Temporal Deep Feature Representation for No-Reference Video Quality Assessment

STDF: Spatio-Temporal Deformable Fusion for Video Quality Enhancement on Embedded Platforms

PTM-VQA: Efficient Video Quality Assessment Leveraging Diverse PreTrained Models from the Wild