Abstract:Over the past two decades, traditional block-based video coding has made remarkable progress and spawned a series of well-known standards such as MPEG-4, H.264/AVC and H.265/HEVC. On the other hand, deep neural networks (DNNs) have shown their powerful capacity for visual content understanding, feature extraction and compact representation. Some previous works have explored the learnt video coding algorithms in an end-to-end manner, which show the great potential compared with traditional methods. In this paper, we propose an end-to-end deep neural video coding framework (NVC), which uses variational autoencoders (VAEs) with joint spatial and temporal prior aggregation (PA) to exploit the correlations in intra-frame pixels, inter-frame motions and inter-frame compensation residuals, respectively. Novel features of NVC include: 1) To estimate and compensate motion over a large range of magnitudes, we propose an unsupervised multiscale motion compensation network (MS-MCN) together with a pyramid decoder in the VAE for coding motion features that generates multiscale flow fields, 2) we design a novel adaptive spatiotemporal context model for efficient entropy coding for motion information, 3) we adopt nonlocal attention modules (NLAM) at the bottlenecks of the VAEs for implicit adaptive feature extraction and activation, leveraging its high transformation capacity and unequal weighting with joint global and local information, and 4) we introduce multi-module optimization and a multi-frame training strategy to minimize the temporal error propagation among P-frames. NVC is evaluated for the low-delay causal settings and compared with H.265/HEVC, H.264/AVC and the other learnt video compression methods following the common test conditions, demonstrating consistent gains across all popular test sequences for both PSNR and MS-SSIM distortion metrics.

End-To-End Compression for Surveillance Video with Unsupervised Foreground-Background Separation

Foreground-Background Parallel Compression with Residual Encoding for Surveillance Video

Multi-scale and Bi-path Method Based on Image Entropy and CNN for Fast CU Partition in VVC

An Efficient Compressive Convolutional Network for Unified Object Detection and Image Compression

Fast Encoding of Surveillance Videos Based on Hevc

A deep learning approach for quality enhancement of surveillance video

Deep Predictive Video Compression Using Mode-Selective Uni- and Bi-Directional Predictions Based on Multi-Frame Hypothesis

The significance of bone marrow involvement in non-Hodgkin's lymphoma: the Eastern Cooperative Oncology Group experience.

A fast background model based surveillance video coding in HEVC

Instance Segmentation Based Background Reference Frame Generation for Surveillance Video Coding

Neural Video Coding Using Multiscale Motion Compensation and Spatiotemporal Context Model

High Efficiency Deep-learning Based Video Compression

A coding unit classification based AVC-to-HEVC transcoding with background modeling for surveillance videos.

A Background Modeling Scheme Based on High Efficiency Motion Classification for Surveillance Video Coding

FVC: An End-to-End Framework Towards Deep Video Compression in Feature Space

Enhanced Surveillance Video Compression with Dual Reference Frames Generation

SSNVC: Single Stream Neural Video Compression with Implicit Temporal Information

High-Efficiency Neural Video Compression via Hierarchical Predictive Learning

Intelligent Analysis Oriented Surveillance Video Coding.

End-to-end Compression Towards Machine Vision: Network Architecture Design and Optimization

M-LVC: Multiple Frames Prediction for Learned Video Compression