Abstract:Inter prediction is an important module in video coding for temporal redundancy removal, where similar reference blocks are searched from previously coded frames and employed to predict the block to be coded. Although traditional video codecs can estimate and compensate for block-level motions, their inter prediction performance is still heavily affected by the remaining inconsistent pixel-wise displacement caused by irregular rotation and deformation. In this paper, we address the problem by proposing a deep frame interpolation network to generate additional reference frames in coding scenarios. First, we summarize the previous adaptive convolutions used for frame interpolation and propose a factorized kernel convolutional network to improve the modeling capacity and simultaneously keep its compact form. Second, to better train this network, multi-domain hierarchical constraints are introduced to regularize the training of our factorized kernel convolutional network. For spatial domain, we use a gradually down-sampled and up-sampled auto-encoder to generate the factorized kernels for frame interpolation at different scales. For quality domain, considering the inconsistent quality of the input frames, the factorized kernel convolution is modulated with quality-related features to learn to exploit more information from high quality frames. For frequency domain, a sum of absolute transformed difference loss that performs frequency transformation is utilized to facilitate network optimization from the view of coding performance. With the well-designed frame interpolation network regularized by multi-domain hierarchical constraints, our method surpasses HEVC on average 6.1% BD-rate saving and up to 11.0% BD-rate saving for the luma component under the random access configuration.

Multiple Resolution Prediction With Deep Up-Sampling for Depth Video Coding

Efficient intra coding algorithm for depth map in 3D video

Convolutional Neural Network Based Up-Sampling for Depth Video Intra Coding

Deep Multi-Domain Prediction for 3D Video Coding.

Efficient depth coding in 3D video to minimize coding bitrate and complexity

A Video Coding System with Spatial-Temporal Down-/Up-Sampling and Super-Resolution Reconstruction

Improved Multi-View Depth Estimation For View Synthesis In 3d Video Coding

Deep Reference Generation with Multi-Domain Hierarchical Constraints for Inter Prediction

Low Complexity Depth Coding Assisted by Coding Information From Color Video

Improved Low-Bitrate HEVC Video Coding Using Deep Learning Based Super-Resolution and Adaptive Block Patching.

Hierarchical Piece-Wise Linear Projections for Efficient Intra-Prediction Coding.

Multi-view Video Coding Based on View Prediction

Multi-Scale Convolutional Neural Network-Based Intra Prediction for Video Coding.

Enhanced Ctu-Level Inter Prediction with Deep Frame Rate Up-Conversion for High Efficiency Video Coding

Reduced Resolution Depth Compression for Multiview Video Plus Depth Coding

Inter Mode Selection for Depth Map Coding in 3D Video.

Deep region segmentation-based intra prediction for depth video coding

A Flexible Lossy Depth Video Coding Scheme Based on Low-rank Tensor Modelling and HEVC Intra Prediction for Free Viewpoint Video

Highly Efficient Multiview Depth Coding Based on Histogram Projection and Allowable Depth Distortion

High-Efficiency Neural Video Compression via Hierarchical Predictive Learning

Deep Inter Prediction Via Pixel-Wise Motion Oriented Reference Generation