Abstract:Decoders play significant roles in recovering scene depths. However, the decoders used in previous works ignore the propagation of multilevel lossless fine-grained information, cannot adaptively capture local and global information in parallel, and cannot perform sufficient global statistical analyses on the final output disparities. In addition, the process of mapping from a low-resolution feature space to a high-resolution feature space is a one-to-many problem that may have multiple solutions. Therefore, the quality of the recovered depth map is low. To this end, we propose a high-quality decoder (HQDec), with which multilevel near-lossless fine-grained information, obtained by the proposed adaptive axial-normalized position-embedded channel attention sampling module (AdaAxialNPCAS), can be adaptively incorporated into a low-resolution feature map with high-level semantics utilizing the proposed adaptive information exchange scheme. In the HQDec, we leverage the proposed adaptive refinement module (AdaRM) to model the local and global dependencies between pixels in parallel and utilize the proposed disparity attention module to model the distribution characteristics of disparity values from a global perspective. To recover fine-grained high-resolution features with maximal accuracy, we adaptively fuse the high-frequency information obtained by constraining the upsampled solution space utilizing the local and global dependencies between pixels into the high-resolution feature map generated from the nonlearning method. Extensive experiments demonstrate that each proposed component improves the quality of the depth estimation results over the baseline results, and the developed approach achieves state-of-the-art results on the KITTI and DDAD datasets. The code and models will be publicly available at \href{<a class="link-external link-https" href="https://github.com/fwucas/HQDec" rel="external noopener nofollow">this https URL</a>}{HQDec}.

Chfnet: a coarse-to-fine hierarchical refinement model for monocular depth estimation

MFF-Net: Towards Efficient Monocular Depth Completion With Multi-Modal Feature Fusion

A Robust Monocular Depth Estimation Framework Based on Light-Weight ERF-Pspnet for Day-Night Driving Scenes

Monocular Depth Estimation Based on Multi-Scale Graph Convolution Networks

A Depth Estimation Framework Based on Unsupervised Learning and Cross-Modal Translation

Monocular depth estimation with hierarchical fusion of dilated CNNs and soft-weighted-sum inference

Self-supervised coarse-to-fine monocular depth estimation using a lightweight attention module

FA-Depth: Toward Fast and Accurate Self-supervised Monocular Depth Estimation

Learning Occlusion-Aware Coarse-to-Fine Depth Map for Self-supervised Monocular Depth Estimation

MSFNet:Multi-scale features network for monocular depth estimation

Pyramid Feature Attention Network for Monocular Depth Prediction

HQDec: Self-Supervised Monocular Depth Estimation Based on a High-Quality Decoder

Fast Monocular Depth Estimation via Side Prediction Aggregation with Continuous Spatial Refinement

HR-Depth: High Resolution Self-Supervised Monocular Depth Estimation

BRNet: Exploring Comprehensive Features for Monocular Depth Estimation.

SAU-Net: Monocular Depth Estimation Combining Multi-Scale Features and Attention Mechanisms

LW-Net: A Lightweight Network for Monocular Depth Estimation

Learning Depth from Single Monocular Images Using Deep Convolutional Neural Fields

A Compromise Principle in Deep Monocular Depth Estimation

Resolution-sensitive self-supervised monocular absolute depth estimation

Structured Attention Guided Convolutional Neural Fields for Monocular Depth Estimation