Abstract:In this article, we tackle the problem of depth estimation from single monocular images. Compared with depth estimation using multiple images such as stereo depth perception, depth from monocular images is much more challenging. Prior work typically focuses on exploiting geometric priors or additional sources of information, most using hand-crafted features. Recently, there is mounting evidence that features from deep convolutional neural networks (CNN) set new records for various vision applications. On the other hand, considering the continuous characteristic of the depth values, depth estimation can be naturally formulated as a continuous conditional random field (CRF) learning problem. Therefore, here we present a deep convolutional neural field model for estimating depths from single monocular images, aiming to jointly explore the capacity of deep CNN and continuous CRF. In particular, we propose a deep structured learning scheme which learns the unary and pairwise potentials of continuous CRF in a unified deep CNN framework. We then further propose an equally effective model based on fully convolutional networks and a novel superpixel pooling method, which is about 10 times faster, to speedup the patch-wise convolutions in the deep model. With this more efficient model, we are able to design deeper networks to pursue better performance. Our proposed method can be used for depth estimation of general scenes with no geometric priors nor any extra information injected. In our case, the integral of the partition function can be calculated in a closed form such that we can exactly solve the log-likelihood maximization. Moreover, solving the inference problem for predicting depths of a test image is highly efficient as closed-form solutions exist. Experiments on both indoor and outdoor scene datasets demonstrate that the proposed method outperforms state-of-the-art depth estimation approaches.

Depth Estimation of Supervised Monocular Images Based on Semantic Segmentation.

A Depth Estimation Framework Based on Unsupervised Learning and Cross-Modal Translation

Monocular Depth Estimation Based on Unsupervised Learning

Binocular Depth Estimation Using Convolutional Neural Network With Siamese Branches.

Monocular image depth estimation using dilated convolution and spatial pyramid polling structure

Monocular Depth Estimation With Affinity, Vertical Pooling, And Label Enhancement

Enhanced Monocular Depth Estimation: A CNN Integrating Semantic Segmentation Embedding And Vanishing Point Detection

Depth Estimation from Monocular Images Using Dilated Convolution and Uncertainty Learning.

Monocular Depth Estimation Based on Dilated Convolutions and Feature Fusion

Depth Monocular Estimation with Attention-based Encoder-Decoder Network from Single Image

Semantic-Guided Representation Enhancement for Self-supervised Monocular Trained Depth Estimation

Attention-Based Monocular Depth Estimation Considering Global and Local Information in Remote Sensing Images

Unsupervised depth estimation from monocular videos with hybrid geometric-refined loss and contextual attention

SDC-Depth: Semantic Divide-and-Conquer Network for Monocular Depth Estimation

Monocular depth estimation with hierarchical fusion of dilated CNNs and soft-weighted-sum inference

A Self-Supervised Monocular Depth Estimation Method Based on High Resolution Convolutional Neural Network

GlobalDepth: Global-Aware Attention Model for Unsupervised Monocular Depth Estimation.

Hierarchical Object Relationship Constrained Monocular Depth Estimation.

Collaborative Deconvolutional Neural Networks for Joint Depth Estimation and Semantic Segmentation

Self-supervised Monocular Depth Estimation with Multi-Scale Feature Fusion

Learning Depth from Single Monocular Images Using Deep Convolutional Neural Fields