Abstract:This work addresses the task of zero-shot monocular depth estimation. A recent advance in this field has been the idea of utilising Text-to-Image foundation models, such as Stable Diffusion. Foundation models provide a rich and generic image representation, and therefore, little training data is required to reformulate them as a depth estimation model that predicts highly-detailed depth maps and has good generalisation capabilities. However, the realisation of this idea has so far led to approaches which are, unfortunately, highly inefficient at test-time due to the underlying iterative denoising process. In this work, we propose a different realisation of this idea and present PrimeDepth, a method that is highly efficient at test time while keeping, or even enhancing, the positive aspects of diffusion-based approaches. Our key idea is to extract from Stable Diffusion a rich, but frozen, image representation by running a single denoising step. This representation, we term preimage, is then fed into a refiner network with an architectural inductive bias, before entering the downstream task. We validate experimentally that PrimeDepth is two orders of magnitude faster than the leading diffusion-based method, Marigold, while being more robust for challenging scenarios and quantitatively marginally superior. Thereby, we reduce the gap to the currently leading data-driven approach, Depth Anything, which is still quantitatively superior, but predicts less detailed depth maps and requires 20 times more labelled data. Due to the complementary nature of our approach, even a simple averaging between PrimeDepth and Depth Anything predictions can improve upon both methods and sets a new state-of-the-art in zero-shot monocular depth estimation. In future, data-driven approaches may also benefit from integrating our preimage.

Refinement of Monocular Depth Maps via Multi-View Differentiable Rendering

V2Depth: Monocular Depth Estimation via Feature-Level Virtual-View Simulation and Refinement

360MonoDepth: High-Resolution 360° Monocular Depth Estimation

Real Time Complete Dense Depth Reconstruction for a Monocular Camera

PatchRefiner: Leveraging Synthetic Data for Real-Domain High-Resolution Monocular Metric Depth Estimation

Crafting Monocular Cues and Velocity Guidance for Self-Supervised Multi-Frame Depth Learning

Monocular Depth Estimation using Diffusion Models

Depth Refinement for Improved Stereo Reconstruction

Double Refinement Network for Efficient Indoor Monocular Depth Estimation

Digging Into Self-Supervised Monocular Depth Estimation

Multi-View Reconstruction using Signed Ray Distance Functions (SRDF)

PrimeDepth: Efficient Monocular Depth Estimation with a Stable Diffusion Preimage

Robust Geometry-Preserving Depth Estimation Using Differentiable Rendering

BetterDepth: Plug-and-Play Diffusion Refiner for Zero-Shot Monocular Depth Estimation

Neural Radiance Field-Inspired Depth Map Refinement for Accurate Multi-View Stereo

Self-Supervised Monocular Depth Estimation with Self-Reference Distillation and Disparity Offset Refinement

3D Hierarchical Refinement and Augmentation for Unsupervised Learning of Depth and Pose From Monocular Video

Monocular Depth Decomposition of Semi-Transparent Volume Renderings

Monocular Depth Estimation with Guidance of Surface Normal Map

Structure-Centric Robust Monocular Depth Estimation via Knowledge Distillation