Abstract:Self-supervised monocular depth estimation can exhibit excellent performance in static environments due to the multi-view consistency assumption during the training process. However, it is hard to maintain depth consistency in dynamic scenes when considering the occlusion problem caused by moving objects. For this reason, we propose a method of self-supervised self-distillation for monocular depth estimation (SS-MDE) in dynamic scenes, where a deep network with a multi-scale decoder and a lightweight pose network are designed to predict depth in a self-supervised manner via the disparity, motion information, and the association between two adjacent frames in the image sequence. Meanwhile, in order to improve the depth estimation accuracy of static areas, the pseudo-depth images generated by the LeReS network are used to provide the pseudo-supervision information, enhancing the effect of depth refinement in static areas. Furthermore, a forgetting factor is leveraged to alleviate the dependency on the pseudo-supervision. In addition, a teacher model is introduced to generate depth prior information, and a multi-view mask filter module is designed to implement feature extraction and noise filtering. This can enable the student model to better learn the deep structure of dynamic scenes, enhancing the generalization and robustness of the entire model in a self-distillation manner. Finally, on four public data datasets, the performance of the proposed SS-MDE method outperformed several state-of-the-art monocular depth estimation techniques, achieving an accuracy (δ1) of 89% while minimizing the error (AbsRel) by 0.102 in NYU-Depth V2 and achieving an accuracy (δ1) of 87% while minimizing the error (AbsRel) by 0.111 in KITTI.

Self-Supervised Human Depth Estimation from Monocular Videos

Monocular Depth Estimation Based on Unsupervised Learning

Region Deformer Networks for Unsupervised Depth Estimation from Unconstrained Monocular Videos

Self-supervised 3D Representation Learning of Dressed Humans from Social Media Videos

3D Object Aided Self-Supervised Monocular Depth Estimation

Unsupervised Scale-Consistent Depth Learning from Video

Self-Supervised Monocular Depth Estimation With Self-Perceptual Anomaly Handling

Towards Practical Consistent Video Depth Estimation.

MDSNet: self-supervised monocular depth estimation for video sequences using self-attention and threshold mask

Unsupervised Monocular Depth Perception: Focusing on Moving Objects

Unsupervised Learning of Depth from Monocular Videos Using 3D-2D Corresponding Constraints

Monocular Depth Estimation via Self-Supervised Self-Distillation

Sparseness Meets Deepness: 3D Human Pose Estimation from Monocular Video

A Lightweight Self-Supervised Training Framework for Monocular Depth Estimation

Embodiment: Self-Supervised Depth Estimation Based on Camera Models

Unsupervised Monocular Depth Learning in Dynamic Scenes

Self-Supervised Learning for Monocular Depth Estimation from Aerial Imagery

RM-Depth: Unsupervised Learning of Recurrent Monocular Depth in Dynamic Scenes

Self-Supervised Monocular Depth Estimation With Multiscale Perception

Self-Supervised Monocular Depth Estimation with Self-Reference Distillation and Disparity Offset Refinement