Abstract:Self-supervised monocular depth estimation can exhibit excellent performance in static environments due to the multi-view consistency assumption during the training process. However, it is hard to maintain depth consistency in dynamic scenes when considering the occlusion problem caused by moving objects. For this reason, we propose a method of self-supervised self-distillation for monocular depth estimation (SS-MDE) in dynamic scenes, where a deep network with a multi-scale decoder and a lightweight pose network are designed to predict depth in a self-supervised manner via the disparity, motion information, and the association between two adjacent frames in the image sequence. Meanwhile, in order to improve the depth estimation accuracy of static areas, the pseudo-depth images generated by the LeReS network are used to provide the pseudo-supervision information, enhancing the effect of depth refinement in static areas. Furthermore, a forgetting factor is leveraged to alleviate the dependency on the pseudo-supervision. In addition, a teacher model is introduced to generate depth prior information, and a multi-view mask filter module is designed to implement feature extraction and noise filtering. This can enable the student model to better learn the deep structure of dynamic scenes, enhancing the generalization and robustness of the entire model in a self-distillation manner. Finally, on four public data datasets, the performance of the proposed SS-MDE method outperformed several state-of-the-art monocular depth estimation techniques, achieving an accuracy (δ1) of 89% while minimizing the error (AbsRel) by 0.102 in NYU-Depth V2 and achieving an accuracy (δ1) of 87% while minimizing the error (AbsRel) by 0.111 in KITTI.

AggNet for Self-supervised Monocular Depth Estimation: Go an Aggressive Step Furthe.

Monocular Depth Estimation Based on Unsupervised Learning

A Depth Estimation Framework Based on Unsupervised Learning and Cross-Modal Translation

MDSNet: self-supervised monocular depth estimation for video sequences using self-attention and threshold mask

Monocular Depth Estimation via Self-Supervised Self-Distillation

MLDA-Net: Multi-Level Dual Attention-Based Network for Self-Supervised Monocular Depth Estimation

MBUDepthNet: Real-Time Unsupervised Monocular Depth Estimation Method for Outdoor Scenes

Self-Supervised Monocular Depth Estimation With Self-Perceptual Anomaly Handling

Self-Supervised Monocular Depth Estimation Based on High-Order Spatial Interactions

MSFNet:Multi-scale features network for monocular depth estimation

Unsupervised Monocular Depth Estimation Based on Hierarchical Feature-Guided Diffusion

Self-supervised monocular depth estimation via joint attention and intelligent mask loss

Frequency-Aware Self-Supervised Monocular Depth Estimation

Digging Into Self-Supervised Monocular Depth Estimation

Unsupervised Monocular Estimation of Depth and Visual Odometry uUsing Attention and Depth-Pose Consistency Loss

AGG-Net: Attention Guided Gated-convolutional Network for Depth Image Completion

Deep Neighbor Layer Aggregation for Lightweight Self-Supervised Monocular Depth Estimation

Detaching and Boosting: Dual Engine for Scale-Invariant Self-Supervised Monocular Depth Estimation

Dual-attention-based semantic-aware self-supervised monocular depth estimation

Towards Loss Balance and Consistent Model in Self-supervised Monocular Depth Estimation