Abstract:Self-supervised monocular depth estimation can exhibit excellent performance in static environments due to the multi-view consistency assumption during the training process. However, it is hard to maintain depth consistency in dynamic scenes when considering the occlusion problem caused by moving objects. For this reason, we propose a method of self-supervised self-distillation for monocular depth estimation (SS-MDE) in dynamic scenes, where a deep network with a multi-scale decoder and a lightweight pose network are designed to predict depth in a self-supervised manner via the disparity, motion information, and the association between two adjacent frames in the image sequence. Meanwhile, in order to improve the depth estimation accuracy of static areas, the pseudo-depth images generated by the LeReS network are used to provide the pseudo-supervision information, enhancing the effect of depth refinement in static areas. Furthermore, a forgetting factor is leveraged to alleviate the dependency on the pseudo-supervision. In addition, a teacher model is introduced to generate depth prior information, and a multi-view mask filter module is designed to implement feature extraction and noise filtering. This can enable the student model to better learn the deep structure of dynamic scenes, enhancing the generalization and robustness of the entire model in a self-distillation manner. Finally, on four public data datasets, the performance of the proposed SS-MDE method outperformed several state-of-the-art monocular depth estimation techniques, achieving an accuracy (δ1) of 89% while minimizing the error (AbsRel) by 0.102 in NYU-Depth V2 and achieving an accuracy (δ1) of 87% while minimizing the error (AbsRel) by 0.111 in KITTI.

Edge Devices Friendly Self-Supervised Monocular Depth Estimation Via Knowledge Distillation.

Monocular Depth Estimation Based on Unsupervised Learning

Evitbins: Edge-Enhanced Vision-Transformer Bins for Monocular Depth Estimation on Edge Devices

Promoting CNNs with Cross-Architecture Knowledge Distillation for Efficient Monocular Depth Estimation

TinyDepth: Lightweight Self-Supervised Monocular Depth Estimation Based on Transformer

Self-supervised Monocular Depth Estimation with Self-Distillation and Dense Skip Connection

Lightweight Monocular Depth Estimation on Edge Devices

Boosting Light-Weight Depth Estimation Via Knowledge Distillation

Monocular Depth Estimation Using Deep Learning: A Review

Structure-Centric Robust Monocular Depth Estimation via Knowledge Distillation

Efficient Monocular Depth Estimation for Edge Devices in Internet of Things

MonoViT: Self-Supervised Monocular Depth Estimation with a Vision Transformer

Deep Neural Networks with Attention Mechanism for Monocular Depth Estimation on Embedded Devices

LUMDE: Light-Weight Unsupervised Monocular Depth Estimation Via Knowledge Distillation

Knowledge Distillation for Fast and Accurate Monocular Depth Estimation on Mobile Devices

Monocular Depth Estimation via Self-Supervised Self-Distillation

Self-Supervised Monocular Depth Estimation with Self-Reference Distillation and Disparity Offset Refinement

AggNet for Self-supervised Monocular Depth Estimation: Go an Aggressive Step Furthe.

Towards a Unified Network for Robust Monocular Depth Estimation: Network Architecture, Training Strategy and Dataset

MonoER - A Edge Refined Self-Supervised Monocular Depth Estimation Method

Crafting Monocular Cues and Velocity Guidance for Self-Supervised Multi-Frame Depth Learning