Attention-Based Monocular Depth Estimation Considering Global and Local Information in Remote Sensing Images

Junwei Lv,Yueting Zhang,Jiayi Guo,Xin Zhao,Ming Gao,Bin Lei
DOI: https://doi.org/10.3390/rs16030585
IF: 5
2024-02-05
Remote Sensing
Abstract:Monocular depth estimation using a single remote sensing image has emerged as a focal point in both remote sensing and computer vision research, proving crucial in tasks such as 3D reconstruction and target instance segmentation. Monocular depth estimation does not require multiple views as references, leading to significant improvements in both time and efficiency. Due to the complexity, occlusion, and uneven depth distribution of remote sensing images, there are currently few monocular depth estimation methods for remote sensing images. This paper proposes an approach to remote sensing monocular depth estimation that integrates an attention mechanism while considering global and local feature information. Leveraging a single remote sensing image as input, the method outputs end-to-end depth estimation for the corresponding area. In the encoder, the proposed method employs a dense neural network (DenseNet) feature extraction module with efficient channel attention (ECA), enhancing the capture of local information and details in remote sensing images. In the decoder stage, this paper proposes a dense atrous spatial pyramid pooling (DenseASPP) module with channel and spatial attention modules, effectively mitigating information loss and strengthening the relationship between the target's position and the background in the image. Additionally, weighted global guidance plane modules are introduced to fuse comprehensive features from different scales and receptive fields, finally predicting monocular depth for remote sensing images. Extensive experiments on the publicly available WHU-OMVS dataset demonstrate that our method yields better depth results in both qualitative and quantitative metrics.
environmental sciences,imaging science & photographic technology,remote sensing,geosciences, multidisciplinary
What problem does this paper attempt to address?
The paper proposes a solution to the problem of monocular depth estimation, especially for depth estimation in remote sensing images. Currently, there are few monocular depth estimation methods for remote sensing images due to their complexity, occlusions, and uneven depth distribution. The paper proposes an attention mechanism that combines global and local information, using a single remote sensing image as input for end-to-end depth estimation. In the encoder stage, the paper uses a DenseNet feature extraction module with Efficient Channel Attention (ECA), which enhances the capture of local information and details in remote sensing images. In the decoder stage, a DenseASPP module with channel and spatial attention is proposed, which reduces information loss and enhances the relationship between target positions and backgrounds in the image. Additionally, a weighted global guidance plane module is introduced to fuse comprehensive features from different scales and receptive fields, to finally predict the monocular depth of remote sensing images. Experiments are conducted on the WHU-OMVS public dataset, and the results show that this method achieves better depth estimation results in both qualitative and quantitative metrics. The objective of the paper is to overcome the challenges in depth estimation for remote sensing images through the improved depth estimation method, in order to improve the accuracy of understanding 3D structures.