BSI-MVS: multi-view stereo network with bidirectional semantic information

Ruiming Jia,Jun Yu,Zhenghui Hu,Fei Yuan
DOI: https://doi.org/10.1038/s41598-024-55612-6
IF: 4.6
2024-03-23
Scientific Reports
Abstract:The basic principle of multi-view stereo (MVS) is to perform 3D reconstruction by extracting depth information from multiple views. Most current SOTA MVS networks are based on Vision Transformer, which usually means expensive computational complexity. To reduce computational complexity and improve depth map accuracy, we propose a MVS network with Bidirectional Semantic Information (BSI-MVS). Firstly, we design a Multi-Level Spatial Pyramid module to generate multiple layers of feature map for extracting multi-scale information. Then we propose a 2D Bidirectional-LSTM module to capture bidirectional semantic information at different time steps in the horizontal and vertical directions, which contains abundant depth information. Finally, cost volumes are built based on various levels of feature maps to optimize the final depth map. We experiment on the DTU and BlendedMVS datasets. The result shows that our network, in terms of overall metrics, surpasses TransMVSNet, CasMVSNet, CVP-MVSNet, and AACVP-MVSNet respectively by 17.84%, 36.42%, 14.96%, and 4.86%, which also shows a noticeable performance enhancement in objective metrics and visualizations.
multidisciplinary sciences
What problem does this paper attempt to address?
The paper attempts to address the issues of high computational complexity and insufficient depth map accuracy in multi-view stereo (MVS) reconstruction. Specifically: 1. **Computational Complexity**: Most state-of-the-art MVS networks are based on Vision Transformer. Although this model performs well in feature extraction, it has high computational complexity, leading to inefficiency when processing high-resolution images and slow convergence speed. 2. **Depth Map Accuracy**: Traditional MVS methods tend to encounter problems such as holes and texture mixing when dealing with complex geometric structures or textureless regions, affecting the reconstruction quality. To solve these problems, the authors propose a new MVS network—BSI-MVS (Bidirectional Semantic Information Multi-View Stereo), with the main innovations including: - **Multi-Scale Spatial Pyramid Module (MLSP)**: Generates multiple levels of feature maps, extracts multi-scale information, and enhances the network's adaptability to different spatial structures. - **Bidirectional LSTM Module (BiLSTM)**: Captures bidirectional semantic information in both horizontal and vertical directions, contains rich depth information, and improves the model's generalization ability and depth map accuracy. - **Cost Volume Construction**: Constructs cost volumes based on feature maps of different levels to optimize the final depth map. With these improvements, BSI-MVS achieves significant enhancements in network performance and depth map accuracy. Experimental results show that BSI-MVS outperforms TransMVSNet, CasMVSNet, CVP-MVSNet, and AACVP-MVSNet on the DTU and BlendedMVS datasets, with overall metrics improved by 17.84%, 36.42%, 14.96%, and 4.86%, respectively.