Abstract:In recent years, convolutional-neural-network based stereo matching methods have achieved significant gains compared to conventional methods in terms of both speed and accuracy. Current state-of-the-art disparity estimation algorithms require many parameters and large amounts of computational resources and are not suited for applications on edge devices. In this paper, we propose an end-to-end light-weight network (LWNet) for fast stereo matching, which consists of an efficient backbone with multi-scale feature fusion for feature extraction, a 3D U-Net aggregation architecture for disparity computation, and color guidance in a 2D convolutional neural network (CNN) for disparity refinement. We adopt MobileNetV2 as an efficient backbone in feature extraction. The channel attention module is applied to improve the representational capacity of features and multi-resolution information is adaptively incorporated into the cost volume via cross-scale connections. In addition, instead of using regular 3D convolutions, we utilize pseudo 3D convolutions in the 3D U-Net architecture to aggregate the cost volume for a better balance between computational cost and accuracy. Further, we introduce a left-right consistency check and color guidance and design a robust disparity refinement network with skip connections and dilated convolutions to capture global context information and further improve disparity-estimation accuracy with little computational cost and memory space. A depth-wise separable convolution is proposed to replace all the standard convolutions in the section of disparity refinement, which can decrease computational complexity and the number of parameters without significant accuracy reduction. Extensive experiments on Scene Flow, KITTI 2015, and KITTI 2012 benchmarks demonstrate that the proposed LWNet achieves competitive accuracy when compared with state-of-the-art stereo matching methods.

DMCNet: Towards Lightweight Volumetric Stereo Matching

Disparity Estimation Using Multilevel and Global Information

Lightweight multi-scale convolutional neural network for real time stereo matching

A Light-Weight Network with Multi-Scale Features Fusion and Color Guidance for Stereo Matching

A Light-Weight Stereo Matching Network Based on Multi-Scale Features Fusion and Robust Disparity Refinement

MSDC-Net: Multi-Scale Dense and Contextual Networks for Automated Disparity Map for Stereo Matching

End-to-End Learning of Multi-scale Convolutional Neural Network for Stereo Matching

Ghost-Stereo: GhostNet-based Cost Volume Enhancement and Aggregation for Stereo Matching Networks

Re-Parameterized Real-Time Stereo Matching Network Based on Mixed Cost Volumes Toward Autonomous Driving

Multi-Dimensional Cooperative Network for Stereo Matching

Multi-Scale Cost Volumes Cascade Network for Stereo Matching

Stereo Matching Using Multi-Level Cost Volume and Multi-Scale Feature Constancy

OD-MVSNet: Omni-dimensional dynamic multi-view stereo network

Exploiting Semantic and Boundary Information for Stereo Matching

DSC-MVSNet: attention aware cost volume regularization based on depthwise separable convolution for multi-view stereo

DCVSMNet: Double Cost Volume Stereo Matching Network

HPA-Net: Hierarchical and Parallel Aggregation Network for Context Learning in Stereo Matching

Multi-scale Cross-form Pyramid Network for Stereo Matching

Stereo Matching with Local Cost Volume Refinement Network

Group-Based Atrous Convolution Stereo Matching Network

Cascaded multi-scale and multi-dimension convolutional neural network for stereo matching.