Abstract:In recent years, convolutional-neural-network based stereo matching methods have achieved significant gains compared to conventional methods in terms of both speed and accuracy. Current state-of-the-art disparity estimation algorithms require many parameters and large amounts of computational resources and are not suited for applications on edge devices. In this paper, we propose an end-to-end light-weight network (LWNet) for fast stereo matching, which consists of an efficient backbone with multi-scale feature fusion for feature extraction, a 3D U-Net aggregation architecture for disparity computation, and color guidance in a 2D convolutional neural network (CNN) for disparity refinement. We adopt MobileNetV2 as an efficient backbone in feature extraction. The channel attention module is applied to improve the representational capacity of features and multi-resolution information is adaptively incorporated into the cost volume via cross-scale connections. In addition, instead of using regular 3D convolutions, we utilize pseudo 3D convolutions in the 3D U-Net architecture to aggregate the cost volume for a better balance between computational cost and accuracy. Further, we introduce a left-right consistency check and color guidance and design a robust disparity refinement network with skip connections and dilated convolutions to capture global context information and further improve disparity-estimation accuracy with little computational cost and memory space. A depth-wise separable convolution is proposed to replace all the standard convolutions in the section of disparity refinement, which can decrease computational complexity and the number of parameters without significant accuracy reduction. Extensive experiments on Scene Flow, KITTI 2015, and KITTI 2012 benchmarks demonstrate that the proposed LWNet achieves competitive accuracy when compared with state-of-the-art stereo matching methods.

CGFNet: 3D Convolution Guided and Multi-scale Volume Fusion Network for fast and robust stereo matching

CGI-Stereo: Accurate and Real-Time Stereo Matching via Context and Geometry Interaction

Fast Multi-Scale Residual Fusion Network for Stereo Matching.

A Light-Weight Network with Multi-Scale Features Fusion and Color Guidance for Stereo Matching

Multi-Scale Cost Volumes Cascade Network for Stereo Matching

UGNet: Uncertainty aware geometry enhanced networks for stereo matching

A Light-Weight Stereo Matching Network Based on Multi-Scale Features Fusion and Robust Disparity Refinement

End-to-End Learning of Multi-scale Convolutional Neural Network for Stereo Matching

Bidirectional Stereo Matching Network With Double Cost Volumes

A Joint 2D-3D Complementary Network for Stereo Matching

A Fast Stereo Matching Network with Multi-Cross Attention

Stereo Matching Using Multi-Level Cost Volume and Multi-Scale Feature Constancy

Multi-Dimensional Cooperative Network for Stereo Matching

Ghost-Stereo: GhostNet-based Cost Volume Enhancement and Aggregation for Stereo Matching Networks

Accuracy and efficiency stereo matching network with adaptive feature modulation

GPDF-Net: geometric prior-guided stereo matching with disparity fusion refinement

MULTI-SCALE CASCADE DISPARITY REFINEMENT STEREO NETWORK

Guided aggregation and disparity refinement for real-time stereo matching

Re-Parameterized Real-Time Stereo Matching Network Based on Mixed Cost Volumes Toward Autonomous Driving

Multi-scale Cross-form Pyramid Network for Stereo Matching

An efficient and accurate multi-level cascaded recurrent network for stereo matching