Abstract:In recent years, convolutional-neural-network based stereo matching methods have achieved significant gains compared to conventional methods in terms of both speed and accuracy. Current state-of-the-art disparity estimation algorithms require many parameters and large amounts of computational resources and are not suited for applications on edge devices. In this paper, we propose an end-to-end light-weight network (LWNet) for fast stereo matching, which consists of an efficient backbone with multi-scale feature fusion for feature extraction, a 3D U-Net aggregation architecture for disparity computation, and color guidance in a 2D convolutional neural network (CNN) for disparity refinement. We adopt MobileNetV2 as an efficient backbone in feature extraction. The channel attention module is applied to improve the representational capacity of features and multi-resolution information is adaptively incorporated into the cost volume via cross-scale connections. In addition, instead of using regular 3D convolutions, we utilize pseudo 3D convolutions in the 3D U-Net architecture to aggregate the cost volume for a better balance between computational cost and accuracy. Further, we introduce a left-right consistency check and color guidance and design a robust disparity refinement network with skip connections and dilated convolutions to capture global context information and further improve disparity-estimation accuracy with little computational cost and memory space. A depth-wise separable convolution is proposed to replace all the standard convolutions in the section of disparity refinement, which can decrease computational complexity and the number of parameters without significant accuracy reduction. Extensive experiments on Scene Flow, KITTI 2015, and KITTI 2012 benchmarks demonstrate that the proposed LWNet achieves competitive accuracy when compared with state-of-the-art stereo matching methods.

Stereo Matching with Local Cost Volume Refinement Network

Disparity Estimation Using Multilevel and Global Information

Stereo Matching Using Multi-Level Cost Volume and Multi-Scale Feature Constancy

Local Similarity Pattern and Cost Self-Reassembling for Deep Stereo Matching Networks

A Light-Weight Stereo Matching Network Based on Multi-Scale Features Fusion and Robust Disparity Refinement

An efficient and accurate multi-level cascaded recurrent network for stereo matching

A Light-Weight Network with Multi-Scale Features Fusion and Color Guidance for Stereo Matching

Edge supervision and multi-scale cost volume for stereo matching

A Fast Stereo Matching Network with Multi-Cross Attention

Deep Stereo Matching With Hysteresis Attention and Supervised Cost Volume Construction

Improved real-time three-dimensional stereo matching with local consistency

Multi-Scale Cost Volumes Cascade Network for Stereo Matching

Better Stereo Matching from Simple Yet Effective Wrangling of Deep Features

Multi-Dimensional Cooperative Network for Stereo Matching

Stereo Matching Method for Remote Sensing Images Based on Attention and Scale Fusion

Re-Parameterized Real-Time Stereo Matching Network Based on Mixed Cost Volumes Toward Autonomous Driving

Stacking Learning with Coalesced Cost Filtering for Accurate Stereo Matching

Adaptive Cost Volume Representation for Unsupervised High-resolution Stereo Matching

Multi-scale Cross-form Pyramid Network for Stereo Matching

Robust Cost Volume Generation Method for Dense Stereo Matching in Endoscopic Scenarios

RUANet: Residual Ultra-Aggregation Network for Stereo Matching