Abstract:In recent years, convolutional-neural-network based stereo matching methods have achieved significant gains compared to conventional methods in terms of both speed and accuracy. Current state-of-the-art disparity estimation algorithms require many parameters and large amounts of computational resources and are not suited for applications on edge devices. In this paper, we propose an end-to-end light-weight network (LWNet) for fast stereo matching, which consists of an efficient backbone with multi-scale feature fusion for feature extraction, a 3D U-Net aggregation architecture for disparity computation, and color guidance in a 2D convolutional neural network (CNN) for disparity refinement. We adopt MobileNetV2 as an efficient backbone in feature extraction. The channel attention module is applied to improve the representational capacity of features and multi-resolution information is adaptively incorporated into the cost volume via cross-scale connections. In addition, instead of using regular 3D convolutions, we utilize pseudo 3D convolutions in the 3D U-Net architecture to aggregate the cost volume for a better balance between computational cost and accuracy. Further, we introduce a left-right consistency check and color guidance and design a robust disparity refinement network with skip connections and dilated convolutions to capture global context information and further improve disparity-estimation accuracy with little computational cost and memory space. A depth-wise separable convolution is proposed to replace all the standard convolutions in the section of disparity refinement, which can decrease computational complexity and the number of parameters without significant accuracy reduction. Extensive experiments on Scene Flow, KITTI 2015, and KITTI 2012 benchmarks demonstrate that the proposed LWNet achieves competitive accuracy when compared with state-of-the-art stereo matching methods.

A Light-Weight Network with Multi-Scale Features Fusion and Color Guidance for Stereo Matching

A Light-Weight Stereo Matching Network Based on Multi-Scale Features Fusion and Robust Disparity Refinement

Disparity Estimation Using Multilevel and Global Information

End-to-End Learning of Multi-scale Convolutional Neural Network for Stereo Matching

Lightweight multi-scale convolutional neural network for real time stereo matching

A Fast Stereo Matching Network with Multi-Cross Attention

Edge supervision and multi-scale cost volume for stereo matching

Stereo Matching Using Multi-Level Cost Volume and Multi-Scale Feature Constancy

A Joint 2D-3D Complementary Network for Stereo Matching

Deeply-fused Attentive Network for Stereo Matching

DMCNet: Towards Lightweight Volumetric Stereo Matching

Improved real-time three-dimensional stereo matching with local consistency

Brain Cholesterol XVIII: Effect of Methylphenidate (Ritalin) on [U-14C] Glucose and [2-3H] Acetate Incorporation

Stereo Matching with Local Cost Volume Refinement Network

CGFNet: 3D Convolution Guided and Multi-scale Volume Fusion Network for fast and robust stereo matching

MSDC-Net: Multi-Scale Dense and Contextual Networks for Automated Disparity Map for Stereo Matching

Multi-scale Cross-form Pyramid Network for Stereo Matching

Multi-Scale Context Attention Network for Stereo Matching

Fast Multi-Scale Residual Fusion Network for Stereo Matching.

Superpixel Guided Network for Three-Dimensional Stereo Matching

Local Similarity Pattern and Cost Self-Reassembling for Deep Stereo Matching Networks