Abstract:Monocular visual odometry (VO) has attracted extensive research attention by providing real-time vehicle motion from cost-effective camera images. However, state-of-the-art optimization-based monocular VO methods suffer from the scale inconsistency problem for long-term predictions. Deep learning has recently been introduced to address this issue by leveraging stereo sequences or ground-truth motions in the training dataset. However, it comes at an additional cost for data collection, and such training data may not be available in all datasets. In this work, we propose VRVO, a novel framework for retrieving the absolute scale from virtual data that can be easily obtained from modern simulation environments, whereas in the real domain no stereo or ground-truth data are required in either the training or inference phases. Specifically, we first train a scale-aware disparity network using both monocular real images and stereo virtual data. The virtual-to-real domain gap is bridged by using an adversarial training strategy to map images from both domains into a shared feature space. The resulting scale-consistent disparities are then integrated with a direct VO system by constructing a virtual stereo objective that ensures the scale consistency over long trajectories. Additionally, to address the suboptimality issue caused by the separate optimization backend and the learning process, we further propose a mutual reinforcement pipeline that allows bidirectional information flow between learning and optimization, which boosts the robustness and accuracy of each other. We demonstrate the effectiveness of our framework on the KITTI and vKITTI2 datasets.

MD2VO: Enhancing Monocular Visual Odometry Through Minimum Depth Difference

Improving Monocular Visual Odometry Using Learned Depth

CodeVIO: Visual-Inertial Odometry with Learned Optimizable Dense Depth

Design of an Enhanced Visual Odometry by Building and Matching Compressive Panoramic Landmarks Online

Self-supervised Visual-LiDAR Odometry with Flip Consistency

DF-VO: What Should Be Learnt for Visual Odometry?

BEV-ODOM: Reducing Scale Drift in Monocular Visual Odometry with BEV Representation

Crafting Monocular Cues and Velocity Guidance for Self-Supervised Multi-Frame Depth Learning

PVO: Panoptic Visual Odometry.

DeepAVO: Efficient Pose Refining with Feature Distilling for Deep Visual Odometry

A self-supervised monocular odometry with visual-inertial and depth representations

A Monocular Visual Odometry Combining Edge Enhance with Deep Learning

Unsupervised Monocular Visual-Inertial Odometry Network

Towards Scale Consistent Monocular Visual Odometry by Learning from the Virtual World

Self-supervised deep monocular visual odometry and depth estimation with observation variation

Approaches, Challenges, and Applications for Deep Visual Odometry: Toward Complicated and Emerging Areas

MS-VRO: A Multi-Stage Visual-Millimeter Wave Radar Fusion Odometry

Approaches, Challenges, and Applications for Deep Visual Odometry: Toward to Complicated and Emerging Areas

Lidar-Monocular Visual Odometry Using Point and Line Features.

DeepVO: A Deep Learning approach for Monocular Visual Odometry

Enhancing Visual Odometry with Estimated Scene Depth: Leveraging RGB-D Data with Deep Learning