Benchmarking Visual-Inertial Deep Multimodal Fusion for Relative Pose Regression and Odometry-aided Absolute Pose Regression

Felix Ott,Nisha Lakshmana Raichur,David Rügamer,Tobias Feigl,Heiko Neumann,Bernd Bischl,Christopher Mutschler

2023-08-04

Abstract:Visual-inertial localization is a key problem in computer vision and robotics applications such as virtual reality, self-driving cars, and aerial vehicles. The goal is to estimate an accurate pose of an object when either the environment or the dynamics are known. Absolute pose regression (APR) techniques directly regress the absolute pose from an image input in a known scene using convolutional and spatio-temporal networks. Odometry methods perform relative pose regression (RPR) that predicts the relative pose from a known object dynamic (visual or inertial inputs). The localization task can be improved by retrieving information from both data sources for a cross-modal setup, which is a challenging problem due to contradictory tasks. In this work, we conduct a benchmark to evaluate deep multimodal fusion based on pose graph optimization and attention networks. Auxiliary and Bayesian learning are utilized for the APR task. We show accuracy improvements for the APR-RPR task and for the RPR-RPR task for aerial vehicles and hand-held devices. We conduct experiments on the EuRoC MAV and PennCOSYVIO datasets and record and evaluate a novel industry dataset.

Computer Vision and Pattern Recognition

What problem does this paper attempt to address?

The paper primarily aims to address key issues in Visual-Inertial Localization, particularly by combining image data and Inertial Measurement Unit (IMU) data to improve the accuracy of object localization. Specifically, the research goal is to estimate the precise position and orientation of an object in a known environment or under dynamic conditions. The paper explores two main tasks: 1. **Absolute Pose Regression (APR)**: Directly regressing the absolute pose from a single image or a set of training images. 2. **Relative Pose Regression (RPR)**: Predicting the relative pose between two known object dynamics (visual or inertial inputs). To improve the performance of these tasks, the research employs multimodal fusion techniques, combining visual information with inertial information. The authors propose a benchmarking framework to evaluate the performance of deep learning-based multimodal fusion methods in visual-inertial localization tasks, including the fusion of absolute pose regression and relative pose regression tasks. Additionally, the paper explores the application of auxiliary learning, Bayesian neural networks, and other techniques to further enhance localization accuracy and quantify model uncertainty. In summary, the problem the paper attempts to solve can be summarized as: How to effectively utilize the complementary characteristics of visual and inertial data to improve the accuracy of localization tasks (especially absolute pose regression and relative pose regression) through multimodal fusion.

Benchmarking Visual-Inertial Deep Multimodal Fusion for Relative Pose Regression and Odometry-aided Absolute Pose Regression

MoreFusion: Multi-object Reasoning for 6D Pose Estimation from Volumetric Fusion

BB-Align: A Lightweight Pose Recovery Framework for Vehicle-to-Vehicle Cooperative Perception

Recurrent Volume-Based 3-D Feature Fusion for Real-Time Multiview Object Pose Estimation.

Recurrent Volume-based 3D Feature Fusion for Real-time Multi-view Object Pose Estimation

Fusing Structure from Motion and Simulation-Augmented Pose Regression from Optical Flow for Challenging Indoor Environments

PRGFlow: Benchmarking SWAP-Aware Unified Deep Visual Inertial Odometry

ViPR: Visual-Odometry-aided Pose Regression for 6DoF Camera Localization

PA-Pose: Partial Point Cloud Fusion Based on Reliable Alignment for 6D Pose Tracking

Beyond Learning: Back to Geometric Essence of Visual Odometry via Fusion-Based Paradigm

Learned Monocular Depth Priors in Visual-Inertial Initialization

Multi-Camera Sensor Fusion for Visual Odometry using Deep Uncertainty Estimation

Pose Estimation Based on Bidirectional Visual–Inertial Odometry with 3D LiDAR (BV-LIO)

Towards Interpretable Camera and LiDAR Data Fusion for Autonomous Ground Vehicles Localisation

Brain-Inspired Visual Odometry: Balancing Speed and Interpretability through a System of Systems Approach

Multi-Sensor Fusion Self-Supervised Deep Odometry and Depth Estimation

Fusing Monocular Images and Sparse IMU Signals for Real-time Human Motion Capture

Efficient 2D-3D Matching for Multi-Camera Visual Localization

Collaborative Learning of Depth Estimation, Visual Odometry and Camera Relocalization from Monocular Videos.

Deep Visual Odometry with Adaptive Memory