DeepInteraction++: Multi-Modality Interaction for Autonomous Driving

Zeyu Yang,Nan Song,Wei Li,Xiatian Zhu,Li Zhang,Philip H.S. Torr

2024-08-15

Abstract:Existing top-performance autonomous driving systems typically rely on the multi-modal fusion strategy for reliable scene understanding. This design is however fundamentally restricted due to overlooking the modality-specific strengths and finally hampering the model performance. To address this limitation, in this work, we introduce a novel modality interaction strategy that allows individual per-modality representations to be learned and maintained throughout, enabling their unique characteristics to be exploited during the whole perception pipeline. To demonstrate the effectiveness of the proposed strategy, we design DeepInteraction++, a multi-modal interaction framework characterized by a multi-modal representational interaction encoder and a multi-modal predictive interaction decoder. Specifically, the encoder is implemented as a dual-stream Transformer with specialized attention operation for information exchange and integration between separate modality-specific representations. Our multi-modal representational learning incorporates both object-centric, precise sampling-based feature alignment and global dense information spreading, essential for the more challenging planning task. The decoder is designed to iteratively refine the predictions by alternately aggregating information from separate representations in a unified modality-agnostic manner, realizing multi-modal predictive interaction. Extensive experiments demonstrate the superior performance of the proposed framework on both 3D object detection and end-to-end autonomous driving tasks. Our code is available at <a class="link-external link-https" href="https://github.com/fudan-zvg/DeepInteraction" rel="external noopener nofollow">this https URL</a>.

Computer Vision and Pattern Recognition

What problem does this paper attempt to address?

The paper aims to address the limitations of multimodal fusion strategies in autonomous driving systems. Specifically: - **Current Issue**: Existing high-performance autonomous driving systems typically rely on multimodal fusion strategies to achieve reliable scene understanding. However, this design inherently overlooks the unique advantages of different modalities (such as LiDAR and cameras), thereby limiting model performance. - **Solution**: The paper proposes a new modality interaction strategy—DeepInteraction++—which allows each modality representation to be independently learned and maintained throughout the perception process, thereby fully leveraging its unique characteristics. - **Methodology**: By designing a multimodal interaction framework, including a multimodal representation interaction encoder and a multimodal prediction interaction decoder, effective information exchange and integration are achieved. The encoder adopts a dual-stream Transformer structure specifically for information exchange and integration; the decoder iteratively aggregates information from different modality representations in a unified manner, achieving multimodal prediction interaction. In summary, the main goal of the paper is to overcome the limitations of existing fusion methods by introducing a new multimodal interaction strategy, thereby improving the overall performance of autonomous driving systems.

DeepInteraction++: Multi-Modality Interaction for Autonomous Driving

DeepInteraction: 3D Object Detection via Modality Interaction

Learning Cross-Modality Interaction for Robust Depth Perception of Autonomous Driving

Multi-Modality Cascaded Fusion Technology for Autonomous Driving

Multi-Modal Sensor Fusion-Based Deep Neural Network for End-to-End Autonomous Driving With Scene Understanding

Multi-modal Integrated Prediction and Decision-making with Adaptive Interaction Modality Explorations

CASIN: Cascading Interaction Network for Robust Depth Sensing with an Auxiliary Task.

DeepDriving: Learning Affordance for Direct Perception in Autonomous Driving

Enhancing 3D object detection through multi-modal fusion for cooperative perception

Driver intention prediction based on multi-dimensional cross-modality information interaction

Multimodal Fusion Using Deep Learning Applied to Driver's Referencing of Outside-Vehicle Objects

A General Framework of Learning Multi-Vehicle Interaction Patterns from Video

Drive Anywhere: Generalizable End-to-end Autonomous Driving with Multi-modal Foundation Models

A General Framework of Learning Multi-Vehicle Interaction Patterns from Videos

Multimodal End-to-End Autonomous Driving

A Camera-Based End-to-End Autonomous Driving Framework Combined with Meta-Based Multi-task Optimization

Spatial-Temporal Multimodal End-to-End Autonomous Driving.

M2DA: Multi-Modal Fusion Transformer Incorporating Driver Attention for Autonomous Driving

Multi-Modal Data-Efficient 3D Scene Understanding for Autonomous Driving