Abstract:Existing top-performance autonomous driving systems typically rely on the multi-modal fusion strategy for reliable scene understanding. This design is however fundamentally restricted due to overlooking the modality-specific strengths and finally hampering the model performance. To address this limitation, in this work, we introduce a novel modality interaction strategy that allows individual per-modality representations to be learned and maintained throughout, enabling their unique characteristics to be exploited during the whole perception pipeline. To demonstrate the effectiveness of the proposed strategy, we design DeepInteraction++, a multi-modal interaction framework characterized by a multi-modal representational interaction encoder and a multi-modal predictive interaction decoder. Specifically, the encoder is implemented as a dual-stream Transformer with specialized attention operation for information exchange and integration between separate modality-specific representations. Our multi-modal representational learning incorporates both object-centric, precise sampling-based feature alignment and global dense information spreading, essential for the more challenging planning task. The decoder is designed to iteratively refine the predictions by alternately aggregating information from separate representations in a unified modality-agnostic manner, realizing multi-modal predictive interaction. Extensive experiments demonstrate the superior performance of the proposed framework on both 3D object detection and end-to-end autonomous driving tasks. Our code is available at <a class="link-external link-https" href="https://github.com/fudan-zvg/DeepInteraction" rel="external noopener nofollow">this https URL</a>.

SDIF-DA: A Shallow-to-Deep Interaction Framework with Data Augmentation for Multi-modal Intent Detection

Sentiment Analysis Using Deep Robust Complementary Fusion of Multi-Features and Multi-Modalities.

A Dynamic Graph Interactive Framework with Label-Semantic Injection for Spoken Language Understanding

AIMDiT: Modality Augmentation and Interaction via Multimodal Dimension Transformation for Emotion Recognition in Conversations

Dual-oriented Disentangled Network with Counterfactual Intervention for Multimodal Intent Detection

Multi-Task Deep Learning for User Intention Understanding in Speech Interaction Systems

DeepInteraction++: Multi-Modality Interaction for Autonomous Driving

Learning Towards Selective Data Augmentation for Dialogue Generation.

DAT: Dialogue-Aware Transformer with Modality-Group Fusion for Human Engagement Estimation

DeepInteraction: 3D Object Detection via Modality Interaction

DeepSpeed-VisualChat: Multi-Round Multi-Image Interleave Chat via Multi-Modal Causal Attention

Data Augmentation Integrating Dialogue Flow and Style to Adapt Spoken Dialogue Systems to Low-Resource User Groups

Exploiting Modality-Specific Features For Multi-Modal Manipulation Detection And Grounding

Deep Multimodal Data Fusion

BlendX: Complex Multi-Intent Detection with Blended Patterns

TOD-DA: Towards Boosting the Robustness of Task-oriented Dialogue Modeling on Spoken Conversations

Continual Generalized Intent Discovery: Marching Towards Dynamic and Open-world Intent Recognition

Enhancing multi-modal fusion in visual dialog via sample debiasing and feature interaction

UniTranSeR: A Unified Transformer Semantic Representation Framework for Multimodal Task-Oriented Dialog System

InMu-Net: Advancing Multi-modal Intent Detection Via Information Bottleneck and Multi-sensory Processing

Promoting Unified Generative Framework with Descriptive Prompts for Joint Multi-Intent Detection and Slot Filling