Abstract:The topic of multimodal conversation systems has recently garnered significant attention across various industries, including travel, retail, and others. While pioneering works in this field have shown promising performance, they often focus solely on context information at the utterance level, overlooking the context-aware dependencies of multimodal semantic elements like words and images. Furthermore, the ordinal information of images, which indicates the relevance between visual context and users’ demands, remains underutilized during the integration of visual content. Additionally, the exploration of how to effectively utilize corresponding attributes provided by users when searching for desired products is still largely unexplored. To address these challenges, we propose a Position-aware Multimodal diAlogue system with semanTic Elements, abbreviated as PMATE. Specifically, to obtain semantic representations at the element-level, we first unfold the multimodal historical utterances and devise a position-aware multimodal element-level encoder. This component considers all images that may be relevant to the current turn and introduces a novel position-aware image selector to choose related images before fusing the information from the two modalities. Finally, we present a knowledge-aware two-stage decoder and an attribute-enhanced image searcher for the tasks of generating textual responses and selecting image responses, respectively. We extensively evaluate our model on two large-scale multimodal dialog datasets, and the results of our experiments demonstrate that our approach outperforms several baseline methods.

Some Can Be Better Than All: Multimodal Star Transformer for Visual Dialog

Multimodal Incremental Transformer with Visual Grounding for Visual Dialogue Generation

SKANet - Structured Knowledge-Aware Network for Visual Dialog.

Hierarchical Vision and Language Transformer for Efficient Visual Dialog

Multimodal Dialogue Generation Based on Transformer and Collaborative Attention

Efficient Attention Mechanism for Visual Dialog that can Handle All the Interactions between Multiple Inputs

Multi-View Attention Network for Visual Dialog

Bridging Text and Video: A Universal Multimodal Transformer for Audio-Visual Scene-Aware Dialog

Modality-Balanced Models for Visual Dialogue

UniTranSeR: A Unified Transformer Semantic Representation Framework for Multimodal Task-Oriented Dialog System

Structure-Aware Multimodal Sequential Learning for Visual Dialog

HVLM: Exploring Human-Like Visual Cognition and Language-Memory Network for Visual Dialog

Multimodal Dialogue Systems via Capturing Context-aware Dependencies and Ordinal Information of Semantic Elements

Vman: visual-modified attention network for multimodal paradigms

Modeling Text-visual Mutual Dependency for Multi-modal Dialog Generation

TransModality: An End2End Fusion Method with Transformer for Multimodal Sentiment Analysis

Multimodal Transformer For Multimodal Machine Translation

Video Dialog Via Progressive Inference and Cross-Transformer.

A Visual Tour Of Current Challenges In Multimodal Language Models

VSET: A MULTIMODAL TRANSFORMER FOR VISUAL SPEECH ENHANCEMENT

Audio Visual Scene-Aware Dialog Generation with Transformer-based Video Representations