Abstract:The topic of multimodal conversation systems has recently garnered significant attention across various industries, including travel, retail, and others. While pioneering works in this field have shown promising performance, they often focus solely on context information at the utterance level, overlooking the context-aware dependencies of multimodal semantic elements like words and images. Furthermore, the ordinal information of images, which indicates the relevance between visual context and users’ demands, remains underutilized during the integration of visual content. Additionally, the exploration of how to effectively utilize corresponding attributes provided by users when searching for desired products is still largely unexplored. To address these challenges, we propose a Position-aware Multimodal diAlogue system with semanTic Elements, abbreviated as PMATE. Specifically, to obtain semantic representations at the element-level, we first unfold the multimodal historical utterances and devise a position-aware multimodal element-level encoder. This component considers all images that may be relevant to the current turn and introduces a novel position-aware image selector to choose related images before fusing the information from the two modalities. Finally, we present a knowledge-aware two-stage decoder and an attribute-enhanced image searcher for the tasks of generating textual responses and selecting image responses, respectively. We extensively evaluate our model on two large-scale multimodal dialog datasets, and the results of our experiments demonstrate that our approach outperforms several baseline methods.

Improving Context Modelling in Multimodal Dialogue Generation

Simple Dialogue System with AUDITED

A Multi-Modal Context Reasoning Approach for Conditional Inference on Joint Textual and Visual Clues

A non-hierarchical attention network with modality dropout for textual response generation in multimodal dialogue systems

Towards Building Large Scale Multimodal Domain-Aware Conversation Systems

Dialogue Generation Model with Hierarchical Encoding and Semantic Segmentation of Dialogue Context

MMDialog: A Large-scale Multi-turn Dialogue Dataset Towards Multi-modal Open-domain Conversation.

Multimodal Dialogue Understanding via Holistic Modeling and Sequence Labeling.

Modality-Balanced Models for Visual Dialogue

Improving Contextual Language Models for Response Retrieval in Multi-Turn Conversation

Contextual Data Augmentation for Task-Oriented Dialog Systems

Self-adaptive Context and Modal-interaction Modeling For Multimodal Emotion Recognition

Multimodal Dialogue Systems via Capturing Context-aware Dependencies and Ordinal Information of Semantic Elements

BI-MDRG: Bridging Image History in Multimodal Dialogue Response Generation

Deep context modeling for multi-turn response selection in dialogue systems

Dual Semantic Knowledge Composed Multimodal Dialog Systems

Saliency infused dialogue response generation: Improving task oriented text generation using feature attribution

DialogMCF: Multimodal Context Flow for Audio Visual Scene-Aware Dialog

S3: A Simple Strong Sample-effective Multimodal Dialog System

Exploring Multi-Modal Representations for Ambiguity Detection & Coreference Resolution in the SIMMC 2.0 Challenge

Multiresolution Recurrent Neural Networks: an Application to Dialogue Response Generation