Abstract:GuessWhich is an engaging visual dialogue game that involves interaction between a Questioner Bot (QBot) and an Answer Bot (ABot) in the context of image-guessing. In this game, QBot's objective is to locate a concealed image solely through a series of visually related questions posed to ABot. However, effectively modeling visually related reasoning in QBot's decision-making process poses a significant challenge. Current approaches either lack visual information or rely on a single real image sampled at each round as decoding context, both of which are inadequate for visual reasoning. To address this limitation, we propose a novel approach that focuses on visually related reasoning through the use of a mental model of the undisclosed image. Within this framework, QBot learns to represent mental imagery, enabling robust visual reasoning by tracking the dialogue state. The dialogue state comprises a collection of representations of mental imagery, as well as representations of the entities involved in the conversation. At each round, QBot engages in visually related reasoning using the dialogue state to construct an internal representation, generate relevant questions, and update both the dialogue state and internal representation upon receiving an answer. Our experimental results on the VisDial datasets (v0.5, 0.9, and 1.0) demonstrate the effectiveness of our proposed model, as it achieves new state-of-the-art performance across all metrics and datasets, surpassing previous state-of-the-art models. Codes and datasets from our experiments are freely available at \href{<a class="link-external link-https" href="https://github.com/xubuvd/GuessWhich" rel="external noopener nofollow">this https URL</a>}.

Generative Visual Dialogue System Via Weighted Likelihood Estimation

Hagan: Hierarchical Attentive Adversarial Learning For Task-Oriented Dialogue System

OpenViDial 2.0: A Larger-Scale, Open-Domain Dialogue Generation Dataset with Visual Contexts

Enhancing Visual Dialog State Tracking through Iterative Object-Entity Alignment in Multi-Round Conversations

Are You Talking to Me? Reasoned Visual Dialog Generation Through Adversarial Learning

DAM: Deliberation, Abandon and Memory Networks for Generating Detailed and Non-repetitive Responses in Visual Dialogue

The Dialog Must Go On: Improving Visual Dialog via Generative Self-Training

Multi-Modal Dialogue State Tracking for Playing GuessWhich Game

Engaging Live Video Comments Generation

DLVGen: A Dual Latent Variable Approach to Personalized Dialogue Generation

Dirichlet Latent Variable Hierarchical Recurrent Encoder-Decoder in Dialogue Generation.

DAL: Dual Adversarial Learning for Dialogue Generation.

Open Domain Dialogue Generation with Latent Images

Recurrent Attention Network with Reinforced Generator for Visual Dialog

Adversarial Learning for Neural Dialogue Generation.

Don't Say That! Making Inconsistent Dialogue Unlikely with Unlikelihood Training

DialogWAE: Multimodal Response Generation with Conditional Wasserstein Auto-Encoder

Multimodal Incremental Transformer with Visual Grounding for Visual Dialogue Generation

Transformer-Based Conditioned Variational Autoencoder for Dialogue Generation

Reasoning Visual Dialogs With Structural And Partial Observations

HVLM: Exploring Human-Like Visual Cognition and Language-Memory Network for Visual Dialog