Abstract:The emerging vision-and-language navigation (VLN) problem aims at learning to navigate an agent to the target location in unseen photo-realistic environments according to the given language instruction. The main challenges of VLN arise mainly from two aspects: first, the agent needs to attend to the meaningful paragraphs of the language instruction corresponding to the dynamically-varying visual environments; second, during the training process, the agent usually imitate the expert demonstrations, i.e., the shortest-path to the target location specified by associated language instructions. Due to the discrepancy of action selection between training and inference, the agent solely on the basis of imitation learning does not perform well. Existing VLN approaches address this issue by sampling the next action from its predicted probability distribution during the training process. This allows the agent to explore diverse routes from the environments, yielding higher success rates. Nevertheless, without being presented with the golden shortest navigation paths during the training process, the agent may arrive at the target location through an unexpected longer route. To overcome these challenges, we design a cross-modal grounding module, which is composed of two complementary attention mechanisms, to equip the agent with a better ability to track the correspondence between the textual and visual modalities. We then propose to recursively alternate the learning schemes of imitation and exploration to narrow the discrepancy between training and inference. We further exploit the advantages of both these two learning schemes via adversarial learning. Extensive experimental results on the Room-to-Room (R2R) benchmark dataset demonstrate that the proposed learning scheme is generalized and complementary to prior arts. Our method performs well against state-of-the-art approaches in terms of effectiveness and efficiency.

Are We There Yet? Learning to Localize in Embodied Instruction Following

Scene-Intuitive Agent for Remote Embodied Visual Grounding

Active Visual Localization in Partially Calibrated Environments.

LACMA: Language-Aligning Contrastive Learning with Meta-Actions for Embodied Instruction Following

Embodied Instruction Following in Unknown Environments

Language-guided Navigation Via Cross-Modal Grounding and Alternate Adversarial Learning

Accessible Instruction-Following Agent

Vision and Language Navigation in the Real World via Online Visual Language Mapping

Where are you? localization from embodied dialog

Alfred: A benchmark for interpreting grounded instructions for everyday tasks

Learning to Act with Affordance-Aware Multimodal Neural SLAM

VLN-Trans: Translator for the Vision and Language Navigation Agent

Deep Active Localization

Embodied Concept Learner: Self-supervised Learning of Concepts and Mapping through Instruction Following

Embodied BERT: A Transformer Model for Embodied, Language-guided Visual Task Completion

Think Holistically, Act Down-to-Earth: A Semantic Navigation Strategy with Continuous Environmental Representation and Multi-step Forward Planning

A Modular Framework for Robot Embodied Instruction Following by Large Language Model

Embodied Learning for Lifelong Visual Perception

Loc4Plan: Locating Before Planning for Outdoor Vision and Language Navigation

Language-guided Robust Navigation for Mobile Robots in Dynamically-changing Environments

Towards Navigation by Reasoning over Spatial Configurations