Abstract:The emerging vision-and-language navigation (VLN) problem aims at learning to navigate an agent to the target location in unseen photo-realistic environments according to the given language instruction. The main challenges of VLN arise mainly from two aspects: first, the agent needs to attend to the meaningful paragraphs of the language instruction corresponding to the dynamically-varying visual environments; second, during the training process, the agent usually imitate the expert demonstrations, i.e., the shortest-path to the target location specified by associated language instructions. Due to the discrepancy of action selection between training and inference, the agent solely on the basis of imitation learning does not perform well. Existing VLN approaches address this issue by sampling the next action from its predicted probability distribution during the training process. This allows the agent to explore diverse routes from the environments, yielding higher success rates. Nevertheless, without being presented with the golden shortest navigation paths during the training process, the agent may arrive at the target location through an unexpected longer route. To overcome these challenges, we design a cross-modal grounding module, which is composed of two complementary attention mechanisms, to equip the agent with a better ability to track the correspondence between the textual and visual modalities. We then propose to recursively alternate the learning schemes of imitation and exploration to narrow the discrepancy between training and inference. We further exploit the advantages of both these two learning schemes via adversarial learning. Extensive experimental results on the Room-to-Room (R2R) benchmark dataset demonstrate that the proposed learning scheme is generalized and complementary to prior arts. Our method performs well against state-of-the-art approaches in terms of effectiveness and efficiency.

Scaling Vision-and-Language Navigation With Offline RL

LangNav: Language as a Perceptual Representation for Navigation

A New Path: Scaling Vision-and-Language Navigation with Synthetic Instructions and Imitation Learning

Continual Vision-and-Language Navigation

Fine-Grained Alignment in Vision-and-Language Navigation through Bayesian Optimization

NavCoT: Boosting LLM-Based Vision-and-Language Navigation via Learning Disentangled Reasoning

Reinforced Cross-Modal Matching and Self-Supervised Imitation Learning for Vision-Language Navigation

Reinforced Vision-and-Language Navigation Based on Historical BERT

Language-guided Navigation Via Cross-Modal Grounding and Alternate Adversarial Learning

Navigating Beyond Instructions: Vision-and-Language Navigation in Obstructed Environments

Vision-Language Navigation Policy Learning and Adaptation

Offline visual representation learning for embodied navigation

Vision-Language Navigation With Self-Supervised Auxiliary Reasoning Tasks

Boosting Efficient Reinforcement Learning for Vision-and-Language Navigation with Open-Sourced LLM

Explore the Potential Performance of Vision-and-Language Navigation Model: a Snapshot Ensemble Method

VLN-Video: Utilizing Driving Videos for Outdoor Vision-and-Language Navigation

Improving Vision-and-Language Navigation with Image-Text Pairs from the Web

Reinforced Structured State-Evolution for Vision-Language Navigation

LM-Nav: Robotic Navigation with Large Pre-Trained Models of Language, Vision, and Action

Mind the Gap: Improving Success Rate of Vision-and-Language Navigation by Revisiting Oracle Success Routes

Language-Conditioned Offline RL for Multi-Robot Navigation