Abstract:Humans can collaborate and complete tasks based on visual signals and instruction from the environment. Training such a robot is difficult especially due to the understanding of the instruction and the complicated environment. Previous instruction-following agents are biased to English-centric corpus, making it unrealizable to be applied to users that use multiple languages or even low-resource languages. Nevertheless, the instruction-following agents are pre-trained in a mode that assumes the user can observe the environment, which limits its accessibility. In this work, we're trying to generalize the success of instruction-following agents to non-English languages with little corpus resources, and improve its intractability and accessibility. We introduce UVLN (Universal Vision-Language Navigation), a novel machine-translation instructional augmented framework for cross-lingual vision-language navigation, with a novel composition of state-of-the-art large language model (GPT3) with the image caption model (BLIP). We first collect a multilanguage vision-language navigation dataset via machine translation. Then we extend the standard VLN training objectives to a multilingual setting via a cross-lingual language encoder. The alignment between different languages is captured through a shared vision and action context via a cross-modal transformer, which encodes the inputs of language instruction, visual observation, and action decision sequences. To improve the intractability, we connect our agent with the large language model that informs the situation and current state to the user and also explains the action decisions. Experiments over Room Across Room Dataset prove the effectiveness of our approach. And the qualitative results show the promising intractability and accessibility of our instruction-following agent.

A New Path: Scaling Vision-and-Language Navigation with Synthetic Instructions and Imitation Learning

Scaling Data Generation in Vision-and-Language Navigation

Improving Vision-and-Language Navigation with Image-Text Pairs from the Web

Learning Vision-and-Language Navigation from YouTube Videos

Learning to Follow and Generate Instructions for Language-Capable Navigation

Vision-and-Language Navigation: Interpreting Visually-Grounded Navigation Instructions in Real Environments

Improving Vision-and-Language Navigation by Generating Future-View Image Semantics

LM-Nav: Robotic Navigation with Large Pre-Trained Models of Language, Vision, and Action

Accessible Instruction-Following Agent

Vision-Language Navigation Policy Learning and Adaptation

Counterfactual Cycle-Consistent Learning for Instruction Following and Generation in Vision-Language Navigation

Navigation Instruction Generation with BEV Perception and Large Language Models

Vision and Language Navigation in the Real World via Online Visual Language Mapping

VLN-Video: Utilizing Driving Videos for Outdoor Vision-and-Language Navigation

$A^2$Nav: Action-Aware Zero-Shot Robot Navigation by Exploiting Vision-and-Language Ability of Foundation Models

LangNav: Language as a Perceptual Representation for Navigation

NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation

Continual Vision-and-Language Navigation

Towards Learning a Generic Agent for Vision-and-Language Navigation via Pre-Training

AIGeN: An Adversarial Approach for Instruction Generation in VLN

SOAT: A Scene- and Object-Aware Transformer for Vision-and-Language Navigation