Abstract:Object Goal Navigation(ObjectNav) is the task that an agent need navigate to an instance of a specific category in an unseen environment through visual observations within limited time steps. This work plays a significant role in enhancing the efficiency of locating specific items in indoor spaces and assisting individuals in completing various tasks, as well as providing support for people with disabilities. To achieve efficient ObjectNav in unfamiliar environments, global perception capabilities, understanding the regularities of space and semantics in the environment layout are significant. In this work, we propose an explicit-prediction method called VLAI that utilizes visual-language alignment information to guide the agent's exploration, unlike previous navigation methods based on frontier potential prediction or egocentric map completion, which only leverage visual observations to construct semantic maps, thus failing to help the agent develop a better global perception. Specifically, when predicting long-term goals, we retrieve previously saved visual observations to obtain visual information around the frontiers based on their position on the incrementally built incomplete semantic map. Then, we apply our designed Chat Describer to this visual information to obtain detailed frontier object descriptions. The Chat Describer, a novel automatic-questioning approach deployed in Visual-to-Language, is composed of Large Language Model(LLM) and the visual-to-language model(VLM), which has visual question-answering functionality. In addition, we also obtain the semantic similarity of target object and frontier object categories. Ultimately, by combining the semantic similarity and the boundary descriptions, the agent can predict the long-term goals more accurately. Our experiments on the Gibson and HM3D datasets reveal that our VLAI approach yields significantly better results compared to earlier methods. The code is released at https://github.com/31539lab/VLAI .

Vision-and-Language Navigation via Latent Semantic Alignment Learning

Cross-modal Semantic Alignment Pre-training for Vision-and-Language Navigation

DELAN: Dual-Level Alignment for Vision-and-Language Navigation by Cross-Modal Contrastive Learning

Self-Supervised 3-D Semantic Representation Learning for Vision-and-Language Navigation

Vision-Language Navigation Policy Learning and Adaptation

Reinforced Cross-Modal Matching and Self-Supervised Imitation Learning for Vision-Language Navigation

Improving Vision-and-Language Navigation by Generating Future-View Image Semantics

Language-guided Navigation Via Cross-Modal Grounding and Alternate Adversarial Learning

Vision Language Navigation with Multi-granularity Observation and Auxiliary Reasoning Tasks

Vision and Language Navigation Using Multi-head Attention Mechanism

Vision-Language Navigation With Self-Supervised Auxiliary Reasoning Tasks

Actional Atomic-Concept Learning for Demystifying Vision-Language Navigation

Self-supervised 3D Semantic Representation Learning for Vision-and-Language Navigation

A Dual Semantic-Aware Recurrent Global-Adaptive Network For Vision-and-Language Navigation

Diagnosing Vision-and-Language Navigation: What Really Matters

Improving Cross-Modal Alignment in Vision Language Navigation via Syntactic Information

Narrowing the Gap between Vision and Action in Navigation

Self-supervised 3D Semantic Representation Learning for Vision-and-Language Navigation

VLAI: Exploration and exploitation based on visual-language aligned information for robotic object goal navigation