Abstract:We study the task of semantic mapping – specifically, an embodied agent (a robot or an egocentric AI assistant) is given a tour of a new environment and asked to build an allocentric top-down semantic map (‘what is where?’) from egocentric observations of an RGB-D camera with known pose (via localization sensors). Importantly, our goal is to build neural episodic memories and spatio-semantic representations of 3D spaces that enable the agent to easily learn subsequent tasks in the same space – navigating to objects seen during the tour (‘Find chair’) or answering questions about the space (‘How many chairs did you see in the house?’). Towards this goal, we present Semantic MapNet (SMNet), which consists of: (1) an Egocentric Visual Encoder that encodes each egocentric RGB-D frame, (2) a Feature Projector that projects egocentric features to appropriate locations on a floor-plan, (3) a Spatial Memory Tensor of size floor-plan length×width×feature-dims that learns to accumulate projected egocentric features, and (4) a Map Decoder that uses the memory tensor to produce semantic top-down maps. SMNet combines the strengths of (known) projective camera geometry and neural representation learning. On the task of semantic mapping in the Matterport3D dataset, SMNet significantly outperforms competitive baselines by 4.01−16.81% (absolute) on mean-IoU and 3.81−19.69% (absolute) on Boundary-F1 metrics. Moreover, we show how to use the spatio-semantic allocentric representations build by SMNet for the task of ObjectNav and Embodied Question Answering. Project page: https://vincentcartillier.github.io/smnet.html.

Semantic MapNet: Building Allocentric Semantic Maps and Representations from Egocentric Views

3D Semantic MapNet: Building Maps for Multi-Object Re-Identification in 3D

Object-aware Semantic Mapping of Indoor Scenes Using Octomap

Learning Navigational Visual Representations with Semantic Map Supervision

Interactive Semantic Map Representation for Skill-Based Visual Object Navigation

MaAST: Map Attention with Semantic Transformersfor Efficient Visual Navigation

Simultaneous Mapping and Target Driven Navigation

Trans4Map: Revisiting Holistic Bird's-Eye-View Mapping from Egocentric Images to Allocentric Semantics with Vision Transformers

Mapping High-level Semantic Regions in Indoor Environments without Object Recognition

SNAP: Self-Supervised Neural Maps for Visual Positioning and Semantic Understanding

An Approach for Construct Semantic Map with Scene Classification and Object Semantic Segmentation

Semantic Map Construction Method Based on Brain-Inspired SLAM

SeMLaPS: Real-time Semantic Mapping with Latent Prior Networks and Quasi-Planar Segmentation

Scalable Spatial Memory for Scene Rendering and Navigation

HDMapNet: A Local Semantic Map Learning and Evaluation Framework.

End-to-End Egospheric Spatial Memory

Neural Semantic Map-Learning for Autonomous Vehicles

Constructing Metric-Semantic Maps using Floor Plan Priors for Long-Term Indoor Localization

Volumetric Instance-Aware Semantic Mapping and 3D Object Discovery

SeanNet: Semantic Understanding Network for Localization Under Object Dynamics