Abstract:We present SLAIM - Simultaneous Localization and Implicit Mapping. We propose a novel coarse-to-fine tracking model tailored for Neural Radiance Field SLAM (NeRF-SLAM) to achieve state-of-the-art tracking performance. Notably, existing NeRF-SLAM systems consistently exhibit inferior tracking performance compared to traditional SLAM algorithms. NeRF-SLAM methods solve camera tracking via image alignment and photometric bundle-adjustment. Such optimization processes are difficult to optimize due to the narrow basin of attraction of the optimization loss in image space (local minima) and the lack of initial correspondences. We mitigate these limitations by implementing a Gaussian pyramid filter on top of NeRF, facilitating a coarse-to-fine tracking optimization strategy. Furthermore, NeRF systems encounter challenges in converging to the right geometry with limited input views. While prior approaches use a Signed-Distance Function (SDF)-based NeRF and directly supervise SDF values by approximating ground truth SDF through depth measurements, this often results in suboptimal geometry. In contrast, our method employs a volume density representation and introduces a novel KL regularizer on the ray termination distribution, constraining scene geometry to consist of empty space and opaque surfaces. Our solution implements both local and global bundle-adjustment to produce a robust (coarse-to-fine) and accurate (KL regularizer) SLAM solution. We conduct experiments on multiple datasets (ScanNet, TUM, Replica) showing state-of-the-art results in tracking and in reconstruction accuracy.

ObjectFusion: Accurate object-level SLAM with neural object priors

ObjectFusion: an Object Detection and Segmentation Framework with RGB-D SLAM and Convolutional Neural Networks

LCP-Fusion: A Neural Implicit SLAM with Enhanced Local Constraints and Computable Prior

Deep-SLAM++: Object-level RGBD SLAM based on class-specific deep shape priors

Contour-SLAM: A Robust Object-Level SLAM Based on Contour Alignment

GeoFusion: Geometric Consistency informed Scene Estimation in Dense Clutter

Toward General Object-level Mapping from Sparse Views with 3D Diffusion Priors

HI-SLAM: Monocular Real-time Dense Mapping with Hybrid Implicit Fields

Improving SLAM Techniques with Integrated Multi-Sensor Fusion for 3D Reconstruction

Co-Fusion: Real-time Segmentation, Tracking and Fusion of Multiple Objects

Open-Fusion: Real-time Open-Vocabulary 3D Mapping and Queryable Scene Representation

SLAIM: Robust Dense Neural SLAM for Online Tracking and Mapping

Compositional Scalable Object SLAM

Object SLAM Based on Spatial Layout and Semantic Consistency

Visual-Inertial Multi-Instance Dynamic SLAM with Object-level Relocalisation

An Object SLAM Framework for Association, Mapping, and High-Level Tasks

vMAP: Vectorised Object Mapping for Neural Field SLAM

Multi-view 3D Object Reconstruction and Uncertainty Modelling with Neural Shape Prior

GO-SLAM: Global Optimization for Consistent 3D Instant Reconstruction

On the Overconfidence Problem in Semantic 3D Mapping

Fusion LiDAR-Inertial-Encoder data for High-Accuracy SLAM