Abstract:Deep neural network models have achieved remarkable progress in 3D scene understanding while trained in the closed-set setting and with full labels. However, the major bottleneck for current 3D recognition approaches is that they do not have the capacity to recognize any unseen novel classes beyond the training categories in diverse kinds of real-world applications. In the meantime, current state-of-the-art 3D scene understanding approaches primarily require high-quality labels to train neural networks, which merely perform well in a fully supervised manner. This work presents a generalized and simple framework for dealing with 3D scene understanding when the labeled scenes are quite limited. To extract knowledge for novel categories from the pre-trained vision-language models, we propose a hierarchical feature-aligned pre-training and knowledge distillation strategy to extract and distill meaningful information from large-scale vision-language models, which helps benefit the open-vocabulary scene understanding tasks. To leverage the boundary information, we propose a novel energy-based loss with boundary awareness benefiting from the region-level boundary predictions. To encourage latent instance discrimination and to guarantee efficiency, we propose the unsupervised region-level semantic contrastive learning scheme for point clouds, using confident predictions of the neural network to discriminate the intermediate feature embeddings at multiple stages. Extensive experiments with both indoor and outdoor scenes demonstrated the effectiveness of our approach in both data-efficient learning and open-world few-shot learning. All codes, models, and data are made publicly available at: <a class="link-external link-https" href="https://drive.google.com/drive/folders/1M58V-PtR8DBEwD296zJkNg_m2qq-MTAP?usp=sharing" rel="external noopener nofollow">this https URL</a>.

Data-efficient 3D instance segmentation by transferring knowledge from synthetic scans

3D Object Segmentation Using Cross-Window Point Transformer with Latent Semantic Boundary Guidance

Learning Semantic Segmentation on Unlabeled Real-World Indoor Point Clouds via Synthetic Data

Generating synthetic photogrammetric data for training deep learning based 3D point cloud segmentation models

3D Semantic Segmentation Using Deep Learning for Large-Scale Indoor Point Cloud

Associate Semantic-Instance Segmentation of 3D Point Clouds Based on Local Feature Extraction

OccuSeg: Occupancy-Aware 3D Instance Segmentation

Instance-Aware Embedding for Point Cloud Instance Segmentation

Cross-modal and Cross-domain Knowledge Transfer for Label-free 3D Segmentation

SA3DIP: Segment Any 3D Instance with Potential 3D Priors

3D Segmentation of Humans in Point Clouds with Synthetic Data

Semantic Labeling and Instance Segmentation of 3D Point Clouds Using Patch Context Analysis and Multiscale Processing

Rethinking 3D LiDAR Point Cloud Segmentation

A review of point cloud segmentation for understanding 3D indoor scenes

Generalized Label-Efficient 3D Scene Parsing via Hierarchical Feature Aligned Pre-Training and Region-Aware Fine-tuning

Automated Semantic Segmentation of Indoor Point Clouds from Close-Range Images with Three-Dimensional Deep Learning

RandomRooms: Unsupervised Pre-training from Synthetic Shapes and Randomized Layouts for 3D Object Detection

Enhancing point cloud semantic segmentation in the data‐scarce domain of industrial plants through synthetic data

LiDAR-Based Real-Time Panoptic Segmentation via Spatiotemporal Sequential Data Fusion

Three-Dimensional Instance Segmentation Using the Generalized Hough Transform and the Adaptive n-Shifted Shuffle Attention

Point Cloud Instance Segmentation of Indoor Scenes Using Learned Pairwise Patch Relations