Abstract:Learning 3D representation plays a critical role in masked autoencoder (MAE) based pre-training methods for point cloud, including single-modal and cross-modal based MAE. Specifically, although cross-modal MAE methods learn strong 3D representations via the auxiliary of other modal knowledge, they often suffer from heavy computational burdens and heavily rely on massive cross-modal data pairs that are often unavailable, which hinders their applications in practice. Instead, single-modal methods with solely point clouds as input are preferred in real applications due to their simplicity and efficiency. However, such methods easily suffer from limited 3D representations with global random mask input. To learn compact 3D representations, we propose a simple yet effective Point Feature Enhancement Masked Autoencoders (Point-FEMAE), which mainly consists of a global branch and a local branch to capture latent semantic features. Specifically, to learn more compact features, a share-parameter Transformer encoder is introduced to extract point features from the global and local unmasked patches obtained by global random and local block mask strategies, followed by a specific decoder to reconstruct. Meanwhile, to further enhance features in the local branch, we propose a Local Enhancement Module with local patch convolution to perceive fine-grained local context at larger scales. Our method significantly improves the pre-training efficiency compared to cross-modal alternatives, and extensive downstream experiments underscore the state-of-the-art effectiveness, particularly outperforming our baseline (Point-MAE) by 5.16%, 5.00%, and 5.04% in three variants of ScanObjectNN, respectively. The code is available at <a class="link-external link-https" href="https://github.com/zyh16143998882/AAAI24-PointFEMAE" rel="external noopener nofollow">this https URL</a>.

PiMAE: Point Cloud and Image Interactive Masked Autoencoders for 3D Object Detection

Masked Autoencoders for Point Cloud Self-supervised Learning.

GD-MAE: Generative Decoder for MAE Pre-training on LiDAR Point Clouds

Point Cloud Self-supervised Learning via 3D to Multi-view Masked Autoencoder

UniM$^2$AE: Multi-modal Masked Autoencoders with Unified 3D Representation for 3D Perception in Autonomous Driving

Point-M2AE: Multi-scale Masked Autoencoders for Hierarchical Point Cloud Pre-training

UniM^2AE: Multi-modal Masked Autoencoders with Unified 3D Representation for 3D Perception in Autonomous Driving

Inter-Modal Masked Autoencoder for Self-Supervised Learning on Point Clouds

LR-MAE: Locate While Reconstructing with Masked Autoencoders for Point Cloud Self-supervised Learning

Joint-MAE: 2D-3D Joint Masked Autoencoders for 3D Point Cloud Pre-training

Towards Compact 3D Representations via Point Feature Enhancement Masked Autoencoders

MultiMAE: Multi-modal Multi-task Masked Autoencoders

BEV-MAE: Bird's Eye View Masked Autoencoders for Point Cloud Pre-training in Autonomous Driving Scenarios

Masked Autoencoders in 3D Point Cloud Representation Learning

PCP-MAE: Learning to Predict Centers for Point Masked Autoencoders

MonoMAE: Enhancing Monocular 3D Detection through Depth-Aware Masked Autoencoders

Masked Autoencoder for Pre-Training on 3D Point Cloud Object Detection

CrossMAE: Cross Modality Masked Autoencoders for Region-Aware Audio-Visual Pretraining

BEV-MAE: Bird's Eye View Masked Autoencoders for Outdoor Point Cloud Pre-training