Abstract:Masked Autoencoders (MAE) have demonstrated promising performance in self-supervised learning for both 2D and 3D computer vision. Nevertheless, existing MAE-based methods still have certain drawbacks. Firstly, the functional decoupling between the encoder and decoder is incomplete, which limits the encoder's representation learning ability. Secondly, downstream tasks solely utilize the encoder, failing to fully leverage the knowledge acquired through the encoder-decoder architecture in the pre-text task. In this paper, we propose Point Regress AutoEncoder (Point-RAE), a new scheme for regressive autoencoders for point cloud self-supervised learning. The proposed method decouples functions between the decoder and the encoder by introducing a mask regressor, which predicts the masked patch representation from the visible patch representation encoded by the encoder and the decoder reconstructs the target from the predicted masked patch representation. By doing so, we minimize the impact of decoder updates on the representation space of the encoder. Moreover, we introduce an alignment constraint to ensure that the representations for masked patches, predicted from the encoded representations of visible patches, are aligned with the masked patch presentations computed from the encoder. To make full use of the knowledge learned in the pre-training stage, we design a new finetune mode for the proposed Point-RAE. Extensive experiments demonstrate that our approach is efficient during pre-training and generalizes well on various downstream tasks. Specifically, our pre-trained models achieve a high accuracy of \textbf{90.28\%} on the ScanObjectNN hardest split and \textbf{94.1\%} accuracy on ModelNet40, surpassing all the other self-supervised learning methods. Our code and pretrained model are public available at: \url{<a class="link-external link-https" href="https://github.com/liuyyy111/Point-RAE" rel="external noopener nofollow">this https URL</a>}.

PatchMixing Masked Autoencoders for 3D Point Cloud Self-Supervised Learning

Masked Autoencoders for Point Cloud Self-supervised Learning.

Masked Autoencoders in 3D Point Cloud Representation Learning

PointPatchMix: Point Cloud Mixing with Patch Scoring

Point‐AGM : Attention Guided Masked Auto‐Encoder for Joint Self‐supervised Learning on Point Clouds

Masked Spatio-Temporal Structure Prediction for Self-supervised Learning on Point Cloud Videos

Point Cloud Self-supervised Learning via 3D to Multi-view Masked Autoencoder

M^3CS: Multi-Target Masked Point Modeling with Learnable Codebook and Siamese Decoders

Inter-Modal Masked Autoencoder for Self-Supervised Learning on Point Clouds

LR-MAE: Locate While Reconstructing with Masked Autoencoders for Point Cloud Self-supervised Learning

GeoMAE: Masked Geometric Target Prediction for Self-supervised Point Cloud Pre-Training

Point-M2AE: Multi-scale Masked Autoencoders for Hierarchical Point Cloud Pre-training

Regress Before Construct: Regress Autoencoder for Point Cloud Self-supervised Learning

Point-LGMask: Local and Global Contexts Embedding for Point Cloud Pre-training with Multi-Ratio Masking

Self-supervised Point Cloud Representation Learning Via Separating Mixed Shapes

Bringing Masked Autoencoders Explicit Contrastive Properties for Point Cloud Self-Supervised Learning

RI-MAE: Rotation-Invariant Masked AutoEncoders for Self-Supervised Point Cloud Representation Learning

Point-MPP: Point Cloud Self-Supervised Learning from Masked Position Prediction

PCP-MAE: Learning to Predict Centers for Point Masked Autoencoders

Towards Compact 3D Representations via Point Feature Enhancement Masked Autoencoders

A Simple Masked Autoencoder Paradigm for Point Cloud