Abstract: Scene flow represents the 3D motion of each point in the scene, which explicitly describes the distance and the direction of each point's movement. Scene flow estimation is used in various applications such as autonomous driving fields, activity recognition, and virtual reality fields. As it is challenging to annotate scene flow with ground truth for real-world data, this leaves no real-world dataset available to provide a large amount of data with ground truth for scene flow estimation. Therefore, many works use synthesized data to pre-train their network and real-world LiDAR data to finetune. Unlike the previous unsupervised learning of scene flow in point clouds, we propose to use odometry information to assist the unsupervised learning of scene flow and use real-world LiDAR data to train our network. Supervised odometry provides more accurate shared cost volume for scene flow. In addition, the proposed network has mask-weighted warp layers to get a more accurate predicted point cloud. The warp operation means applying an estimated pose transformation or scene flow to a source point cloud to obtain a predicted point cloud and is the key to refining scene flow from coarse to fine. When performing warp operations, the points in different states use different weights for the pose transformation and scene flow transformation. We classify the states of points as static, dynamic, and occluded, where the static masks are used to divide static and dynamic points, and the occlusion masks are used to divide occluded points. The mask-weighted warp layer indicates that static masks and occlusion masks are used as weights when performing warp operations. Our designs are proved to be effective in ablation experiments. The experiment results show the promising prospect of an odometry-assisted unsupervised learning method for 3D scene flow in real-world data.

Integrating Semantic Segmentation Model for Self-Supervised Scene Flow Estimation Via Cross Task Distillation

Unsupervised Learning of Scene Flow Estimation Fusing with Local Rigidity.

Efficient Semantic Segmentation Via Self-Attention and Self-Distillation

Self-Supervised Monocular Scene Flow Estimation

X-Distill: Improving Self-Supervised Monocular Depth via Cross-Task Distillation

SeFlow: A Self-Supervised Scene Flow Method in Autonomous Driving

Adversarial Self-Supervised Scene Flow Estimation

Do not trust the neighbors! Adversarial Metric Learning for Self-Supervised Scene Flow Estimation

RAFT-MSF: Self-Supervised Monocular Scene Flow Using Recurrent Optimizer

Learning By Analogy: Reliable Supervision From Transformations For Unsupervised Optical Flow Estimation

Unsupervised Learning of 3D Scene Flow from Monocular Camera

Cross-Image Distillation for Semi-Supervised Semantic Segmentation

Hierarchical Attention Learning of Scene Flow in 3D Point Clouds

SemHint-MD: Learning from Noisy Semantic Labels for Self-Supervised Monocular Depth Estimation

Unsupervised Learning of 3D Scene Flow with 3D Odometry Assistance

Hidden Gems: 4D Radar Scene Flow Learning Using Cross-Modal Supervision

SSFlowNet: Semi-supervised Scene Flow Estimation On Point Clouds With Pseudo Label

Learning Graph-Based Representations for Scene Flow Estimation

Monocular Depth Estimation via Self-Supervised Self-Distillation

Semantic Flow: Learning Semantic Field of Dynamic Scenes from Monocular Videos