Abstract: Scene flow represents the 3D motion of each point in the scene, which explicitly describes the distance and the direction of each point's movement. Scene flow estimation is used in various applications such as autonomous driving fields, activity recognition, and virtual reality fields. As it is challenging to annotate scene flow with ground truth for real-world data, this leaves no real-world dataset available to provide a large amount of data with ground truth for scene flow estimation. Therefore, many works use synthesized data to pre-train their network and real-world LiDAR data to finetune. Unlike the previous unsupervised learning of scene flow in point clouds, we propose to use odometry information to assist the unsupervised learning of scene flow and use real-world LiDAR data to train our network. Supervised odometry provides more accurate shared cost volume for scene flow. In addition, the proposed network has mask-weighted warp layers to get a more accurate predicted point cloud. The warp operation means applying an estimated pose transformation or scene flow to a source point cloud to obtain a predicted point cloud and is the key to refining scene flow from coarse to fine. When performing warp operations, the points in different states use different weights for the pose transformation and scene flow transformation. We classify the states of points as static, dynamic, and occluded, where the static masks are used to divide static and dynamic points, and the occlusion masks are used to divide occluded points. The mask-weighted warp layer indicates that static masks and occlusion masks are used as weights when performing warp operations. Our designs are proved to be effective in ablation experiments. The experiment results show the promising prospect of an odometry-assisted unsupervised learning method for 3D scene flow in real-world data.

Learning Rigidity in Dynamic Scenes with a Moving Camera for 3D Motion Field Estimation

Unsupervised Learning of Scene Flow Estimation Fusing with Local Rigidity.

EffiScene: Efficient Per-Pixel Rigidity Inference for Unsupervised Joint Learning of Optical Flow, Depth, Camera Pose and Motion Segmentation

Layered RGBD Scene Flow Estimation with Global Non-rigid Local Rigid Assumption

Weakly Supervised Learning of Rigid 3D Scene Flow

Camera Pose Estimation in Dynamic Scenes with Background Tracking

EMR-MSF: Self-Supervised Recurrent Monocular Scene Flow Exploiting Ego-Motion Rigidity

Self-Supervised 3D Scene Flow Estimation and Motion Prediction using Local Rigidity Prior

Robust Real-time RGB-D Visual Odometry in Dynamic Environments via Rigid Motion Model

Motion Detection for Rapidly Moving Cameras in Fully 3D Scenes

Unsupervised Learning of 3D Scene Flow with 3D Odometry Assistance

Learning Residual Flow as Dynamic Motion from Stereo Videos

Scene Motion Decomposition for Learnable Visual Odometry

Dense Depth Estimation of a Complex Dynamic Scene without Explicit 3D Motion Estimation

Region Deformer Networks for Unsupervised Depth Estimation from Unconstrained Monocular Videos

RigidFlow: Self-Supervised Scene Flow Learning on Point Clouds by Local Rigidity Prior

Self-Supervised Learning of Non-Rigid Residual Flow and Ego-Motion

Motion Rectification Network for Unsupervised Learning of Monocular Depth and Camera Motion

Camera Motion Estimation from RGB-D-Inertial Scene Flow

ScaleFlow++: Robust and Accurate Estimation of 3D Motion from Video

R3D3: Dense 3D Reconstruction of Dynamic Scenes from Multiple Cameras