You Only Label Once: 3D Box Adaptation from Point Cloud to Image with Semi-Supervised Learning

Shi Jieqi,Li Peiliang,Chen Xiaozhi,Shen Shaojie
DOI: https://doi.org/10.1109/lra.2023.3310433
IF: 5.2
2023-01-01
IEEE Robotics and Automation Letters
Abstract:The image-based 3D object detection task expects that the predicted 3D bounding box has a “tightness” projection (also referred to as cuboid) to facilitate 2D-based training, which fits the object contour well on the image while remaining reasonable on the 3D space. These requirements bring significant challenges to the annotation. Projecting the Lidar-labeled 3D boxes to the image leads to non-trivial misalignment, while directly drawing a cuboid on the image cannot access the original 3D information. In this work, we propose a learning-based 3D box adaptation approach that automatically adjusts minimum parameters of the 360 $^{\circ }$ Lidar 3D bounding box to fit the image appearance of panoramic cameras perfectly. With only a few 2D boxes annotation as guidance during the training phase, our network can produce accurate image-level cuboid annotations with 3D properties from Lidar boxes. We call our method “you only label once”, which means labeling on the point cloud once and automatically adapting to all surrounding cameras. Our refinement balances the accuracy and efficiency well and dramatically reduces the labeling effort for accurate cuboid annotation. Extensive experiments on the public Waymo and NuScenes datasets show that our method can produce human-level cuboid annotation on the image without manual adjustment and can accelerate monocular-3D training tasks.
What problem does this paper attempt to address?