Orthographic Feature Transform for Monocular 3D Object Detection

Thomas Roddick,Alex Kendall,Roberto Cipolla

DOI: https://doi.org/10.48550/arXiv.1811.08188

2018-11-20

Abstract:3D object detection from monocular images has proven to be an enormously challenging task, with the performance of leading systems not yet achieving even 10\% of that of LiDAR-based counterparts. One explanation for this performance gap is that existing systems are entirely at the mercy of the perspective image-based representation, in which the appearance and scale of objects varies drastically with depth and meaningful distances are difficult to infer. In this work we argue that the ability to reason about the world in 3D is an essential element of the 3D object detection task. To this end, we introduce the orthographic feature transform, which enables us to escape the image domain by mapping image-based features into an orthographic 3D space. This allows us to reason holistically about the spatial configuration of the scene in a domain where scale is consistent and distances between objects are meaningful. We apply this transformation as part of an end-to-end deep learning architecture and achieve state-of-the-art performance on the KITTI 3D object benchmark.\footnote{We will release full source code and pretrained models upon acceptance of this manuscript for publication.

Computer Vision and Pattern Recognition

What problem does this paper attempt to address?

The problem that this paper attempts to solve is the problem of 3D object detection from monocular images. Specifically, the author points out that the performance of current 3D object detection methods based on monocular images is far lower than that of methods based on LiDAR. The main reason is that the appearance and scale of objects in monocular images will change significantly with the change of distance, and it is difficult to directly infer meaningful distances. To solve this problem, the paper introduces a new method - Orthographic Feature Transform (OFT). This method can map image - based features to an orthogonal 3D space, so as to perform reasoning in a space where the scale is consistent and the distances between objects are meaningful. In this way, the method proposed in the paper has reached the state - of - the - art level among monocular methods in the KITTI 3D object detection benchmark test. The main contributions of the paper include: 1. Introducing the Orthographic Feature Transform (OFT), which can map perspective image features to an orthogonal bird - eye - view representation and use integral images to achieve fast average pooling. 2. Describing a deep - learning architecture for predicting 3D bounding boxes from monocular RGB images. 3. Emphasizing the importance of 3D reasoning in object detection tasks. Through these contributions, the paper aims to improve the performance of 3D object detection based on monocular images. Especially in applications such as autonomous driving, the cost and redundancy of monocular image sensors make it an important choice.

Orthographic Feature Transform for Monocular 3D Object Detection

Dynamic Depth Fusion and Transformation for Monocular 3D Object Detection.

Kinematic 3D Object Detection in Monocular Video

Progressive Coordinate Transforms for Monocular 3D Object Detection

OCM3D: Object-Centric Monocular 3D Object Detection

DETR3D: 3D Object Detection from Multi-view Images via 3D-to-2D Queries

Efficient Feature Aggregation and Scale-Aware Regression for Monocular 3D Object Detection

Learning 2D to 3D Lifting for Object Detection in 3D for Autonomous Vehicles

Disentangling Monocular 3D Object Detection

Depth-Vision-Decoupled Transformer With Cascaded Group Convolutional Attention for Monocular 3-D Object Detection

MonoDTR: Monocular 3D Object Detection with Depth-Aware Transformer

RTM3D: Real-Time Monocular 3D Detection from Object Keypoints for Autonomous Driving

3D Street Object Detection from Monocular Images Using Deep Learning and Depth Information

OcTr: Octree-based Transformer for 3D Object Detection

Perspective-aware Convolution for Monocular 3D Object Detection

Monocular 3D Object Detection: An Extrinsic Parameter Free Approach

Transformation-Equivariant 3D Object Detection for Autonomous Driving

Monocular 3D Object Detection Leveraging Accurate Proposals and Shape Reconstruction

Ground-aware Monocular 3D Object Detection for Autonomous Driving

3D Object Recognition By Corresponding and Quantizing Neural 3D Scene Representations

Training an Open-Vocabulary Monocular 3D Object Detection Model without 3D Data