Abstract:Despite radar's popularity in the automotive industry, for fusion-based 3D object detection, most existing works focus on LiDAR and camera fusion. In this paper, we propose TransCAR, a Transformer-based Camera-And-Radar fusion solution for 3D object detection. Our TransCAR consists of two modules. The first module learns 2D features from surround-view camera images and then uses a sparse set of 3D object queries to index into these 2D features. The vision-updated queries then interact with each other via transformer self-attention layer. The second module learns radar features from multiple radar scans and then applies transformer decoder to learn the interactions between radar features and vision-updated queries. The cross-attention layer within the transformer decoder can adaptively learn the soft-association between the radar features and vision-updated queries instead of hard-association based on sensor calibration only. Finally, our model estimates a bounding box per query using set-to-set Hungarian loss, which enables the method to avoid non-maximum suppression. TransCAR improves the velocity estimation using the radar scans without temporal information. The superior experimental results of our TransCAR on the challenging nuScenes datasets illustrate that our TransCAR outperforms state-of-the-art Camera-Radar fusion-based 3D object detection approaches.

What problem does this paper attempt to address?

The paper attempts to address the problem of 3D object detection in autonomous driving systems using radar and camera fusion. Although radar is very popular in the automotive industry, most existing 3D object detection research focuses primarily on the fusion of LiDAR and cameras. Radar data is sparser compared to LiDAR point clouds and lacks height information, making it challenging to directly estimate and classify 3D bounding boxes from radar data. However, radar performs well under adverse weather and lighting conditions, can accurately measure the radial velocity of objects, and is cost-effective. The paper proposes a new method called TransCAR, which is based on the Transformer framework and learns the interaction between camera and radar features through self-attention and cross-attention mechanisms. Specifically, TransCAR consists of two modules: the first module extracts 2D features from surround-view images and uses a set of sparse 3D object queries to index these 2D features; the second module learns radar features from multiple radar scans and applies a Transformer decoder to learn the interaction between radar features and vision-updated queries. This method can adaptively learn soft associations, thereby avoiding the limitations brought by hard associations based solely on sensor calibration. Experimental results show that TransCAR outperforms existing radar-camera fusion methods on the nuScenes dataset, demonstrating significant improvements in 3D object detection tasks across various distance ranges. Particularly, the improvements are most notable in the 20 to 40 meters range. Additionally, for objects without radar echoes, TransCAR can still maintain baseline performance, while for objects with radar echoes, it significantly enhances detection performance.

TransCAR: Transformer-based Camera-And-Radar Fusion for 3D Object Detection

CenterTransFuser: radar point cloud and visual information fusion for 3D object detection

Fusing LiDAR and Radar with Pillars Attention for 3D Object Detection

RCTrans: Radar-Camera Transformer via Radar Densifier and Sequential Decoder for 3D Object Detection

TransFusion: Multi-Modal Robust Fusion for 3D Object Detection in Foggy Weather Based on Spatial Vision Transformer

Radar-camera Fusion for 3D Object Detection with Aggregation Transformer

CRAFT: Camera-Radar 3D Object Detection with Spatio-Contextual Fusion Transformer

CR-DINO: A Novel Camera-Radar Fusion 2D Object Detection Model Based On Transformer

Radar and Camera Fusion for Multi-Task Sensing in Autonomous Driving

BEV-Radar: Bidirectional Radar-Camera Fusion for 3D Object Detection

Radar-Image Fusion Transformer for Object Detection with Uncertainty Estimation

CRT-Fusion: Camera, Radar, Temporal Fusion Using Motion Information for 3D Object Detection

TransFusion: Robust LiDAR-Camera Fusion for 3D Object Detection with Transformers

TransRadar: Adaptive-Directional Transformer for Real-Time Multi-View Radar Semantic Segmentation

RCFusion: Fusing 4-D Radar and Camera with Bird's-Eye View Features for 3-D Object Detection.

Radar Voxel Fusion for 3D Object Detection

RCDPT: Radar-Camera fusion Dense Prediction Transformer

CramNet: Camera-Radar Fusion with Ray-Constrained Cross-Attention for Robust 3D Object Detection

RCM-Fusion: Radar-Camera Multi-Level Fusion for 3D Object Detection

Radar-Camera Sensor Fusion for Joint Object Detection and Distance Estimation in Autonomous Vehicles

HVDetFusion: A Simple and Robust Camera-Radar Fusion Framework