Abstract:LiDAR-based sparse 3-D object detection plays a crucial role in autonomous driving applications due to its computational efficiency advantages. Existing methods either use the features of a single central voxel as an object proxy or treat an aggregated cluster of foreground points as an object proxy. However, the former cannot aggregate contextual information, resulting in insufficient information expression in object proxies. The latter relies on multistage pipelines and auxiliary tasks, which reduce the inference speed. To maintain the efficiency of the sparse framework while fully aggregating contextual information, in this work, we propose SparseDet that designs sparse queries as object proxies. It introduces two key modules: the local multiscale feature aggregation (LMFA) module and the global feature aggregation (GFA) module, aiming to fully capture the contextual information, thereby enhancing the ability of the proxies to represent objects. The LMFA module achieves feature fusion across different scales for sparse key voxels via coordinate transformations and using nearest neighbor relationships to capture object-level details and local contextual information, whereas the GFA module uses self-attention mechanisms to selectively aggregate the features of the key voxels across the entire scene for capturing scene-level contextual information. Experiments on nuScenes and KITTI demonstrate the effectiveness of our method. Specifically, SparseDet surpasses the previous best sparse detector VoxelNeXt (a typical method using voxels as object proxies) by 2.2% mean average precision (mAP) with 13.5 frames/s on nuScenes and outperforms VoxelNeXt by 1.12% on hard level tasks with 17.9 frames/s on KITTI. What is more, not only the mAP of SparseDet exceeds that of FSDV2 (a classical method using clusters of foreground points as object proxies) but also its inference speed is 1.3 times faster than FSDV2 on the nuScenes test set. The code has been released in https://github.com/liulin813/SparseDet.git.

SparseFormer: Detecting Objects in HRW Shots Via Sparse Vision Transformer

Bridging the Gap Between Object Detection in Close-Up and High-Resolution Wide Shots

Vision Transformers for Single Image Dehazing

SSF: Sparse Point Cloud Object Detection Based on Self-Adaptive Voxel Encoding and Focal-Sparse Convolution

SWFormer: Sparse Window Transformer for 3D Object Detection in Point Clouds

SparseFormer: Sparse Visual Recognition via Limited Latent Tokens

SparseFusion: Efficient Sparse Multi-Modal Fusion Framework for Long-Range 3D Perception

SparseDet: A Simple and Effective Framework for Fully Sparse LiDAR-based 3D Object Detection

SparseDet: A Simple and Effective Framework for Fully Sparse LiDAR-Based 3-D Object Detection

Super Sparse 3D Object Detection

Embracing Single Stride 3D Object Detector with Sparse Transformer

Vision Transformer with Sparse Scan Prior

Few-Shot Object Detection with Sparse Context Transformers

Dynamic Spatial Sparsification for Efficient Vision Transformers and Convolutional Neural Networks

DS-Trans: A 3D Object Detection Method Based on a Deformable Spatiotemporal Transformer for Autonomous Vehicles

High-Resolution Network with Transformer Embedding Parallel Detection for Small Object Detection in Optical Remote Sensing Images

Sparse4D v2: Recurrent Temporal Fusion with Sparse Model

SparseFormer: Sparse Transformer Network for Point Cloud Classification

HRFormer: High-Resolution Transformer for Dense Prediction

DSVT: Dynamic Sparse Voxel Transformer with Rotated Sets

DFS-DETR: Detailed-Feature-Sensitive Detector for Small Object Detection in Aerial Images Using Transformer