Abstract:Striking a balance between precision and efficiency presents a prominent challenge in the bird's-eye-view (BEV) 3D object detection. Although previous camera-based BEV methods achieved remarkable performance by incorporating long-term temporal information, most of them still face the problem of low efficiency. One potential solution is knowledge distillation. Existing distillation methods only focus on reconstructing spatial features, while overlooking temporal knowledge. To this end, we propose TempDistiller, a Temporal knowledge Distiller, to acquire long-term memory from a teacher detector when provided with a limited number of frames. Specifically, a reconstruction target is formulated by integrating long-term temporal knowledge through self-attention operation applied to feature teachers. Subsequently, novel features are generated for masked student features via a generator. Ultimately, we utilize this reconstruction target to reconstruct the student features. In addition, we also explore temporal relational knowledge when inputting full frames for the student model. We verify the effectiveness of the proposed method on the nuScenes benchmark. The experimental results show our method obtain an enhancement of +1.6 mAP and +1.1 NDS compared to the baseline, a speed improvement of approximately 6 FPS after compressing temporal knowledge, and the most accurate velocity estimation.

What problem does this paper attempt to address?

This paper attempts to solve the problem of balancing accuracy and efficiency in bird - eye - view (BEV) 3D object detection. Although previous camera - based BEV methods have achieved significant performance improvements through fusing long - term temporal information, most methods still face the problem of inefficiency. The paper proposes a temporal knowledge distillation method named TempDistiller, which aims to obtain long - term memory from the teacher detector, and this can be achieved even when a limited number of frames are provided. Specifically, the long - term temporal knowledge in the teacher features is integrated through self - attention operations to form a reconstruction target, and new features are generated for the masked parts in the student features by the generator, and finally the student features are reconstructed according to the reconstruction target. In addition, when the complete frames are input, the paper also explores the temporal relationship knowledge to further improve the speed and accuracy, especially in speed estimation. The main contributions of the paper include: - Proposing the first method to transfer temporal knowledge in 3D object detection, and exploring how to learn temporal knowledge and its relationship information. - Providing a novel perspective. Through knowledge distillation to handle long - term temporal fusion, the student detector can obtain long - term temporal knowledge from the teacher detector even when the number of input frames is reduced. - The experimental results show that compared with the baseline, this method improves 1.6 mAP and 1.1 NDS, the inference speed is increased by about 6 FPS, and at the same time, the most accurate speed estimation is achieved. These contributions not only solve the efficiency problems existing in the existing methods when dealing with long - term temporal information, but also provide a new direction for future 3D object detection research.

Distilling Temporal Knowledge with Masked Feature Reconstruction for 3D Object Detection

Distilling Focal Knowledge from Imperfect Expert for 3D Object Detection

Structured Knowledge Distillation Towards Efficient and Compact Multi-View 3D Detection

SimDistill: Simulated Multi-Modal Distillation for BEV 3D Object Detection

DMKD: Improving Feature-based Knowledge Distillation for Object Detection Via Dual Masking Augmentation

UniDistill: A Universal Cross-Modality Knowledge Distillation Framework for 3D Object Detection in Bird's-Eye View

BEVDistill: Cross-Modal BEV Distillation for Multi-View 3D Object Detection

Distilling Object Detectors with Global Knowledge

Towards Efficient 3D Object Detection with Knowledge Distillation

Representation Disparity-aware Distillation for 3D Object Detection

SAM-Guided Masked Token Prediction for 3D Scene Understanding

TKD: Temporal Knowledge Distillation for Active Perception

Leveraging Vision-Centric Multi-Modal Expertise for 3D Object Detection

Attention-Based Depth Distillation with 3D-Aware Positional Encoding for Monocular 3D Object Detection

PointDistiller: Structured Knowledge Distillation Towards Efficient and Compact 3D Detection

AMD: Adaptive Masked Distillation for Object Detection

LabelDistill: Label-guided Cross-modal Knowledge Distillation for Camera-based 3D Object Detection

FSD-BEV: Foreground Self-Distillation for Multi-view 3D Object Detection

Temporal Enhanced Training of Multi-view 3D Object Detector via Historical Object Prediction

Cyclic Refiner: Object-Aware Temporal Representation Learning for Multi-view 3D Detection and Tracking

Distilling Object Detectors With Fine-Grained Feature Imitation