Abstract:Although recent Siamese network-based trackers have achieved impressive perceptual accuracy for single object tracking in LiDAR point clouds, they usually utilized heavy correlation operations to capture category-level characteristics only, and overlook the inherent merit of arbitrariness in contrast to multiple object tracking. In this work, we propose a radically novel one-stream network with the strength of the instance-level encoding, which avoids the correlation operations occurring in previous Siamese network, thus considerably reducing the computational effort. In particular, the proposed method mainly consists of a Template-aware Transformer Module (TTM) and a Multi-scale Feature Aggregation (MFA) module capable of fusing spatial and semantic information. The TTM stitches the specified template and the search region together and leverages an attention mechanism to establish the information flow, breaking the previous pattern of independent extraction-and-correlation. As a result, this module makes it possible to directly generate template-aware features that are suitable for the arbitrary and continuously changing nature of the target, enabling the model to deal with unseen categories. In addition, the MFA is proposed to make spatial and semantic information complementary to each other, which is characterized by reverse directional feature propagation that aggregates information from shallow to deep layers. Extensive experiments on KITTI and nuScenes demonstrate that our method has achieved considerable performance not only for class-specific tracking but also for class-agnostic tracking with less computation and higher efficiency.

CenterTube: Tracking Multiple 3D Objects with 4D Tubelets in Dynamic Point Clouds

Exploit the Connectivity: Multi-Object Tracking with TrackletNet

Exploit Spatiotemporal Contextual Information for 3D Single Object Tracking Via Memory Networks

Object-Level Pseudo-3D Lifting for Distance-Aware Tracking

InterTrack: Interaction Transformer for 3D Multi-Object Tracking

Chained-Tracker: Chaining Paired Attentive Regression Results for End-to-End Joint Multiple-Object Detection and Tracking

Tracking Objects as Points

DeepPCT: Single Object Tracking in Dynamic Point Cloud Sequences

3D Multi-Object Tracking in Point Clouds Based on Prediction Confidence-Guided Data Association

PointTrackNet: An End-to-End Network For 3-D Object Detection and Tracking From Point Clouds

EasyTrack: Efficient and Compact One-stream 3D Point Clouds Tracker

ByteTrackV2: 2D and 3D Multi-Object Tracking by Associating Every Detection Box

MotionTrack: Learning Robust Short-term and Long-term Motions for Multi-Object Tracking

TrackNet: Simultaneous Object Detection and Tracking and Its Application in Traffic Video Analysis

TransCenter: Transformers With Dense Representations for Multiple-Object Tracking

MFITrack: Multi-Frame Integration Strategy for Enhanced Motion-Centric Single Object Tracking

Beyond 3D Siamese Tracking: A Motion-Centric Paradigm for 3D Single Object Tracking in Point Clouds

OST: Efficient One-stream Network for 3D Single Object Tracking in Point Clouds

TPTrack: Strengthening tracking-by-detection methods from tracklet processing perspectives

MCTrack: A Unified 3D Multi-Object Tracking Framework for Autonomous Driving