Multiple Object Tracking as ID Prediction

Ruopeng Gao,Yijun Zhang,Limin Wang

2024-03-25

Abstract:In Multiple Object Tracking (MOT), tracking-by-detection methods have stood the test for a long time, which split the process into two parts according to the definition: object detection and association. They leverage robust single-frame detectors and treat object association as a post-processing step through hand-crafted heuristic algorithms and surrogate tasks. However, the nature of heuristic techniques prevents end-to-end exploitation of training data, leading to increasingly cumbersome and challenging manual modification while facing complicated or novel scenarios. In this paper, we regard this object association task as an End-to-End in-context ID prediction problem and propose a streamlined baseline called MOTIP. Specifically, we form the target embeddings into historical trajectory information while considering the corresponding IDs as in-context prompts, then directly predict the ID labels for the objects in the current frame. Thanks to this end-to-end process, MOTIP can learn tracking capabilities straight from training data, freeing itself from burdensome hand-crafted algorithms. Without bells and whistles, our method achieves impressive state-of-the-art performance in complex scenarios like DanceTrack and SportsMOT, and it performs competitively with other transformer-based methods on MOT17. We believe that MOTIP demonstrates remarkable potential and can serve as a starting point for future research. The code is available at

Computer Science

What problem does this paper attempt to address?

### Problems the Paper Attempts to Solve The paper aims to address the object association problem in Multiple Object Tracking (MOT). Traditional tracking methods are usually divided into two stages: object detection and association. These methods rely on manually designed heuristic algorithms to accomplish the object association task, but this approach becomes cumbersome and difficult to adjust in complex or novel situations. The paper proposes a new perspective, treating object association as an end-to-end ID prediction problem, and introduces a concise baseline method called MOTIP. Specifically, MOTIP utilizes historical trajectory information to directly predict the ID labels of objects in the current frame, thereby avoiding traditional manually designed algorithms. This approach can better learn tracking strategies from training data and performs well in complex scenarios (such as DanceTrack and SportsMOT). Additionally, MOTIP is competitive on the MOT17 dataset and demonstrates potential for future research. In this way, MOTIP simplifies the traditional tracking pipeline, achieves an end-to-end learning process, and achieves excellent results in multiple benchmarks.

Multiple Object Tracking as ID Prediction

Split and Connect: A Universal Tracklet Booster for Multi-Object Tracking

Exploit the Connectivity: Multi-Object Tracking with TrackletNet

APPTracker: Improving Tracking Multiple Objects in Low-Frame-Rate Videos

TransLink: Transformer-Based Embedding for Tracklets’ Global Link

MAT: Motion-Aware Multi-Object Tracking

MotionTrack: Learning Motion Predictor for Multiple Object Tracking

Object-Level Pseudo-3D Lifting for Distance-Aware Tracking

Towards Real-Time Multi-Object Tracking

Chained-Tracker: Chaining Paired Attentive Regression Results for End-to-End Joint Multiple-Object Detection and Tracking

MotionTrack: Learning Robust Short-term and Long-term Motions for Multi-Object Tracking

MOTR: End-to-End Multiple-Object Tracking with Transformer

ETTrack: Enhanced Temporal Motion Predictor for Multi-Object Tracking

TR-MOT: Multi-Object Tracking by Reference

Learning a Proposal Classifier for Multiple Object Tracking

Multi-object tracking algorithm based on interactive attention network and adaptive trajectory reconnection

Multi-object tracking with Siamese-RPN and adaptive matching strategy

Multi-object Tracking via Discriminative Embeddings for the Internet of Things

MeMOTR: Long-Term Memory-Augmented Transformer for Multi-Object Tracking

PatchTrack: Multiple Object Tracking Using Frame Patches

DiffMOT: A Real-time Diffusion-based Multiple Object Tracker with Non-linear Prediction