Abstract:This paper proposes a lightweight object tracking algorithm based on the transformer architecture. The joint attention module is introduced to leverage spatiotemporal context information, enhancing feature extraction capabilities. In addition, to address occlusion, the following two strategies are adopted. A position encoding generator module has been added to the transformer structure to obtain position discrimination. A dynamic template update strategy is added to increase template reliability. These two strategies greatly improve algorithm robustness while reducing computational requirements. At present, the multi‐object tracking method based on transformer generally uses its powerful self‐attention mechanism and global modelling ability to improve the accuracy of object tracking. However, most existing methods excessively rely on hardware devices, leading to an inconsistency between accuracy and speed in practical applications. Therefore, a lightweight transformer joint position awareness algorithm is proposed to solve the above problems. Firstly, a joint attention module to enhance the ShuffleNet V2 network is proposed. This module comprises the spatio‐temporal pyramid module and the convolutional block attention module. The spatio‐temporal pyramid module fuses multi‐scale features to capture information on different spatial and temporal scales. The convolutional block attention module aggregates channel and spatial dimension information to enhance the representation ability of the model. Then, a position encoding generator module and a dynamic template update strategy are proposed to solve the occlusion. Group convolution is adopted in the input sequence through position encoding generator module, with each convolution group responsible for handling the relative positional relationships of a specific range. In order to improve the reliability of the template, dynamic template update strategy is used to update the template at the appropriate time. The effectiveness of the approach is validated on the MOT16, MOT17, and MOT20 datasets.

TransMOT: Spatial-Temporal Graph Transformer for Multiple Object Tracking

Exploit the Connectivity: Multi-Object Tracking with TrackletNet

TransLink: Transformer-Based Embedding for Tracklets’ Global Link

Exploit the Connectivity

MOTR: End-to-End Multiple-Object Tracking with Transformer

Joint Spatial-Temporal and Appearance Modeling with Transformer for Multiple Object Tracking

Split and Connect: A Universal Tracklet Booster for Multi-Object Tracking

PuTR: A Pure Transformer for Decoupled and Online Multi-Object Tracking

MeMOTR: Long-Term Memory-Augmented Transformer for Multi-Object Tracking

TransCenter: Transformers With Dense Representations for Multiple-Object Tracking

FastTrackTr:Towards Fast Multi-Object Tracking with Transformers

Transformer-Based Multiple-Object Tracking via Anchor-Based-Query and Template Matching

InterTrack: Interaction Transformer for 3D Multi-Object Tracking

STMT: Spatio-temporal memory transformer for multi-object tracking

Chained-Tracker: Chaining Paired Attentive Regression Results for End-to-End Joint Multiple-Object Detection and Tracking

A transformer‐based lightweight method for multiple‐object tracking

TrackFormer: Multi-Object Tracking with Transformers

MAT: Motion-Aware Multi-Object Tracking

Spatio-Temporal Point Process for Multiple Object Tracking

Transformer Network for Multi-Person Tracking and Re-Identification in Unconstrained Environment

STMMOT: Advancing multi-object tracking through spatiotemporal memory networks and multi-scale attention pyramids