Abstract:Convolutional neural networks (CNNs) have been successfully applied to the single target tracking task in recent years. Generally, training a deep CNN model requires numerous labeled training samples, and the number and quality of these samples directly affect the representational capability of the trained model. However, this approach is restrictive in practice, because manually labeling such a large number of training samples is time-consuming and prohibitively expensive. In this article, we propose an active learning method for deep visual tracking, which selects and annotates the unlabeled samples to train the deep CNN model. Under the guidance of active learning, the tracker based on the trained deep CNN model can achieve competitive tracking performance while reducing the labeling cost. More specifically, to ensure the diversity of selected samples, we propose an active learning method based on multiframe collaboration to select those training samples that should be and need to be annotated. Meanwhile, considering the representativeness of these selected samples, we adopt a nearest-neighbor discrimination method based on the average nearest-neighbor distance to screen isolated samples and low-quality samples. Therefore, the training samples' subset selected based on our method requires only a given budget to maintain the diversity and representativeness of the entire sample set. Furthermore, we adopt a Tversky loss to improve the bounding box estimation of our tracker, which can ensure that the tracker achieves more accurate target states. Extensive experimental results confirm that our active-learning-based tracker (ALT) achieves competitive tracking accuracy and speed compared with state-of-the-art trackers on the seven most challenging evaluation benchmarks. Project website: https://sites.google.com/view/altrack/.

Visual Tracking Using Online Deep Reinforcement Learning with Heatmap

Real-time visual tracking by deep reinforced decision making

Deep Reinforcement Learning With Iterative Shift For Visual Tracking

Dynamical Hyperparameter Optimization Via Deep Reinforcement Learning in Tracking

End-to-end Active Object Tracking Via Reinforcement Learning

Tracking as Online Decision-Making: Learning a Policy from Streaming Videos with Reinforcement Learning

Meta-Tracker: Fast and Robust Online Adaptation for Visual Object Trackers

End-to-End Active Object Tracking and Its Real-World Deployment Via Reinforcement Learning

Empowering Embodied Visual Tracking with Visual Foundation Models and Offline RL

Revisiting Jump-Diffusion Process for Visual Tracking: A Reinforcement Learning Approach

Learning Motion-Aware Policies for Robust Visual Tracking

Learning a Deep Compact Image Representation for Visual Tracking

Recursive Least-Squares Estimator-Aided Online Learning for Visual Tracking

Deep visual tracking: Review and experimental comparison

AEVRNet: Adaptive Exploration Network with Variance Reduced Optimization for Visual Tracking

Deep Reinforcement Learning-Based DQN Agent Algorithm for Visual Object Tracking in a Virtual Environmental Simulation

Learning reinforced attentional representation for end-to-end visual tracking

Enhancing Continuous Control of Mobile Robots for End-to-End Visual Active Tracking

Robust Visual Tracking Method via Deep Learning

Hierarchical Tracking by Reinforcement Learning-Based Searching and Coarse-to-Fine Verifying

Active Learning for Deep Visual Tracking