Abstract:Human video instance segmentation plays an important role in computer understanding of human activities and is widely used in video processing, video surveillance, and human modeling in virtual reality. Most current VIS methods are based on Mask-RCNN framework, where the target appearance and motion information for data matching will increase computational cost and have an impact on segmentation real-time performance; on the other hand, the existing datasets for VIS focus less on all the people appearing in the video. In this paper, to solve the problems, we develop a new method for human video instance segmentation based on single-stage detector. To tracking the instance across the video, we have adopted data association strategy for matching the same instance in the video sequence, where we jointly learn target instance appearances and their affinities in a pair of video frames in an end-to-end fashion. We have also adopted the centroid sampling strategy for enhancing the embedding extraction ability of instance, which is to bias the instance position to the inside of each instance mask with heavy overlap condition. As a result, even there exists a sudden change in the character activity, the instance position will not move out of the mask, so that the problem that the same instance is represented by two different instances can be alleviated. Finally, we collect PVIS dataset by assembling several video instance segmentation datasets to fill the gap of the current lack of datasets dedicated to human video segmentation. Extensive simulations based on such dataset has been conduct. Simulation results verify the effectiveness and efficiency of the proposed work.

Self-supervised Multi-view Multi-Human Association and Tracking

Unveiling the Power of Self-supervision for Multi-view Multi-human Association and Tracking

Beyond Traditional Driving Scenes: A Robotic-Centric Paradigm for 2D+3D Human Tracking Using Siamese Transformer Network

Multi-View Multi-Human Association With Deep Assignment Network

Multiple Human Association and Tracking from Egocentric and Complementary Top Views

Multi-person Multi-Camera Tracking for Live Stream Videos Based on Improved Motion Model and Matching Cascade

Benchmarking the Complementary-View Multi-human Association and Tracking

Multi-object tracking via discriminative appearance modeling.

MvHAAN: multi-view hierarchical attention adversarial network for person re-identification

Multiple Human Association between Top and Horizontal Views by Matching Subjects' Spatial Distributions

Deep Human-Interaction and Association by Graph-Based Learning for Multiple Object Tracking in the Wild

GMT: A Robust Global Association Model for Multi-Target Multi-Camera Tracking

Object Tracking with Multi-View Support Vector Machines.

Multiple Human Tracking Based on Multi-view Upper-Body Detection and Discriminative Learning

Special Issue on Visual Tracking

Hierarchical Multi-Supervision Multi-Interaction Graph Attention Network for Multi-Camera Pedestrian Trajectory Prediction

Human Instance Segmentation and Tracking via Data Association and Single-stage Detector

MSA-MOT: Multi-Stage Association for 3D Multimodality Multi-Object Tracking

Learning a Neural Association Network for Self-supervised Multi-Object Tracking

MHSAN: Multi-view hierarchical self-attention network for 3D shape recognition

Spatial-Temporal Multi-level Association for Video Object Segmentation