Abstract:In the booming video era, video segmentation attracts increasing research attention in the multimedia community. Semi-supervised video object segmentation (VOS) aims at segmenting objects in all target frames of a video, given annotated object masks of reference frames. Most existing methods build pixel-wise reference-target correlations and then perform pixel-wise tracking to obtain target masks. Due to neglecting object-level cues, pixel-level approaches make the tracking vulnerable to perturbations, and even indiscriminate among similar objects. Towards robust VOS, the key insight is to calibrate the representation and mask of each specific object to be expressive and discriminative. Accordingly, we propose a new deep network, which can adaptively construct object representations and calibrate object masks to achieve stronger robustness. First, we construct the object representations by applying an adaptive object proxy (AOP) aggregation method, where the proxies represent arbitrary-shaped segments at multi-levels for reference. Then, prototype masks are initially generated from the reference-target correlations based on AOP. Afterwards, such proto-masks are further calibrated through network modulation, conditioning on the object proxy representations. We consolidate this conditional mask calibration process in a progressive manner, where the object representations and proto-masks evolve to be discriminative iteratively. Extensive experiments are conducted on the standard VOS benchmarks, YouTube-VOS-18/19 and DAVIS-17. Our model achieves the state-of-the-art performance among existing published works, and also exhibits superior robustness against perturbations. Our project repo is at https://github.com/JerryX1110/Robust-Video-Object-Segmentation

Learning to Segment Video Object with Accurate Boundaries.

Delving Deeper into Mask Utilization in Video Object Segmentation

Fast Real-Time Video Object Segmentation with a Tangled Memory Network

Learning Quality-aware Dynamic Memory for Video Object Segmentation

Fast Video Object Segmentation Via Dynamic Targeting Network

Spatiotemporal Graph Neural Network Based Mask Reconstruction for Video Object Segmentation

Rethinking Video Segmentation with Masked Video Consistency: Did the Model Learn as Intended?

Attention-guided Temporally Coherent Video Object Matting

Fast and Accurate Online Video Object Segmentation Via Tracking Parts.

Towards Robust Video Object Segmentation with Adaptive Object Calibration

Fusion target attention mask generation network for video segmentation

Supervised Edge Attention Network for Accurate Image Instance Segmentation

CenterMask: Real-Time Anchor-Free Instance Segmentation

CRNet: Collaborative Refinement Network for Self-Supervised Video Object Segmentation

Boundary-Aware Network for Fast and High-Accuracy Portrait Segmentation

MacFormer: Semantic Segmentation with Fine Object Boundaries

Unified Mask Embedding and Correspondence Learning for Self-Supervised Video Segmentation

BorderPointsMask: One-stage Instance Segmentation with Boundary Points Representation.

MSN: Efficient Online Mask Selection Network for Video Instance Segmentation

Spatial Feature Calibration and Temporal Fusion for Effective One-stage Video Instance Segmentation