Abstract:Video text spotting (VTS) aims to simultaneously localize, recognize and track text instances in videos. To address the limited recognition capability of end-to-end methods, recent methods track the zero-shot results of state-of-the-art image text spotters directly, and achieve impressive performance. However, owing to the domain gap between different datasets, these methods usually obtain limited tracking trajectories on extreme dataset. Fine-tuning transformer-based text spotters on specific datasets could yield performance enhancements, albeit at the expense of considerable training resources. In this paper, we propose a Language Collaboration and Glyph Perception Model, termed LOGO, an innovative framework designed to enhance the performance of conventional text spotters. To achieve this goal, we design a language synergy classifier (LSC) to explicitly discern text instances from background noise in the recognition stage. Specially, the language synergy classifier can output text content or background code based on the legibility of text regions, thus computing language scores. Subsequently, fusion scores are computed by taking the average of detection scores and language scores, and are utilized to re-score the detection results before tracking. By the re-scoring mechanism, the proposed LSC facilitates the detection of low-resolution text instances while filtering out text-like regions. Moreover, the glyph supervision is introduced to enhance the recognition accuracy of noisy text regions. In addition, we propose the visual position mixture module, which can merge the position information and visual features efficiently, and acquire more discriminative tracking features. Extensive experiments on public benchmarks validate the effectiveness of the proposed method.

Towards Accurate Video Text Spotting with Text-wise Semantic Reasoning

A new video text detection method.

You Only Recognize Once: Towards Fast Video Text Spotting

Boosting Video-Text Retrieval with Explicit High-Level Semantics

Real-time End-to-End Video Text Spotter with Contrastive Representation Learning

LOGO: Video Text Spotting with Language Collaboration and Glyph Perception Model

VimTS: A Unified Video and Image Text Spotter for Enhancing the Cross-domain Generalization

FREE: A Fast and Robust End-to-End Video Text Spotter

Video Text Tracking With a Spatio-Temporal Complementary Model

A Bilingual, OpenWorld Video Text Dataset and End-to-end Video Text Spotter with Transformer

Contrastive Learning of Semantic and Visual Representations for Text Tracking

AE TextSpotter: Learning Visual and Linguistic Representation for Ambiguous Text Spotting

SPTS v2: Single-Point Scene Text Spotting

End-to-End Video Text Spotting with Transformer

Learning Spatial-Semantic Features for Robust Video Object Segmentation

ASTS: A Unified Framework for Arbitrary Shape Text Spotting.

Semantic Lens: Instance-Centric Semantic Alignment for Video Super-Resolution

SwinTextSpotter v2: Towards Better Synergy for Scene Text Spotting

ESTextSpotter: Towards Better Scene Text Spotting with Explicit Synergy in Transformer

TextFormer: A Query-based End-to-End Text Spotter with Mixed Supervision