Abstract:Video text spotting (VTS) aims to simultaneously localize, recognize and track text instances in videos. To address the limited recognition capability of end-to-end methods, recent methods track the zero-shot results of state-of-the-art image text spotters directly, and achieve impressive performance. However, owing to the domain gap between different datasets, these methods usually obtain limited tracking trajectories on extreme dataset. Fine-tuning transformer-based text spotters on specific datasets could yield performance enhancements, albeit at the expense of considerable training resources. In this paper, we propose a Language Collaboration and Glyph Perception Model, termed LOGO, an innovative framework designed to enhance the performance of conventional text spotters. To achieve this goal, we design a language synergy classifier (LSC) to explicitly discern text instances from background noise in the recognition stage. Specially, the language synergy classifier can output text content or background code based on the legibility of text regions, thus computing language scores. Subsequently, fusion scores are computed by taking the average of detection scores and language scores, and are utilized to re-score the detection results before tracking. By the re-scoring mechanism, the proposed LSC facilitates the detection of low-resolution text instances while filtering out text-like regions. Moreover, the glyph supervision is introduced to enhance the recognition accuracy of noisy text regions. In addition, we propose the visual position mixture module, which can merge the position information and visual features efficiently, and acquire more discriminative tracking features. Extensive experiments on public benchmarks validate the effectiveness of the proposed method.

Multi-modal Sign Language Spotting by Multi/One-Shot Learning.

SLV: Spatial Likelihood Voting for Weakly Supervised Object Detection

Sign Language Video Retrieval with Free-Form Textual Queries

Watch, read and lookup: learning to spot signs from multiple supervisors

Mlts: A Multi-Language Scene Text Spotter

Skeleton Aware Multi-modal Sign Language Recognition

OS-MSL: One Stage Multimodal Sequential Link Framework for Scene Segmentation and Classification

StepNet: Spatial-temporal Part-aware Network for Isolated Sign Language Recognition

Multi-Modal Zero-Shot Sign Language Recognition

LOGO: Video Text Spotting with Language Collaboration and Glyph Perception Model

DualSign: Semi-Supervised Sign Language Production with Balanced Multi-Modal Multi-Task Dual Transformation

Multi-View Spatial-Temporal Network for Continuous Sign Language Recognition

Multi-Stream Keypoint Attention Network for Sign Language Recognition and Translation

Sign Language Recognition with Multi-modal Features.

Natural Language-Assisted Sign Language Recognition

Sign Language Recognition via Skeleton-Aware Multi-Model Ensemble

Spatial-Temporal Multi-Cue Network for Continuous Sign Language Recognition

SignVTCL: Multi-Modal Continuous Sign Language Recognition Enhanced by Visual-Textual Contrastive Learning

Improving Continuous Sign Language Recognition with Adapted Image Models

Using an LLM to Turn Sign Spottings into Spoken Language Sentences