Abstract:Context-aware methods have achieved remarkable advancements in supervised scene text recognition by leveraging semantic priors from words. Considering the heterogeneity of text and background in STR, we propose that such contextual priors can be reinterpreted as the relations between textual elements, serving as effective self-supervised labels for representation learning. However, textual relations are restricted to the finite size of the dataset due to lexical dependencies, which causes over-fitting problem, thus compromising the representation quality. To address this, our work introduces a unified framework of Relational Contrastive Learning and Masked Image Modeling for STR (RCMSTR), which explicitly models the enriched textual relations. For the RCL branch, we first introduce the relational rearrangement module to cultivate new relations on the fly. Based on this, we further conduct relational contrastive learning to model the intra- and inter-hierarchical relations for frames, sub-words and words. On the other hand, MIM can naturally boost the context information via masking, where we find that the block masking strategy is more effective for STR. For the effective integration of RCL and MIM, we also introduce a novel decoupling design aimed at mitigating the impact of masked images on contrastive learning. Additionally, to enhance the compatibility of MIM with CNNs, we propose the adoption of sparse convolutions and directly sharing the weights with dense convolutions in training. The proposed RCMSTR demonstrates superior performance in various evaluation protocols for different STR-related downstream tasks, outperforming the existing state-of-the-art self-supervised STR techniques. Ablation studies and qualitative experimental results further validate the effectiveness of our method. The code and pre-trained models will be available at <a class="link-external link-https" href="https://github.com/ThunderVVV/RCMSTR" rel="external noopener nofollow">this https URL</a> .

Masked Text Modeling: A Self-Supervised Pre-training Method for Scene Text Detection

Leveraging Text Localization for Scene Text Removal via Text-aware Masked Image Modeling

Turning a CLIP Model into a Scene Text Detector

Self-Supervised Pre-training with Symmetric Superimposition Modeling for Scene Text Recognition

Mask is All You Need: Rethinking Mask R-CNN for Dense and Arbitrary-Shaped Scene Text Detection

Maskstr: Guide Scene Text Recognition Models with Masking

Incorporating Self-attention Mechanism and Multi-task Learning into Scene Text Detection

Mask Scene Text Recognizer

Relational Contrastive Learning and Masked Image Modeling for Scene Text Recognition

TAMC: Textual Alignment and Masked Consistency for Open-Vocabulary 3D Scene Understanding

MaskOCR: Text Recognition with Masked Encoder-Decoder Pretraining

StrucTexTv2: Masked Visual-Textual Prediction for Document Image Pre-training

Self-supervised Scene Text Segmentation with Object-centric Layered Representations Augmented by Text Regions

OPMP: An Omnidirectional Pyramid Mask Proposal Network for Arbitrary-Shape Scene Text Detection

MaskCLIP: Masked Self-Distillation Advances Contrastive Language-Image Pretraining

Texts As Lines: Text Detection with Weak Supervision

Class-Aware Mask-guided feature refinement for scene text recognition

Real-Time Text Detection with Similar Mask in Traffic, Industrial, and Natural Scenes

Train No Evil: Selective Masking for Task-Guided Pre-Training

MTA-CLIP: Language-Guided Semantic Segmentation with Mask-Text Alignment

Dual Masked Modeling for Weakly-Supervised Temporal Boundary Discovery