Abstract:Context-aware methods have achieved remarkable advancements in supervised scene text recognition by leveraging semantic priors from words. Considering the heterogeneity of text and background in STR, we propose that such contextual priors can be reinterpreted as the relations between textual elements, serving as effective self-supervised labels for representation learning. However, textual relations are restricted to the finite size of the dataset due to lexical dependencies, which causes over-fitting problem, thus compromising the representation quality. To address this, our work introduces a unified framework of Relational Contrastive Learning and Masked Image Modeling for STR (RCMSTR), which explicitly models the enriched textual relations. For the RCL branch, we first introduce the relational rearrangement module to cultivate new relations on the fly. Based on this, we further conduct relational contrastive learning to model the intra- and inter-hierarchical relations for frames, sub-words and words. On the other hand, MIM can naturally boost the context information via masking, where we find that the block masking strategy is more effective for STR. For the effective integration of RCL and MIM, we also introduce a novel decoupling design aimed at mitigating the impact of masked images on contrastive learning. Additionally, to enhance the compatibility of MIM with CNNs, we propose the adoption of sparse convolutions and directly sharing the weights with dense convolutions in training. The proposed RCMSTR demonstrates superior performance in various evaluation protocols for different STR-related downstream tasks, outperforming the existing state-of-the-art self-supervised STR techniques. Ablation studies and qualitative experimental results further validate the effectiveness of our method. The code and pre-trained models will be available at <a class="link-external link-https" href="https://github.com/ThunderVVV/RCMSTR" rel="external noopener nofollow">this https URL</a> .

CLIP4STR: A Simple Baseline for Scene Text Recognition with Pre-trained Vision-Language Model

CLIP-Llama: A New Approach for Scene Text Recognition with a Pre-Trained Vision-Language Model and a Pre-Trained Language Model

Symmetrical Linguistic Feature Distillation with CLIP for Scene Text Recognition

SVIPTR: Fast and Efficient Scene Text Recognition with Vision Permutable Extractor

Decoder Pre-Training with only Text for Scene Text Recognition

Multi-Granularity Prediction for Scene Text Recognition

SpaceCLIP: A Vision-Language Pretraining Framework With Spatial Reconstruction On Text

SVTR: Scene Text Recognition with a Single Visual Model

ViTSTR-Transducer: Cross-Attention-Free Vision Transformer Transducer for Scene Text Recognition

Flexible scene text recognition based on dual attention mechanism

Relational Contrastive Learning and Masked Image Modeling for Scene Text Recognition

Linguistic More: Taking a Further Step Toward Efficient and Accurate Scene Text Recognition.

Multi-modal In-Context Learning Makes an Ego-evolving Scene Text Recognizer

How Much Can CLIP Benefit Vision-and-Language Tasks?

Turning a CLIP Model into a Scene Text Detector

Multi-Granularity Prediction with Learnable Fusion for Scene Text Recognition

Multimodal Visual-Semantic Representations Learning for Scene Text Recognition

Self-Calibrated CLIP for Training-Free Open-Vocabulary Segmentation

Scene Text Recognition with Self-supervised Contrastive Predictive Coding

ViTEraser: Harnessing the Power of Vision Transformers for Scene Text Removal with SegMIM Pretraining

CLIP-ViP: Adapting Pre-trained Image-Text Model to Video-Language Representation Alignment