Abstract:Arbitrary-shaped text detection in natural images is a challenging task due to the complexity of the background and the diversity of text properties. The difficulty lies in two aspects: accurate separation of adjacent texts and sufficient text feature representation. To handle these problems, we consider text detection as instance segmentation and propose a novel text detection framework, which jointly learns semantic segmentation and a pixel affinity pyramid in a unified fully convolutional network. Specifically, the pixel affinity pyramid is proposed to encode multi-scale instance affiliation relationships of pixels, which is not only robust to varying shapes of text but also provides an accurate boundary description for separating closely located texts. In the inference phase, a simple but effective post-processing is presented to reconstruct text instances from the semantic segmentation results under the guidance of the learned pixel affinity pyramid, achieving good accuracy and efficiency. Furthermore, to enhance the representation of text features in the neural network, two modules — the Region Enhancement Module (REM) and Attentional Fusion Module (AFM) — are proposed. The REM models the semantic correlations of regional features to enhance the features from the text area, which effectively suppresses false-positive detection. The AFM adaptively fuses multi-scale textual information through an attention mechanism to obtain abundant text semantic features, which benefits multi-sized text detection. Extensive ablation experiments are conducted demonstrating the effectiveness of the REM and AFM. Evaluation results on standard benchmarks, including Total-Text, ICDAR2015, SCUT-CTW1500, and MSRA-TD500, show that our method surpasses most existing text detectors and achieves state-of-the-art performance, denoting its superior capability in detecting arbitrary-shaped texts.

Learning and Fusing Multi-Scale Representations for Accurate Arbitrary-Shaped Scene Text Recognition.

A Multi-Scale Natural Scene Text Detection Method Based on Attention Feature Extraction and Cascade Feature Fusion

Accurate Scene Text Detection Via Scale-Aware Data Augmentation and Shape Similarity Constraint

MASTER: Multi-Aspect Non-local Network for Scene Text Recognition

Scene Text Recognition from Two-Dimensional Perspective

A Feasible Framework for Arbitrary-Shaped Scene Text Recognition

Multi-Spectral Fusion Based Approach for Arbitrarily Oriented Scene Text Detection in Video Images

Residual Dual Scale Scene Text Spotting by Fusing Bottom-Up and Top-Down Processing

Real-Time Scene Text Detection With Differentiable Binarization and Adaptive Scale Fusion

Multi-orientation scene text detection with scale-guided regression

A New Perspective for Flexible Feature Gathering in Scene Text Recognition Via Character Anchor Pooling

Region-aware Arbitrary-shaped Text Detection with Progressive Fusion

Transforming Scene Text Detection and Recognition: A Multi-Scale End-to-End Approach With Transformer Framework

Robust Scene Text Recognition Through Adaptive Image Enhancement

Scale based region growing for scene text detection.

Scale-Residual Learning Network for Scene Text Detection

Efficient Neural Network for Text Recognition in Natural Scenes Based on End-to-End Multi-Scale Attention Mechanism

Class-Aware Mask-guided feature refinement for scene text recognition

Learning Pixel Affinity Pyramid for Arbitrary-Shaped Text Detection

A holistic representation guided attention network for scene text recognition

ASTS: A Unified Framework for Arbitrary Shape Text Spotting.