Abstract:Recently, scene text detection has become an active research topic in computer vision and document analysis, because of its great importance and significant challenge. However, vast majority of the existing methods detect text within local regions, typically through extracting character, word or line level candidates followed by candidate aggregation and false positive elimination, which potentially exclude the effect of wide-scope and long-range contextual cues in the scene. To take full advantage of the rich information available in the whole natural image, we propose to localize text in a holistic manner, by casting scene text detection as a semantic segmentation problem. The proposed algorithm directly runs on full images and produces global, pixel-wise prediction maps, in which detections are subsequently formed. To better make use of the properties of text, three types of information regarding text region, individual characters and their relationship are estimated, with a single Fully Convolutional Network (FCN) model. With such predictions of text properties, the proposed algorithm can simultaneously handle horizontal, multi-oriented and curved text in real-world natural images. The experiments on standard benchmarks, including ICDAR 2013, ICDAR 2015 and MSRA-TD500, demonstrate that the proposed algorithm substantially outperforms previous state-of-the-art approaches. Moreover, we report the first baseline result on the recently-released, large-scale dataset COCO-Text.

Fast and Robust Text Detection in MOOCs Videos

A new video text detection method.

A Novel Approach to Text Detection and Extraction from Videos by Discriminative Features and Density

Robust Text Detection in Natural Scene Images

Learning Markov Clustering Networks for Scene Text Detection

A Robust Color-Independent Text Detection Method from Complex Videos

A Text Detection System for Natural Scenes with Convolutional Feature Learning and Cascaded Classification

Fast and robust text detection in images and video frames

A Multi-stage Method for Chinese Text Detection in News Videos

Video text detection and segmentation for optical character recognition

Detecting both superimposed and scene text with multiple languages and multiple alignments in video

Scene Text Detection via Holistic, Multi-Channel Prediction

Look More Than Once: An Accurate Detector for Text of Arbitrary Shapes

Real-Time Text Detection and Recognition

Fast and Effective Text Detection.

OPMP: An Omnidirectional Pyramid Mask Proposal Network for Arbitrary-Shape Scene Text Detection

Intelligent Detection Method of English Text in Natural Scenes in Video

A Robust Symmetry-Based Method For Scene/Video Text Detection Through Neural Network

Automatic video superimposed text detection based on Nonsubsampled Contourlet Transform

A Robust Hierarchical Detection Method for Scene Text Based on Convolutional Neural Networks

Multi-oriented Scene Text Detection via Corner Localization and Region Segmentation