Abstract:Optical Character Recognition (OCR) on historical printings is a challenging task mainly due to the complexity of the layout and the highly variant typography. Nevertheless, in the last few years great progress has been made in the area of historical OCR, resulting in several powerful open-source tools for preprocessing, layout recognition and segmentation, character recognition and post-processing. The drawback of these tools often is their limited applicability by non-technical users like humanist scholars and in particular the combined use of several tools in a workflow. In this paper we present an open-source OCR software called OCR4all, which combines state-of-the-art OCR components and continuous model training into a comprehensive workflow. A comfortable GUI allows error corrections not only in the final output, but already in early stages to minimize error propagations. Further on, extensive configuration capabilities are provided to set the degree of automation of the workflow and to make adaptations to the carefully selected default parameters for specific printings, if necessary. Experiments showed that users with minimal or no experience were able to capture the text of even the earliest printed books with manageable effort and great quality, achieving excellent character error rates (CERs) below 0.5%. The fully automated application on 19th century novels showed that OCR4all can considerably outperform the commercial state-of-the-art tool ABBYY Finereader on moderate layouts if suitably pretrained mixed OCR models are available. The architecture of OCR4all allows the easy integration (or substitution) of newly developed tools for its main components by standardized interfaces like PageXML, thus aiming at continual higher automation for historical printings.

Survey of Post-OCR Processing Approaches

Survey of Post-OCR Processing Approaches

A survey of modern optical character recognition techniques

Reference-Based Post-OCR Processing with LLM for Diacritic Languages

A Novel Pipeline for Improving Optical Character Recognition through Post-processing Using Natural Language Processing

Advanced Digital Image Processing Technique based Optical Character Recognition of Scanned Document

Postprocessing Algorithm for the Optical Recognition of Degraded Characters

OCR accuracy improvement on document images through a novel pre-processing approach

Toward the Optimized Crowdsourcing Strategy for OCR Post-Correction

OCR Post-Processing Error Correction Algorithm using Google Online Spelling Suggestion

A Tool for Facilitating OCR Postediting in Historical Documents

Rerunning OCR: A Machine Learning Approach to Quality Assessment and Enhancement Prediction

Estimating Post-OCR Denoising Complexity on Numerical Texts

OCR4all -- An Open-Source Tool Providing a (Semi-)Automatic OCR Workflow for Historical Printings

Neural Machine Translation with BERT for Post-OCR Error Detection and Correction

At the frontiers of OCR

OCR Post Correction for Endangered Language Texts

Optimization of Image Processing Algorithms for Character Recognition in Cultural Typewritten Documents

OCR Result Optimization Based on Pattern Matching.

Implementation of OCR using Convolutional Neural Network (CNN): A Survey

OCR Context-Sensitive Error Correction Based on Google Web 1T 5-Gram Data Set