Abstract:Over the past few decades, large archives of paper-based historical documents, such as books and newspapers, have been digitized using the Optical Character Recognition (OCR) technology. Unfortunately, this broadly used technology is error-prone, especially when an OCRed document was written hundreds of years ago. Neural networks have shown great success in solving various text processing tasks, including OCR post-correction. The main disadvantage of using neural networks for historical corpora is the lack of sufficiently large training datasets they require to learn from, especially for morphologically rich languages like Hebrew. Moreover, it is not clear what are the optimal structure and values of hyperparameters (predefined parameters) of neural networks for OCR error correction in Hebrew due to its unique features. Furthermore, languages change across genres and periods. These changes may affect the accuracy of OCR post-correction neural network models. To overcome these challenges, we developed a new multi-phase method for generating artificial training datasets with OCR errors and hyperparameters’ optimization for building an effective neural network for OCR post-correction in Hebrew. To evaluate the proposed approach, a series of experiments using several literary Hebrew corpora from various periods and genres were conducted. The obtained results demonstrate that (1) training a network on texts from a similar period dramatically improves the network's ability to fix OCR errors, (2) using the proposed error injection algorithm, based on character-level period-specific errors, minimizes the need for manually corrected data and improves the network accuracy by 9%, (3) the optimized network design improves the accuracy by 3% compared to the state-of-the-art network, and (4) the constructed optimized network outperforms neural machine translation models and industry-leading spellcheckers. The proposed methodology may have practical implications for digital humanities projects that aim to search and analyze OCRed documents in Hebrew and potentially other morphologically rich languages.

RoundTripOCR: A Data Generation Technique for Enhancing Post-OCR Error Correction in Low-Resource Devanagari Languages

Upcycle Your OCR: Reusing OCRs for Post-OCR Text Correction in Romanised Sanskrit

OCR Post Correction for Endangered Language Texts

A Novel Pipeline for Improving Optical Character Recognition through Post-processing Using Natural Language Processing

Advancing Post-OCR Correction: A Comparative Study of Synthetic Data

Neural OCR Post-Hoc Correction of Historical Corpora

A Novel Approach to Skew-Detection and Correction of English Alphabets for OCR

A Cost Efficient Approach to Correct OCR Errors in Large Document Collections

Vartani Spellcheck -- Automatic Context-Sensitive Spelling Correction of OCR-generated Hindi Text Using BERT and Levenshtein Distance

Leveraging Text Repetitions and Denoising Autoencoders in OCR Post-correction

Reference-Based Post-OCR Processing with LLM for Diacritic Languages

Optical Text Recognition in Nepali and Bengali: A Transformer-based Approach

Large Synthetic Data from the arXiv for OCR Post Correction of Historic Scientific Articles

OCR Improves Machine Translation for Low-Resource Languages

CNN-Bidirectional LSTM Based Optical Character Recognition of Sanskrit Manuscripts : A Comprehensive Systematic Literature Review

OCR Post-Processing Error Correction Algorithm using Google Online Spelling Suggestion

Statistical Learning for OCR Text Correction

Making Old Kurdish Publications Processable by Augmenting Available Optical Character Recognition Engines

Toward a Period-specific Optimized Neural Network for OCR Error Correction of Historical Hebrew Texts

Handwritten OCR for Indic Scripts: A Comprehensive Overview of Machine Learning and Deep Learning Techniques

Toward the Optimized Crowdsourcing Strategy for OCR Post-Correction