Abstract:Optical character recognition for the English text may be considered one of the most important research topics, whether, printed or handwritten. Although excellent results have been reached in the English text, there is a lack of this type of research in the Arabic text. This is because of the nature of the Arabic alphabet, and the multiplicity of forms of the same letter. Arabic handwritten character recognition (AHCR) systems involve several issues, and challenges from finding a suitable, and public Arabic handwritten text dataset phase to recognition, and classification phase passing through segmentation, and feature extraction phases. The paper objectives are: Firstly, a large, and complex Arabic handwritten characters’ dataset (HMBD) is presented for training, testing, and validation phases, as well as, discussing its collection, preparation, cleaning, and preprocessing. Secondly, we introduce a deep learning (DL) system with two convolutional neural network (CNN) architectures (named HMB1 and HMB2); with the appliance of optimization, regularization, and dropout techniques. This system can serve as a baseline for future research on handwritten Arabic text. Different performance metrics were calculated such as accuracy, recall, precision, and F1. 16 experiments were applied to the described system using HMBD, and another two datasets: CMATER, and AIA9k. Experiments’ results were captured and compared to study the effects of weight initializers, optimizers, data augmentation, and regularization on overfitting, and accuracy. He Uniform weight initializer and AdaDelta optimizer reported the highest accuracies. Data augmentation showed an improvement in the accuracies. HMB1 reported testing accuracy of 98.4% with 865,840 records using augmentation on HMBD. CMATER and AIA9k datasets were used for validating the generalization. Data augmentation was applied, and the best results were 100%, and 99.0% for testing accuracies, respectively. A cross-over validation between the described architectures, and a previous state-of-the-art architecture, and dataset was performed in two phases. First, the previous control architecture cannot generalize for the presented dataset in the current study. Second, the study described architectures generalize for the control dataset, with higher accuracies (97.3%, and 96.8% for HMB1, and HMB2, respectively), than the reported accuracy in the selected control study.

A self-attention-based deep architecture for online handwriting recognition

On the Use of Attention Mechanism in a Seq2Seq Based Approach for Off-Line Handwritten Digit String Recognition

An End-to-End, Segmentation-Free, Arabic Handwritten Recognition Model on KHATT

The State of the Art Recognize in Arabic Script through Combination of Online and Offline

Improving Accuracy and Explainability of Online Handwriting Recognition

An Effective Approach to Offline Arabic Handwriting Recognition

Refocus attention span networks for handwriting line recognition

Neural Computing for Online Arabic Handwriting Character Recognition using Hard Stroke Features Mining

Digitizing Handwriting with a Sensor Pen: A Writer-Independent Recognizer

Towards an IMU-based Pen Online Handwriting Recognizer

CNN-BiLSTM model for English Handwriting Recognition: Comprehensive Evaluation on the IAM Dataset

A Transformer-based Approach for Arabic Offline Handwritten Text Recognition

Handwritten Amharic Word Recognition With Additive Attention Mechanism

Handwriting quality analysis using online-offline models

ADOCRNet: A Deep Learning OCR for Arabic Documents Recognition

Arabic Handwritten Text for Person Biometric Identification: A Deep Learning Approach

Multi-layer Recurrent Neural Network Based Offline Arabic Handwriting Recognition

A new Arabic handwritten character recognition deep learning system (AHCR-DLS)

Novel Deep Convolutional Neural Network-Based Contextual Recognition of Arabic Handwritten Scripts

On usage of an end-to-end deep neural architecture for handwritten digit string recognition