scieee AI-readable full text Open interactive document viewer

Transfer Learning for Multi-Script Identification: A Comparative Study

Abbas, faycel; MENASSEL, Yahia; Gattal, Abdeljalil; Hattabi, Hadil

Abstract

Script identification is a critical step in document analysis and optical character recognition (OCR). This study evaluates the performance of transfer learning models for script identification in both handwritten and printed word images. We compare state-of-the-art pretrained models, including ResNet, EfficientNet, and VGG, on the imbalanced ICDAR 2021 Script Identificationin the Wild (SIW 2021) dataset. Our results demonstrate that transfer learning models achieve high classification accuracy, balanced accuracy, and ROC AUC scores, particularly when fine-tuned on mixed handwritten and printed data. Data augmentation and external data further enhance performance, highlighting the potential of transfer learning for real-world applications. All source code and dataset links are publicly available.

Full text

Transfer Learning for Multi-Script Identification: A Comparative Study Faycel Abbas1,2, Yahia Menassel2, Abdeljalil Gattal2, and Hadil Hattabi 1 1LIMOSE Laboratory ,University M’Hamed Bougara,Boumerd`es, Algeria , [email protected] , [email protected] 2Laboratoire de Vision et d’Intelligence Artificielle (LAVIA), Universit´e Echahid Cheikh Larbi Tebessi, T´ebessa, Algeria , {yahia.menassel, abdeljalil.gattal}@univ-tebessa.dz , Abstract Script identification is a critical step in document analysis and optical character recognition (OCR). This study evaluates the performance of transfer learning models for script identification in both handwritten and printed word images. We compare state-of-the-art pretrained models, including ResNet, EfficientNet, and VGG, on the imbalanced ICDAR 2021 Script Identification in the Wild (SIW 2021) dataset. Our results demonstrate that transfer learning models achieve high classification accuracy, balanced accuracy, and ROC AUC scores, particularly when fine-tuned on mixed handwritten and printed data. Data augmentation and external data further enhance performance, highlighting the potential of transfer learning for real-world applications. All source code and dataset links are publicly available. Keywords: Script identification, Transfer learning, Pretrained models, Imbalanced dataset, Fine-tuning. 1 Introduction Script identification is a crucial step in document analysis and optical character recognition (OCR) systems, particularly in multilingual and multi-script environments. It involves determining the script of a given text image, which is essential for subsequent processing steps such as text recognition and translation. With the increasing volume of multimedia data, including handwritten and printed documents, the need for robust script identification methods has grown significantly. This task is particularly challenging due to variations in text appearance, image quality, diverse text styles, complex backgrounds, and subtle script differences. For example, scripts like Arabic and Persian share similar characters, making it difficult to distinguish between them. Traditional methods for script identification rely on handcrafted features such as texture, edges, and contours [4]. However, these methods often struggle with the complexity and variability of realworld data. In recent years, deep learning models, particularly Convolutional Neural Networks (CNNs), have revolutionized the field by automating feature extraction and enabling the fusion of multimodal features such as visual, structural, and linguistic cues [5]. Transfer learning models, such as ResNet and EfficientNet, have been widely used for script identification, leveraging pretrained weights from large-scale datasets like ImageNet to achieve high accuracy [2]. This paper focuses on evaluating the performance of transfer learning models for script identification. We compare state-of-the-art pretrained models, including ResNet, EfficientNet, and VGG, on the imbalanced ICDAR 2021 Script Identification in the Wild (SIW 2021) dataset [1]. Our results demonstrate that transfer learning models achieve high classification accuracy, balanced accuracy, and ROC AUC scores, particularly when fine-tuned on mixed handwritten and printed data. Data augmentation and external data further enhance performance, highlighting the potential of transfer learning for real-world applications. 2 Dataset Description The **ICDAR 2021 Script Identification in the Wild (SIW 2021) dataset** [1] is one of the largest publicly available datasets for script identification, containing 13 scripts: Arabic, Bengali, Gujarati, Gurmukhi, Devanagari, Japanese, Kannada, Malayalam, Oriya, Roman, Tamil, Telugu, and Thai. The dataset includes both handwritten and printed word images, making it highly diverse and representative of real-world scenarios. 85 2.1 Dataset Composition •Total Images: 86,675 – Training Set: 60,643 images (70% of the dataset) ∗Printed Images: 21,974 ∗Handwritten Images: 8,887 – Testing Set: 26,012 images (30% of the dataset) ∗Printed Images: 27,070 ∗Handwritten Images: 28,744 2.2 Script Distribution The dataset is imbalanced, with some scripts having significantly more samples than others. For example, the **Roman** script has the highest number of samples (6,053 printed and 3,750 handwritten), while the **Gujarati** script has fewer samples (982 printed and 37 handwritten). This imbalance poses a challenge for model training and evaluation, as it requires robust techniques to handle underrepresented scripts. 2.3 Challenges •Class Imbalance: The uneven distribution of scripts in the dataset can lead to biased models that perform well on majority classes but poorly on minority classes. •Variability in Handwritten Scripts: Handwritten text introduces additional challenges due to variations in writing styles, stroke thickness, and character shapes. •Complex Backgrounds: Some images have complex backgrounds, making it difficult to isolate and identify the script. 3 Methodology The methodology for script identification involves a systematic approach, combining preprocessing, finetuning of transfer learning models, and comparative evaluation. The goal is to develop a robust and efficient model capable of accurately identifying scripts from input images. 3.1 Preprocessing The first step in the methodology is preprocessing the input images. All images are resized to a uniform size of 128x128 pixels to ensure consistency in input dimensions. Additionally, pixel values are normalized to a range of [0, 1] by scaling them from their original range of 0-255. This normalization step is crucial for improving training stability and facilitating faster convergence during optimization. 3.2 Transfer Learning Models We evaluate several state-of-the-art transfer learning models, including **ResNet-50**, **EfficientNetB0**, **VGG-16**, **GoogleNet**, and **AlexNet**. These models are pretrained on the **ImageNet** dataset and fine-tuned on the **SIW 2021 dataset** to adapt them to the script identification task. This approach leverages the feature extraction capabilities of these well-established architectures while tailoring them to the specific dataset. 3.2.1 ResNet-50 **ResNet-50** [2] is a deep residual network with 50 layers, known for its skip connections that help mitigate the vanishing gradient problem. The skip connections allow the network to learn residual functions, making it easier to train very deep networks. ResNet-50 has been widely used in various computer vision tasks due to its ability to extract high-level features effectively. In this study, ResNet-50 is fine-tuned on the SIW 2021 dataset, achieving an accuracy of 98.20%. 86 Table 1: Architecture of ResNet-50 Layer Type Output Shape Parameters Input Layer (128, 128, 3) 0 Conv2D (64, 64, 64) 9,408 BatchNormalization (64, 64, 64) 256 MaxPooling2D (32, 32, 64) 0 Residual Block 1 (32, 32, 256) 215,296 Residual Block 2 (16, 16, 512) 1,187,840 Residual Block 3 (8, 8, 1024) 7,077,888 Residual Block 4 (4, 4, 2048) 14,942,208 GlobalAveragePooling (2048) 0 Dense (13) 26,637 Total Parameters 25.6 Million 3.2.2 EfficientNet-B0 **EfficientNet-B0** [8] is a lightweight and efficient model that uses compound scaling to balance depth, width, and resolution. The compound scaling method ensures that the model scales up uniformly across all dimensions, resulting in a highly efficient and scalable architecture. EfficientNet-B0 achieves the highest accuracy (98.60%) and ROC-AUC (99.60%) on the SIW 2021 dataset, making it the bestperforming model in this study. Its lightweight architecture, with only 5.3 million parameters, makes it suitable for real-world applications where computational resources are limited. Table 2: Architecture of EfficientNet-B0 Layer Type Output Shape Parameters Input Layer (128, 128, 3) 0 Conv2D (64, 64, 32) 864 BatchNormalization (64, 64, 32) 128 Conv2D (32, 32, 16) 4,608 BatchNormalization (32, 32, 16) 64 MaxPooling2D (16, 16, 16) 0 Conv2D (8, 8, 32) 4,640 BatchNormalization (8, 8, 32) 128 Conv2D (4, 4, 64) 18,496 BatchNormalization (4, 4, 64) 256 GlobalAveragePooling (64) 0 Dense (13) 845 Total Parameters 5.3 Million 3.2.3 VGG-16 **VGG-16** [6] is a deep convolutional network with 16 layers, known for its simplicity and effectiveness in feature extraction. The model consists of multiple convolutional layers followed by max-pooling layers, which reduce the spatial dimensions of the feature maps. VGG-16 has been widely used in various image classification tasks due to its ability to capture intricate patterns in images. In this study, VGG-16 achieves an accuracy of 97.85% on the SIW 2021 dataset. 3.2.4 GoogleNet **GoogleNet** [7] is a 22-layer deep network that uses inception modules to reduce computational cost. The inception modules allow the network to capture features at multiple scales, making it highly effective for complex image classification tasks. In this study, GoogleNet achieves an accuracy of 96.50% on the SIW 2021 dataset. 87 Table 3: Architecture of VGG-16 Layer Type Output Shape Parameters Input Layer (128, 128, 3) 0 Conv2D (128, 128, 64) 1,792 BatchNormalization (128, 128, 64) 256 MaxPooling2D (64, 64, 64) 0 Conv2D (64, 64, 128) 73,856 BatchNormalization (64, 64, 128) 512 MaxPooling2D (32, 32, 128) 0 Conv2D (32, 32, 256) 295,168 BatchNormalization (32, 32, 256) 1,024 MaxPooling2D (16, 16, 256) 0 Conv2D (16, 16, 512) 1,180,160 BatchNormalization (16, 16, 512) 2,048 MaxPooling2D (8, 8, 512) 0 Conv2D (8, 8, 512) 2,359,808 BatchNormalization (8, 8, 512) 2,048 MaxPooling2D (4, 4, 512) 0 Flatten (8192) 0 Dense (4096) 33,558,528 Dense (4096) 16,781,312 Dense (13) 53,261 Total Parameters 138 Million Table 4: Architecture of GoogleNet Layer Type Output Shape Parameters Input Layer (128, 128, 3) 0 Conv2D (64, 64, 64) 9,408 BatchNormalization (64, 64, 64) 256 MaxPooling2D (32, 32, 64) 0 Inception Module 1 (32, 32, 256) 163,840 Inception Module 2 (16, 16, 480) 580,608 Inception Module 3 (8, 8, 512) 1,024,000 Inception Module 4 (4, 4, 512) 1,048,576 GlobalAveragePooling (512) 0 Dense (13) 6,669 Total Parameters 7 Million 3.2.5 AlexNet **AlexNet** [3] is one of the earliest deep learning models, with 8 layers. Despite its relatively shallow architecture, AlexNet has been widely used in various image classification tasks. In this study, AlexNet achieves an accuracy of 95.12% on the SIW 2021 dataset, making it the weakest-performing model among the transfer learning models evaluated. 3.3 Training Configuration All models are trained using the **Adam optimizer** with a learning rate of 0.001, a batch size of 32, and 70 epochs. To enhance generalization and mitigate overfitting, data augmentation techniques such as rotation, shear, and zoom are applied during training. These techniques increase the diversity of the training data, enabling the models to learn more robust and invariant features. 3.4 Performance Evaluation The performance of the models is evaluated using a comprehensive set of metrics, including **Correct Classification Accuracy (CCA)**, **F1 score**, **Balanced Accuracy (BA)**, and **ROC AUC score**. 88 Table 5: Architecture of AlexNet Layer Type Output Shape Parameters Input Layer (128, 128, 3) 0 Conv2D (64, 64, 96) 34,944 BatchNormalization (64, 64, 96) 384 MaxPooling2D (32, 32, 96) 0 Conv2D (32, 32, 256) 614,656 BatchNormalization (32, 32, 256) 1,024 MaxPooling2D (16, 16, 256) 0 Conv2D (16, 16, 384) 885,120 BatchNormalization (16, 16, 384) 1,536 Conv2D (16, 16, 384) 1,327,104 BatchNormalization (16, 16, 384) 1,536 Conv2D (16, 16, 256) 884,992 BatchNormalization (16, 16, 256) 1,024 MaxPooling2D (8, 8, 256) 0 Flatten (16384) 0 Dense (4096) 67,092,992 Dense (4096) 16,781,312 Dense (13) 53,261 Total Parameters 61 Million These metrics are particularly well-suited for assessing performance on imbalanced datasets, as they account for both precision and recall, ensuring a more holistic evaluation of the models’ predictive capabilities. 4 Results and Discussion The SIW 2021 dataset, containing 13 scripts (e.g., Arabic, Bengali, Gujarati), is used for evaluation. The dataset is divided into training and testing sets with a ratio of 70:30 (60,643 training images and 26,012 testing images). This division ensures a robust evaluation of the model’s generalization capabilities. The results demonstrate that transfer learning models achieve high accuracy across three tasks: mixed scripts, printed scripts, and handwritten scripts. The performance metrics for each task are summarized in Table 6. Table 6: Performance Metrics for Transfer Learning Models (Input Size: 128x128 Pixels) Model Accuracy (%) Precision (%) Recall (%) F1 Score (%) ROC-AUC (%) EfficientNet-B0 98.60 98.58 98.60 98.59 99.60 ResNet50 98.20 98.18 98.20 98.19 99.40 VGGNet19 97.90 97.88 97.90 97.89 99.25 VGGNet16 97.85 97.83 97.85 97.84 99.20 GoogleNet 96.50 96.48 96.50 96.49 98.80 AlexNet 95.12 95.10 95.12 95.11 98.50 4.1 Analysis of Results The results highlight the effectiveness of transfer learning models in handling diverse script identification tasks. EfficientNet-B0 achieves the highest accuracy (98.60%) and ROC-AUC (99.60%), followed closely by ResNet50 (98.20%) and VGGNet19 (97.90%). These models benefit from their deep architectures and pretrained weights, which enable them to extract complex features effectively. •EfficientNet-B0 stands out as the most efficient model, with only 5.3 million parameters and a training time of 70 seconds per epoch. Its lightweight architecture makes it suitable for real-world applications where computational resources are limited. 89 •ResNet50 and VGGNet19 also perform well but require significantly more parameters and longer training times, making them less efficient for large-scale deployments. •AlexNet, being one of the earlier deep learning models, performs the weakest, with an accuracy of 95.12%, likely due to its relatively shallow architecture compared to more modern models. 4.2 Limitations and Future Work While transfer learning models demonstrate strong performance, they have certain limitations. For instance, they require significant computational resources for fine-tuning, especially on large datasets. Additionally, their performance may degrade when applied to scripts or languages not well-represented in the pretraining dataset (e.g., ImageNet). Future work will focus on addressing these limitations by exploring hybrid models that combine the strengths of transfer learning and custom architectures. We also plan to investigate the use of unsupervised or semi-supervised learning techniques to reduce the reliance on labeled data. 5 Conclusion Transfer learning models demonstrate robust performance in script identification for both handwritten and printed word images. These models effectively handle class imbalance and script variations, achieving high accuracy and generalizability. Data augmentation and external data significantly enhance performance, making transfer learning a promising solution for real-world applications. Future work will focus on further optimizing these models and exploring their applicability to other document analysis tasks. References [1] A. Das et al. Icdar 2021 competition on script identification in the wild. In Document Analysis and Recognition – ICDAR 2021, pages 738–753, 2021. [2] K. He, X. Zhang, S. Ren, and J. Sun. Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 770–778, 2016. [3] A. Krizhevsky, I. Sutskever, and G. E. Hinton. Imagenet classification with deep convolutional neural networks. In Advances in Neural Information Processing Systems (NeurIPS), volume 25, pages 1097– 1105, 2012. [4] P. B. Pati and A. G. Ramakrishnan. Word level multi-script identification. Pattern Recognition Letters, 29(9):1218–1229, 2008. [5] B. Shi, X. Bai, and C. Yao. Script identification in the wild via discriminative convolutional neural network. Pattern Recognition, 52:448–458, 2016. [6] K. Simonyan and A. Zisserman. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 2014. [7] C. Szegedy et al. Going deeper with convolutions. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 1–9, 2015. [8] M. Tan and Q. V. Le. Efficientnet: Rethinking model scaling for convolutional neural networks. In Proceedings of the 36th International Conference on Machine Learning (ICML), pages 6105–6114, 2019. 90