scieee AI-readable full text Open interactive document viewer

PraNet-based Gastrointestinal Polyp Segmentation on Heterogeneous Datasets: Study of Augmentation Impact on Performance

Cristea, Daniela-Maria; Sima, Ioan; Sınitın, Vladimir; Iantovics, Laszlo Barna

Abstract

Deep Learning (DL) in the healthcare is exemplified by specific applications such as colorectal polyp segmentationusing the Parallel Reverse Attention Network (PraNet) model. The main objective of this research consistedon investigating how data augmentation impacts the performance of the PraNet model across heterogeneousdatasets. It is proposed a data augmentation methodology with properties to conserve the original image quality.Seven public polyp image datasets with a unified normalization pipeline and specific augmentation techniqueswere applied to expand the training data. Four experiments were conducted: (1) baseline training on each normalizeddataset; (2) training on each dataset with augmentation; (3) combined training on all normalized datasets,tested on an independent dataset to assess generalization; and (4) combined training on all normalized plusaugmented data, tested on the same independent set to examine augmentation-induced bias. The results showthat the specific augmentation substantially improves in-domain segmentation accuracy under five metrics Dice,IoU, MAE, S-measure and E-measure. Average Dice score gains of 0.20–0.25 on small datasets, and trainingon the aggregated multi-source data yields strong generalization (Dice ≈ 0.81) on the independent set. Whenaugmentation is added to combined training, performance on the independent test slightly decreases, the Dicescore decreases from 0.81 to 0.77, suggesting a slight decrease in segmentation performance under varying testconditions, indicating potential overfitting or distribution bias introduced by excessive augmentation. Combiningmultiple datasets resulted in a model with strong generalization to an independent dataset, where real datadiversity can effectively limit the bias.

Full text

ScienceDirect Available online at www.sciencedirect.com Procedia Computer Science 270 (2025) 4155–4164 1877-0509 © 2025 The Authors. Published by Elsevier B.V. This is an open access article under the CC BY-NC-ND license (https://creativecommons.org/licenses/by-nc-nd/4.0) Peer-review under responsibility of the scientific committee of the KES International. 10.1016/j.procs.2025.09.540 Available online at www.sciencedirect.com Procedia Computer Science 00 (2025) 000–000 www.elsevier.com/locate/procedia 29th International Conference on Knowledge-Based and Intelligent Information & Engineering Systems (KES 2025) PraNet-based Gastrointestinal Polyp Segmentation on Heterogeneous Datasets: Study of Augmentation Impact on Performance Daniela-Maria Cristeaa,c,∗, Ioan Simab,S ˆ ınit ,ˆ ın Vladimira, Laszlo Barna Iantovicsd aUniversity ’1 Decembrie 1918’ of Alba Iulia, 510009, Romania; [email protected], [email protected] bBabes-Bolyai University, Computer Science Department, Cluj-Napoca, 43017-6221, Romania; [email protected]o cDoctoral School of Letters, Humanities and Applied Sciences, George Emil Palade University of Medicine, Pharmacy, Sciences and Technology of Targu Mures, 540142, Romania; [email protected]o dElectrical Engineering and Information Technology Department, George Emil Palade University of Medicine, Pharmacy, Sciences and Technology of Targu Mures, 540142, Romania; [email protected]o Abstract Deep Learning (DL) in the healthcare is exemplified by specific applications such as colorectal polyp segmentation using the Parallel Reverse Attention Network (PraNet) model. The main objective of this research consisted on investigating how data augmentation impacts the performance of the PraNet model across heterogeneous datasets. It is proposed a data augmentation methodology with properties to conserve the original image quality. Seven public polyp image datasets with a unified normalization pipeline and specific augmentation techniques were applied to expand the training data. Four experiments were conducted: (1) baseline training on each normalized dataset; (2) training on each dataset with augmentation; (3) combined training on all normalized datasets, tested on an independent dataset to assess generalization; and (4) combined training on all normalized plus augmented data, tested on the same independent set to examine augmentation-induced bias. The results show that the specific augmentation substantially improves in-domain segmentation accuracy under five metrics Dice, IoU, MAE, S-measure and E-measure. Average Dice score gains of 0.20–0.25 on small datasets, and training on the aggregated multi-source data yields strong generalization (Dice ≈0.81) on the independent set. When augmentation is added to combined training, performance on the independent test slightly decreases, the Dice score decreases from 0.81 to 0.77, suggesting a slight decrease in segmentation performance under varying test conditions, indicating potential overfitting or distribution bias introduced by excessive augmentation. Combining multiple datasets resulted in a model with strong generalization to an independent dataset, where real data diversity can effectively limit the bias. ©2025 The Authors. Published by Elsevier B.V. This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-ncnd/4.0/) Peer-review under responsibility of the scientific committee of the KES International. Keywords: Gastrointestinal Polyps; Medical Image Augmentation and Segmentation; Deep Learning; Colorectal Cancer. 1877-0509 ©2025 The Authors. Published by Elsevier B.V. This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/) Peer-review under responsibility of the scientific committee of the KES International. 10.1016/j.procs.2025.09.540 1877-0509 © 2025 The Authors. Published by Elsevier B.V. This is an open access article under the CC BY-NC-ND license (https://creativecommons.org/licenses/by-nc-nd/4.0) Peer-review under responsibility of the scientific committee of the KES International. Available online at www.sciencedirect.com Procedia Computer Science 00 (2025) 000–000 www.elsevier.com/locate/procedia 29th International Conference on Knowledge-Based and Intelligent Information & Engineering Systems (KES 2025) PraNet-based Gastrointestinal Polyp Segmentation on Heterogeneous Datasets: Study of Augmentation Impact on Performance Daniela-Maria Cristeaa,c,∗, Ioan Simab,S ˆ ınit ,ˆ ın Vladimira, Laszlo Barna Iantovicsd aUniversity ’1 Decembrie 1918’ of Alba Iulia, 510009, Romania; [email protected], [email protected] bBabes-Bolyai University, Computer Science Department, Cluj-Napoca, 43017-6221, Romania; [email protected]o cDoctoral School of Letters, Humanities and Applied Sciences, George Emil Palade University of Medicine, Pharmacy, Sciences and Technology of Targu Mures, 540142, Romania; [email protected]o dElectrical Engineering and Information Technology Department, George Emil Palade University of Medicine, Pharmacy, Sciences and Technology of Targu Mures, 540142, Romania; barna.ianto[email protected]o Abstract Deep Learning (DL) in the healthcare is exemplified by specific applications such as colorectal polyp segmentation using the Parallel Reverse Attention Network (PraNet) model. The main objective of this research consisted on investigating how data augmentation impacts the performance of the PraNet model across heterogeneous datasets. It is proposed a data augmentation methodology with properties to conserve the original image quality. Seven public polyp image datasets with a unified normalization pipeline and specific augmentation techniques were applied to expand the training data. Four experiments were conducted: (1) baseline training on each normalized dataset; (2) training on each dataset with augmentation; (3) combined training on all normalized datasets, tested on an independent dataset to assess generalization; and (4) combined training on all normalized plus augmented data, tested on the same independent set to examine augmentation-induced bias. The results show that the specific augmentation substantially improves in-domain segmentation accuracy under five metrics Dice, IoU, MAE, S-measure and E-measure. Average Dice score gains of 0.20–0.25 on small datasets, and training on the aggregated multi-source data yields strong generalization (Dice ≈0.81) on the independent set. When augmentation is added to combined training, performance on the independent test slightly decreases, the Dice score decreases from 0.81 to 0.77, suggesting a slight decrease in segmentation performance under varying test conditions, indicating potential overfitting or distribution bias introduced by excessive augmentation. Combining multiple datasets resulted in a model with strong generalization to an independent dataset, where real data diversity can effectively limit the bias. ©2025 The Authors. Published by Elsevier B.V. This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-ncnd/4.0/) Peer-review under responsibility of the scientific committee of the KES International. Keywords: Gastrointestinal Polyps; Medical Image Augmentation and Segmentation; Deep Learning; Colorectal Cancer. 1877-0509 ©2025 The Authors. Published by Elsevier B.V. This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/) Peer-review under responsibility of the scientific committee of the KES International. 4156 Daniela-Maria Cristea et al. / Procedia Computer Science 270 (2025) 4155–4164 2D.-M. Cristea et al. /Procedia Computer Science 00 (2025) 000–000 1. Introduction Colorectal cancer (CRC) is a leading cause of cancer-related deaths, often developing from precancerous polyps found during colonoscopy. Accurate polyp segmentation in colonoscopy images is therefore vital for diagnosis and guiding polyp removal, as it can improve adenoma detection rates and assist in complete resection [2]. However, models frequently encounter challenges such as limited annotated data, variations in imaging conditions across datasets, and domain shift between training and real-world clinical environments. In recent years, deep learning models have achieved noteworthy success in this task, greatly outperforming traditional methods in accuracy and efficiency [8,2]. One such model, introduced in MICCAI [5], specifically tackled the challenges of polyp images segmentation from endoscopic imagery – the variability in polyp size/appearance and ambiguous boundaries – by combining a partial decoder and reverse attention modules. Parallel Reverse Attention Network (PraNet) demonstrated state-of-the-art performance on multiple polyp datasets and has become a widely-used baseline for this problem [8]. Despite these advances, limitations remain. Many polyp segmentation models are trained and validated on the same single-center datasets [3,2]. There is a recognized need for heterogeneous or multi-center data to ensure robustness [2]. However, merging datasets from different sources introduces variability in image resolution, aspect ratio, and illumination, necessitating careful data normalization. Standardizing image size and intensity can reduce such dataset biases so that models learn consistent features. Furthermore, the lack of annotated medical images (some polyp datasets contain only tens of training samples) often leads to high variance in model performance. Data augmentation is a common strategy to address this issue by synthetically increasing data diversity [9]. Augmentation can act as a regularizer to reduce overfitting, but if applied inappropriately, it may introduce unrealistic samples or biases that do not translate to real-world improvement [9,4]. Indeed, recent studies have started to examine how augmentation might inadvertently create fake features that a model can exploit, negatively influence generalization when those augmented artifacts are absent in deployment data [5]. In the context of polyp segmentation, augmentation is widely used (e.g., horizontal flipping is almost universally applied), yet the community lacks a systematic analysis of its true impact on both in-domain and cross-domain performance. An additional objective of our research was to ensure the reproducibility of our experiments. We achieve this by utilizing publicly available code (the PraNet repository) and providing full details of our data preprocessing scripts, training procedure, and computing environment. By sharing precise normalization operations, augmentation parameters, and hardware/software specifications, we enable others researchers to replicate and extend our study. The main contributions of the research to the state-of-the-art are: (1) Increased Baseline Performance with Normalized Data: Establish benchmark segmentation results on each dataset after applying a uniform normalization pipeline, to quantify the effect of standardizing data pre-training. (2) Increased Effect of Augmentation: Evaluate how integrating on-the-fly augmentation per dataset influences model performance, by training PraNet on augmented versions of each dataset and comparing against the baseline. ∗Corresponding author: Daniela-Maria Cristea Daniela-Maria Cristea et al. / Procedia Computer Science 270 (2025) 4155–4164 4157 2D.-M. Cristea et al. /Procedia Computer Science 00 (2025) 000–000 1. Introduction Colorectal cancer (CRC) is a leading cause of cancer-related deaths, often developing from precancerous polyps found during colonoscopy. Accurate polyp segmentation in colonoscopy images is therefore vital for diagnosis and guiding polyp removal, as it can improve adenoma detection rates and assist in complete resection [2]. However, models frequently encounter challenges such as limited annotated data, variations in imaging conditions across datasets, and domain shift between training and real-world clinical environments. In recent years, deep learning models have achieved noteworthy success in this task, greatly outperforming traditional methods in accuracy and efficiency [8,2]. One such model, introduced in MICCAI [5], specifically tackled the challenges of polyp images segmentation from endoscopic imagery – the variability in polyp size/appearance and ambiguous boundaries – by combining a partial decoder and reverse attention modules. Parallel Reverse Attention Network (PraNet) demonstrated state-of-the-art performance on multiple polyp datasets and has become a widely-used baseline for this problem [8]. Despite these advances, limitations remain. Many polyp segmentation models are trained and validated on the same single-center datasets [3,2]. There is a recognized need for heterogeneous or multi-center data to ensure robustness [2]. However, merging datasets from different sources introduces variability in image resolution, aspect ratio, and illumination, necessitating careful data normalization. Standardizing image size and intensity can reduce such dataset biases so that models learn consistent features. Furthermore, the lack of annotated medical images (some polyp datasets contain only tens of training samples) often leads to high variance in model performance. Data augmentation is a common strategy to address this issue by synthetically increasing data diversity [9]. Augmentation can act as a regularizer to reduce overfitting, but if applied inappropriately, it may introduce unrealistic samples or biases that do not translate to real-world improvement [9,4]. Indeed, recent studies have started to examine how augmentation might inadvertently create fake features that a model can exploit, negatively influence generalization when those augmented artifacts are absent in deployment data [5]. In the context of polyp segmentation, augmentation is widely used (e.g., horizontal flipping is almost universally applied), yet the community lacks a systematic analysis of its true impact on both in-domain and cross-domain performance. An additional objective of our research was to ensure the reproducibility of our experiments. We achieve this by utilizing publicly available code (the PraNet repository) and providing full details of our data preprocessing scripts, training procedure, and computing environment. By sharing precise normalization operations, augmentation parameters, and hardware/software specifications, we enable others researchers to replicate and extend our study. The main contributions of the research to the state-of-the-art are: (1) Increased Baseline Performance with Normalized Data: Establish benchmark segmentation results on each dataset after applying a uniform normalization pipeline, to quantify the effect of standardizing data pre-training. (2) Increased Effect of Augmentation: Evaluate how integrating on-the-fly augmentation per dataset influences model performance, by training PraNet on augmented versions of each dataset and comparing against the baseline. ∗Corresponding author: Daniela-Maria Cristea D.-M. Cristea et al. /Procedia Computer Science 00 (2025) 000–000 3 (3) In-depth Cross-Dataset Bias Analysis: Investigate model generalizability bias by training PraNet on a combined pool of normalized datasets and testing on an independent dataset not utilized in training, resulting in a broad training distribution that transfers to a new clinical data source. (4) Improved Augmentation-Induced Bias Check: Using the same independent test, assess whether adding augmentation to the combined training (effectively doubling the training set with augmented samples) improves generalization or introduces a performance decrease, indicating overfitting to synthetic patterns or other biases. In summary, this work systematically dissects the role of normalization and augmentation in polyp segmentation, contributing insights on how to balance data diversity and realism to train models that generalize well across heterogeneous medical datasets. We discuss these findings in the context of recent literature on medical image segmentation and data augmentation, noting that while augmentation reduce inadequate availability of data necessary for informed decision-making and overfitting, it must be carefully tuned to avoid diminishing returns. Finally, we emphasize reproducibility by detailing our preprocessing scripts and system configuration, providing a foundation for future extensions on other datasets, advanced augmentations, and bias reduction strategies. PraNet-based models have demonstrated promising results in the segmentation of colorectal polyps from endoscopic imagery. By evaluating performance across normalised and augmented heterogeneous datasets, it becomes possible to assess both the strengths and the limitations of such approaches within realistic clinical scenarios. 2. Materials and Methods 2.1. Datasets and Data preprocessing A collection of seven public polyp segmentation datasets was leveraged, representing heterogeneous sources (i.e., different hospitals, imaging devices and patient populations). These include the CVC subset series and Kvasir datasets all of them splited in 80% train and 20% test for first 2 experiments: CVC-300[16]: 60 colonoscopy images with corresponding polyp masks; CVC-ClinicDB[7] with 62 images; CVC-ColonDB[6]with 380 images; ETIS-LaribPolypDB[15]with 196 images; Kvasir[14]with 100 polyp images; Kvasir-SEG[13]: 1000 annotated polyp images; Kvasir-Sessile with 196 images. Image Resizing with Padding. Each image and its corresponding binary mask were resized to a target size of 512×512 pixels while preserving the original aspect ratio. The appropriate scaling factor was computed, and the remaining area was padded with a constant black background (pixel value 0). By padding rather than stretching, geometric distortions of polyps are avoided. In practice, smaller images (e.g., 384 ×288) are centered in a 512 ×512 frame and non-square images are letterboxed. Identical resizing and padding were applied to both images and masks to maintain alignment. Intensity Normalization. After resizing, image pixel values were converted from the 8-bit (0-255) range to a 0.0-1.0 floating-point range by simple linear scaling. This normalization equalizes brightness across images acquired under different lighting conditions or via different endoscopic devices. Although the normalized images were eventually stored as 8-bit (0-255) PNGs for compatibility, the underlying pixel intensity distribution is rendered consistent across datasets. Similarly, the masks, originally binary or grayscale, were normalized and saved, yielding binary masks in {0,255}after round-trip conversion. This standardizes the mask format despite variations in the original datasets. No additional filtering techniques were applied; the focus was on geometric and scale normalization. The outcome is a set of normalized datasets where all images share a standard size and general intensity scale, making them ready for training. This consistent normalization helps reduce inter-dataset variation originating solely from technical factors, so that any subsequent performance differences are more likely to reflect differences in training data content. 4158 Daniela-Maria Cristea et al. / Procedia Computer Science 270 (2025) 4155–4164 4D.-M. Cristea et al. /Procedia Computer Science 00 (2025) 000–000 2.2. Data Augmentation To create augmented training samples instantly, a particular augmentation module was suggested and created. The approach was intended to introduce realistic modifications while maintaining the fundamental polyp structures. In Fig. 1, a set of random transformations was created using the Albumentations library and applied to each image-mask pair with predetermined probabilities. Fig. 1: Example augmented images: (a) CVC-300 panel; (b) Kvasir panel. Geometric Transforms. -Random rotations: Up to ±15◦applied with a probability(p) by 0.8. -Flips: Horizontal and vertical flips (each with p=0.5). -Scaling: Slight scaling (zoom in/out by up to 2%). -Affine Transform: A small random affine transformation (translation up to 3–5% of the image size and an additional ±10◦rotation). Photometric Transforms. -Brightness and Contrast Adjustments: Applied with p=0.5. -Gamma Correction Variation: Gamma values varied by approximately ±10% with p=0.3. -Gaussian Noise Injection: Injects noise between 1 and 3 in pixel values, with p≈0.5. -Custom Brightness Perturbation: Included in some runs to mimic illumination changes. Random Cropping. With a 0.5 probability, a random 256 ×256 patch was extracted from the 512 ×512 image and subsequently resized back to 512 ×512. The corresponding mask patch was taken from the same location. This operation provides a zoomed-in view of a random region, helping the model to learn from partial polyp views and different scales. The sequence of augmentation operations was composed in such a way that each training image can undergo a unique series of transformations per epoch, thereby greatly expanding the diversity of the training data. To preserve spatial correspondence, augmentations were applied synchronously to images and masks. Operations that might alter the mask content, such as elastic deformations or random erasing, were deliberately avoided.By combining these operations, our augmentation pipeline introduces substantial variability – different orientations, positions, brightness levels, and partial views – while keeping the polyp anatomy intact. During training, augmented samples are generated on the fly rather than presaved, meaning the model effectively never sees the same image twice. The intent was to improve the model’s robustness to real-world variation. At the same time, we aimed to avoid overly aggressive augmentation that could produce implausible images (for instance, we did not fill up with noise or warp the image unnaturally). This balanced approach aligns with recommended practices for medical image augmentation, where the augmentations should reflect realistic alterations [9]. In the later experiments, we specifically examine whether this augmentation strategy indeed improved performance and whether it might have any unintended side effects on generalization. Daniela-Maria Cristea et al. / Procedia Computer Science 270 (2025) 4155–4164 4159 4D.-M. Cristea et al. /Procedia Computer Science 00 (2025) 000–000 2.2. Data Augmentation To create augmented training samples instantly, a particular augmentation module was suggested and created. The approach was intended to introduce realistic modifications while maintaining the fundamental polyp structures. In Fig. 1, a set of random transformations was created using the Albumentations library and applied to each image-mask pair with predetermined probabilities. Fig. 1: Example augmented images: (a) CVC-300 panel; (b) Kvasir panel. Geometric Transforms. -Random rotations: Up to ±15◦applied with a probability(p) by 0.8. -Flips: Horizontal and vertical flips (each with p=0.5). -Scaling: Slight scaling (zoom in/out by up to 2%). -Affine Transform: A small random affine transformation (translation up to 3–5% of the image size and an additional ±10◦rotation). Photometric Transforms. -Brightness and Contrast Adjustments: Applied with p=0.5. -Gamma Correction Variation: Gamma values varied by approximately ±10% with p=0.3. -Gaussian Noise Injection: Injects noise between 1 and 3 in pixel values, with p≈0.5. -Custom Brightness Perturbation: Included in some runs to mimic illumination changes. Random Cropping. With a 0.5 probability, a random 256 ×256 patch was extracted from the 512 ×512 image and subsequently resized back to 512 ×512. The corresponding mask patch was taken from the same location. This operation provides a zoomed-in view of a random region, helping the model to learn from partial polyp views and different scales. The sequence of augmentation operations was composed in such a way that each training image can undergo a unique series of transformations per epoch, thereby greatly expanding the diversity of the training data. To preserve spatial correspondence, augmentations were applied synchronously to images and masks. Operations that might alter the mask content, such as elastic deformations or random erasing, were deliberately avoided.By combining these operations, our augmentation pipeline introduces substantial variability – different orientations, positions, brightness levels, and partial views – while keeping the polyp anatomy intact. During training, augmented samples are generated on the fly rather than presaved, meaning the model effectively never sees the same image twice. The intent was to improve the model’s robustness to real-world variation. At the same time, we aimed to avoid overly aggressive augmentation that could produce implausible images (for instance, we did not fill up with noise or warp the image unnaturally). This balanced approach aligns with recommended practices for medical image augmentation, where the augmentations should reflect realistic alterations [9]. In the later experiments, we specifically examine whether this augmentation strategy indeed improved performance and whether it might have any unintended side effects on generalization. D.-M. Cristea et al. /Procedia Computer Science 00 (2025) 000–000 5 There was no overlap between training and testing samples within any dataset. All training images were normalized as described above. Summary tables (see Results section) detail the dataset sizes and baseline performance. In addition, an independent side dataset, the BKAI-IGH NeoPolyp dataset (1000 polyp images and masks), was used exclusively for external validation in Exp 3 and 4. This dataset originates from a different institution and serves as an unseen domain to evaluate model bias and generalization. Four experiments were designed to address the study objectives: Experiment 1 (Exp 1): Baseline (Normalized only). A PraNet model was trained on each dataset using normalized images without augmentation. This provided per-dataset baseline performance (assessed by Dice, IoU, etc.). Experiment 2 (Exp 2): Augmented (Normalized plus Augmented per dataset). A PraNet model was trained with augmentation enabled. Here, the training set was effectively doubled (each epoch saw both original and augmented variants), allowing the effect of data augmentation on segmentation performance to be isolated. Experiment 3 (Exp 3): Combined Training (All Normalized), Cross-Dataset Testing. All normalized training images (a total of 1994) were pooled to train a single PraNet model. No augmentation was applied. Evaluation was performed on the BKAI-IGH NeoPolyp test set. Experiment 4 (Exp 4): Combined Training (Normalized plus Augmented), Cross-Dataset Testing. The combined training set was augmented (doubling the 1994 training samples to 3988), and the model was evaluated on the BKAI-IGH NeoPolyp dataset. This experiment helps determine whether augmentation improves or harms generalization. Across all experiments, evaluation metrics included the Dice coefficient, Intersection over Union (IoU), Mean Absolute Error (MAE) of the prediction mask, S-measure (structure measure), Emeasure (enhanced-alignment measure). These metrics are consistent [8], providing a comprehensive view of segmentation quality. 2.3. System Configuration and Training Details Experiments were conducted on a computer running Microsoft Windows 11 Pro (64-bit, version 10.0.22631) equipped with an Intel Core i7-13700H CPU, an NVIDIA RTX A1000 Laptop GPU (6 GB VRAM), and 64 GB of RAM. Deep learning code was written in Python and executed within a dedicated virtual environment. Key libraries included: PyTorch 2.5.1 (with CUDA 12.1 support), Torchvision 0.20.1,Albumentations 1.3.x (for data augmentation), OpenCV (cv2 for image I/O and resizing), and NumPy (for array operations). The adapted program code was based on the official PraNet repository [8], employing a Res2Net backbone with reverse attention modules. For every experiment, training was executed using the PraNet training script (MyTrain.py) with the appropriate paths and parameters. Training was capped at a maximum of 20 epochs per dataset. The batch size was chosen based on the dataset size and GPU memory constraints. The optimizer was Stochastic Gradient Descent (SGD) with momentum and a learning rate of 1 ×10−4with poly decay. No extensive hyperparameter tuning was performed, so differences in performance are attributable to variations in data composition and augmentation. Model checkpoints were saved during training to ensure accurate tracking of dataset usage. For testing, the PraNet testing script (MyTest.py) was used, with test images resized to 352 ×352 as recommended. The experimental design progressed from controlled single-dataset evaluations (Exp 1 and 2) to a comprehensive multi-dataset approach (Exp 3 and 4), both with and without augmentation. This systematic approach allows for a detailed assessment of normalization baselines, augmentation effects, and any augmentation-induced biases when generalizing to unseen clinical data. 4160 Daniela-Maria Cristea et al. / Procedia Computer Science 270 (2025) 4155–4164 6D.-M. Cristea et al. /Procedia Computer Science 00 (2025) 000–000 3. Results The performance evaluation incorporated several key metrics and criteria as previously established in [10,12,11]. The segmentation performance is reported under each experimental setting using five evaluation metrics. The key comparative findings are highlighted for baseline versus augmented training and combined training versus independent testing. Experiment 1. Baseline (Normalized Data) Table 1presents the baseline results for the PraNet model trained on each dataset using only normalized images (i.e., without augmentation). Performance varies widely across datasets. For example, on the small CVC-ClinicDB, the model achieved an average (AVG) Dice of 0.511, whereas on the larger Kvasir dataset, Dice reached 0.903. Larger datasets such as Kvasir and Kvasir-SEG yielded higher Dice and IoU values. The S-measure exceeds 0.915 for ColonDB and Kvasir-SEG, whereas for CVC-300 and ClinicDB it is lower (0.539–0.582), indicating more fragmented predictions. These outcomes align with our expectation that datasets with easier polyp instances (e.g., ETIS-LaribPolypDB and ColonDB) yield better segmentation (Dice ∼0.716–0.768), whereas challenging cases like those in ClinicDB result in lower Dice scores. Table 1: Exp 1: PraNet trained on normalized datasets (no augmentation) Dataset Avg Dice Avg IoU Avg MAE Avg S Avg E CVC-300 0.588 0.497 0.081 0.539 0.533 CVC-ClinicDB 0.511 0.387 0.071 0.582 0.575 CVC-ColonDB 0.768 0.666 0.013 0.915 0.898 ETIS-LaribPolypDB 0.716 0.636 0.027 0.701 0.676 Kvasir 0.903 0.832 0.040 0.780 0.936 Kvasir-SEG 0.891 0.833 0.029 0.944 0.937 Kvasir-Sessile 0.816 0.717 0.034 0.882 0.893 Experiment 2. Augmented Training per Dataset Table 2shows the performance metrics when training with both normalized and augmented images. Augmentation provides a notable boost in performance across nearly every dataset. For instance, CVC-300’s Dice improves from 0.588 to 0.806 and CVC-ClinicDB from 0.511 to 0.754. Even highperforming datasets (Kvasir and Kvasir-SEG) revise small improvements (e.g., Dice increases from 0.903 to 0.926 in Kvasir). MAE consistently decreased, indicating fewer pixel-level errors. Table 2: Exp 2: PraNet trained on each datasets normalized & augmented images (test sets remain unchanged). Dataset Avg Dice Avg IoU Avg MAE Avg S Avg E CVC-300 0.806 0.692 0.015 0.793 0.843 CVC-ClinicDB 0.754 0.632 0.021 0.784 0.816 CVC-ColonDB 0.856 0.782 0.006 0.975 0.959 ETIS-LaribPolypDB 0.899 0.825 0.003 0.981 0.978 Kvasir 0.926 0.872 0.032 0.880 0.956 Kvasir-SEG 0.920 0.870 0.020 0.956 0.957 Kvasir-Sessile 0.871 0.796 0.025 0.944 0.938 Experiment 3. Combined Training (All Normalized), Tested on Independent Data. After establishing per-dataset gains, we trained a single model on the entire collection of normalized data (1994 images aggregated) and tested on the BKAI-IGH NeoPolyp dataset, which consists Daniela-Maria Cristea et al. / Procedia Computer Science 270 (2025) 4155–4164 4161 6D.-M. Cristea et al. /Procedia Computer Science 00 (2025) 000–000 3. Results The performance evaluation incorporated several key metrics and criteria as previously established in [10,12,11]. The segmentation performance is reported under each experimental setting using five evaluation metrics. The key comparative findings are highlighted for baseline versus augmented training and combined training versus independent testing. Experiment 1. Baseline (Normalized Data) Table 1presents the baseline results for the PraNet model trained on each dataset using only normalized images (i.e., without augmentation). Performance varies widely across datasets. For example, on the small CVC-ClinicDB, the model achieved an average (AVG) Dice of 0.511, whereas on the larger Kvasir dataset, Dice reached 0.903. Larger datasets such as Kvasir and Kvasir-SEG yielded higher Dice and IoU values. The S-measure exceeds 0.915 for ColonDB and Kvasir-SEG, whereas for CVC-300 and ClinicDB it is lower (0.539–0.582), indicating more fragmented predictions. These outcomes align with our expectation that datasets with easier polyp instances (e.g., ETIS-LaribPolypDB and ColonDB) yield better segmentation (Dice ∼0.716–0.768), whereas challenging cases like those in ClinicDB result in lower Dice scores. Table 1: Exp 1: PraNet trained on normalized datasets (no augmentation) Dataset Avg Dice Avg IoU Avg MAE Avg S Avg E CVC-300 0.588 0.497 0.081 0.539 0.533 CVC-ClinicDB 0.511 0.387 0.071 0.582 0.575 CVC-ColonDB 0.768 0.666 0.013 0.915 0.898 ETIS-LaribPolypDB 0.716 0.636 0.027 0.701 0.676 Kvasir 0.903 0.832 0.040 0.780 0.936 Kvasir-SEG 0.891 0.833 0.029 0.944 0.937 Kvasir-Sessile 0.816 0.717 0.034 0.882 0.893 Experiment 2. Augmented Training per Dataset Table 2shows the performance metrics when training with both normalized and augmented images. Augmentation provides a notable boost in performance across nearly every dataset. For instance, CVC-300’s Dice improves from 0.588 to 0.806 and CVC-ClinicDB from 0.511 to 0.754. Even highperforming datasets (Kvasir and Kvasir-SEG) revise small improvements (e.g., Dice increases from 0.903 to 0.926 in Kvasir). MAE consistently decreased, indicating fewer pixel-level errors. Table 2: Exp 2: PraNet trained on each datasets normalized & augmented images (test sets remain unchanged). Dataset Avg Dice Avg IoU Avg MAE Avg S Avg E CVC-300 0.806 0.692 0.015 0.793 0.843 CVC-ClinicDB 0.754 0.632 0.021 0.784 0.816 CVC-ColonDB 0.856 0.782 0.006 0.975 0.959 ETIS-LaribPolypDB 0.899 0.825 0.003 0.981 0.978 Kvasir 0.926 0.872 0.032 0.880 0.956 Kvasir-SEG 0.920 0.870 0.020 0.956 0.957 Kvasir-Sessile 0.871 0.796 0.025 0.944 0.938 Experiment 3. Combined Training (All Normalized), Tested on Independent Data. After establishing per-dataset gains, we trained a single model on the entire collection of normalized data (1994 images aggregated) and tested on the BKAI-IGH NeoPolyp dataset, which consists D.-M. Cristea et al. /Procedia Computer Science 00 (2025) 000–000 7 of 1000 images from a different source. The performance of this model on the side dataset is shown in Table 3 (left column). The combined model achieved an Average Dice of 0.809 and Average IoU of 0.730 on the BKAI-IGH dataset. These are fairly high scores, especially considering that BKAI is a completely independent dataset. For comparison, these metrics are on par with the model’s performance on some of the known datasets: e.g., 0.809 Dice is similar to what the model got on KvasirSessile in Exp 2 (0.871) or Kvasir-SEG in Exp 1 (0.891). It indicates that by training on a diverse set of sources, the PraNet model learned a representation that generalizes to new images reasonably well. The S-measure of 0.964 is very high, suggesting the model can detect the salient polyp regions in the BKAI images with nearly perfect consistency. The MAE is 0.017, meaning on average only 1.76% of pixels were misclassified, which is low for an unseen dataset. Overall, these results demonstrate successful generalization: the combined training has mitigated bias toward any single source and created a robust model. From a bias analysis perspective, we note that the combined model (Exp 3) did not overfit to the training datasets – if it had, we would expect much lower scores on BKAI-IGH. Experiment 4. Combined Training (Normalized +Augmented), Tested on Independent Data The training set was augmented (doubling the sample size to 3,988), and the model was tested on the BKAI-IGH dataset. Surprisingly, the augmented combined model resulted in a decrease in performance (e.g., Average Dice fell to 0.766 and IoU to 0.688). The increase in MAE and decrease in Eand S-measures suggest that, while augmentation benefits in-domain performance (as in Exp 2), it may introduce a distribution mismatch when the test data originates from a different clinical source. The results are pivotal, whereas in the in-domain cases, augmentation helped; in this cross-domain Table 3: Cross-dataset evaluation on BKAI-IGH NeoPolyp dataset. Left: Combined training on normalized data (Exp 3). Right: Combined training on normalized +augmented data (Exp 4). Metric Exp3 (Combined Norm) Exp4 (Combined Norm+Aug) Average Dice 0.809 0.766 Average IoU 0.730 0.688 Average MAE 0.017 0.036 Average S-measure 0.964 0.939 Average E-measure 0.909 0.874 case, augmentation hindered performance slightly. The Exp 4 model likely had a better fit on the training data (according to Exp 2 augmentation improves in-domain scores), but those gains did not carry over to the external dataset. Several interpretations are possible: (a) The augmented combined model might be overfitting to the peculiarities of the augmented images (for example, it may rely on patterns introduced by augmentation that are absent in BKAI images). (b) The augmentation effectively altered the training distribution, possibly reducing the emphasis on some real variations present in BKAI. Since our augmentation included random cropping and brightness changes, the model might have learned to expect these, whereas the BKAI dataset might have different characteristics. In contrast, the non-augmented combined model trained purely on real images from various datasets, which might be closer to the real images in BKAI, hence it performed slightly better. Nonetheless, even the augmented model’s performance is weak – Dice 0.766 is still respectable – but it is worse than the non-augmented model on the same test, pointing to a form of augmentationinduced bias/overfit. This finding aligns with reports in the literature that augmentation needs to be carefully tuned; if it diverges too much from the target domain’s characteristics, it might not yield additional benefit [4]. Here we have a concrete example: adding all those random transformations during training made the model a bit less effective on normal-looking images of BKAI. To ensure this difference is not due to randomness, recall that we kept all other factors the same. The only difference between Exp 3 and Exp 4 models was the presence of augmented images in training. The 4162 Daniela-Maria Cristea et al. / Procedia Computer Science 270 (2025) 4155–4164 8D.-M. Cristea et al. /Procedia Computer Science 00 (2025) 000–000 models went through the same number of training epochs, etc. The consistent decrease across all metrics (Dice, IoU, E, S, MAE decreasing gives confidence that this is a real effect, not noise. Augmentation significantly improved segmentation performance on the same distribution as the training data (Exp 1 compared to Exp 2), fulfilling our expectation that it makes the model generalize to similar images and overcome small dataset limitations. However, when evaluating on a different distribution (BKAI) and using a broadly trained model, we found that the model trained without augmentation generalized better than the one with augmentation (Exp 3 compared to Exp 4). This suggests that in our case, the diversity provided by combining many datasets was sufficient for generalization, and additional synthetic diversity from augmentation might have started to introduce irrelevant variations that did not match the new dataset, slightly confusing the model. For completeness, we note that the training time for the combined models was longer (Exp 4 took nearly ∼15 hours while Exp 3 took 8 hours). But since all models converged within 20 epochs, training length did not impact the results beyond providing context that augmentation increases training effort. Finally, our results constitute empirical evidence regarding augmentation-induced bias: we have quantified a case where more training data (via augmentation) did not equal better performance on an independent test. Note: All experiments were performed using the PraNet model architecture[8]. The BKAI-IGH NeoPolyp dataset is an independent dataset used solely for evaluating generalizability. 4. Discussion Normalization and Baseline Performance. By resizing and scaling all images uniformly, were removed simple disparities between datasets, allowing PraNet to train effectively on a combined dataset in Exp 3. The baseline results (Exp 1) show that even without any augmentation, PraNet achieved good accuracy on each dataset. This indicates that our normalization did not distort the data; if anything, it likely improved PraNet’s performance by providing images at the resolution the network expects and by regularizing image sizes. The combined training (Exp 3) was feasible only because the data were homogenized in format; without normalization (e.g., mixing batches of 512 × 512 with 1080 ×1920 images), training on mixed data would have been problematic. The successful generalization of the combined model (Dice ∼0.81 on BKAI) underscores that normalization enabled the model to benefit from heterogeneous content without being distracted by heterogeneity of form. In practical terms, this means that researchers can confidently merge datasets after applying such preprocessing. Our work thus reinforces a commonly assumed—but rarely quantified—point: data normalization is a crucial step for training robust medical image models across different sources. Impact of Data Augmentation on In-domain Performance. Data augmentation largely improves in-domain segmentation performance by reducing overfitting and enhancing feature learning. Exp 2 demonstrated clear gains in Dice, IoU, and other metrics for every dataset when augmentation was applied. This is in line with many studies that report augmentation as a means to improve model generalization to similar data [9]. Particularly on smaller datasets, augmentation had an outsized impact. For instance, the increase from 0.51 to 0.75 Dice on CVC-ClinicDB is likely because augmentation presented variations of the 49 training images that covered scenarios appearing in the 13 test images (e.g., a rotated training image might match an unusual polyp orientation in test, or a darkened image might mirror a low-brightness test image). Additionally, the substantial reduction in MAE indicates that the augmented model produced nearly realistic masks for many images. From a feature learning perspective, augmentations such as flips and rotations enforce equivariance the model learns that a polyp is a polyp regardless of orientation. Likewise, brightness shifts teach the model to rely on shape and texture rather than absolute intensity. These effects collectively result in more robust feature representations and better performance on test sets drawn from the same distribution. Benefits of Combining Multiple Datasets. Combining multiple datasets results in a model with strong generalization to an independent dataset, where real data diversity can effectively limit bias. Daniela-Maria Cristea et al. / Procedia Computer Science 270 (2025) 4155–4164 4163 8D.-M. Cristea et al. /Procedia Computer Science 00 (2025) 000–000 models went through the same number of training epochs, etc. The consistent decrease across all metrics (Dice, IoU, E, S, MAE decreasing gives confidence that this is a real effect, not noise. Augmentation significantly improved segmentation performance on the same distribution as the training data (Exp 1 compared to Exp 2), fulfilling our expectation that it makes the model generalize to similar images and overcome small dataset limitations. However, when evaluating on a different distribution (BKAI) and using a broadly trained model, we found that the model trained without augmentation generalized better than the one with augmentation (Exp 3 compared to Exp 4). This suggests that in our case, the diversity provided by combining many datasets was sufficient for generalization, and additional synthetic diversity from augmentation might have started to introduce irrelevant variations that did not match the new dataset, slightly confusing the model. For completeness, we note that the training time for the combined models was longer (Exp 4 took nearly ∼15 hours while Exp 3 took 8 hours). But since all models converged within 20 epochs, training length did not impact the results beyond providing context that augmentation increases training effort. Finally, our results constitute empirical evidence regarding augmentation-induced bias: we have quantified a case where more training data (via augmentation) did not equal better performance on an independent test. Note: All experiments were performed using the PraNet model architecture[8]. The BKAI-IGH NeoPolyp dataset is an independent dataset used solely for evaluating generalizability. 4. Discussion Normalization and Baseline Performance. By resizing and scaling all images uniformly, were removed simple disparities between datasets, allowing PraNet to train effectively on a combined dataset in Exp 3. The baseline results (Exp 1) show that even without any augmentation, PraNet achieved good accuracy on each dataset. This indicates that our normalization did not distort the data; if anything, it likely improved PraNet’s performance by providing images at the resolution the network expects and by regularizing image sizes. The combined training (Exp 3) was feasible only because the data were homogenized in format; without normalization (e.g., mixing batches of 512 × 512 with 1080 ×1920 images), training on mixed data would have been problematic. The successful generalization of the combined model (Dice ∼0.81 on BKAI) underscores that normalization enabled the model to benefit from heterogeneous content without being distracted by heterogeneity of form. In practical terms, this means that researchers can confidently merge datasets after applying such preprocessing. Our work thus reinforces a commonly assumed—but rarely quantified—point: data normalization is a crucial step for training robust medical image models across different sources. Impact of Data Augmentation on In-domain Performance. Data augmentation largely improves in-domain segmentation performance by reducing overfitting and enhancing feature learning. Exp 2 demonstrated clear gains in Dice, IoU, and other metrics for every dataset when augmentation was applied. This is in line with many studies that report augmentation as a means to improve model generalization to similar data [9]. Particularly on smaller datasets, augmentation had an outsized impact. For instance, the increase from 0.51 to 0.75 Dice on CVC-ClinicDB is likely because augmentation presented variations of the 49 training images that covered scenarios appearing in the 13 test images (e.g., a rotated training image might match an unusual polyp orientation in test, or a darkened image might mirror a low-brightness test image). Additionally, the substantial reduction in MAE indicates that the augmented model produced nearly realistic masks for many images. From a feature learning perspective, augmentations such as flips and rotations enforce equivariance the model learns that a polyp is a polyp regardless of orientation. Likewise, brightness shifts teach the model to rely on shape and texture rather than absolute intensity. These effects collectively result in more robust feature representations and better performance on test sets drawn from the same distribution. Benefits of Combining Multiple Datasets. Combining multiple datasets results in a model with strong generalization to an independent dataset, where real data diversity can effectively limit bias. D.-M. Cristea et al. /Procedia Computer Science 00 (2025) 000–000 9 In Exp 3, the model trained on seven different datasets performed well on the independent BKAI dataset. Each training dataset contributed with unique examples. The model that train on all these varied examples generalized to BKAI, which likely contains a mix of these characteristics. BKAI NeoPolyp is a multi-center dataset itself, so a broadly trained model is necessary. Achieving a Dice of ∼0.81 on BKAI is impressive given that no BKAI images were used during training. For context, a model trained solely on one dataset (e.g., CVC-ColonDB) would likely perform much worse on BKAI due to domain gaps. This result is in line with findings from multi-center studies [2] and the design of larger datasets such as PolypGen [3]. Balancing Generalization Improvements and Augmentation Bias. The improvements observed in Exp 2 are a result of improvements in segmenting the same distribution. However, the decrease seen in Exp 4 suggests that the augmented training distribution over-optimized for synthetic cues at the expense of the target domain a form of domain shift. The Exp 3 model’s Dice of 0.80 on BKAI might be considered the ceiling achievable with our architecture, given the training diversity, whereas the Exp 4 model’s 0.76 represents an approximate 5.4% absolute decrease. This decrease is critical in medical imaging [17], where a few percent may lead to clinically significant errors, such as missing a small polyp. This controlled comparison demonstrates that such performance decreases are attributable to augmentation bias rather than model saturation (data with the same content). Comparison with Related Work. Our positive results with in-domain augmentation (Exp 2) reinforce that augmentation benefits limited medical datasets [9]. Meanwhile, the strong performance of the combined model (Exp 3) supports the view that training with data from diverse sources is critical for combating domain bias [3,2]. Our observation of augmentation-induced bias aligns with Alomar et al. [4]. In this study, focused on evaluating PraNet-based colorectal and gastrointestinal polyp image segmentation across normalised and augmented heterogeneous datasets, generative AI was intentionally not employed. This decision ensures that model performance is evaluated strictly on real data, avoiding entirely synthetic artefacts like unnatural-looking textures. 5. Conclusion Our findings can be summarized as follows: Data augmentation. Augmented models achieved higher Dice and IoU scores on the test sets drawn from the same distribution as the training data. This confirms that augmentations such as rotations, flips, and intensity jitter reduce overfitting and help the model learn a more complete concept of a polyp, regardless of minor transformations. Strong generalization to an independent dataset was reached by training on a combination of normalized datasets (BKAI, Dice ≈0.81, IoU ≈0.73). This finding shows how crucial data diversity is to overcoming dataset bias. The model acquired a clear representation of polyp traits by combining results from several centers, which translated to data that was not visible. When augmentation was introduced to the combined training, performance on the independent test somewhat declined, resulting in augmentation-induced bias. This suggests that certain augmentations introduced a distribution supervised to overfitting on synthetic patterns that did not generalize well. Research provides a detailed analysis of how data normalization and augmentation influence polyp segmentation outcomes. We demonstrated that normalization is indispensable and that augmentation, while powerful, must be applied accordingly, especially in multi-center training scenarios. Our work not only aids in the development of more polyp segmentation models but also serves as a methodological reference for addressing data processing in medical image segmentation [1]. Limitations and Future Directions. However, some limitations exist in our study that compares four scenarios using PraNet. First, we employed only a single segmentation architecture; alternative architectures (e.g., Transformer-based models) may respond differently to augmentation. Second, our