Full text
Transfer Learning with Enhanced Models for Skin Cancer Detection: A Comprehensive Evaluation of Transfer Learning and Data Augmentation on the ISIC 2020 Dataset SAADNA Yassmina1, MEZZOUDJ Saliha2, and KHELIFA Meriem3 1Lastic Laboratory, Department of Mathematics and Computer Science, University of Batna 2 Mostefa Ben Boula¨ıd, Batna, Algeria, [email protected] 2Faculty of Sciences, Department of Computer Science, University of Algiers 1, Algiers, Algeria, [email protected] 3Department of Computer Science and Information Technology, University of Kasdi Merbah Ouargla, Ouargla, Algeria Abstract Skin cancer, encompassing melanoma and non-melanoma variants, remains a prevalent global malignancy, necessitating timely detection to enhance patient outcomes. This study employs transfer learning with pre-trained convolutional neural networks (CNNs)—VGG16, DenseNet121, ResNet50, and InceptionV3—to classify skin lesions as benign or malignant using the ISIC 2020 dataset of 17,755 dermoscopic images. We evaluated baseline models, data augmentation effects, and enhanced architectures with additional trainable layers. Enhanced DenseNet121 achieved superior performance, with 96.45% accuracy, 96.32% precision, 96.69% recall, and a 96.50% F1-score. Data augmentation, however, reduced accuracy, underscoring its context-specific limitations. These findings highlight the efficacy of enhanced transfer learning for automated skin cancer diagnostics, offering a scalable, precise solution. Keywords: Skin Cancer, Melanoma, Transfer Learning, Deep Learning, Convolutional Neural Networks, ISIC 2020 1 Introduction Skin cancer ranks among the most prevalent malignancies globally, with an estimated 1.5 million new cases annually, posing a significant public health challenge [22,18]. Melanoma, though less common, is the deadliest form due to its metastatic potential, while non-melanoma variants, such as basal cell and squamous cell carcinomas, contribute substantially to morbidity, particularly in fair-skinned populations exposed to ultraviolet radiation [12]. Early detection is paramount, as it can increase five-year survival rates for melanoma from 25% in advanced stages to over 95% when identified early [18]. However, conventional diagnostic methods—visual inspection, dermoscopy, and histopathology—are time-consuming, subjective, and reliant on expert dermatologists, who are often scarce, especially in resource-limited regions [7]. This gap underscores the urgent need for automated, accurate, and accessible diagnostic tools to bridge disparities in skin cancer care. Deep learning, particularly convolutional neural networks (CNNs), has emerged as a transformative approach in medical imaging, offering the potential to automate skin cancer detection with high precision [11]. Yet, training CNNs from scratch requires vast labeled datasets, a resource rarely available in medical contexts due to data scarcity and annotation challenges [17]. Transfer learning addresses this limitation by adapting pre-trained CNNs, originally trained on large-scale natural image datasets like ImageNet, to specialized medical tasks [15]. Despite its success, challenges persist, including domain shifts between natural and dermoscopic images and the inconsistent performance of data augmentation, which can degrade rather than enhance diagnostic accuracy [8]. These issues highlight the need for innovative strategies to optimize transfer learning for skin cancer detection. This study tackles these challenges by investigating transfer learning with enhanced models for classifying skin lesions as benign or malignant using the ISIC 2020 dataset, comprising 17,755 dermoscopic images. We systematically compare four pre-trained CNNs—VGG16, DenseNet121, ResNet50, and InceptionV3—across three configurations: baseline transfer learning, data augmentation, and enhanced models with additional trainable layers. Our primary contribution is the development of an enhanced 2
DenseNet121 model that achieves a state-of-the-art accuracy of 96.45%, surpassing baseline performance (94%) and demonstrating robustness against data augmentation’s limitations. Key strengths include the model’s ability to leverage dense connectivity for superior feature extraction, its adaptability to dermoscopic-specific features through architectural enhancements, and its potential for clinical deployment in resource-constrained settings. By providing a comprehensive evaluation, including training/validation curves and confusion matrices, and a novel enhancement strategy, this work advances the field of AI-driven skin cancer diagnostics, offering a scalable, high-precision solution that rivals expertlevel accuracy. 2 Related Works The integration of artificial intelligence into healthcare has reshaped diagnostics, with deep learning demonstrating exceptional proficiency in image-based disease detection. In dermatology, convolutional neural networks (CNNs) have shown promise in identifying subtle lesion patterns, driven by seminal works and ongoing advancements in transfer learning. Esteva et al. (2017) set a benchmark by achieving dermatologist-level accuracy (91%) using a fine-tuned InceptionV3 model on a dataset of 129,450 images, establishing the feasibility of AI-driven skin cancer detection. This work underscored transfer learning’s potential to adapt pre-trained models, trained on large-scale datasets like ImageNet, to medical imaging tasks with limited data. Subsequent studies have expanded this paradigm. Brinker et al. (2019) pitted a CNN against 136 dermatologists, achieving 89% accuracy on dermoscopic images, highlighting AI’s competitive edge. Haenssle et al. (2018) reported a CNN’s superior performance (95% sensitivity) over 58 dermatologists, reinforcing clinical relevance. These efforts often leverage datasets like HAM10000, introduced by Tschandl et al. (2018) with 10,015 multi-source dermoscopic images, and ISIC 2020, a standardized benchmark for lesion classification. However, challenges persist, including domain shifts between natural and dermoscopic images and variable data augmentation efficacy. Recent advancements have focused on optimizing transfer learning for skin cancer detection, particularly with ISIC 2020. Nawaz et al. (2022) combined deep learning with fuzzy k-means clustering, achieving accuracies of 95.4%, 93.1%, and 95.6% across ISBI-2016, ISIC-2017, and PH2 datasets, respectively. Rashid et al. (2022) utilized MobileNetV2 with data augmentation, reporting 98.2% accuracy on ISIC 2020, emphasizing lightweight models. Lee et al. (2023) compared CNNs (e.g., ResNet, Inception) on ISIC 2020, achieving up to 95% accuracy, providing a direct benchmark. Within Artificial Intelligence in Medicine, Khan et al. (2023) enhanced CNNs with optimization techniques, achieving 98.2% accuracy on ISIC 2020, while Hosny et al. (2022) developed a hybrid CNN model, reporting 96% accuracy with multi-source integration. Additional AIM contributions include Mishra et al. (2022), who employed transfer learning with DenseNet for melanoma detection, achieving 94.8% accuracy on ISIC 2019, and Zhang et al. (2021), who introduced a multi-task CNN framework for skin lesion segmentation and classification, yielding 92.5% accuracy. Cicalese et al. (2023) further advanced diagnostics with a generative adversarial network (GAN) to synthesize dermoscopic images, improving classification by 3% over baseline CNNs on ISIC data. Data augmentation’s role remains debated. Johnson et al. (2022) explored domain-specific augmentation, achieving 93% accuracy by preserving lesion features, contrasting with generic methods’ limitations, as reviewed by Shorten and Khoshgoftaar (2019). Zhang et al. (2019) introduced attention mechanisms, improving classification to 94% accuracy. Architectural enhancements have also progressed, with Huang et al.’s (2017) DenseNet introducing dense connectivity, He et al.’s (2016) ResNet addressing gradient issues, Szegedy et al.’s (2016) InceptionV3 optimizing multi-scale extraction, and Simonyan and Zisserman’s (2014) VGG16 emphasizing depth. In AIM, Li et al. (2020) applied transfer learning to retinal imaging, adapting CNNs for disease classification with 95.2% accuracy, paralleling dermoscopic efforts. Dosovitskiy et al. (2021) proposed Vision Transformers, suggesting future directions. Our work extends these efforts by enhancing DenseNet121 with trainable layers, addressing domain shifts and outperforming baseline models, contributing to AI-driven skin cancer diagnostics. 3
3 Methods 3.1 Dataset Robust datasets underpin deep learning efficacy in medical imaging [2]. We utilized a subset of the ISIC 2020 dataset, ”ISIC2020 60 40,” comprising 17,755 dermoscopic images sourced from Kaggle. The dataset was split into training (10,653 images: 5,400 benign, 5,253 malignant) and testing (7,103 images: 3,600 benign, 3,502 malignant) sets, with images resized to 256 ×256 ×3 pixels for computational efficiency. Other datasets considered include HAM10000 (10,015 images across seven classes) [21] and PH2 (200 RGB dermoscopic images) [13]. ISIC 2020 was selected for its scale and binary focus, aligning with clinical diagnostic needs. 3.2 Pre-trained Models Four CNNs were selected for their architectural diversity and proven efficacy in transfer learning [15]: •VGG16: Features 16 layers with 3 ×3 filters, pre-trained on ImageNet, emphasizing depth for detailed feature extraction [19]. •DenseNet121: Comprises 121 layers with dense connectivity, pre-trained on ImageNet, enhancing feature reuse and efficiency [6]. •ResNet50: Employs 50 layers with residual connections, pre-trained on ImageNet, mitigating gradient issues in deep networks [5]. •InceptionV3: Utilizes 42 layers with multi-scale convolutions, pre-trained on ImageNet, for robust feature capture [20]. ImageNet, with 1.2 million natural images across 1,000 classes, provided initial weights [17]. 3.3 Experimental Design Experiments were conducted using Python 3.8 with Keras on Google Colab. Three configurations were assessed: •Baseline: Simple transfer learning, freezing pre-trained layers and adding a softmax output layer. •Data Augmentation: Applied via Keras’ ImageDataGenerator using pixel normalization (1./255), 90°rotation, 0.2 width/height shift, 0.2 shear, 0.2 zoom, and horizontal flip. •Enhanced Models: Augmented pre-trained bases with trainable convolutional (ReLU activation) and dense (dropout) layers, enabling task-specific adaptation beyond final-layer retraining. Hyperparameters (Table 1) were optimized for convergence and performance evaluation. Table 1: Experimental hyperparameters. Parameter Value Loss Function Multi-class cross-entropy Activation Functions ReLU (hidden), Softmax (output) Optimizer Adam Batch Size 128 Epochs 40 Metrics Accuracy, Precision, Recall, F1-Score, Confusion Matrix 3.4 Evaluation Metrics Performance was quantified using accuracy, precision, recall, F1-score, and confusion matrices to assess classification across benign and malignant classes. Training and validation curves were analyzed to evaluate convergence and generalization. 4
4 Results 4.1 Baseline Models Baseline models established initial performance: DenseNet121 achieved 94% accuracy (15% loss), followed by InceptionV3 (88%, 26%), VGG16 (85%, 35%), and ResNet50 (76%, 52%). Confusion matrices (Figure 1) showed DenseNet121 misclassified 135 benign and 240 malignant lesions, outperforming others (e.g., ResNet50: 446 benign, 1,233 malignant). Training and validation curves (Figure 2) indicated stable convergence for DenseNet121, with minimal overfitting compared to ResNet50’s higher validation loss. Detailed metrics are in Table 2. Table 2: Baseline model performance metrics. Model F1-Score Precision Recall VGG16 0.8574 0.8314 0.8850 ResNet50 0.7898 0.8761 0.9300 DenseNet121 0.9487 0.9352 0.9625 InceptionV3 0.8875 0.9176 0.8594 Figure 1: Confusion matrices for baseline models. 4.2 Data Augmentation Data augmentation reduced accuracy: DenseNet121 dropped to 82% (40% loss), InceptionV3 to 77% (46%), VGG16 to 79% (48%), and ResNet50 to 69% (60%). Confusion matrices (Figure 3) revealed increased errors (e.g., DenseNet121: 125 benign, 1,163 malignant), reflecting disrupted feature recognition. Training curves (Figure 4) showed higher volatility and loss divergence, indicating poor generalization. Metrics are in Table 3. Table 3: Performance metrics with data augmentation. Model F1-Score Precision Recall VGG16 0.7737 0.8582 0.7044 ResNet50 0.8122 0.8524 0.4764 DenseNet121 0.8436 0.7492 0.9653 InceptionV3 0.7857 0.7585 0.8150 5
Figure 2: Training and validation curves for baseline models. 4.3 Enhanced Models Enhanced models improved outcomes: DenseNet121 reached 96.45% accuracy (10% loss), VGG16 93.49% (18%), ResNet50 94.07% (20%), and InceptionV3 93.68% (22%). Confusion matrices (Figure 5) showed DenseNet121 misclassified only 119 benign and 133 malignant lesions, minimizing errors. Training and validation curves (Figure 6) exhibited smooth convergence and low loss, confirming robust generalization. Metrics are in Table 4. Table 4: Enhanced model performance metrics. Model Accuracy F1-Score Precision Recall VGG16 93.49% 93.65% 92.52% 94.81% ResNet50 94.07% 94.08% 95.19% 93.00% DenseNet121 96.45% 96.50% 96.32% 96.69% InceptionV3 93.68% 93.55% 96.93% 90.39% 6
Figure 3: Confusion matrices with data augmentation. 5 Discussion Enhanced DenseNet121’s standout performance (96.45% accuracy) underscores the value of architectural augmentation in transfer learning, as evidenced by Table 5. Baseline accuracy (94%) dropped to 82% with augmentation, rebounding to 96.45% with enhancements—a pattern mirrored across models (e.g., VGG16: 85% to 79% to 93.49%). Training curves (Figures 2,4,6) and confusion matrices (Figures 1,3, 5) reveal why: baseline models leveraged pre-trained weights well, augmentation disrupted key features, and enhancements restored and refined them. The consistent drop in accuracy across all models with data augmentation—from 94% to 82% for DenseNet121—likely stems from the disruption of critical dermoscopic features like lesion asymmetry and border irregularity, essential for malignancy detection in the ISIC 2020 dataset. Geometric transformations such as 90◦rotation and horizontal flipping, applied to the 10,653 training images, may have altered these diagnostic markers, misaligning them with the dataset’s centered lesion patterns and causing a surge in false negatives (e.g., 1,163/3,502 for DenseNet121). Given the dataset’s size and diversity, these augmentations introduced noise rather than beneficial variance, a contrast to their efficacy in natural image tasks. Pre-trained models, initialized on ImageNet, struggled to adapt to these distortions, as dermoscopic images demand specific feature preservation unlike the broader textures of natural scenes. Similar performance declines with augmentation have been observed by Pooch et al. [14] found reduced accuracy in chest radiograph classification due to domain shifts, while Chlap et al. [1] noted degraded radiotherapy model outcomes from excessive geometric changes, also Johnson et al [9] underscoring the need for domain-specific strategies in medical imaging. Table 5: Comparative accuracy across configurations. Model Baseline Accuracy Augmented Accuracy Enhanced Accuracy VGG16 85% 79% 93.49% ResNet50 76% 69% 94.07% DenseNet121 94% 82% 96.45% InceptionV3 88% 77% 93.68% DenseNet121’s edge lies in its dense connectivity [6], where each layer accesses all prior outputs, fostering feature reuse (e.g., edges from early layers inform deeper lesion pattern detection). This contrasts with VGG16’s linear depth, which redundantly relearns features, or ResNet50’s residuals, which mitigate gradients but lack DenseNet’s efficiency (fewer parameters: ∼7M vs. ResNet50’s ∼25M). Added convolutional and dense layers with ReLU and dropout further tuned this advantage, adapting ImageNet-derived filters to dermoscopic specifics—likely prioritizing irregular borders or pigment variations over generic 7
Figure 4: Training and validation curves with data augmentation. textures. The drop from 252 total errors (baseline) to 252 (enhanced) reflects this, with false negatives halving (240 to 133), vital for avoiding missed diagnoses. Augmentation’s failure (82% accuracy) likely stems from altering medically significant features—e.g., flipping a lesion might obscure asymmetry, a malignancy marker [9]. This contrasts with natural image tasks where such distortions aid robustness, highlighting a domain mismatch. Enhanced models counter this by learning task-specific filters, evidenced by tighter training/validation alignment (Figure 6) and a 2.45% accuracy gain over baseline—statistically notable given the 7,103-image test set (approximate 95% confidence interval: ±0.8%). Compared to prior work, 96.45% approaches Rashid et al.’s 98.2% with MobileNetV2 [16] and Khan et al.’s 98.2% with optimized CNNs [10], surpassing Esteva et al.’s 91% [3]. Unlike MobileNetV2’s lightweight focus, DenseNet121 balances complexity and precision, suiting clinical deployment where sensitivity (96.69%) outweighs speed. Misclassification analysis suggests most errors are false positives (119 benign), tolerable in screening as they trigger further checks, unlike false negatives (133 malignant), which risk delayed treatment—still, a 3.80% miss rate rivals expert dermatologists (e.g., Haenssle et al.’s 95% sensitivity [4]). Limitations include ISIC 2020’s binary focus—multi-class datasets like HAM10000 could test generalization—and augmentation’s context-specific failure, warranting tailored strategies [9]. 8
Figure 5: Confusion matrices for enhanced models. 6 Conclusion This study demonstrates the efficacy of transfer learning with enhanced convolutional neural networks for skin cancer detection, achieving 96.45% accuracy, 96.32% precision, 96.69% recall, and a 96.50% F1-score with DenseNet121 on the ISIC 2020 dataset of 17,755 dermoscopic images. Incorporating trainable layers improved performance over baseline transfer learning (94% to 96.45%), harnessing dense connectivity to optimize feature reuse, gradient propagation, and adaptation of pre-trained weights to dermoscopic characteristics, including irregular borders and pigment variations. This enhancement yielded a false negative rate of 3.80% (133/3,502), critical for early malignancy identification, and a false positive rate of 3.31% (119/3,600), supporting its utility in clinical screening workflows requiring subsequent validation. In contrast, data augmentation reduced accuracy (94% to 82% for DenseNet121), exposing its limitations in medical imaging contexts. Geometric transformations—rotation, flipping, shifting, shearing, and zooming—altered diagnostic features such as asymmetry and border irregularity, increasing false negatives to 33.21% (1,163/3,502) and disrupting training stability. This divergence from its benefits in natural image domains underscores a domain-specific mismatch, with the 10,653 training images proving sufficient for baseline generalization without augmentation. These findings affirm a scalable, high-sensitivity approach for automated skin cancer detection, particularly valuable in resource-constrained environments. Future investigations should prioritize domainadapted augmentation strategies, such as color-based adjustments, assess multi-class classification on diverse datasets, and evaluate advanced architectures or multi-modal inputs integrating dermoscopy with patient data to refine diagnostic accuracy further. This work advances the integration of AI into clinical dermatology by optimizing model architecture while highlighting the need for tailored data preprocessing. References [1] Phillip Chlap, Hang Min, Nicholas Vandenberg, Jason Dowling, Lois Holloway, and Annette Haworth. A review of medical image data augmentation techniques for deep learning applications. Journal of Medical Imaging and Radiation Oncology, 65(5):545–563, 2021. [2] N. C. Codella et al. Skin lesion analysis toward melanoma detection 2018: A challenge hosted by isic. arXiv preprint arXiv:1902.03368, 2019. 9
Figure 6: Training and validation curves for enhanced models. [3] A. Esteva et al. Dermatologist-level classification of skin cancer with deep neural networks. Nature, 542(7639):115–118, 2017. [4] H. A. Haenssle et al. Man against machine: Diagnostic performance of a deep learning convolutional neural network for dermoscopic melanoma recognition. Annals of Oncology, 29(8):1836–1842, 2018. [5] K. He et al. Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 770–778, 2016. [6] G. Huang et al. Densely connected convolutional networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 4700–4708, 2017. [7] International Dermoscopy Society. International dermoscopy society consensus – terminology for dermoscopy. Journal of the American Academy of Dermatology, 85(3):678–687, 2021. 10