Explainable Deep Learning for Computer-Aided Skin Cancer Detection Using CNNs and Vision Transformers
Eirini Karantina, Antreas Kantaros, Grigoris Nikolaou, Nikolaos Laskaris, Paraskevi ZachariaEarly and accurate detection of skin cancer, particularly melanoma, remains a critical challenge in computer-aided diagnosis, motivating the development of reliable and interpretable machine learning solutions. This study presents a comparative algorithmic analysis of deep learning models for automated skin cancer detection using dermoscopic images. Specifically, convolutional neural networks (CNNs) and Vision Transformers (ViTs) are implemented within a unified framework, employing transfer learning and standardized preprocessing techniques on a benchmark dataset. The proposed methodology incorporates data augmentation and class imbalance handling strategies, while model performance is evaluated using clinically relevant metrics, including accuracy, precision, recall, F1-score, and area under the ROC curve. In addition, explainability techniques such as Grad-CAM and attention visualization are employed to enhance model interpretability, and decision threshold analysis is conducted to assess trade-offs between sensitivity and specificity in melanoma detection. Experimental results demonstrate that CNN-based architectures achieve robust performance in capturing local spatial features, while transformer-based models provide competitive results through global contextual representation. However, variations are observed in model calibration and false-negative rates, which are critical for clinical deployment. Overall, the findings highlight the importance of combining algorithmic performance with interpretability and threshold optimization to support reliable and clinically meaningful computer-aided diagnosis systems.