Dear Editor, Automated image classification by machine vision is evolving, largely because of the increasing efficiency of neural networks with special architectures such as convolutional neural networks (CNNs).1 A recent study comparing board‐certified dermatologists with CNNs suggested that CNNs may achieve diagnostic accuracies similar to those of human experts.2 We demonstrated that medical students without prior knowledge of dermatoscopy learn equally well from an analytical or heuristic teaching approach.3 During heuristic teaching, representative images of each class are presented together with the correct diagnosis without any further explanations, mimicking training of a neural network. In this experiment we trained a CNN (pretrained for general‐purpose image classification) and a group of previously untrained medical students to diagnose pigmented skin lesions on dermatoscopic images, and compared their diagnostic performance after training. We invited final‐year medical students without prior knowledge of dermatoscopy to participate in a 1‐h training session.3 During this training, the students saw dermatoscopic images of 298 pigmented lesions: 40 basal cell carcinomas (BCCs), seven dermatofibromas, 62 melanomas, 129 melanocytic naevi, 38 seborrhoeic keratoses and 22 vascular lesions. The instructor mentioned the diagnosis but did not explain any diagnostic features. After the training session the students had to rate a test set of 50 randomly selected pigmented lesions including 10 melanomas, 10 BCCs, 14 naevi, nine seborrhoeic keratoses and five vascular lesions or dermatofibromas. The students were asked to make a specific diagnosis and to indicate the probability of malignancy (scale from 1 to 5). No additional information (age, site or medical history) was provided. All images from the students' training session were also used to retrain the last layer of the ‘GoogLeNet Inception v3' neural network,4 without any kind of test‐set augmentation (4000 epochs, learning rate 0·001, batch size 50). We used the retrained network on the same test set as the students. The diagnosis was considered to indicate a malignant lesion if the top class of the softmax outputs was BCC or melanoma. The softmax outputs indicate probabilities for each class and were used to calculate receiver operating characteristics (ROC) curves. We calculated the areas under the curves (AUCs) and their confidence intervals (CIs) with the package pROC,5 and used epiR (in R, https://www.r-project.org) for the calculation of diagnostic values. We assessed the inter‐rater agreement by calculating Cohen's kappa for each class and each reader pair using the irr package in R. The study was approved by the institutional review board of the Medical University of Vienna (no. 1285/2014). Twenty‐seven students (mean age 24·9 ± 2·3 years, 74% female) completed the training session and evaluated the test set. The retrained neural network (R‐NN) correctly identified malignant lesions (BCC and melanoma) with a sensitivity of 90% (95% CI 68–99) and a specificity of 71% (95% CI 51–87). The sensitivity was slightly higher than that achieved by the students (mean 86%, 95% CI 83–88) but the R‐NN specificity was slightly lower than that of the students (mean 79%, 95% CI 74–83). For the R‐NN, the AUC for correctly predicting malignancy was 0·91 (95% CI 0·81–0·97), which was similar to the AUC of the pooled student ratings of 0·85 (95% CI 0·87–0·91, P = 0·64; Fig. 1a). (a) Diagnostic accuracy. The retrained neural network (R‐NN) and students (pooled ratings) showed similar diagnostic accuracy in detecting skin cancer (basal cell carcinoma and melanoma) as measured by the areas under the receiver operating characteristic curves. (b) Misclassified lesions. Correct (black) and incorrect (red) specific diagnoses of every case (rows 1–48) by every student (columns 1–37) and the R‐NN, and the mean correct ratings of the students. A striking similarity is seen between misdiagnosed lesions of the students and the R‐NN. The R‐NN made a correct specific diagnosis in 69% of lesions. Two of 10 melanomas were misclassified as naevi and two as BCC. Cases misdiagnosed by the R‐NN were also misdiagnosed by 46% of students. Of the 13 cases that were misdiagnosed by more than half of the students, nine were also misclassified by the R‐NN (Fig. 1b). Agreement on specific diagnoses between the R‐NN and students was moderate (mean Cohen's kappa 0·57, 95% CI 0·52–0·61) and similar to the average agreement between students. We show that an R‐NN achieves a diagnostic accuracy similar to that of medical students under comparable training conditions. The neural network (GoogLeNet Inception v3)4 was pretrained on more than 1 million images for the purpose of general image classification but was not pretrained to diagnose skin cancer. During training, students and the R‐NN had to extract features that discriminate benign from malignant pigmented skin lesions. Given the small training set, the students and the R‐NN completed the task surprisingly well with similar diagnostic performance. We also showed that misclassified lesions overlapped and shared morphological features with other classes. It seems that the R‐NN and medical students extracted similar morphological features. This supports the assumption that the pretrained architecture of the neural network is able to imitate the visual cortex of the human brain on a higher level, although further studies are needed to confirm this. Funding sources: none. Conflicts of interest: none declared.
No takes yet. Share an insight, caveat, or question.
Tschandl et al. (2017) studied this question.
Synapse has enriched 3 closely related papers on similar clinical questions. Consider them for comparative context: