Original Research A benchmark for neural network robustness in skin cancer classification

Maron, Roman C. and Schlager, Justin G. and Haggenmuller, Sarah and von Kalle, Christof and Utikal, Jochen S. and Meier, Friedegund and Gellrich, Frank F. and Hobelsberger, Sarah and Hauschild, Axel and French, Lars and Heinzerling, Lucie and Schlaak, Max and Ghoreschi, Kamran and Hilke, Franz J. and Poch, Gabriela and Heppt, Markus and Berking, Carola and Haferkamp, Sebastian and Sondermann, Wiebke and Schadendorf, Dirk and Schilling, Bastian and Goebeler, Matthias and Krieghoff-Henning, Eva and Hekler, Achim and Frohling, Stefan and Lipka, Daniel B. and Kather, Jakob N. and Brinker, Titus J. (2021) Original Research A benchmark for neural network robustness in skin cancer classification. EUROPEAN JOURNAL OF CANCER, 155. pp. 191-199. ISSN 0959-8049, 1879-0852

Full text not available from this repository. (Request a copy)

Abstract

Background: One prominent application for deep learning-based classifiers is skin cancer classification on dermoscopic images. However, classifier evaluation is often limited to holdout data which can mask common shortcomings such as susceptibility to confounding factors. To increase clinical applicability, it is necessary to thoroughly evaluate such classifiers on out-of-distribution (OOD) data. Objective: The objective of the study was to establish a dermoscopic skin cancer benchmark in which classifier robustness to OOD data can be measured. Methods: Using a proprietary dermoscopic image database and a set of image transformations, we create an OOD robustness benchmark and evaluate the robustness of four different convolutional neural network (CNN) architectures on it. Results: The benchmark contains three data sets-Skin Archive Munich (SAM), SAM -corrupted (SAM-C) and SAM-perturbed (SAM-P)-and is publicly available for download. To maintain the benchmark's OOD status, ground truth labels are not provided and test results should be sent to us for assessment. The SAM data set contains 319 unmodified and biopsy-verified dermoscopic melanoma (n = 194) and nevus (n = 125) images. SAM-C and SAM-P contain images from SAM which were artificially modified to test a classifier against low-quality inputs and to measure its prediction stability over small image changes, respectively. All four CNNs showed susceptibility to corruptions and perturbations. Conclusions: This benchmark provides three data sets which allow for OOD testing of binary skin cancer classifiers. Our classifier performance confirms the shortcomings of CNNs and provides a frame of reference. Altogether, this benchmark should facilitate a more thorough evaluation process and thereby enable the development of more robust skin cancer classifiers. 2021 The Author(s). Published by Elsevier Ltd. This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/).

Item Type:	Article
Uncontrolled Keywords:	IMAGE CLASSIFICATION; LEVEL CLASSIFICATION; DERMATOLOGISTS; SUPERIOR; Benchmarking; Dermatology; Skin neoplasms; Melanoma; Nevus; Artificial intelligence; Deep learning
Subjects:	600 Technology > 610 Medical sciences Medicine
Divisions:	Medicine > Lehrstuhl für Dermatologie und Venerologie
Depositing User:	Dr. Gernot Deinzer
Date Deposited:	27 Sep 2022 09:55
Last Modified:	27 Sep 2022 09:55
URI:	https://pred.uni-regensburg.de/id/eprint/48058

Actions (login required)

View Item