The advent of large–scale pre–training has produced powerful computer vision models, yet their effectiveness as feature extractors on highly specialized, out–of–distribution domains remains a critical issue. To fill this gap, we systematically compare the feature extraction capabilities of different, state–of–the–art architectures on the challenging task of dermoscopic skin–lesion classification—a domain substantially different from the natural images used during pre–training. We evaluate four distinct backbone architectures for feature extraction, encompassing both unimodal and multimodal models. Our selection includes EfficientNet B0 and Vision Transformer as representative unimodal models, along with Stable Diffusion Model and CLIP as state–of–the–art multimodal models. Using a subset of the ISIC archive, we assess the quality of extracted features via both linear probing and classification with a Multi Layer Perceptron. Our experiments suggest that all architectures produce highly effective features, achieving competitive performance. Notably, multimodal models, although pre–trained for tasks other than image classification, extract features that are remarkably competitive in this context. This work can serve as a comparative guide for the selection of feature extractors when tackling classification in specialized domains, such as medical imaging. The results highlight not only the generalizability of modern architectures, but also the surprising versatility of multimodal models as powerful feature extractors for interdisciplinary tasks.

Andreini, P., Bonechi, S., Bianchini, M., Scarselli, F., Manni, F., Quercioli, M., et al. (2026). A Comparison on Cross–Domain Representation Learning for Skin–Lesion Classification. In M.G. Monica Bianchini (a cura di), Artificial Life and Evolutionary Computation - Proceedings of WIVACE 2025 (pp. 47-60). Springer Nature.

A Comparison on Cross–Domain Representation Learning for Skin–Lesion Classification

Paolo Andreini;Monica Bianchini;Franco Scarselli;Francesco Manni;Matilde Quercioli;Barbara Toniella Corradini
2026-01-01

Abstract

The advent of large–scale pre–training has produced powerful computer vision models, yet their effectiveness as feature extractors on highly specialized, out–of–distribution domains remains a critical issue. To fill this gap, we systematically compare the feature extraction capabilities of different, state–of–the–art architectures on the challenging task of dermoscopic skin–lesion classification—a domain substantially different from the natural images used during pre–training. We evaluate four distinct backbone architectures for feature extraction, encompassing both unimodal and multimodal models. Our selection includes EfficientNet B0 and Vision Transformer as representative unimodal models, along with Stable Diffusion Model and CLIP as state–of–the–art multimodal models. Using a subset of the ISIC archive, we assess the quality of extracted features via both linear probing and classification with a Multi Layer Perceptron. Our experiments suggest that all architectures produce highly effective features, achieving competitive performance. Notably, multimodal models, although pre–trained for tasks other than image classification, extract features that are remarkably competitive in this context. This work can serve as a comparative guide for the selection of feature extractors when tackling classification in specialized domains, such as medical imaging. The results highlight not only the generalizability of modern architectures, but also the surprising versatility of multimodal models as powerful feature extractors for interdisciplinary tasks.
2026
Andreini, P., Bonechi, S., Bianchini, M., Scarselli, F., Manni, F., Quercioli, M., et al. (2026). A Comparison on Cross–Domain Representation Learning for Skin–Lesion Classification. In M.G. Monica Bianchini (a cura di), Artificial Life and Evolutionary Computation - Proceedings of WIVACE 2025 (pp. 47-60). Springer Nature.
File in questo prodotto:
Non ci sono file associati a questo prodotto.

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11365/1324634
 Attenzione

Attenzione! I dati visualizzati non sono stati sottoposti a validazione da parte dell'ateneo