The advent of large–scale pre–training has produced powerful computer vision models, yet their effectiveness as feature extractors on highly specialized, out–of–distribution domains remains a critical issue. To fill this gap, we systematically compare the feature extraction capabilities of different, state–of–the–art architectures on the challenging task of dermoscopic skin–lesion classification—a domain substantially different from the natural images used during pre–training. We evaluate four distinct backbone architectures for feature extraction, encompassing both unimodal and multimodal models. Our selection includes EfficientNet B0 and Vision Transformer as representative unimodal models, along with Stable Diffusion Model and CLIP as state–of–the–art multimodal models. Using a subset of the ISIC archive, we assess the quality of extracted features via both linear probing and classification with a Multi Layer Perceptron. Our experiments suggest that all architectures produce highly effective features, achieving competitive performance. Notably, multimodal models, although pre–trained for tasks other than image classification, extract features that are remarkably competitive in this context. This work can serve as a comparative guide for the selection of feature extractors when tackling classification in specialized domains, such as medical imaging. The results highlight not only the generalizability of modern architectures, but also the surprising versatility of multimodal models as powerful feature extractors for interdisciplinary tasks.
Andreini, P., Bonechi, S., Bianchini, M., Scarselli, F., Manni, F., Quercioli, M., et al. (2027). A Comparison on Cross–Domain Representation Learning for Skin–Lesion Classification. In Artificial Life and Evolutionary Computation. WIVACE 2025 (pp.47-60). Cham : Springer [10.1007/978-3-032-33185-4_4].
A Comparison on Cross–Domain Representation Learning for Skin–Lesion Classification
Bonechi, Simone;Bianchini, Monica;Scarselli, Franco;
2027-01-01
Abstract
The advent of large–scale pre–training has produced powerful computer vision models, yet their effectiveness as feature extractors on highly specialized, out–of–distribution domains remains a critical issue. To fill this gap, we systematically compare the feature extraction capabilities of different, state–of–the–art architectures on the challenging task of dermoscopic skin–lesion classification—a domain substantially different from the natural images used during pre–training. We evaluate four distinct backbone architectures for feature extraction, encompassing both unimodal and multimodal models. Our selection includes EfficientNet B0 and Vision Transformer as representative unimodal models, along with Stable Diffusion Model and CLIP as state–of–the–art multimodal models. Using a subset of the ISIC archive, we assess the quality of extracted features via both linear probing and classification with a Multi Layer Perceptron. Our experiments suggest that all architectures produce highly effective features, achieving competitive performance. Notably, multimodal models, although pre–trained for tasks other than image classification, extract features that are remarkably competitive in this context. This work can serve as a comparative guide for the selection of feature extractors when tackling classification in specialized domains, such as medical imaging. The results highlight not only the generalizability of modern architectures, but also the surprising versatility of multimodal models as powerful feature extractors for interdisciplinary tasks.| File | Dimensione | Formato | |
|---|---|---|---|
|
WIVACEAndreini.pdf
non disponiibile
Tipologia:
PDF editoriale
Licenza:
NON PUBBLICO - Accesso privato/ristretto
Dimensione
899.81 kB
Formato
Adobe PDF
|
899.81 kB | Adobe PDF | Visualizza/Apri Richiedi una copia |
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/11365/1324634
