The classification of natural products (NPs) remains a central challenge in cheminformatics, due to their structural complexity and the limited availability of curated data. In this study, we explore whether general-purpose language models—pre-trained on non-chemical corpora—can achieve performance comparable to models explicitly trained on molecular data. We evaluate six models, including ChemBERTa-zinc-base-v1, ChemBERTa-77M-MLM, LLaMA, GPT-2, MPNet, and MiniLM, across two established datasets (REFMET and NPClassifier), using both unsupervised clustering (K-Means) and supervised classification (logistic regression, MLP). In addition, we assess model robustness under input degradation, introducing systematic SMILES occlusion and additive embedding noise. Our results show that general-purpose models can match or even surpass domain-specific alternatives, not only in clean classification scenarios but also under perturbation. Notably, general-purpose models demonstrate strong resilience and generalization, underscoring the transferability of domain-agnostic architectures to molecular tasks. These findings support the use of general-purpose language models as reliable tools for natural product classification, highlighting their potential to address complex challenges in cheminformatics through large-scale pre-trained architectures. Project page: https://github.com/bcorrad/MetaboLM25.
Corradini, B.T., Bianchini, M., Scarselli, F., Nigi, L., Sali, V., Prete, A.L. (2027). Comparing General and Domain-Specific Pre-trained Language Models for Natural Product Classification. In Artificial Life and Evolutionary Computation - Proceedings of WIVACE 2025 (pp.61-74). Cham : Springer [10.1007/978-3-032-33185-4_5].
Comparing General and Domain-Specific Pre-trained Language Models for Natural Product Classification
Bianchini, Monica;Scarselli, Franco;Prete, Alessia Lucia
2027-01-01
Abstract
The classification of natural products (NPs) remains a central challenge in cheminformatics, due to their structural complexity and the limited availability of curated data. In this study, we explore whether general-purpose language models—pre-trained on non-chemical corpora—can achieve performance comparable to models explicitly trained on molecular data. We evaluate six models, including ChemBERTa-zinc-base-v1, ChemBERTa-77M-MLM, LLaMA, GPT-2, MPNet, and MiniLM, across two established datasets (REFMET and NPClassifier), using both unsupervised clustering (K-Means) and supervised classification (logistic regression, MLP). In addition, we assess model robustness under input degradation, introducing systematic SMILES occlusion and additive embedding noise. Our results show that general-purpose models can match or even surpass domain-specific alternatives, not only in clean classification scenarios but also under perturbation. Notably, general-purpose models demonstrate strong resilience and generalization, underscoring the transferability of domain-agnostic architectures to molecular tasks. These findings support the use of general-purpose language models as reliable tools for natural product classification, highlighting their potential to address complex challenges in cheminformatics through large-scale pre-trained architectures. Project page: https://github.com/bcorrad/MetaboLM25.| File | Dimensione | Formato | |
|---|---|---|---|
|
WIVACECorradini.pdf
non disponiibile
Tipologia:
PDF editoriale
Licenza:
NON PUBBLICO - Accesso privato/ristretto
Dimensione
5.41 MB
Formato
Adobe PDF
|
5.41 MB | Adobe PDF | Visualizza/Apri Richiedi una copia |
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/11365/1324635
