Background: Transcriptomic biomarker discovery often fails to produce reproducible gene signatures across independent cohorts due to model-specific biases and dataset heterogeneity. While single-algorithm approaches may perform well on training data, they frequently fail to generalize effectively. Ensemble methods have proven effective in general machine learning applications, yet their systematic integration for consensus-based feature prioritization remains underexplored in transcriptomics. Results: We developed TENTACLES (Transcriptomic ExplorationTool through Aggregation of Classifiers), an open-source modular framework for robust biomarker discovery through multi-algorithm consensus. The tool is an open-source R package that integrates up to 15 supervised learning algorithms and 6 unsupervised clustering methods. The tool utilizes a modular architecture to automate data preprocessing, multi-algorithm feature prioritization, and cross-cohort validation. By aggregating variable importance across multiple models, TENTACLES identifies gene signatures resilient to algorithm-specific biases. We validated the framework using Crohn’s disease as a high-heterogeneity case study across 689 samples from four independent publicly available RNA-seq cohorts. TENTACLES identified a 28-gene consensus panel that achieved superior cross-cohort generalizability compared to single-algorithm-derived signatures and conventional differential expression methods while using, compared to the latter, 95% fewer features. This signature was further refined to a minimal 5-gene core that maintained robust discriminatory power in completely unsupervised validation. These results confirm the tool’s ability to extract stable biological signals from complex, noisy datasets. Conclusions: TENTACLES provides a scalable, disease-agnostic solution for identifying minimal reproducible gene signatures from heterogeneous transcriptomic data. By bridging the gap between complex ensemble modeling and practical biomarker discovery, the software could serve as a versatile resource for researchers aiming to derive reproducible biomarkers across diverse disease contexts.

Montesi, G., Mouta, G.D.S., Novedrati, M., Cunha, A.F., Fuschi, A., Lucchesi, S., et al. (2026). TENTACLES: a consensus machine learning tool for robust biomarker discovery in heterogeneous data. BIODATA MINING, 19(1) [10.1186/s13040-026-00568-8].

TENTACLES: a consensus machine learning tool for robust biomarker discovery in heterogeneous data

Montesi, Giorgio;Mouta, Gabriel Dos Santos;Novedrati, Maria;Lucchesi, Simone;Sonnati, Chiara;Ciabattini, Annalisa;Santoro, Francesco;Medaglini, Donata;
2026-01-01

Abstract

Background: Transcriptomic biomarker discovery often fails to produce reproducible gene signatures across independent cohorts due to model-specific biases and dataset heterogeneity. While single-algorithm approaches may perform well on training data, they frequently fail to generalize effectively. Ensemble methods have proven effective in general machine learning applications, yet their systematic integration for consensus-based feature prioritization remains underexplored in transcriptomics. Results: We developed TENTACLES (Transcriptomic ExplorationTool through Aggregation of Classifiers), an open-source modular framework for robust biomarker discovery through multi-algorithm consensus. The tool is an open-source R package that integrates up to 15 supervised learning algorithms and 6 unsupervised clustering methods. The tool utilizes a modular architecture to automate data preprocessing, multi-algorithm feature prioritization, and cross-cohort validation. By aggregating variable importance across multiple models, TENTACLES identifies gene signatures resilient to algorithm-specific biases. We validated the framework using Crohn’s disease as a high-heterogeneity case study across 689 samples from four independent publicly available RNA-seq cohorts. TENTACLES identified a 28-gene consensus panel that achieved superior cross-cohort generalizability compared to single-algorithm-derived signatures and conventional differential expression methods while using, compared to the latter, 95% fewer features. This signature was further refined to a minimal 5-gene core that maintained robust discriminatory power in completely unsupervised validation. These results confirm the tool’s ability to extract stable biological signals from complex, noisy datasets. Conclusions: TENTACLES provides a scalable, disease-agnostic solution for identifying minimal reproducible gene signatures from heterogeneous transcriptomic data. By bridging the gap between complex ensemble modeling and practical biomarker discovery, the software could serve as a versatile resource for researchers aiming to derive reproducible biomarkers across diverse disease contexts.
2026
Montesi, G., Mouta, G.D.S., Novedrati, M., Cunha, A.F., Fuschi, A., Lucchesi, S., et al. (2026). TENTACLES: a consensus machine learning tool for robust biomarker discovery in heterogeneous data. BIODATA MINING, 19(1) [10.1186/s13040-026-00568-8].
File in questo prodotto:
File Dimensione Formato  
s13040-026-00568-8.pdf

accesso aperto

Descrizione: Articolo
Tipologia: PDF editoriale
Licenza: Creative commons
Dimensione 3.49 MB
Formato Adobe PDF
3.49 MB Adobe PDF Visualizza/Apri

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11365/1325274