Biologically informed neural networks are increasingly adopted in bioinformatics under the premise that embedding biological knowledge into model architectures yields more accurate and interpretable predictions. This approach has driven a growing literature of pathway-informed models aiming to move beyond black-box learning by explicitly encoding biological structure. However, it remains unclear whether these models exploit biological knowledge or instead benefit from a different inductive bias. Here, we systematically investigate this question across 29 state-of-the-art pathway-informed neural networks by explicitly decoupling biological annotations from network architecture. For each evaluable model, we implement a structure-matched randomization protocol, in which pathway annotations are replaced with random associations while preserving sparsity and architectural constraints, allowing for a direct comparison under controlled conditions. Across multiple prediction tasks, datasets, and evaluation metrics, the randomized models consistently match or outperform their biologically informed counterparts. Moreover, pathway-informed models show no systematic advantage in interpretability: randomized models recover disease-associated biomarkers with comparable accuracy and yield highly correlated feature rankings. Our results reveal that the performance gains commonly attributed to biological pathway integration arise predominantly from sparsity-induced regularization rather than from biological knowledge itself. We provide a general evaluation workflow to test whether biological priors contribute predictive information beyond sparsity, offering practical guidance for the development of biology-aware neural networks. The code implementing the proposed methodology is available on GitHub at https://github.com/compbiomed-unito/Pathway_Randomization.

Sparsity is all you need: rethinking biologically informed neural networks

Caranzano, Isabella
First
;
Pancotti, Corrado;Rollo, Cesare;Sartori, Flavio;Fariselli, Piero
;
Sanavia, Tiziana
Last
2026-01-01

Abstract

Biologically informed neural networks are increasingly adopted in bioinformatics under the premise that embedding biological knowledge into model architectures yields more accurate and interpretable predictions. This approach has driven a growing literature of pathway-informed models aiming to move beyond black-box learning by explicitly encoding biological structure. However, it remains unclear whether these models exploit biological knowledge or instead benefit from a different inductive bias. Here, we systematically investigate this question across 29 state-of-the-art pathway-informed neural networks by explicitly decoupling biological annotations from network architecture. For each evaluable model, we implement a structure-matched randomization protocol, in which pathway annotations are replaced with random associations while preserving sparsity and architectural constraints, allowing for a direct comparison under controlled conditions. Across multiple prediction tasks, datasets, and evaluation metrics, the randomized models consistently match or outperform their biologically informed counterparts. Moreover, pathway-informed models show no systematic advantage in interpretability: randomized models recover disease-associated biomarkers with comparable accuracy and yield highly correlated feature rankings. Our results reveal that the performance gains commonly attributed to biological pathway integration arise predominantly from sparsity-induced regularization rather than from biological knowledge itself. We provide a general evaluation workflow to test whether biological priors contribute predictive information beyond sparsity, offering practical guidance for the development of biology-aware neural networks. The code implementing the proposed methodology is available on GitHub at https://github.com/compbiomed-unito/Pathway_Randomization.
2026
27
4
1
15
biologically informed neural networks; cancer genomics; model randomization analysis; multi-omics integration; pathway-based deep learning
Caranzano, Isabella; Pancotti, Corrado; Rollo, Cesare; Sartori, Flavio; Liò, Pietro; Fariselli, Piero; Sanavia, Tiziana
File in questo prodotto:
File Dimensione Formato  
bbag425.pdf

Accesso aperto

Tipo di file: PDF EDITORIALE
Dimensione 2.43 MB
Formato Adobe PDF
2.43 MB Adobe PDF Visualizza/Apri

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/2318/2162230
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? 0
social impact