CINECA IRIS Institutional Research Information System

Genome interpretation (GI) encompasses the computational attempts to model the relationship between genotype and phenotype with the goal of understanding how the first leads to the second. While traditional approaches have focused on sub-problems such as predicting the effect of single nucleotide variants or finding genetic associations, recent advances in neural networks (NNs) have made it possible to develop end-to-end GI models that take genomic data as input and predict phenotypes as output. However, technical and modeling issues still need to be fixed for these models to be effective, including the widespread underdetermination of genomic datasets, making them unsuitable for training large, overfitting-prone, NNs. Here we propose novel GI models to address this issue, exploring the use of two types of transfer learning approaches and proposing a novel Biologically Meaningful Sparse NN layer specifically designed for end-to-end GI. Our models predict the leaf and seed ionome in A.thaliana, obtaining comparable results to our previous over-parameterized model while reducing the number of parameters by 8.8 folds. We also investigate how the effect of population stratification influences the evaluation of the performances, highlighting how it leads to (1) an instance of the Simpson's Paradox, and (2) model generalization limitations.

Biologically meaningful genome interpretation models to address data underdetermination for the leaf and seed ionome prediction in Arabidopsis thaliana

Passemiers, Antoine;Verplaetse, Nora;Corso, Massimiliano;Ferrero-Serrano, Ángel;Nazzicari, Nelson;Biscarini, Filippo;Fariselli, Piero;Moreau, Yves

2024-01-01

Abstract

Genome interpretation (GI) encompasses the computational attempts to model the relationship between genotype and phenotype with the goal of understanding how the first leads to the second. While traditional approaches have focused on sub-problems such as predicting the effect of single nucleotide variants or finding genetic associations, recent advances in neural networks (NNs) have made it possible to develop end-to-end GI models that take genomic data as input and predict phenotypes as output. However, technical and modeling issues still need to be fixed for these models to be effective, including the widespread underdetermination of genomic datasets, making them unsuitable for training large, overfitting-prone, NNs. Here we propose novel GI models to address this issue, exploring the use of two types of transfer learning approaches and proposing a novel Biologically Meaningful Sparse NN layer specifically designed for end-to-end GI. Our models predict the leaf and seed ionome in A.thaliana, obtaining comparable results to our previous over-parameterized model while reducing the number of parameters by 8.8 folds. We also investigate how the effect of population stratification influences the evaluation of the performances, highlighting how it leads to (1) an instance of the Simpson's Paradox, and (2) model generalization limitations.

Scheda breve

Scheda completa

Scheda completa (DC)

	Anno
	
				2024
			
	Lingua di pubblicazione
	
				Inglese
			
	Codice ISI WoS
	
				WOS:001248258900040
			
	Codice Scopus
	
				2-s2.0-85195534007
			
	Referee
	
				Esperti anonimi
			
	Titolo rivista
	
				SCIENTIFIC REPORTS
			
	N. Volume
	
				14
			
	Fascicolo
	
				1
			
	Pagine (da)
	
				1
			
	Pagine (a)
	
				11
			
	Numero di pagine totale
	
				11
			
	DOI
	
				https://dx.doi.org/10.1038/s41598-024-63855-6
			
	Parole Chiave
	
				Arabidopsis thaliana; Deep learning; Ionome prediction; Genomic prediction; Genome interpretation
			
	Coautori affiliati a enti stranieri
	
				sì
			
	Nazione dell'ente di affiliazione
	
				FRANCIA
STATI UNITI D'AMERICA
BELGIO
			
	Prodotto conforme al Regolamento di Ateneo sull'accesso aperto?
	
				1 – prodotto con  file in versione Open Access (allegherò il file al passo 6 - Carica)
			
	Tipologia sito docente
	
				262
			
	Numero autori
	
				9
			
	Tutti gli autori
	
						Raimondi, Daniele; Passemiers, Antoine; Verplaetse, Nora; Corso, Massimiliano; Ferrero-Serrano, Ángel; Nazzicari, Nelson; Biscarini, Filippo; Farisell...espandi
						
	Tipologia
	
				info:eu-repo/semantics/article
			
	Fulltext
	
				open
			
	Tipologia
	
				03-CONTRIBUTO IN RIVISTA::03A-Articolo su Rivista
			
	Appare nelle tipologie:
	
				03A-Articolo su Rivista

File in questo prodotto:

File	Dimensione	Formato
Raimondi_SciRep_2024_s41598-024-63855-6.pdf Accesso aperto Tipo di file: PDF EDITORIALE Dimensione 3.71 MB Formato Adobe PDF Visualizza/Apri	3.71 MB	Adobe PDF	Visualizza/Apri

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/2318/2028718

Citazioni

ND

0

0

social impact