The rapid expansion of high-throughput omics analyses has revolutionized the field of biomedicine, yet the reliability of computational models analyzing such data remains limited by biases, noisy signals, and suboptimal validation practices. This thesis advances the theoretical and methodological foundations of machine learning for computational biology and precision oncology, focusing on two central challenges: (1) robust evaluation of Drug Response Prediction (DRP) models, and (2) accurate quantification of residual batch effects in single-cell RNA sequencing (scRNA-seq). In the first part of this thesis, we focus on DRP, a central task in precision oncology aimed at predicting cancer cell line responses to therapeutic compounds. We introduce our NxtDRP model, which integrates non-linear entity–relation data fusion, and use it as a flexible framework to better understand the behavior and limitations of DRP systems. We show that typical benchmarks overestimate model generalizability due to latent dataset biases and inadequate management of validation splits. We illustrate that global performance metrics computation, regardless of a model’s assumptions, may actually be “gamed” by models that learn spurious correlations, a phenomenon we refer to as specification gaming. To address the problem, we propose a reliable benchmarking paradigm for validating DRP models that requires novel Aggregation Strategies for metrics computation, including Global, Fixed-Drug, as well as Fixed-Cell Line, that decouple overall generalizability across various drugs as well as cell lines. Further, we indicate that IC50-derived target labels continue to be biased by experimental design factors such as maximum concentrations tested, proposing a paradigm shift towards instead utilizing Area Under the Dose–Response Curve as the target prediction label for DRP. The second part of this thesis focuses on batch effect correction in scRNA-seq data, with the goal of assessing how effectively current methods remove technical variation. To address this problem, we introduce the Batch Probing Score (BPS), a supervised metric that quantifies residual batch signal by directly testing whether batch labels remain predictable after correction. Unlike widely used unsupervised integration metrics, which rely on geometric properties of low-dimensional embeddings, BPS provides a more direct and conservative estimate of technical bias. Across multiple real datasets and controlled synthetic experiments, we show that existing evaluation metrics often lack sensitivity and specificity, leading to overly optimistic assessments of batch correction quality. In contrast, BPS consistently detects machine-learning actionable batch signal even when standard metrics suggest successful integration. Our results indicate that most batch correction methods substantially reduce apparent batch structure but leave residual signal exploitable by downstream models, with deep learning approaches generally performing better than classical methods. Moreover, BPS offers an interpretable framework to identify genes driving residual technical variation and to estimate the potential impact of batch effects on downstream analyses. Together, these works provide an in-depth analysis of how validation procedures are crucial in bioinformatics and computational biology, and how they could and should be improved. This thesis contributes to the development of more trustworthy, generalizable, and biologically relevant machine learning models for precision medicine
Reliable Inference in Computational Biomedicine: Evaluation Protocols for Omics-Based Models(2026 Jul 23).
Reliable Inference in Computational Biomedicine: Evaluation Protocols for Omics-Based Models
CODICE', FRANCESCO
2026-07-23
Abstract
The rapid expansion of high-throughput omics analyses has revolutionized the field of biomedicine, yet the reliability of computational models analyzing such data remains limited by biases, noisy signals, and suboptimal validation practices. This thesis advances the theoretical and methodological foundations of machine learning for computational biology and precision oncology, focusing on two central challenges: (1) robust evaluation of Drug Response Prediction (DRP) models, and (2) accurate quantification of residual batch effects in single-cell RNA sequencing (scRNA-seq). In the first part of this thesis, we focus on DRP, a central task in precision oncology aimed at predicting cancer cell line responses to therapeutic compounds. We introduce our NxtDRP model, which integrates non-linear entity–relation data fusion, and use it as a flexible framework to better understand the behavior and limitations of DRP systems. We show that typical benchmarks overestimate model generalizability due to latent dataset biases and inadequate management of validation splits. We illustrate that global performance metrics computation, regardless of a model’s assumptions, may actually be “gamed” by models that learn spurious correlations, a phenomenon we refer to as specification gaming. To address the problem, we propose a reliable benchmarking paradigm for validating DRP models that requires novel Aggregation Strategies for metrics computation, including Global, Fixed-Drug, as well as Fixed-Cell Line, that decouple overall generalizability across various drugs as well as cell lines. Further, we indicate that IC50-derived target labels continue to be biased by experimental design factors such as maximum concentrations tested, proposing a paradigm shift towards instead utilizing Area Under the Dose–Response Curve as the target prediction label for DRP. The second part of this thesis focuses on batch effect correction in scRNA-seq data, with the goal of assessing how effectively current methods remove technical variation. To address this problem, we introduce the Batch Probing Score (BPS), a supervised metric that quantifies residual batch signal by directly testing whether batch labels remain predictable after correction. Unlike widely used unsupervised integration metrics, which rely on geometric properties of low-dimensional embeddings, BPS provides a more direct and conservative estimate of technical bias. Across multiple real datasets and controlled synthetic experiments, we show that existing evaluation metrics often lack sensitivity and specificity, leading to overly optimistic assessments of batch correction quality. In contrast, BPS consistently detects machine-learning actionable batch signal even when standard metrics suggest successful integration. Our results indicate that most batch correction methods substantially reduce apparent batch structure but leave residual signal exploitable by downstream models, with deep learning approaches generally performing better than classical methods. Moreover, BPS offers an interpretable framework to identify genes driving residual technical variation and to estimate the potential impact of batch effects on downstream analyses. Together, these works provide an in-depth analysis of how validation procedures are crucial in bioinformatics and computational biology, and how they could and should be improved. This thesis contributes to the development of more trustworthy, generalizable, and biologically relevant machine learning models for precision medicine| File | Dimensione | Formato | |
|---|---|---|---|
|
Tesi-Codice-Francesco.pdf
Accesso aperto
Descrizione: Tesi
Dimensione
8.31 MB
Formato
Adobe PDF
|
8.31 MB | Adobe PDF | Visualizza/Apri |
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.



