Deep learning has become a cornerstone of computer vision applications such as image classification, object detection, and image segmentation across domains ranging from medical imaging to autonomous systems. However, deep learning demands large-scale annotated datasets for model training, which are often expensive, time-consuming, or impractical to acquire and annotate. This thesis investigates synthetic data as both a practical tool and a methodological framework for improving deep learning in data-scarce scenarios. The work was carried out within an industrial PhD in collaboration with Synesthesia srl SB and reflects a dual academic and industrial research agenda: both tracks share a common focus on understanding how synthetic data can support and enhance the training of deep learning models when real annotated data is limited or costly. The first part of this thesis explores synthetic data along three complementary directions from an academic perspective. Low-data regimes are first investigated through a case study of Byzantine seals, in which synthetic data are used to support object detection and classification. To bridge the gap between synthetic and real data, style transfer techniques are applied and analyzed, providing insight into how domain adaptation strategies affect generalization and how the synthetic-to-real domain gap can be systematically reduced. The approach is then extended to structured visual domains by introducing a synthetic light-field (LFs) generation pipeline for learning-based intra-prediction and compression. In this context, synthetic data serves as a controllable source of supervision, enabling fine-grained modeling of spatial and angular consistency. Finally, synthetic data serve as a fully controllable environment for studying learning dynamics by introducing neural velocity, a novel metric that characterizes the evolution of learned representations across training epochs. Computed on an auxiliary synthetic dataset, generated independently of both the training and validation data, neural velocity enables convergence detection, early stopping, and adaptive hyperparameter tuning without requiring annotated validation data. The framework is further extended to federated learning, where neural velocity helps address data heterogeneity and non-independent and identically distributed (non-IID) data distributions, supporting more informed and efficient distributed training. The second part of this thesis builds upon the theoretical framework described above and presents a synthetic data generation pipeline for object detection in industrial environments where annotated data is scarce or costly to obtain. Beyond its design and implementation, the pipeline’s key components are analyzed to identify the factors that most significantly affect performance and deployability. Based on these findings, practical guidelines are provided for designing effective synthetic data pipelines in data-scarce industrial settings
Beyond Real Data: Synthetic Data for Scalable and Efficient Deep Learning(2026 Jul 14).
Beyond Real Data: Synthetic Data for Scalable and Efficient Deep Learning
DALMASSO, GIANLUCA
2026-07-14
Abstract
Deep learning has become a cornerstone of computer vision applications such as image classification, object detection, and image segmentation across domains ranging from medical imaging to autonomous systems. However, deep learning demands large-scale annotated datasets for model training, which are often expensive, time-consuming, or impractical to acquire and annotate. This thesis investigates synthetic data as both a practical tool and a methodological framework for improving deep learning in data-scarce scenarios. The work was carried out within an industrial PhD in collaboration with Synesthesia srl SB and reflects a dual academic and industrial research agenda: both tracks share a common focus on understanding how synthetic data can support and enhance the training of deep learning models when real annotated data is limited or costly. The first part of this thesis explores synthetic data along three complementary directions from an academic perspective. Low-data regimes are first investigated through a case study of Byzantine seals, in which synthetic data are used to support object detection and classification. To bridge the gap between synthetic and real data, style transfer techniques are applied and analyzed, providing insight into how domain adaptation strategies affect generalization and how the synthetic-to-real domain gap can be systematically reduced. The approach is then extended to structured visual domains by introducing a synthetic light-field (LFs) generation pipeline for learning-based intra-prediction and compression. In this context, synthetic data serves as a controllable source of supervision, enabling fine-grained modeling of spatial and angular consistency. Finally, synthetic data serve as a fully controllable environment for studying learning dynamics by introducing neural velocity, a novel metric that characterizes the evolution of learned representations across training epochs. Computed on an auxiliary synthetic dataset, generated independently of both the training and validation data, neural velocity enables convergence detection, early stopping, and adaptive hyperparameter tuning without requiring annotated validation data. The framework is further extended to federated learning, where neural velocity helps address data heterogeneity and non-independent and identically distributed (non-IID) data distributions, supporting more informed and efficient distributed training. The second part of this thesis builds upon the theoretical framework described above and presents a synthetic data generation pipeline for object detection in industrial environments where annotated data is scarce or costly to obtain. Beyond its design and implementation, the pipeline’s key components are analyzed to identify the factors that most significantly affect performance and deployability. Based on these findings, practical guidelines are provided for designing effective synthetic data pipelines in data-scarce industrial settings| File | Dimensione | Formato | |
|---|---|---|---|
|
Tesi-Dalmasso-Gianluca.pdf
Accesso aperto
Descrizione: Tesi
Dimensione
45.34 MB
Formato
Adobe PDF
|
45.34 MB | Adobe PDF | Visualizza/Apri |
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.



