Co-clustering is a powerful data mining technique designed to extract summary information from data matrices by simultaneously grouping rows and columns. By exploiting latent relationships between objects and their attributes, co-clustering provides a compact representation of the data, identifying coherent groups of similar instances and their interactions with clusters of correlated features. This approach has demonstrated effectiveness across a wide range of real-world applications, including gene expression analysis, text mining, recommender systems, market basket analysis, and community detection. In this dissertation, we address the problem of co-clustering in sparse, high-dimensional contingency tables, with a particular focus on algorithmic fairness. Although fairness has been extensively studied in supervised learning and clustering, it has received only limited attention in co-clustering. Ensuring fairness is crucial when co-clustering methods are employed in decision-making processes involving sensitive attributes such as gender or ethnicity, since resulting co-clusters may have under- or over-represented social groups, potentially leading to discrimination against minorities. To mitigate such risks, we introduce a formal definition of fair co-clustering, where a co-clustering algorithm ensures group fairness if each co-cluster preserves the same proportion of protected groups as in the overall dataset — capturing the concept of balance in the context of fair clustering. Building on this definition, we propose a novel fair co-clustering algorithm, called Fair-$\tau$CC, based on an associative measure derived from Goodman–Kruskal's $\tau$, which exhibits favorable convergence properties. The algorithm is parameterless (it does not require specifying the final number of row and column clusters) and achieves a good trade-off between fairness and clustering quality. Moreover, Fair-$\tau$CC supports fairness constraints involving two or more protected groups and can enforce fairness not only across row clusters but also across column clusters, when protected group information for the columns is also available.

Fair Associative Co-Clustering(2026 Jul 23).

Fair Associative Co-Clustering

PEIRETTI, Federico
2026-07-23

Abstract

Co-clustering is a powerful data mining technique designed to extract summary information from data matrices by simultaneously grouping rows and columns. By exploiting latent relationships between objects and their attributes, co-clustering provides a compact representation of the data, identifying coherent groups of similar instances and their interactions with clusters of correlated features. This approach has demonstrated effectiveness across a wide range of real-world applications, including gene expression analysis, text mining, recommender systems, market basket analysis, and community detection. In this dissertation, we address the problem of co-clustering in sparse, high-dimensional contingency tables, with a particular focus on algorithmic fairness. Although fairness has been extensively studied in supervised learning and clustering, it has received only limited attention in co-clustering. Ensuring fairness is crucial when co-clustering methods are employed in decision-making processes involving sensitive attributes such as gender or ethnicity, since resulting co-clusters may have under- or over-represented social groups, potentially leading to discrimination against minorities. To mitigate such risks, we introduce a formal definition of fair co-clustering, where a co-clustering algorithm ensures group fairness if each co-cluster preserves the same proportion of protected groups as in the overall dataset — capturing the concept of balance in the context of fair clustering. Building on this definition, we propose a novel fair co-clustering algorithm, called Fair-$\tau$CC, based on an associative measure derived from Goodman–Kruskal's $\tau$, which exhibits favorable convergence properties. The algorithm is parameterless (it does not require specifying the final number of row and column clusters) and achieves a good trade-off between fairness and clustering quality. Moreover, Fair-$\tau$CC supports fairness constraints involving two or more protected groups and can enforce fairness not only across row clusters but also across column clusters, when protected group information for the columns is also available.
23-lug-2026
38
INFORMATICA
CAPECCHI, Sara
PENSA, Ruggero Gaetano
BIOGLIO, Livio
File in questo prodotto:
File Dimensione Formato  
Tesi-Peiretti-Federico.pdf

Accesso aperto

Descrizione: Tesi
Dimensione 18.69 MB
Formato Adobe PDF
18.69 MB Adobe PDF Visualizza/Apri

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/2318/2153570
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
social impact