Topological Quality of Subsets via Persistence Matching Diagrams

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Torras-Casas, Álvaro, Paluzo-Hidalgo, Eduardo, Gonzalez-Diaz, Rocio
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909328030564352
author Torras-Casas, Álvaro
Paluzo-Hidalgo, Eduardo
Gonzalez-Diaz, Rocio
author_facet Torras-Casas, Álvaro
Paluzo-Hidalgo, Eduardo
Gonzalez-Diaz, Rocio
contents Data quality is crucial for the successful training, generalization and performance of machine learning models. We propose to measure the quality of a subset concerning the dataset it represents, using topological data analysis techniques. Specifically, we define the persistence matching diagram, a topological invariant derived from combining embeddings with persistent homology. We provide an algorithm to compute it using minimum spanning trees. Also, the invariant allows us to understand whether the subset ``represents well" the clusters from the larger dataset or not, and we also use it to estimate bounds for the Hausdorff distance between the subset and the complete dataset. In particular, this approach enables us to explain why the chosen subset is likely to result in poor performance of a supervised learning model.
format Preprint
id arxiv_https___arxiv_org_abs_2306_02411
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Topological Quality of Subsets via Persistence Matching Diagrams
Torras-Casas, Álvaro
Paluzo-Hidalgo, Eduardo
Gonzalez-Diaz, Rocio
Algebraic Topology
Artificial Intelligence
Machine Learning
Data quality is crucial for the successful training, generalization and performance of machine learning models. We propose to measure the quality of a subset concerning the dataset it represents, using topological data analysis techniques. Specifically, we define the persistence matching diagram, a topological invariant derived from combining embeddings with persistent homology. We provide an algorithm to compute it using minimum spanning trees. Also, the invariant allows us to understand whether the subset ``represents well" the clusters from the larger dataset or not, and we also use it to estimate bounds for the Hausdorff distance between the subset and the complete dataset. In particular, this approach enables us to explain why the chosen subset is likely to result in poor performance of a supervised learning model.
title Topological Quality of Subsets via Persistence Matching Diagrams
topic Algebraic Topology
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2306.02411