Topological Quality of Subsets via Persistence Matching Diagrams
Fuente:
arXiv
Salvato in:
| Autori principali: | , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866909328030564352 |
|---|---|
| author | Torras-Casas, Álvaro Paluzo-Hidalgo, Eduardo Gonzalez-Diaz, Rocio |
| author_facet | Torras-Casas, Álvaro Paluzo-Hidalgo, Eduardo Gonzalez-Diaz, Rocio |
| contents | Data quality is crucial for the successful training, generalization and performance of machine learning models. We propose to measure the quality of a subset concerning the dataset it represents, using topological data analysis techniques. Specifically, we define the persistence matching diagram, a topological invariant derived from combining embeddings with persistent homology. We provide an algorithm to compute it using minimum spanning trees. Also, the invariant allows us to understand whether the subset ``represents well" the clusters from the larger dataset or not, and we also use it to estimate bounds for the Hausdorff distance between the subset and the complete dataset. In particular, this approach enables us to explain why the chosen subset is likely to result in poor performance of a supervised learning model. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2306_02411 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | Topological Quality of Subsets via Persistence Matching Diagrams Torras-Casas, Álvaro Paluzo-Hidalgo, Eduardo Gonzalez-Diaz, Rocio Algebraic Topology Artificial Intelligence Machine Learning Data quality is crucial for the successful training, generalization and performance of machine learning models. We propose to measure the quality of a subset concerning the dataset it represents, using topological data analysis techniques. Specifically, we define the persistence matching diagram, a topological invariant derived from combining embeddings with persistent homology. We provide an algorithm to compute it using minimum spanning trees. Also, the invariant allows us to understand whether the subset ``represents well" the clusters from the larger dataset or not, and we also use it to estimate bounds for the Hausdorff distance between the subset and the complete dataset. In particular, this approach enables us to explain why the chosen subset is likely to result in poor performance of a supervised learning model. |
| title | Topological Quality of Subsets via Persistence Matching Diagrams |
| topic | Algebraic Topology Artificial Intelligence Machine Learning |
| url | https://arxiv.org/abs/2306.02411 |