Assessing and improving reliability of neighbor embedding methods: a map-continuity perspective

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Liu, Zhexuan, Ma, Rong, Zhong, Yiqiao
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908292276551680
author Liu, Zhexuan
Ma, Rong
Zhong, Yiqiao
author_facet Liu, Zhexuan
Ma, Rong
Zhong, Yiqiao
contents Visualizing high-dimensional data is essential for understanding biomedical data and deep learning models. Neighbor embedding methods, such as t-SNE and UMAP, are widely used but can introduce misleading visual artifacts. We find that the manifold learning interpretations from many prior works are inaccurate and that the misuse stems from a lack of data-independent notions of embedding maps, which project high-dimensional data into a lower-dimensional space. Leveraging the leave-one-out principle, we introduce LOO-map, a framework that extends embedding maps beyond discrete points to the entire input space. We identify two forms of map discontinuity that distort visualizations: one exaggerates cluster separation and the other creates spurious local structures. As a remedy, we develop two types of point-wise diagnostic scores to detect unreliable embedding points and improve hyperparameter selection, which are validated on datasets from computer vision and single-cell omics.
format Preprint
id arxiv_https___arxiv_org_abs_2410_16608
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Assessing and improving reliability of neighbor embedding methods: a map-continuity perspective
Liu, Zhexuan
Ma, Rong
Zhong, Yiqiao
Methodology
Machine Learning
Computation
62-08
Visualizing high-dimensional data is essential for understanding biomedical data and deep learning models. Neighbor embedding methods, such as t-SNE and UMAP, are widely used but can introduce misleading visual artifacts. We find that the manifold learning interpretations from many prior works are inaccurate and that the misuse stems from a lack of data-independent notions of embedding maps, which project high-dimensional data into a lower-dimensional space. Leveraging the leave-one-out principle, we introduce LOO-map, a framework that extends embedding maps beyond discrete points to the entire input space. We identify two forms of map discontinuity that distort visualizations: one exaggerates cluster separation and the other creates spurious local structures. As a remedy, we develop two types of point-wise diagnostic scores to detect unreliable embedding points and improve hyperparameter selection, which are validated on datasets from computer vision and single-cell omics.
title Assessing and improving reliability of neighbor embedding methods: a map-continuity perspective
topic Methodology
Machine Learning
Computation
62-08
url https://arxiv.org/abs/2410.16608