Tree-Wasserstein Distance for High Dimensional Data with a Latent Feature Hierarchy

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Lin, Ya-Wei Eileen, Coifman, Ronald R., Mishne, Gal, Talmon, Ronen
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866910841875464192
author Lin, Ya-Wei Eileen
Coifman, Ronald R.
Mishne, Gal
Talmon, Ronen
author_facet Lin, Ya-Wei Eileen
Coifman, Ronald R.
Mishne, Gal
Talmon, Ronen
contents Finding meaningful distances between high-dimensional data samples is an important scientific task. To this end, we propose a new tree-Wasserstein distance (TWD) for high-dimensional data with two key aspects. First, our TWD is specifically designed for data with a latent feature hierarchy, i.e., the features lie in a hierarchical space, in contrast to the usual focus on embedding samples in hyperbolic space. Second, while the conventional use of TWD is to speed up the computation of the Wasserstein distance, we use its inherent tree as a means to learn the latent feature hierarchy. The key idea of our method is to embed the features into a multi-scale hyperbolic space using diffusion geometry and then present a new tree decoding method by establishing analogies between the hyperbolic embedding and trees. We show that our TWD computed based on data observations provably recovers the TWD defined with the latent feature hierarchy and that its computation is efficient and scalable. We showcase the usefulness of the proposed TWD in applications to word-document and single-cell RNA-sequencing datasets, demonstrating its advantages over existing TWDs and methods based on pre-trained models.
format Preprint
id arxiv_https___arxiv_org_abs_2410_21107
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Tree-Wasserstein Distance for High Dimensional Data with a Latent Feature Hierarchy
Lin, Ya-Wei Eileen
Coifman, Ronald R.
Mishne, Gal
Talmon, Ronen
Machine Learning
Finding meaningful distances between high-dimensional data samples is an important scientific task. To this end, we propose a new tree-Wasserstein distance (TWD) for high-dimensional data with two key aspects. First, our TWD is specifically designed for data with a latent feature hierarchy, i.e., the features lie in a hierarchical space, in contrast to the usual focus on embedding samples in hyperbolic space. Second, while the conventional use of TWD is to speed up the computation of the Wasserstein distance, we use its inherent tree as a means to learn the latent feature hierarchy. The key idea of our method is to embed the features into a multi-scale hyperbolic space using diffusion geometry and then present a new tree decoding method by establishing analogies between the hyperbolic embedding and trees. We show that our TWD computed based on data observations provably recovers the TWD defined with the latent feature hierarchy and that its computation is efficient and scalable. We showcase the usefulness of the proposed TWD in applications to word-document and single-cell RNA-sequencing datasets, demonstrating its advantages over existing TWDs and methods based on pre-trained models.
title Tree-Wasserstein Distance for High Dimensional Data with a Latent Feature Hierarchy
topic Machine Learning
url https://arxiv.org/abs/2410.21107