Distributional Reduction: Unifying Dimensionality Reduction and Clustering with Gromov-Wasserstein

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Van Assel, Hugues, Vincent-Cuaz, Cédric, Courty, Nicolas, Flamary, Rémi, Frossard, Pascal, Vayer, Titouan
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913914439073792
author Van Assel, Hugues
Vincent-Cuaz, Cédric
Courty, Nicolas
Flamary, Rémi
Frossard, Pascal
Vayer, Titouan
author_facet Van Assel, Hugues
Vincent-Cuaz, Cédric
Courty, Nicolas
Flamary, Rémi
Frossard, Pascal
Vayer, Titouan
contents Unsupervised learning aims to capture the underlying structure of potentially large and high-dimensional datasets. Traditionally, this involves using dimensionality reduction (DR) methods to project data onto lower-dimensional spaces or organizing points into meaningful clusters (clustering). In this work, we revisit these approaches under the lens of optimal transport and exhibit relationships with the Gromov-Wasserstein problem. This unveils a new general framework, called distributional reduction, that recovers DR and clustering as special cases and allows addressing them jointly within a single optimization problem. We empirically demonstrate its relevance to the identification of low-dimensional prototypes representing data at different scales, across multiple image and genomic datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2402_02239
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Distributional Reduction: Unifying Dimensionality Reduction and Clustering with Gromov-Wasserstein
Van Assel, Hugues
Vincent-Cuaz, Cédric
Courty, Nicolas
Flamary, Rémi
Frossard, Pascal
Vayer, Titouan
Machine Learning
Unsupervised learning aims to capture the underlying structure of potentially large and high-dimensional datasets. Traditionally, this involves using dimensionality reduction (DR) methods to project data onto lower-dimensional spaces or organizing points into meaningful clusters (clustering). In this work, we revisit these approaches under the lens of optimal transport and exhibit relationships with the Gromov-Wasserstein problem. This unveils a new general framework, called distributional reduction, that recovers DR and clustering as special cases and allows addressing them jointly within a single optimization problem. We empirically demonstrate its relevance to the identification of low-dimensional prototypes representing data at different scales, across multiple image and genomic datasets.
title Distributional Reduction: Unifying Dimensionality Reduction and Clustering with Gromov-Wasserstein
topic Machine Learning
url https://arxiv.org/abs/2402.02239