Navigating Perplexity: A linear relationship with the data set size in t-SNE embeddings

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Skrodzki, Martin, Chaves-de-Plaza, Nicolas F., Höllt, Thomas, Eisemann, Elmar, Hildebrandt, Klaus
Format: Preprint
Veröffentlicht: 2023
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866929615909421056
author Skrodzki, Martin
Chaves-de-Plaza, Nicolas F.
Höllt, Thomas
Eisemann, Elmar
Hildebrandt, Klaus
author_facet Skrodzki, Martin
Chaves-de-Plaza, Nicolas F.
Höllt, Thomas
Eisemann, Elmar
Hildebrandt, Klaus
contents Widely used pipelines for analyzing high-dimensional data utilize two-dimensional visualizations. These are created, for instance, via t-distributed stochastic neighbor embedding (t-SNE). A crucial element of the t-SNE embedding procedure is the perplexity hyperparameter. That is because the embedding structure varies when perplexity is changed. A suitable perplexity choice depends on the data set and the intended usage for the embedding. Therefore, perplexity is often chosen based on heuristics, intuition, and prior experience. This paper uncovers a linear relationship between perplexity and the data set size. Namely, we show that embeddings remain structurally consistent across data set samples when perplexity is adjusted accordingly. Qualitative and quantitative experimental results support these findings. This informs the visualization process, guiding the user when choosing a perplexity value. Finally, we outline several applications for the visualization of high-dimensional data via t-SNE based on this linear relationship.
format Preprint
id arxiv_https___arxiv_org_abs_2308_15513
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Navigating Perplexity: A linear relationship with the data set size in t-SNE embeddings
Skrodzki, Martin
Chaves-de-Plaza, Nicolas F.
Höllt, Thomas
Eisemann, Elmar
Hildebrandt, Klaus
Machine Learning
Artificial Intelligence
Quantitative Methods
Widely used pipelines for analyzing high-dimensional data utilize two-dimensional visualizations. These are created, for instance, via t-distributed stochastic neighbor embedding (t-SNE). A crucial element of the t-SNE embedding procedure is the perplexity hyperparameter. That is because the embedding structure varies when perplexity is changed. A suitable perplexity choice depends on the data set and the intended usage for the embedding. Therefore, perplexity is often chosen based on heuristics, intuition, and prior experience. This paper uncovers a linear relationship between perplexity and the data set size. Namely, we show that embeddings remain structurally consistent across data set samples when perplexity is adjusted accordingly. Qualitative and quantitative experimental results support these findings. This informs the visualization process, guiding the user when choosing a perplexity value. Finally, we outline several applications for the visualization of high-dimensional data via t-SNE based on this linear relationship.
title Navigating Perplexity: A linear relationship with the data set size in t-SNE embeddings
topic Machine Learning
Artificial Intelligence
Quantitative Methods
url https://arxiv.org/abs/2308.15513