A Survey: Potential Dimensionality Reduction Methods

Fuente: arXiv
Saved in:
Bibliographic Details
Main Author: Chang, Yuan-chin Ivan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915156912504832
author Chang, Yuan-chin Ivan
author_facet Chang, Yuan-chin Ivan
contents Dimensionality reduction is a fundamental technique in machine learning and data analysis, enabling efficient representation and visualization of high-dimensional data. This paper explores five key methods: Principal Component Analysis (PCA), Kernel PCA (KPCA), Sparse Kernel PCA, t-Distributed Stochastic Neighbor Embedding (t-SNE), and Uniform Manifold Approximation and Projection (UMAP). PCA provides a linear approach to capturing variance, whereas KPCA and Sparse KPCA extend this concept to non-linear structures using kernel functions. Meanwhile, t-SNE and UMAP focus on preserving local relationships, making them effective for data visualization. Each method is examined in terms of its mathematical formulation, computational complexity, strengths, and limitations. The trade-offs between global structure preservation, computational efficiency, and interpretability are discussed to guide practitioners in selecting the appropriate technique based on their application needs.
format Preprint
id arxiv_https___arxiv_org_abs_2502_11036
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Survey: Potential Dimensionality Reduction Methods
Chang, Yuan-chin Ivan
Other Statistics
Dimensionality reduction is a fundamental technique in machine learning and data analysis, enabling efficient representation and visualization of high-dimensional data. This paper explores five key methods: Principal Component Analysis (PCA), Kernel PCA (KPCA), Sparse Kernel PCA, t-Distributed Stochastic Neighbor Embedding (t-SNE), and Uniform Manifold Approximation and Projection (UMAP). PCA provides a linear approach to capturing variance, whereas KPCA and Sparse KPCA extend this concept to non-linear structures using kernel functions. Meanwhile, t-SNE and UMAP focus on preserving local relationships, making them effective for data visualization. Each method is examined in terms of its mathematical formulation, computational complexity, strengths, and limitations. The trade-offs between global structure preservation, computational efficiency, and interpretability are discussed to guide practitioners in selecting the appropriate technique based on their application needs.
title A Survey: Potential Dimensionality Reduction Methods
topic Other Statistics
url https://arxiv.org/abs/2502.11036