Principal nested spheres for high-dimensional data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Monem, Mymuna, Dryden, Ian L., George, Florence
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909898424451072
author Monem, Mymuna
Dryden, Ian L.
George, Florence
author_facet Monem, Mymuna
Dryden, Ian L.
George, Florence
contents The method of Principal Nested Spheres (PNS) is a non-linear dimension reduction technique for spherical data. The method is a backwards fitting procedure, starting with fitting a high-dimensional sphere and then successively reducing dimension at each stage. After reviewing the PNS method in detail, we introduce some new methods for model selection at each stage between great and small subspheres, based on the Kolmogorov-Smirnov test, a variance test and a likelihood ratio test. The current PNS fitting method is slow for high-dimensional spherical data, and so we introduce a fast PNS method which involves an initial principal components analysis decomposition to select a basis for lower dimensional PNS. A new visual method called the PNS biplot is introduced for examining the effects of the original variables on the PNS, and this involves procedures for back-fitting from the PNS scores back to the original variables. The methodology is illustrated with two high-dimensional datasets from cancer research: Melanoma proteomics data with 500 variables and 205 patients, and a Pan Cancer dataset with 12,478 genes and 300 patients. In both applications the PNS biplot is used to select variables for effective classification.
format Preprint
id arxiv_https___arxiv_org_abs_2511_08398
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Principal nested spheres for high-dimensional data
Monem, Mymuna
Dryden, Ian L.
George, Florence
Methodology
62R30
The method of Principal Nested Spheres (PNS) is a non-linear dimension reduction technique for spherical data. The method is a backwards fitting procedure, starting with fitting a high-dimensional sphere and then successively reducing dimension at each stage. After reviewing the PNS method in detail, we introduce some new methods for model selection at each stage between great and small subspheres, based on the Kolmogorov-Smirnov test, a variance test and a likelihood ratio test. The current PNS fitting method is slow for high-dimensional spherical data, and so we introduce a fast PNS method which involves an initial principal components analysis decomposition to select a basis for lower dimensional PNS. A new visual method called the PNS biplot is introduced for examining the effects of the original variables on the PNS, and this involves procedures for back-fitting from the PNS scores back to the original variables. The methodology is illustrated with two high-dimensional datasets from cancer research: Melanoma proteomics data with 500 variables and 205 patients, and a Pan Cancer dataset with 12,478 genes and 300 patients. In both applications the PNS biplot is used to select variables for effective classification.
title Principal nested spheres for high-dimensional data
topic Methodology
62R30
url https://arxiv.org/abs/2511.08398