Optimal Test-Data Piling in HDLSS Classification with Covariance Heterogeneity

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kim, Taehyun, Ahn, Jeongyoun, Jung, Sungkyu
Format: Preprint
Published: 2022
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912528296050688
author Kim, Taehyun
Ahn, Jeongyoun
Jung, Sungkyu
author_facet Kim, Taehyun
Ahn, Jeongyoun
Jung, Sungkyu
contents This work addresses a longstanding question in high-dimensional linear classification: Is perfect classification achievable in heterogeneous covariance structures? We focus on the phenomenon of data piling, where projected data points collapse onto discrete values. We provide a comprehensive characterization of two distinct types of data piling. The first type of data piling refers to the phenomenon where projecting the training data onto a certain direction yields exactly two distinct values-one for each class. This occurs universally when the data dimension $p$ exceeds the sample size $n$. The second type concerns independent test data and arises asymptotically as $p \to \infty$ with fixed $n$. While previous work established the existence of such double data piling under homogeneously spiked covariance structures using negatively ridged classifiers, our analysis extends to the more general and realistic case of heterogeneous covariance. We identify an optimal direction among all piling directions that maximizes the separation between test data piles, which is called the Second Maximal Data Piling direction. An algorithm based on data splitting is proposed to compute this direction using only training data. Our analysis reveals a key insight: the main obstacle to discovering this direction is the imbalance of the tail eigenvalues, rather than differences in spike count, spike magnitude, or the alignment of leading eigenspaces. Extensive simulations confirm our theoretical results and demonstrate the effectiveness of the proposed classifier across a wide range of high-dimensional scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2211_15562
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle Optimal Test-Data Piling in HDLSS Classification with Covariance Heterogeneity
Kim, Taehyun
Ahn, Jeongyoun
Jung, Sungkyu
Statistics Theory
This work addresses a longstanding question in high-dimensional linear classification: Is perfect classification achievable in heterogeneous covariance structures? We focus on the phenomenon of data piling, where projected data points collapse onto discrete values. We provide a comprehensive characterization of two distinct types of data piling. The first type of data piling refers to the phenomenon where projecting the training data onto a certain direction yields exactly two distinct values-one for each class. This occurs universally when the data dimension $p$ exceeds the sample size $n$. The second type concerns independent test data and arises asymptotically as $p \to \infty$ with fixed $n$. While previous work established the existence of such double data piling under homogeneously spiked covariance structures using negatively ridged classifiers, our analysis extends to the more general and realistic case of heterogeneous covariance. We identify an optimal direction among all piling directions that maximizes the separation between test data piles, which is called the Second Maximal Data Piling direction. An algorithm based on data splitting is proposed to compute this direction using only training data. Our analysis reveals a key insight: the main obstacle to discovering this direction is the imbalance of the tail eigenvalues, rather than differences in spike count, spike magnitude, or the alignment of leading eigenspaces. Extensive simulations confirm our theoretical results and demonstrate the effectiveness of the proposed classifier across a wide range of high-dimensional scenarios.
title Optimal Test-Data Piling in HDLSS Classification with Covariance Heterogeneity
topic Statistics Theory
url https://arxiv.org/abs/2211.15562