Efficient Canonical Correlation Analysis with Sparsity

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Zixuan, Tuzhilina, Elena, Donnat, Claire
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918092821495808
author Wu, Zixuan
Tuzhilina, Elena
Donnat, Claire
author_facet Wu, Zixuan
Tuzhilina, Elena
Donnat, Claire
contents In high-dimensional settings, Canonical Correlation Analysis (CCA) often fails, and existing sparse methods force an untenable choice between computational speed and statistical rigor. This work introduces a fast and provably consistent sparse CCA algorithm (ECCAR) that resolves this trade-off. We formulate CCA as a high-dimensional reduced-rank regression problem, which allows us to derive consistent estimators with high-probability error bounds without relying on computationally expensive techniques like Fantope projections. The resulting algorithm is scalable, projection-free, and significantly faster than its competitors. We validate our method through extensive simulations and demonstrate its power to uncover reliable and interpretable associations in two complex biological datasets, as well as in an ML interpretability task. Our work makes sparse CCA a practical and trustworthy tool for large-scale multimodal data analysis. A companion R package has been made available.
format Preprint
id arxiv_https___arxiv_org_abs_2507_11160
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Efficient Canonical Correlation Analysis with Sparsity
Wu, Zixuan
Tuzhilina, Elena
Donnat, Claire
Methodology
In high-dimensional settings, Canonical Correlation Analysis (CCA) often fails, and existing sparse methods force an untenable choice between computational speed and statistical rigor. This work introduces a fast and provably consistent sparse CCA algorithm (ECCAR) that resolves this trade-off. We formulate CCA as a high-dimensional reduced-rank regression problem, which allows us to derive consistent estimators with high-probability error bounds without relying on computationally expensive techniques like Fantope projections. The resulting algorithm is scalable, projection-free, and significantly faster than its competitors. We validate our method through extensive simulations and demonstrate its power to uncover reliable and interpretable associations in two complex biological datasets, as well as in an ML interpretability task. Our work makes sparse CCA a practical and trustworthy tool for large-scale multimodal data analysis. A companion R package has been made available.
title Efficient Canonical Correlation Analysis with Sparsity
topic Methodology
url https://arxiv.org/abs/2507.11160