Two approaches to multiple canonical correlation analysis for repeated measures data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Górecki, Tomasz, Krzyśko, Mirosław, Gnettner, Felix, Kokoszka, Piotr
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912666251952128
author Górecki, Tomasz
Krzyśko, Mirosław
Gnettner, Felix
Kokoszka, Piotr
author_facet Górecki, Tomasz
Krzyśko, Mirosław
Gnettner, Felix
Kokoszka, Piotr
contents In classical canonical correlation analysis (CCA), the goal is to determine the linear transformations of two random vectors into two new random variables that are most strongly correlated. Canonical variables are pairs of these new random variables, while canonical correlations are correlations between these pairs. In this paper, we propose and study two generalizations of this classical method: (1) Instead of two random vectors we study more complex data structures that appear in important applications. In these structures, there are $L$ features, each described by $p_l$ scalars, $1 \le l \le L$. We observe $n$ such objects over $T$ time points. We derive a suitable analog of the CCA for such data. Our approach relies on embeddings into Reproducing Kernel Hilbert Spaces, and covers several related data structures as well. (2) We develop an analogous approach for multidimensional random processes. In this case, the experimental units are multivariate continuous, square-integrable functions over a given interval. These functions are modeled as elements of a Hilbert space, so in this case, we define the multiple functional canonical correlation analysis, MFCCA. We justify our approaches by their application to two data sets and suitable large sample theory. We derive consistency rates for the related transformation and correlation estimators, and show that it is possible to relax two common assumptions on the compactness of the underlying cross-covariance operators and the independence of the data.
format Preprint
id arxiv_https___arxiv_org_abs_2510_04457
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Two approaches to multiple canonical correlation analysis for repeated measures data
Górecki, Tomasz
Krzyśko, Mirosław
Gnettner, Felix
Kokoszka, Piotr
Methodology
Statistics Theory
Applications
Machine Learning
62H20 (Primary) 62G05, 62G20, 62R07, 68T05 (Secondary)
In classical canonical correlation analysis (CCA), the goal is to determine the linear transformations of two random vectors into two new random variables that are most strongly correlated. Canonical variables are pairs of these new random variables, while canonical correlations are correlations between these pairs. In this paper, we propose and study two generalizations of this classical method: (1) Instead of two random vectors we study more complex data structures that appear in important applications. In these structures, there are $L$ features, each described by $p_l$ scalars, $1 \le l \le L$. We observe $n$ such objects over $T$ time points. We derive a suitable analog of the CCA for such data. Our approach relies on embeddings into Reproducing Kernel Hilbert Spaces, and covers several related data structures as well. (2) We develop an analogous approach for multidimensional random processes. In this case, the experimental units are multivariate continuous, square-integrable functions over a given interval. These functions are modeled as elements of a Hilbert space, so in this case, we define the multiple functional canonical correlation analysis, MFCCA. We justify our approaches by their application to two data sets and suitable large sample theory. We derive consistency rates for the related transformation and correlation estimators, and show that it is possible to relax two common assumptions on the compactness of the underlying cross-covariance operators and the independence of the data.
title Two approaches to multiple canonical correlation analysis for repeated measures data
topic Methodology
Statistics Theory
Applications
Machine Learning
62H20 (Primary) 62G05, 62G20, 62R07, 68T05 (Secondary)
url https://arxiv.org/abs/2510.04457