On the Structural Dimension of Sliced Inverse Regression
Fuente:
arXiv
Guardado en:
| Autores principales: | , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866914866634162176 |
|---|---|
| author | Huang, Dongming Tian, Songtao Lin, Qian |
| author_facet | Huang, Dongming Tian, Songtao Lin, Qian |
| contents | In this work, we address the longstanding puzzle that Sliced Inverse Regression (SIR) often performs poorly for sufficient dimension reduction when the structural dimension $d$ (the dimension of the central space) exceeds 4. We first show that in the multiple index model $Y=f( \mathbf{P} \boldsymbol{X})+ε$ where $\boldsymbol{X}$ is a $p$-standard normal vector, $ε$ is an independent noise, and $\mathbf{P}$ is a projection operator from $\mathbb R^{p}$ to $\mathbb R^{d}$, if the link function $f$ follows the law of a Gaussian process, then with high probability, the $d$-th eigenvalue $λ_{d}$ of $\mathrm{Cov}\left[\mathbb{E}(\boldsymbol{X}\mid Y)\right]$ satisfies $λ_{d}\leq C e^{-θd}$ for some positive constants $C$ and $θ$. We then focus on the low signal regime where $λ_{d}$ can be arbitrarily small and not larger than $d^{-8.1}$, and prove that the minimax risk of estimating the central space is lower bounded by $\frac{dp}{nλ_{d}}$. Combining these two results, we provide a convincing explanation for the poor performance of SIR when $d$ is large, a phenomenon that has perplexed researchers for nearly three decades. The technical tools developed here may be of independent interest for studying other sufficient dimension reduction methods. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2305_04340 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | On the Structural Dimension of Sliced Inverse Regression Huang, Dongming Tian, Songtao Lin, Qian Statistics Theory 62J02 (Primary) 62H12, 62C20 (Secondary) In this work, we address the longstanding puzzle that Sliced Inverse Regression (SIR) often performs poorly for sufficient dimension reduction when the structural dimension $d$ (the dimension of the central space) exceeds 4. We first show that in the multiple index model $Y=f( \mathbf{P} \boldsymbol{X})+ε$ where $\boldsymbol{X}$ is a $p$-standard normal vector, $ε$ is an independent noise, and $\mathbf{P}$ is a projection operator from $\mathbb R^{p}$ to $\mathbb R^{d}$, if the link function $f$ follows the law of a Gaussian process, then with high probability, the $d$-th eigenvalue $λ_{d}$ of $\mathrm{Cov}\left[\mathbb{E}(\boldsymbol{X}\mid Y)\right]$ satisfies $λ_{d}\leq C e^{-θd}$ for some positive constants $C$ and $θ$. We then focus on the low signal regime where $λ_{d}$ can be arbitrarily small and not larger than $d^{-8.1}$, and prove that the minimax risk of estimating the central space is lower bounded by $\frac{dp}{nλ_{d}}$. Combining these two results, we provide a convincing explanation for the poor performance of SIR when $d$ is large, a phenomenon that has perplexed researchers for nearly three decades. The technical tools developed here may be of independent interest for studying other sufficient dimension reduction methods. |
| title | On the Structural Dimension of Sliced Inverse Regression |
| topic | Statistics Theory 62J02 (Primary) 62H12, 62C20 (Secondary) |
| url | https://arxiv.org/abs/2305.04340 |