Penalized Quasi-likelihood for High-dimensional Longitudinal Data via Within-cluster Resampling

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Ma, Yue, Wang, Haofeng, Jiang, Xuejun
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913632895369216
author Ma, Yue
Wang, Haofeng
Jiang, Xuejun
author_facet Ma, Yue
Wang, Haofeng
Jiang, Xuejun
contents The generalized estimating equation (GEE) method is a popular tool for longitudinal data analysis. However, GEE produces biased estimates when the outcome of interest is associated with cluster size, a phenomenon known as informative cluster size (ICS). In this study, we address this issue by formulating the impact of ICS and proposing an integrated approach to mitigate its effects. Our method combines the concept of within-cluster resampling with a penalized quasi-likelihood framework applied to each resampled dataset, ensuring consistency in model selection and estimation. To aggregate the estimators from the resampled datasets, we introduce a penalized mean regression technique, resulting in a final estimator that improves true positive discovery rates while reducing false positives. Simulation studies and an application to yeast cell-cycle gene expression data demonstrate the excellent performance of the proposed penalized quasi-likelihood method via within-cluster resampling.
format Preprint
id arxiv_https___arxiv_org_abs_2501_01021
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Penalized Quasi-likelihood for High-dimensional Longitudinal Data via Within-cluster Resampling
Ma, Yue
Wang, Haofeng
Jiang, Xuejun
Methodology
Computation
The generalized estimating equation (GEE) method is a popular tool for longitudinal data analysis. However, GEE produces biased estimates when the outcome of interest is associated with cluster size, a phenomenon known as informative cluster size (ICS). In this study, we address this issue by formulating the impact of ICS and proposing an integrated approach to mitigate its effects. Our method combines the concept of within-cluster resampling with a penalized quasi-likelihood framework applied to each resampled dataset, ensuring consistency in model selection and estimation. To aggregate the estimators from the resampled datasets, we introduce a penalized mean regression technique, resulting in a final estimator that improves true positive discovery rates while reducing false positives. Simulation studies and an application to yeast cell-cycle gene expression data demonstrate the excellent performance of the proposed penalized quasi-likelihood method via within-cluster resampling.
title Penalized Quasi-likelihood for High-dimensional Longitudinal Data via Within-cluster Resampling
topic Methodology
Computation
url https://arxiv.org/abs/2501.01021