A Weighted U Statistic for Genetic Association Analyses of Sequencing Data
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2015
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915446966452224 |
|---|---|
| author | Wei, Changshuai Li, Ming He, Zihuai Vsevolozhskaya, Olga Schaid, Daniel J. Lu, Qing |
| author_facet | Wei, Changshuai Li, Ming He, Zihuai Vsevolozhskaya, Olga Schaid, Daniel J. Lu, Qing |
| contents | With advancements in next generation sequencing technology, a massive amount of sequencing data are generated, offering a great opportunity to comprehensively investigate the role of rare variants in the genetic etiology of complex diseases. Nevertheless, this poses a great challenge for the statistical analysis of high-dimensional sequencing data. The association analyses based on traditional statistical methods suffer substantial power loss because of the low frequency of genetic variants and the extremely high dimensionality of the data. We developed a weighted U statistic, referred to as WU-seq, for the high-dimensional association analysis of sequencing data. Based on a non-parametric U statistic, WU-SEQ makes no assumption of the underlying disease model and phenotype distribution, and can be applied to a variety of phenotypes. Through simulation studies and an empirical study, we showed that WU-SEQ outperformed a commonly used SKAT method when the underlying assumptions were violated (e.g., the phenotype followed a heavy-tailed distribution). Even when the assumptions were satisfied, WU-SEQ still attained comparable performance to SKAT. Finally, we applied WU-seq to sequencing data from the Dallas Heart Study (DHS), and detected an association between ANGPTL 4 and very low density lipoprotein cholesterol. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_1505_01204 |
| institution | arXiv |
| publishDate | 2015 |
| record_format | arxiv |
| spellingShingle | A Weighted U Statistic for Genetic Association Analyses of Sequencing Data Wei, Changshuai Li, Ming He, Zihuai Vsevolozhskaya, Olga Schaid, Daniel J. Lu, Qing Methodology Artificial Intelligence Machine Learning Quantitative Methods With advancements in next generation sequencing technology, a massive amount of sequencing data are generated, offering a great opportunity to comprehensively investigate the role of rare variants in the genetic etiology of complex diseases. Nevertheless, this poses a great challenge for the statistical analysis of high-dimensional sequencing data. The association analyses based on traditional statistical methods suffer substantial power loss because of the low frequency of genetic variants and the extremely high dimensionality of the data. We developed a weighted U statistic, referred to as WU-seq, for the high-dimensional association analysis of sequencing data. Based on a non-parametric U statistic, WU-SEQ makes no assumption of the underlying disease model and phenotype distribution, and can be applied to a variety of phenotypes. Through simulation studies and an empirical study, we showed that WU-SEQ outperformed a commonly used SKAT method when the underlying assumptions were violated (e.g., the phenotype followed a heavy-tailed distribution). Even when the assumptions were satisfied, WU-SEQ still attained comparable performance to SKAT. Finally, we applied WU-seq to sequencing data from the Dallas Heart Study (DHS), and detected an association between ANGPTL 4 and very low density lipoprotein cholesterol. |
| title | A Weighted U Statistic for Genetic Association Analyses of Sequencing Data |
| topic | Methodology Artificial Intelligence Machine Learning Quantitative Methods |
| url | https://arxiv.org/abs/1505.01204 |