A near-exact linear mixed model for genome-wide association studies

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pu, Zhibin, Ge, Shufei, Wang, Shijia
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915433561456640
author Pu, Zhibin
Ge, Shufei
Wang, Shijia
author_facet Pu, Zhibin
Ge, Shufei
Wang, Shijia
contents Linear mixed models (LMM) are widely adopted in genome-wide association studies (GWAS) to account for population stratification and cryptic relatedness. However, the parameter estimation of LMMs imposes substantial computational burdens due to large-scale operations on genetic similarity matrices (GSM). We introduced the near-exact linear mixed model (NExt-LMM), a novel LMM framework that overcomes critical computational bottlenecks in GWAS through the following key innovations. Firstly, we exploit the inherent low-rank structure of the GSM iteratively with the Hierarchical Off-Diagonal Low-Rank (HODLR) format, which is much faster than traditional decomposition methods. Secondly, we leverage the HODLR-approximated GSM to dramatically accelerate the further maximum likelihood estimation with the shared heritability ratios. Moreover, we establish rigorous error bounds for the NExt-LMM estimator, proving that Kullback-Leibler divergence between the approximated and exact estimators can be arbitrarily small. Consequently, our proposed dual approach accelerates inference of LMMs while guaranteeing low approximation errors. We use numerical experiments to demonstrate that the NExt-LMM significantly improves inference efficiency compared to existing methods. We develop a Python package that is available at https://github.com/ZhibinPU/NExt-LMM.
format Preprint
id arxiv_https___arxiv_org_abs_2508_05278
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A near-exact linear mixed model for genome-wide association studies
Pu, Zhibin
Ge, Shufei
Wang, Shijia
Computation
82-10, 62-08
G.1.2; G.1.6; J.3
Linear mixed models (LMM) are widely adopted in genome-wide association studies (GWAS) to account for population stratification and cryptic relatedness. However, the parameter estimation of LMMs imposes substantial computational burdens due to large-scale operations on genetic similarity matrices (GSM). We introduced the near-exact linear mixed model (NExt-LMM), a novel LMM framework that overcomes critical computational bottlenecks in GWAS through the following key innovations. Firstly, we exploit the inherent low-rank structure of the GSM iteratively with the Hierarchical Off-Diagonal Low-Rank (HODLR) format, which is much faster than traditional decomposition methods. Secondly, we leverage the HODLR-approximated GSM to dramatically accelerate the further maximum likelihood estimation with the shared heritability ratios. Moreover, we establish rigorous error bounds for the NExt-LMM estimator, proving that Kullback-Leibler divergence between the approximated and exact estimators can be arbitrarily small. Consequently, our proposed dual approach accelerates inference of LMMs while guaranteeing low approximation errors. We use numerical experiments to demonstrate that the NExt-LMM significantly improves inference efficiency compared to existing methods. We develop a Python package that is available at https://github.com/ZhibinPU/NExt-LMM.
title A near-exact linear mixed model for genome-wide association studies
topic Computation
82-10, 62-08
G.1.2; G.1.6; J.3
url https://arxiv.org/abs/2508.05278