Incorporating LLM Embeddings for Variation Across the Human Genome

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Niu, Hongqian, Bryan, Jordan, Williams, Jacob, Zhou, Hufeng, Zhang, Haoyu, Li, Xihao, Li, Didong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915900525903872
author Niu, Hongqian
Bryan, Jordan
Williams, Jacob
Zhou, Hufeng
Zhang, Haoyu
Li, Xihao
Li, Didong
author_facet Niu, Hongqian
Bryan, Jordan
Williams, Jacob
Zhou, Hufeng
Zhang, Haoyu
Li, Xihao
Li, Didong
contents Recent advances in large language model (LLM) embeddings have enabled powerful representations for biological data, but most applications to date focus on gene-level information. We present one of the first systematic frameworks to generate genetic variant-level embeddings across the entire human genome. Using curated annotations from FAVOR, ClinVar, and the GWAS Catalog, we construct functional text descriptions for 8.9 billion possible variants and generated embeddings at three scales: 1.5 million HapMap3/MEGA variants, 90 million imputed UK Biobank (UKB) variants, and 9 billion all possible variants. Embeddings were produced using general purpose models including both OpenAI's text-embedding-3-large and the open-source Qwen3-Embedding-0.6B models. Baseline quality control experiments demonstrate high predictive accuracy for variant-level properties, validating the embeddings as structured representations of genomic variation. We further apply them to real-world embedding-augmented genetic risk predictions that demonstrate the performance of using LLM embeddings in polygenic risk score (PRS) style predictions over the UK Biobank cohort data. These resources, publicly available on Hugging Face, provide a foundation for advancing large-scale genomic discovery and precision medicine.
format Preprint
id arxiv_https___arxiv_org_abs_2509_20702
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Incorporating LLM Embeddings for Variation Across the Human Genome
Niu, Hongqian
Bryan, Jordan
Williams, Jacob
Zhou, Hufeng
Zhang, Haoyu
Li, Xihao
Li, Didong
Applications
Artificial Intelligence
Genomics
Recent advances in large language model (LLM) embeddings have enabled powerful representations for biological data, but most applications to date focus on gene-level information. We present one of the first systematic frameworks to generate genetic variant-level embeddings across the entire human genome. Using curated annotations from FAVOR, ClinVar, and the GWAS Catalog, we construct functional text descriptions for 8.9 billion possible variants and generated embeddings at three scales: 1.5 million HapMap3/MEGA variants, 90 million imputed UK Biobank (UKB) variants, and 9 billion all possible variants. Embeddings were produced using general purpose models including both OpenAI's text-embedding-3-large and the open-source Qwen3-Embedding-0.6B models. Baseline quality control experiments demonstrate high predictive accuracy for variant-level properties, validating the embeddings as structured representations of genomic variation. We further apply them to real-world embedding-augmented genetic risk predictions that demonstrate the performance of using LLM embeddings in polygenic risk score (PRS) style predictions over the UK Biobank cohort data. These resources, publicly available on Hugging Face, provide a foundation for advancing large-scale genomic discovery and precision medicine.
title Incorporating LLM Embeddings for Variation Across the Human Genome
topic Applications
Artificial Intelligence
Genomics
url https://arxiv.org/abs/2509.20702