Integrating Large Language Models for Genetic Variant Classification

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Boulaimen, Youssef, Fossi, Gabriele, Outemzabet, Leila, Jeanray, Nathalie, Levenets, Oleksandr, Gerart, Stephane, Vachenc, Sebastien, Raieli, Salvatore, Giemza, Joanna
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909380334583808
author Boulaimen, Youssef
Fossi, Gabriele
Outemzabet, Leila
Jeanray, Nathalie
Levenets, Oleksandr
Gerart, Stephane
Vachenc, Sebastien
Raieli, Salvatore
Giemza, Joanna
author_facet Boulaimen, Youssef
Fossi, Gabriele
Outemzabet, Leila
Jeanray, Nathalie
Levenets, Oleksandr
Gerart, Stephane
Vachenc, Sebastien
Raieli, Salvatore
Giemza, Joanna
contents The classification of genetic variants, particularly Variants of Uncertain Significance (VUS), poses a significant challenge in clinical genetics and precision medicine. Large Language Models (LLMs) have emerged as transformative tools in this realm. These models can uncover intricate patterns and predictive insights that traditional methods might miss, thus enhancing the predictive accuracy of genetic variant pathogenicity. This study investigates the integration of state-of-the-art LLMs, including GPN-MSA, ESM1b, and AlphaMissense, which leverage DNA and protein sequence data alongside structural insights to form a comprehensive analytical framework for variant classification. Our approach evaluates these integrated models using the well-annotated ProteinGym and ClinVar datasets, setting new benchmarks in classification performance. The models were rigorously tested on a set of challenging variants, demonstrating substantial improvements over existing state-of-the-art tools, especially in handling ambiguous and clinically uncertain variants. The results of this research underline the efficacy of combining multiple modeling approaches to significantly refine the accuracy and reliability of genetic variant classification systems. These findings support the deployment of these advanced computational models in clinical environments, where they can significantly enhance the diagnostic processes for genetic disorders, ultimately pushing the boundaries of personalized medicine by offering more detailed and actionable genetic insights.
format Preprint
id arxiv_https___arxiv_org_abs_2411_05055
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Integrating Large Language Models for Genetic Variant Classification
Boulaimen, Youssef
Fossi, Gabriele
Outemzabet, Leila
Jeanray, Nathalie
Levenets, Oleksandr
Gerart, Stephane
Vachenc, Sebastien
Raieli, Salvatore
Giemza, Joanna
Genomics
Artificial Intelligence
Machine Learning
68T07 (Primary) 92D20, 92C40 (Secondary)
I.2.7; J.3
The classification of genetic variants, particularly Variants of Uncertain Significance (VUS), poses a significant challenge in clinical genetics and precision medicine. Large Language Models (LLMs) have emerged as transformative tools in this realm. These models can uncover intricate patterns and predictive insights that traditional methods might miss, thus enhancing the predictive accuracy of genetic variant pathogenicity. This study investigates the integration of state-of-the-art LLMs, including GPN-MSA, ESM1b, and AlphaMissense, which leverage DNA and protein sequence data alongside structural insights to form a comprehensive analytical framework for variant classification. Our approach evaluates these integrated models using the well-annotated ProteinGym and ClinVar datasets, setting new benchmarks in classification performance. The models were rigorously tested on a set of challenging variants, demonstrating substantial improvements over existing state-of-the-art tools, especially in handling ambiguous and clinically uncertain variants. The results of this research underline the efficacy of combining multiple modeling approaches to significantly refine the accuracy and reliability of genetic variant classification systems. These findings support the deployment of these advanced computational models in clinical environments, where they can significantly enhance the diagnostic processes for genetic disorders, ultimately pushing the boundaries of personalized medicine by offering more detailed and actionable genetic insights.
title Integrating Large Language Models for Genetic Variant Classification
topic Genomics
Artificial Intelligence
Machine Learning
68T07 (Primary) 92D20, 92C40 (Secondary)
I.2.7; J.3
url https://arxiv.org/abs/2411.05055