Leveraging genomic deep learning models for the prediction of non-coding variant effects

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kathail, Pooja, Bajwa, Ayesha, Ioannidis, Nilah M.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911281253974016
author Kathail, Pooja
Bajwa, Ayesha
Ioannidis, Nilah M.
author_facet Kathail, Pooja
Bajwa, Ayesha
Ioannidis, Nilah M.
contents Characterizing non-coding variant function remains an important challenge in human genetics. Genomic deep learning models have emerged as a promising approach to enable in silico prediction of variant effects. These include supervised sequence-to-activity models, which predict molecular phenotypes such as genome-wide chromatin states or gene expression levels directly from DNA sequence, and self-supervised genomic language models. Here, we review progress in leveraging these models for non-coding variant effect prediction. We describe practical considerations for making such predictions and categorize the types of ground truth data used to evaluate variant effect predictions, providing insight into the settings in which current models are most useful. Our Review highlights key considerations for practitioners and opportunities for improvement in model development and evaluation.
format Preprint
id arxiv_https___arxiv_org_abs_2411_11158
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Leveraging genomic deep learning models for the prediction of non-coding variant effects
Kathail, Pooja
Bajwa, Ayesha
Ioannidis, Nilah M.
Genomics
Characterizing non-coding variant function remains an important challenge in human genetics. Genomic deep learning models have emerged as a promising approach to enable in silico prediction of variant effects. These include supervised sequence-to-activity models, which predict molecular phenotypes such as genome-wide chromatin states or gene expression levels directly from DNA sequence, and self-supervised genomic language models. Here, we review progress in leveraging these models for non-coding variant effect prediction. We describe practical considerations for making such predictions and categorize the types of ground truth data used to evaluate variant effect predictions, providing insight into the settings in which current models are most useful. Our Review highlights key considerations for practitioners and opportunities for improvement in model development and evaluation.
title Leveraging genomic deep learning models for the prediction of non-coding variant effects
topic Genomics
url https://arxiv.org/abs/2411.11158