BEND: Benchmarking DNA Language Models on biologically meaningful tasks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Marin, Frederikke Isa, Teufel, Felix, Horlacher, Marc, Madsen, Dennis, Pultz, Dennis, Winther, Ole, Boomsma, Wouter
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917634563375104
author Marin, Frederikke Isa
Teufel, Felix
Horlacher, Marc
Madsen, Dennis
Pultz, Dennis
Winther, Ole
Boomsma, Wouter
author_facet Marin, Frederikke Isa
Teufel, Felix
Horlacher, Marc
Madsen, Dennis
Pultz, Dennis
Winther, Ole
Boomsma, Wouter
contents The genome sequence contains the blueprint for governing cellular processes. While the availability of genomes has vastly increased over the last decades, experimental annotation of the various functional, non-coding and regulatory elements encoded in the DNA sequence remains both expensive and challenging. This has sparked interest in unsupervised language modeling of genomic DNA, a paradigm that has seen great success for protein sequence data. Although various DNA language models have been proposed, evaluation tasks often differ between individual works, and might not fully recapitulate the fundamental challenges of genome annotation, including the length, scale and sparsity of the data. In this study, we introduce BEND, a Benchmark for DNA language models, featuring a collection of realistic and biologically meaningful downstream tasks defined on the human genome. We find that embeddings from current DNA LMs can approach performance of expert methods on some tasks, but only capture limited information about long-range features. BEND is available at https://github.com/frederikkemarin/BEND.
format Preprint
id arxiv_https___arxiv_org_abs_2311_12570
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle BEND: Benchmarking DNA Language Models on biologically meaningful tasks
Marin, Frederikke Isa
Teufel, Felix
Horlacher, Marc
Madsen, Dennis
Pultz, Dennis
Winther, Ole
Boomsma, Wouter
Genomics
Machine Learning
The genome sequence contains the blueprint for governing cellular processes. While the availability of genomes has vastly increased over the last decades, experimental annotation of the various functional, non-coding and regulatory elements encoded in the DNA sequence remains both expensive and challenging. This has sparked interest in unsupervised language modeling of genomic DNA, a paradigm that has seen great success for protein sequence data. Although various DNA language models have been proposed, evaluation tasks often differ between individual works, and might not fully recapitulate the fundamental challenges of genome annotation, including the length, scale and sparsity of the data. In this study, we introduce BEND, a Benchmark for DNA language models, featuring a collection of realistic and biologically meaningful downstream tasks defined on the human genome. We find that embeddings from current DNA LMs can approach performance of expert methods on some tasks, but only capture limited information about long-range features. BEND is available at https://github.com/frederikkemarin/BEND.
title BEND: Benchmarking DNA Language Models on biologically meaningful tasks
topic Genomics
Machine Learning
url https://arxiv.org/abs/2311.12570