GenomeQA: Benchmarking General Large Language Models for Genome Sequence Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Long, Weicai, Hou, Yusen, Feng, Junning, Su, Houcheng, Yang, Shuo, Xie, Donglin, Zhang, Yanlin
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914452000997376
author Long, Weicai
Hou, Yusen
Feng, Junning
Su, Houcheng
Yang, Shuo
Xie, Donglin
Zhang, Yanlin
author_facet Long, Weicai
Hou, Yusen
Feng, Junning
Su, Houcheng
Yang, Shuo
Xie, Donglin
Zhang, Yanlin
contents Large Language Models (LLMs) are increasingly adopted as conversational assistants in genomics, where they are mainly used to reason over biological knowledge, annotations, and analysis outputs through natural language interfaces. However, existing benchmarks either focus on specialized DNA models trained for sequence prediction or evaluate biological knowledge using text-only questions, leaving the behavior of general-purpose LLMs when directly exposed to raw genome sequences underexplored. We introduce GenomeQA, a benchmark designed to provide a controlled evaluation setting for general-purpose LLMs on sequence-based genome inference tasks. GenomeQA comprises 5,200 samples drawn from multiple biological databases, with sequence lengths ranging from 6 to 1,000 base pairs (bp), spanning six task families: Enhancer and Promoter Identification, Splice Site Identification, Taxonomic Classification, Histone Mark Prediction, Transcription Factor Binding Site Prediction, and TF Motif Prediction. Across six frontier LLMs, we find that models consistently outperform random baselines and can exploit local sequence signals such as GC content and short motifs, while performance degrades on tasks that require more indirect or multi-step inference over sequence patterns. GenomeQA establishes a diagnostic benchmark for studying and improving the use of general-purpose LLMs on raw genomic sequences.
format Preprint
id arxiv_https___arxiv_org_abs_2604_05774
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle GenomeQA: Benchmarking General Large Language Models for Genome Sequence Understanding
Long, Weicai
Hou, Yusen
Feng, Junning
Su, Houcheng
Yang, Shuo
Xie, Donglin
Zhang, Yanlin
Genomics
Computation and Language
Large Language Models (LLMs) are increasingly adopted as conversational assistants in genomics, where they are mainly used to reason over biological knowledge, annotations, and analysis outputs through natural language interfaces. However, existing benchmarks either focus on specialized DNA models trained for sequence prediction or evaluate biological knowledge using text-only questions, leaving the behavior of general-purpose LLMs when directly exposed to raw genome sequences underexplored. We introduce GenomeQA, a benchmark designed to provide a controlled evaluation setting for general-purpose LLMs on sequence-based genome inference tasks. GenomeQA comprises 5,200 samples drawn from multiple biological databases, with sequence lengths ranging from 6 to 1,000 base pairs (bp), spanning six task families: Enhancer and Promoter Identification, Splice Site Identification, Taxonomic Classification, Histone Mark Prediction, Transcription Factor Binding Site Prediction, and TF Motif Prediction. Across six frontier LLMs, we find that models consistently outperform random baselines and can exploit local sequence signals such as GC content and short motifs, while performance degrades on tasks that require more indirect or multi-step inference over sequence patterns. GenomeQA establishes a diagnostic benchmark for studying and improving the use of general-purpose LLMs on raw genomic sequences.
title GenomeQA: Benchmarking General Large Language Models for Genome Sequence Understanding
topic Genomics
Computation and Language
url https://arxiv.org/abs/2604.05774