GQVis: A Dataset of Genomics Data Questions and Visualizations for Generative AI

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Walters, Skylar Sargent, Valderrama, Arthea, Smits, Thomas C., Kouřil, David, Nguyen, Huyen N., L'Yi, Sehi, Lange, Devin, Gehlenborg, Nils
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914095963308032
author Walters, Skylar Sargent
Valderrama, Arthea
Smits, Thomas C.
Kouřil, David
Nguyen, Huyen N.
L'Yi, Sehi
Lange, Devin
Gehlenborg, Nils
author_facet Walters, Skylar Sargent
Valderrama, Arthea
Smits, Thomas C.
Kouřil, David
Nguyen, Huyen N.
L'Yi, Sehi
Lange, Devin
Gehlenborg, Nils
contents Data visualization is a fundamental tool in genomics research, enabling the exploration, interpretation, and communication of complex genomic features. While machine learning models show promise for transforming data into insightful visualizations, current models lack the training foundation for domain-specific tasks. In an effort to provide a foundational resource for genomics-focused model training, we present a framework for generating a dataset that pairs abstract, low-level questions about genomics data with corresponding visualizations. Building on prior work with statistical plots, our approach adapts to the complexity of genomics data and the specialized representations used to depict them. We further incorporate multiple linked queries and visualizations, along with justifications for design choices, figure captions, and image alt-texts for each item in the dataset. We use genomics data retrieved from three distinct genomics data repositories (4DN, ENCODE, Chromoscope) to produce GQVis: a dataset consisting of 1.14 million single-query data points, 628k query pairs, and 589k query chains. The GQVis dataset and generation code are available at https://huggingface.co/datasets/HIDIVE/GQVis and https://github.com/hms-dbmi/GQVis-Generation.
format Preprint
id arxiv_https___arxiv_org_abs_2510_13816
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GQVis: A Dataset of Genomics Data Questions and Visualizations for Generative AI
Walters, Skylar Sargent
Valderrama, Arthea
Smits, Thomas C.
Kouřil, David
Nguyen, Huyen N.
L'Yi, Sehi
Lange, Devin
Gehlenborg, Nils
Genomics
Artificial Intelligence
Human-Computer Interaction
Machine Learning
Data visualization is a fundamental tool in genomics research, enabling the exploration, interpretation, and communication of complex genomic features. While machine learning models show promise for transforming data into insightful visualizations, current models lack the training foundation for domain-specific tasks. In an effort to provide a foundational resource for genomics-focused model training, we present a framework for generating a dataset that pairs abstract, low-level questions about genomics data with corresponding visualizations. Building on prior work with statistical plots, our approach adapts to the complexity of genomics data and the specialized representations used to depict them. We further incorporate multiple linked queries and visualizations, along with justifications for design choices, figure captions, and image alt-texts for each item in the dataset. We use genomics data retrieved from three distinct genomics data repositories (4DN, ENCODE, Chromoscope) to produce GQVis: a dataset consisting of 1.14 million single-query data points, 628k query pairs, and 589k query chains. The GQVis dataset and generation code are available at https://huggingface.co/datasets/HIDIVE/GQVis and https://github.com/hms-dbmi/GQVis-Generation.
title GQVis: A Dataset of Genomics Data Questions and Visualizations for Generative AI
topic Genomics
Artificial Intelligence
Human-Computer Interaction
Machine Learning
url https://arxiv.org/abs/2510.13816