MetaHQ: Harmonized, high-quality metadata annotations of public omics samples and studies

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hicks, Parker, Valtadoros, Lydia E, Mancuso, Christopher A, Alquadoomi, Faisal, Johnson, Kayla A, Sundar, Sneha, Krishnan, Arjun
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917337211338752
author Hicks, Parker
Valtadoros, Lydia E
Mancuso, Christopher A
Alquadoomi, Faisal
Johnson, Kayla A
Sundar, Sneha
Krishnan, Arjun
author_facet Hicks, Parker
Valtadoros, Lydia E
Mancuso, Christopher A
Alquadoomi, Faisal
Johnson, Kayla A
Sundar, Sneha
Krishnan, Arjun
contents Public omics databases like the Gene Expression Omnibus and the Sequence Read Archive offer substantial opportunities for data reuse to address novel biomedical questions. However, it is still difficult to find samples and studies of interest since they are described by free-text metadata and lack standardized annotations. To address this issue, multiple research groups have undertaken curation efforts to add standardized annotations to large collections of these data, but these annotations are fragmented across online resources and are stored in different formats subject to varying standardization criteria, hindering the integration of annotations across sources. We developed MetaHQ to harmonize and distribute standardized metadata for public omics samples. MetaHQ comprises a database with nearly 200,000 annotations from 13 sources and a user-friendly command-line interface (CLI) to query the database and retrieve annotations. The MetaHQ CLI is deployed as a Python Package on PyPI at https://pypi.org/project/metahq-cli that accesses the MetaHQ database available at https://doi.org/10.5281/zenodo.18462463. Project source code and documentation are available at https://github.com/krishnanlab/meta-hq.
format Preprint
id arxiv_https___arxiv_org_abs_2602_07805
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle MetaHQ: Harmonized, high-quality metadata annotations of public omics samples and studies
Hicks, Parker
Valtadoros, Lydia E
Mancuso, Christopher A
Alquadoomi, Faisal
Johnson, Kayla A
Sundar, Sneha
Krishnan, Arjun
Genomics
Public omics databases like the Gene Expression Omnibus and the Sequence Read Archive offer substantial opportunities for data reuse to address novel biomedical questions. However, it is still difficult to find samples and studies of interest since they are described by free-text metadata and lack standardized annotations. To address this issue, multiple research groups have undertaken curation efforts to add standardized annotations to large collections of these data, but these annotations are fragmented across online resources and are stored in different formats subject to varying standardization criteria, hindering the integration of annotations across sources. We developed MetaHQ to harmonize and distribute standardized metadata for public omics samples. MetaHQ comprises a database with nearly 200,000 annotations from 13 sources and a user-friendly command-line interface (CLI) to query the database and retrieve annotations. The MetaHQ CLI is deployed as a Python Package on PyPI at https://pypi.org/project/metahq-cli that accesses the MetaHQ database available at https://doi.org/10.5281/zenodo.18462463. Project source code and documentation are available at https://github.com/krishnanlab/meta-hq.
title MetaHQ: Harmonized, high-quality metadata annotations of public omics samples and studies
topic Genomics
url https://arxiv.org/abs/2602.07805