UniEntrezDB: Large-scale Gene Ontology Annotation Dataset and Evaluation Benchmarks with Unified Entrez Gene Identifiers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Miao, Yuwei, Guo, Yuzhi, Ma, Hehuan, Yan, Jingquan, Jiang, Feng, An, Weizhi, Gao, Jean, Huang, Junzhou
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916528649142272
author Miao, Yuwei
Guo, Yuzhi
Ma, Hehuan
Yan, Jingquan
Jiang, Feng
An, Weizhi
Gao, Jean
Huang, Junzhou
author_facet Miao, Yuwei
Guo, Yuzhi
Ma, Hehuan
Yan, Jingquan
Jiang, Feng
An, Weizhi
Gao, Jean
Huang, Junzhou
contents Gene studies are crucial for fields such as protein structure prediction, drug discovery, and cancer genomics, yet they face challenges in fully utilizing the vast and diverse information available. Gene studies require clean, factual datasets to ensure reliable results. Ontology graphs, neatly organized domain terminology graphs, provide ideal sources for domain facts. However, available gene ontology annotations are currently distributed across various databases without unified identifiers for genes and gene products. To address these challenges, we introduce Unified Entrez Gene Identifier Dataset and Benchmarks (UniEntrezDB), the first systematic effort to unify large-scale public Gene Ontology Annotations (GOA) from various databases using unique gene identifiers. UniEntrezDB includes a pre-training dataset and four downstream tasks designed to comprehensively evaluate gene embedding performance from gene, protein, and cell levels, ultimately enhancing the reliability and applicability of LLMs in gene research and other professional settings.
format Preprint
id arxiv_https___arxiv_org_abs_2412_12688
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle UniEntrezDB: Large-scale Gene Ontology Annotation Dataset and Evaluation Benchmarks with Unified Entrez Gene Identifiers
Miao, Yuwei
Guo, Yuzhi
Ma, Hehuan
Yan, Jingquan
Jiang, Feng
An, Weizhi
Gao, Jean
Huang, Junzhou
Databases
Gene studies are crucial for fields such as protein structure prediction, drug discovery, and cancer genomics, yet they face challenges in fully utilizing the vast and diverse information available. Gene studies require clean, factual datasets to ensure reliable results. Ontology graphs, neatly organized domain terminology graphs, provide ideal sources for domain facts. However, available gene ontology annotations are currently distributed across various databases without unified identifiers for genes and gene products. To address these challenges, we introduce Unified Entrez Gene Identifier Dataset and Benchmarks (UniEntrezDB), the first systematic effort to unify large-scale public Gene Ontology Annotations (GOA) from various databases using unique gene identifiers. UniEntrezDB includes a pre-training dataset and four downstream tasks designed to comprehensively evaluate gene embedding performance from gene, protein, and cell levels, ultimately enhancing the reliability and applicability of LLMs in gene research and other professional settings.
title UniEntrezDB: Large-scale Gene Ontology Annotation Dataset and Evaluation Benchmarks with Unified Entrez Gene Identifiers
topic Databases
url https://arxiv.org/abs/2412.12688