VitaGraph: Building a Knowledge Graph for Biologically Relevant Learning Tasks

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Madeddu, Francesco, Testa, Lucia, De Carlo, Gianluca, Pieroni, Michele, Mastropietro, Andrea, Anagnostopoulos, Aris, Tieri, Paolo, Barbarossa, Sergio
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918022219825152
author Madeddu, Francesco
Testa, Lucia
De Carlo, Gianluca
Pieroni, Michele
Mastropietro, Andrea
Anagnostopoulos, Aris
Tieri, Paolo
Barbarossa, Sergio
author_facet Madeddu, Francesco
Testa, Lucia
De Carlo, Gianluca
Pieroni, Michele
Mastropietro, Andrea
Anagnostopoulos, Aris
Tieri, Paolo
Barbarossa, Sergio
contents The intrinsic complexity of human biology presents ongoing challenges to scientific understanding. Researchers collaborate across disciplines to expand our knowledge of the biological interactions that define human life. AI methodologies have emerged as powerful tools across scientific domains, particularly in computational biology, where graph data structures effectively model biological entities such as protein-protein interaction (PPI) networks and gene functional networks. Those networks are used as datasets for paramount network medicine tasks, such as gene-disease association prediction, drug repurposing, and polypharmacy side effect studies. Reliable predictions from machine learning models require high-quality foundational data. In this work, we present a comprehensive multi-purpose biological knowledge graph constructed by integrating and refining multiple publicly available datasets. Building upon the Drug Repurposing Knowledge Graph (DRKG), we define a pipeline tasked with a) cleaning inconsistencies and redundancies present in DRKG, b) coalescing information from the main available public data sources, and c) enriching the graph nodes with expressive feature vectors such as molecular fingerprints and gene ontologies. Biologically and chemically relevant features improve the capacity of machine learning models to generate accurate and well-structured embedding spaces. The resulting resource represents a coherent and reliable biological knowledge graph that serves as a state-of-the-art platform to advance research in computational biology and precision medicine. Moreover, it offers the opportunity to benchmark graph-based machine learning and network medicine models on relevant tasks. We demonstrate the effectiveness of the proposed dataset by benchmarking it against the task of drug repurposing, PPI prediction, and side-effect prediction, modeled as link prediction problems.
format Preprint
id arxiv_https___arxiv_org_abs_2505_11185
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VitaGraph: Building a Knowledge Graph for Biologically Relevant Learning Tasks
Madeddu, Francesco
Testa, Lucia
De Carlo, Gianluca
Pieroni, Michele
Mastropietro, Andrea
Anagnostopoulos, Aris
Tieri, Paolo
Barbarossa, Sergio
Machine Learning
The intrinsic complexity of human biology presents ongoing challenges to scientific understanding. Researchers collaborate across disciplines to expand our knowledge of the biological interactions that define human life. AI methodologies have emerged as powerful tools across scientific domains, particularly in computational biology, where graph data structures effectively model biological entities such as protein-protein interaction (PPI) networks and gene functional networks. Those networks are used as datasets for paramount network medicine tasks, such as gene-disease association prediction, drug repurposing, and polypharmacy side effect studies. Reliable predictions from machine learning models require high-quality foundational data. In this work, we present a comprehensive multi-purpose biological knowledge graph constructed by integrating and refining multiple publicly available datasets. Building upon the Drug Repurposing Knowledge Graph (DRKG), we define a pipeline tasked with a) cleaning inconsistencies and redundancies present in DRKG, b) coalescing information from the main available public data sources, and c) enriching the graph nodes with expressive feature vectors such as molecular fingerprints and gene ontologies. Biologically and chemically relevant features improve the capacity of machine learning models to generate accurate and well-structured embedding spaces. The resulting resource represents a coherent and reliable biological knowledge graph that serves as a state-of-the-art platform to advance research in computational biology and precision medicine. Moreover, it offers the opportunity to benchmark graph-based machine learning and network medicine models on relevant tasks. We demonstrate the effectiveness of the proposed dataset by benchmarking it against the task of drug repurposing, PPI prediction, and side-effect prediction, modeled as link prediction problems.
title VitaGraph: Building a Knowledge Graph for Biologically Relevant Learning Tasks
topic Machine Learning
url https://arxiv.org/abs/2505.11185