Integrating Heterogeneous Gene Expression Data through Knowledge Graphs for Improving Diabetes Prediction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sousa, Rita T., Paulheim, Heiko
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916219640086528
author Sousa, Rita T.
Paulheim, Heiko
author_facet Sousa, Rita T.
Paulheim, Heiko
contents Diabetes is a worldwide health issue affecting millions of people. Machine learning methods have shown promising results in improving diabetes prediction, particularly through the analysis of diverse data types, namely gene expression data. While gene expression data can provide valuable insights, challenges arise from the fact that the sample sizes in expression datasets are usually limited, and the data from different datasets with different gene expressions cannot be easily combined. This work proposes a novel approach to address these challenges by integrating multiple gene expression datasets and domain-specific knowledge using knowledge graphs, a unique tool for biomedical data integration. KG embedding methods are then employed to generate vector representations, serving as inputs for a classifier. Experiments demonstrated the efficacy of our approach, revealing improvements in diabetes prediction when integrating multiple gene expression datasets and domain-specific knowledge about protein functions and interactions.
format Preprint
id arxiv_https___arxiv_org_abs_2404_14970
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Integrating Heterogeneous Gene Expression Data through Knowledge Graphs for Improving Diabetes Prediction
Sousa, Rita T.
Paulheim, Heiko
Machine Learning
J.3
Diabetes is a worldwide health issue affecting millions of people. Machine learning methods have shown promising results in improving diabetes prediction, particularly through the analysis of diverse data types, namely gene expression data. While gene expression data can provide valuable insights, challenges arise from the fact that the sample sizes in expression datasets are usually limited, and the data from different datasets with different gene expressions cannot be easily combined. This work proposes a novel approach to address these challenges by integrating multiple gene expression datasets and domain-specific knowledge using knowledge graphs, a unique tool for biomedical data integration. KG embedding methods are then employed to generate vector representations, serving as inputs for a classifier. Experiments demonstrated the efficacy of our approach, revealing improvements in diabetes prediction when integrating multiple gene expression datasets and domain-specific knowledge about protein functions and interactions.
title Integrating Heterogeneous Gene Expression Data through Knowledge Graphs for Improving Diabetes Prediction
topic Machine Learning
J.3
url https://arxiv.org/abs/2404.14970