Multi-omics data integration for early diagnosis of hepatocellular carcinoma (HCC) using machine learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Spooner, Annette, Moridani, Mohammad Karimi, Safarchi, Azadeh, Maher, Salim, Vafaee, Fatemeh, Zekry, Amany, Sowmya, Arcot
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909581956874240
author Spooner, Annette
Moridani, Mohammad Karimi
Safarchi, Azadeh
Maher, Salim
Vafaee, Fatemeh
Zekry, Amany
Sowmya, Arcot
author_facet Spooner, Annette
Moridani, Mohammad Karimi
Safarchi, Azadeh
Maher, Salim
Vafaee, Fatemeh
Zekry, Amany
Sowmya, Arcot
contents The complementary information found in different modalities of patient data can aid in more accurate modelling of a patient's disease state and a better understanding of the underlying biological processes of a disease. However, the analysis of multi-modal, multi-omics data presents many challenges, including high dimensionality and varying size, statistical distribution, scale and signal strength between modalities. In this work we compare the performance of a variety of ensemble machine learning algorithms that are capable of late integration of multi-class data from different modalities. The ensemble methods and their variations tested were i) a voting ensemble, with hard and soft vote, ii) a meta learner, iii) a multi-modal Adaboost model using a hard vote, a soft vote and a meta learner to integrate the modalities on each boosting round, the PB-MVBoost model and a novel application of a mixture of experts model. These were compared to simple concatenation as a baseline. We examine these methods using data from an in-house study on hepatocellular carcinoma (HCC), along with four validation datasets on studies from breast cancer and irritable bowel disease (IBD). Using the area under the receiver operating curve as a measure of performance we develop models that achieve a performance value of up to 0.85 and find that two boosted methods, PB-MVBoost and Adaboost with a soft vote were the overall best performing models. We also examine the stability of features selected, and the size of the clinical signature determined. Finally, we provide recommendations for the integration of multi-modal multi-class data.
format Preprint
id arxiv_https___arxiv_org_abs_2409_13791
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Multi-omics data integration for early diagnosis of hepatocellular carcinoma (HCC) using machine learning
Spooner, Annette
Moridani, Mohammad Karimi
Safarchi, Azadeh
Maher, Salim
Vafaee, Fatemeh
Zekry, Amany
Sowmya, Arcot
Machine Learning
Artificial Intelligence
The complementary information found in different modalities of patient data can aid in more accurate modelling of a patient's disease state and a better understanding of the underlying biological processes of a disease. However, the analysis of multi-modal, multi-omics data presents many challenges, including high dimensionality and varying size, statistical distribution, scale and signal strength between modalities. In this work we compare the performance of a variety of ensemble machine learning algorithms that are capable of late integration of multi-class data from different modalities. The ensemble methods and their variations tested were i) a voting ensemble, with hard and soft vote, ii) a meta learner, iii) a multi-modal Adaboost model using a hard vote, a soft vote and a meta learner to integrate the modalities on each boosting round, the PB-MVBoost model and a novel application of a mixture of experts model. These were compared to simple concatenation as a baseline. We examine these methods using data from an in-house study on hepatocellular carcinoma (HCC), along with four validation datasets on studies from breast cancer and irritable bowel disease (IBD). Using the area under the receiver operating curve as a measure of performance we develop models that achieve a performance value of up to 0.85 and find that two boosted methods, PB-MVBoost and Adaboost with a soft vote were the overall best performing models. We also examine the stability of features selected, and the size of the clinical signature determined. Finally, we provide recommendations for the integration of multi-modal multi-class data.
title Multi-omics data integration for early diagnosis of hepatocellular carcinoma (HCC) using machine learning
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2409.13791