A technical review of multi-omics data integration methods: from classical statistical to deep generative approaches

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Baião, Ana R., Cai, Zhaoxiang, Poulos, Rebecca C, Robinson, Phillip J., Reddel, Roger R, Zhong, Qing, Vinga, Susana, Gonçalves, Emanuel
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912209450303488
author Baião, Ana R.
Cai, Zhaoxiang
Poulos, Rebecca C
Robinson, Phillip J.
Reddel, Roger R
Zhong, Qing
Vinga, Susana
Gonçalves, Emanuel
author_facet Baião, Ana R.
Cai, Zhaoxiang
Poulos, Rebecca C
Robinson, Phillip J.
Reddel, Roger R
Zhong, Qing
Vinga, Susana
Gonçalves, Emanuel
contents The rapid advancement of high-throughput sequencing and other assay technologies has resulted in the generation of large and complex multi-omics datasets, offering unprecedented opportunities for advancing precision medicine strategies. However, multi-omics data integration presents significant challenges due to the high dimensionality, heterogeneity, experimental gaps, and frequency of missing values across data types. Computational methods have been developed to address these issues, employing statistical and machine learning approaches to uncover complex biological patterns and provide deeper insights into our understanding of disease mechanisms. Here, we comprehensively review state-of-the-art multi-omics data integration methods with a focus on deep generative models, particularly variational autoencoders (VAEs) that have been widely used for data imputation and augmentation, joint embedding creation, and batch effect correction. We explore the technical aspects of loss functions and regularisation techniques including adversarial training, disentanglement and contrastive learning. Moreover, we discuss recent advancements in foundation models and the integration of emerging data modalities, while describing the current limitations and outlining future directions for enhancing multi-modal methodologies in biomedical research.
format Preprint
id arxiv_https___arxiv_org_abs_2501_17729
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A technical review of multi-omics data integration methods: from classical statistical to deep generative approaches
Baião, Ana R.
Cai, Zhaoxiang
Poulos, Rebecca C
Robinson, Phillip J.
Reddel, Roger R
Zhong, Qing
Vinga, Susana
Gonçalves, Emanuel
Quantitative Methods
The rapid advancement of high-throughput sequencing and other assay technologies has resulted in the generation of large and complex multi-omics datasets, offering unprecedented opportunities for advancing precision medicine strategies. However, multi-omics data integration presents significant challenges due to the high dimensionality, heterogeneity, experimental gaps, and frequency of missing values across data types. Computational methods have been developed to address these issues, employing statistical and machine learning approaches to uncover complex biological patterns and provide deeper insights into our understanding of disease mechanisms. Here, we comprehensively review state-of-the-art multi-omics data integration methods with a focus on deep generative models, particularly variational autoencoders (VAEs) that have been widely used for data imputation and augmentation, joint embedding creation, and batch effect correction. We explore the technical aspects of loss functions and regularisation techniques including adversarial training, disentanglement and contrastive learning. Moreover, we discuss recent advancements in foundation models and the integration of emerging data modalities, while describing the current limitations and outlining future directions for enhancing multi-modal methodologies in biomedical research.
title A technical review of multi-omics data integration methods: from classical statistical to deep generative approaches
topic Quantitative Methods
url https://arxiv.org/abs/2501.17729