Multi-View Variational Autoencoder for Missing Value Imputation in Untargeted Metabolomics

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Chen, Su, Kuan-Jui, Wu, Chong, Cao, Xuewei, Sha, Qiuying, Li, Wu, Luo, Zhe, Qin, Tian, Qiu, Chuan, Zhao, Lan Juan, Liu, Anqi, Jiang, Lindong, Zhang, Xiao, Shen, Hui, Zhou, Weihua, Deng, Hong-Wen
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929272888754176
author Zhao, Chen
Su, Kuan-Jui
Wu, Chong
Cao, Xuewei
Sha, Qiuying
Li, Wu
Luo, Zhe
Qin, Tian
Qiu, Chuan
Zhao, Lan Juan
Liu, Anqi
Jiang, Lindong
Zhang, Xiao
Shen, Hui
Zhou, Weihua
Deng, Hong-Wen
author_facet Zhao, Chen
Su, Kuan-Jui
Wu, Chong
Cao, Xuewei
Sha, Qiuying
Li, Wu
Luo, Zhe
Qin, Tian
Qiu, Chuan
Zhao, Lan Juan
Liu, Anqi
Jiang, Lindong
Zhang, Xiao
Shen, Hui
Zhou, Weihua
Deng, Hong-Wen
contents Background: Missing data is a common challenge in mass spectrometry-based metabolomics, which can lead to biased and incomplete analyses. The integration of whole-genome sequencing (WGS) data with metabolomics data has emerged as a promising approach to enhance the accuracy of data imputation in metabolomics studies. Method: In this study, we propose a novel method that leverages the information from WGS data and reference metabolites to impute unknown metabolites. Our approach utilizes a multi-view variational autoencoder to jointly model the burden score, polygenetic risk score (PGS), and linkage disequilibrium (LD) pruned single nucleotide polymorphisms (SNPs) for feature extraction and missing metabolomics data imputation. By learning the latent representations of both omics data, our method can effectively impute missing metabolomics values based on genomic information. Results: We evaluate the performance of our method on empirical metabolomics datasets with missing values and demonstrate its superiority compared to conventional imputation techniques. Using 35 template metabolites derived burden scores, PGS and LD-pruned SNPs, the proposed methods achieved R^2-scores > 0.01 for 71.55% of metabolites. Conclusion: The integration of WGS data in metabolomics imputation not only improves data completeness but also enhances downstream analyses, paving the way for more comprehensive and accurate investigations of metabolic pathways and disease associations. Our findings offer valuable insights into the potential benefits of utilizing WGS data for metabolomics data imputation and underscore the importance of leveraging multi-modal data integration in precision medicine research.
format Preprint
id arxiv_https___arxiv_org_abs_2310_07990
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Multi-View Variational Autoencoder for Missing Value Imputation in Untargeted Metabolomics
Zhao, Chen
Su, Kuan-Jui
Wu, Chong
Cao, Xuewei
Sha, Qiuying
Li, Wu
Luo, Zhe
Qin, Tian
Qiu, Chuan
Zhao, Lan Juan
Liu, Anqi
Jiang, Lindong
Zhang, Xiao
Shen, Hui
Zhou, Weihua
Deng, Hong-Wen
Genomics
Information Retrieval
Machine Learning
Applications
Background: Missing data is a common challenge in mass spectrometry-based metabolomics, which can lead to biased and incomplete analyses. The integration of whole-genome sequencing (WGS) data with metabolomics data has emerged as a promising approach to enhance the accuracy of data imputation in metabolomics studies. Method: In this study, we propose a novel method that leverages the information from WGS data and reference metabolites to impute unknown metabolites. Our approach utilizes a multi-view variational autoencoder to jointly model the burden score, polygenetic risk score (PGS), and linkage disequilibrium (LD) pruned single nucleotide polymorphisms (SNPs) for feature extraction and missing metabolomics data imputation. By learning the latent representations of both omics data, our method can effectively impute missing metabolomics values based on genomic information. Results: We evaluate the performance of our method on empirical metabolomics datasets with missing values and demonstrate its superiority compared to conventional imputation techniques. Using 35 template metabolites derived burden scores, PGS and LD-pruned SNPs, the proposed methods achieved R^2-scores > 0.01 for 71.55% of metabolites. Conclusion: The integration of WGS data in metabolomics imputation not only improves data completeness but also enhances downstream analyses, paving the way for more comprehensive and accurate investigations of metabolic pathways and disease associations. Our findings offer valuable insights into the potential benefits of utilizing WGS data for metabolomics data imputation and underscore the importance of leveraging multi-modal data integration in precision medicine research.
title Multi-View Variational Autoencoder for Missing Value Imputation in Untargeted Metabolomics
topic Genomics
Information Retrieval
Machine Learning
Applications
url https://arxiv.org/abs/2310.07990