Metadata-guided Feature Disentanglement for Functional Genomics

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rakowski, Alexander, Monti, Remo, Huryn, Viktoriia, Lemanczyk, Marta, Ohler, Uwe, Lippert, Christoph
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929363804487680
author Rakowski, Alexander
Monti, Remo
Huryn, Viktoriia
Lemanczyk, Marta
Ohler, Uwe
Lippert, Christoph
author_facet Rakowski, Alexander
Monti, Remo
Huryn, Viktoriia
Lemanczyk, Marta
Ohler, Uwe
Lippert, Christoph
contents With the development of high-throughput technologies, genomics datasets rapidly grow in size, including functional genomics data. This has allowed the training of large Deep Learning (DL) models to predict epigenetic readouts, such as protein binding or histone modifications, from genome sequences. However, large dataset sizes come at a price of data consistency, often aggregating results from a large number of studies, conducted under varying experimental conditions. While data from large-scale consortia are useful as they allow studying the effects of different biological conditions, they can also contain unwanted biases from confounding experimental factors. Here, we introduce Metadata-guided Feature Disentanglement (MFD) - an approach that allows disentangling biologically relevant features from potential technical biases. MFD incorporates target metadata into model training, by conditioning weights of the model output layer on different experimental factors. It then separates the factors into disjoint groups and enforces independence of the corresponding feature subspaces with an adversarially learned penalty. We show that the metadata-driven disentanglement approach allows for better model introspection, by connecting latent features to experimental factors, without compromising, or even improving performance in downstream tasks, such as enhancer prediction, or genetic variant discovery. The code for our implemementation is available at https://github.com/HealthML/MFD
format Preprint
id arxiv_https___arxiv_org_abs_2405_19057
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Metadata-guided Feature Disentanglement for Functional Genomics
Rakowski, Alexander
Monti, Remo
Huryn, Viktoriia
Lemanczyk, Marta
Ohler, Uwe
Lippert, Christoph
Genomics
With the development of high-throughput technologies, genomics datasets rapidly grow in size, including functional genomics data. This has allowed the training of large Deep Learning (DL) models to predict epigenetic readouts, such as protein binding or histone modifications, from genome sequences. However, large dataset sizes come at a price of data consistency, often aggregating results from a large number of studies, conducted under varying experimental conditions. While data from large-scale consortia are useful as they allow studying the effects of different biological conditions, they can also contain unwanted biases from confounding experimental factors. Here, we introduce Metadata-guided Feature Disentanglement (MFD) - an approach that allows disentangling biologically relevant features from potential technical biases. MFD incorporates target metadata into model training, by conditioning weights of the model output layer on different experimental factors. It then separates the factors into disjoint groups and enforces independence of the corresponding feature subspaces with an adversarially learned penalty. We show that the metadata-driven disentanglement approach allows for better model introspection, by connecting latent features to experimental factors, without compromising, or even improving performance in downstream tasks, such as enhancer prediction, or genetic variant discovery. The code for our implemementation is available at https://github.com/HealthML/MFD
title Metadata-guided Feature Disentanglement for Functional Genomics
topic Genomics
url https://arxiv.org/abs/2405.19057