A Markov Random Field Multi-Modal Variational AutoEncoder

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Oubari, Fouad, Baha, Mohamed El, Meunier, Raphael, Décatoire, Rodrigue, Mougeot, Mathilde
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866910818092711936
author Oubari, Fouad
Baha, Mohamed El
Meunier, Raphael
Décatoire, Rodrigue
Mougeot, Mathilde
author_facet Oubari, Fouad
Baha, Mohamed El
Meunier, Raphael
Décatoire, Rodrigue
Mougeot, Mathilde
contents Recent advancements in multimodal Variational AutoEncoders (VAEs) have highlighted their potential for modeling complex data from multiple modalities. However, many existing approaches use relatively straightforward aggregating schemes that may not fully capture the complex dynamics present between different modalities. This work introduces a novel multimodal VAE that incorporates a Markov Random Field (MRF) into both the prior and posterior distributions. This integration aims to capture complex intermodal interactions more effectively. Unlike previous models, our approach is specifically designed to model and leverage the intricacies of these relationships, enabling a more faithful representation of multimodal data. Our experiments demonstrate that our model performs competitively on the standard PolyMNIST dataset and shows superior performance in managing complex intermodal dependencies in a specially designed synthetic dataset, intended to test intricate relationships.
format Preprint
id arxiv_https___arxiv_org_abs_2408_09576
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Markov Random Field Multi-Modal Variational AutoEncoder
Oubari, Fouad
Baha, Mohamed El
Meunier, Raphael
Décatoire, Rodrigue
Mougeot, Mathilde
Machine Learning
Recent advancements in multimodal Variational AutoEncoders (VAEs) have highlighted their potential for modeling complex data from multiple modalities. However, many existing approaches use relatively straightforward aggregating schemes that may not fully capture the complex dynamics present between different modalities. This work introduces a novel multimodal VAE that incorporates a Markov Random Field (MRF) into both the prior and posterior distributions. This integration aims to capture complex intermodal interactions more effectively. Unlike previous models, our approach is specifically designed to model and leverage the intricacies of these relationships, enabling a more faithful representation of multimodal data. Our experiments demonstrate that our model performs competitively on the standard PolyMNIST dataset and shows superior performance in managing complex intermodal dependencies in a specially designed synthetic dataset, intended to test intricate relationships.
title A Markov Random Field Multi-Modal Variational AutoEncoder
topic Machine Learning
url https://arxiv.org/abs/2408.09576