MMoE: Enhancing Multimodal Models with Mixtures of Multimodal Interaction Experts

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Yu, Haofei, Qi, Zhengyang, Jang, Lawrence, Salakhutdinov, Ruslan, Morency, Louis-Philippe, Liang, Paul Pu
Formato: Preprint
Publicado: 2023
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910619914993664
author Yu, Haofei
Qi, Zhengyang
Jang, Lawrence
Salakhutdinov, Ruslan
Morency, Louis-Philippe
Liang, Paul Pu
author_facet Yu, Haofei
Qi, Zhengyang
Jang, Lawrence
Salakhutdinov, Ruslan
Morency, Louis-Philippe
Liang, Paul Pu
contents Advances in multimodal models have greatly improved how interactions relevant to various tasks are modeled. Today's multimodal models mainly focus on the correspondence between images and text, using this for tasks like image-text matching. However, this covers only a subset of real-world interactions. Novel interactions, such as sarcasm expressed through opposing spoken words and gestures or humor expressed through utterances and tone of voice, remain challenging. In this paper, we introduce an approach to enhance multimodal models, which we call Multimodal Mixtures of Experts (MMoE). The key idea in MMoE is to train separate expert models for each type of multimodal interaction, such as redundancy present in both modalities, uniqueness in one modality, or synergy that emerges when both modalities are fused. On a sarcasm detection task (MUStARD) and a humor detection task (URFUNNY), we obtain new state-of-the-art results. MMoE is also able to be applied to various types of models to gain improvement.
format Preprint
id arxiv_https___arxiv_org_abs_2311_09580
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle MMoE: Enhancing Multimodal Models with Mixtures of Multimodal Interaction Experts
Yu, Haofei
Qi, Zhengyang
Jang, Lawrence
Salakhutdinov, Ruslan
Morency, Louis-Philippe
Liang, Paul Pu
Computation and Language
Advances in multimodal models have greatly improved how interactions relevant to various tasks are modeled. Today's multimodal models mainly focus on the correspondence between images and text, using this for tasks like image-text matching. However, this covers only a subset of real-world interactions. Novel interactions, such as sarcasm expressed through opposing spoken words and gestures or humor expressed through utterances and tone of voice, remain challenging. In this paper, we introduce an approach to enhance multimodal models, which we call Multimodal Mixtures of Experts (MMoE). The key idea in MMoE is to train separate expert models for each type of multimodal interaction, such as redundancy present in both modalities, uniqueness in one modality, or synergy that emerges when both modalities are fused. On a sarcasm detection task (MUStARD) and a humor detection task (URFUNNY), we obtain new state-of-the-art results. MMoE is also able to be applied to various types of models to gain improvement.
title MMoE: Enhancing Multimodal Models with Mixtures of Multimodal Interaction Experts
topic Computation and Language
url https://arxiv.org/abs/2311.09580