Debiased Multimodal Understanding for Human Language Sequences

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Xu, Zhi, Yang, Dingkang, Li, Mingcheng, Wang, Yuzheng, Chen, Zhaoyu, Chen, Jiawei, Wei, Jinjie, Zhang, Lihua
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913610198941696
author Xu, Zhi
Yang, Dingkang
Li, Mingcheng
Wang, Yuzheng
Chen, Zhaoyu
Chen, Jiawei
Wei, Jinjie
Zhang, Lihua
author_facet Xu, Zhi
Yang, Dingkang
Li, Mingcheng
Wang, Yuzheng
Chen, Zhaoyu
Chen, Jiawei
Wei, Jinjie
Zhang, Lihua
contents Human multimodal language understanding (MLU) is an indispensable component of expression analysis (e.g., sentiment or humor) from heterogeneous modalities, including visual postures, linguistic contents, and acoustic behaviours. Existing works invariably focus on designing sophisticated structures or fusion strategies to achieve impressive improvements. Unfortunately, they all suffer from the subject variation problem due to data distribution discrepancies among subjects. Concretely, MLU models are easily misled by distinct subjects with different expression customs and characteristics in the training data to learn subject-specific spurious correlations, limiting performance and generalizability across new subjects. Motivated by this observation, we introduce a recapitulative causal graph to formulate the MLU procedure and analyze the confounding effect of subjects. Then, we propose SuCI, a simple yet effective causal intervention module to disentangle the impact of subjects acting as unobserved confounders and achieve model training via true causal effects. As a plug-and-play component, SuCI can be widely applied to most methods that seek unbiased predictions. Comprehensive experiments on several MLU benchmarks clearly show the effectiveness of the proposed module.
format Preprint
id arxiv_https___arxiv_org_abs_2403_05025
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Debiased Multimodal Understanding for Human Language Sequences
Xu, Zhi
Yang, Dingkang
Li, Mingcheng
Wang, Yuzheng
Chen, Zhaoyu
Chen, Jiawei
Wei, Jinjie
Zhang, Lihua
Artificial Intelligence
Human multimodal language understanding (MLU) is an indispensable component of expression analysis (e.g., sentiment or humor) from heterogeneous modalities, including visual postures, linguistic contents, and acoustic behaviours. Existing works invariably focus on designing sophisticated structures or fusion strategies to achieve impressive improvements. Unfortunately, they all suffer from the subject variation problem due to data distribution discrepancies among subjects. Concretely, MLU models are easily misled by distinct subjects with different expression customs and characteristics in the training data to learn subject-specific spurious correlations, limiting performance and generalizability across new subjects. Motivated by this observation, we introduce a recapitulative causal graph to formulate the MLU procedure and analyze the confounding effect of subjects. Then, we propose SuCI, a simple yet effective causal intervention module to disentangle the impact of subjects acting as unobserved confounders and achieve model training via true causal effects. As a plug-and-play component, SuCI can be widely applied to most methods that seek unbiased predictions. Comprehensive experiments on several MLU benchmarks clearly show the effectiveness of the proposed module.
title Debiased Multimodal Understanding for Human Language Sequences
topic Artificial Intelligence
url https://arxiv.org/abs/2403.05025