Mixed Signals: Understanding Model Disagreement in Multimodal Empathy Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Srikanth, Maya, Chen, Run, Hirschberg, Julia
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909897289891840
author Srikanth, Maya
Chen, Run
Hirschberg, Julia
author_facet Srikanth, Maya
Chen, Run
Hirschberg, Julia
contents Multimodal models play a key role in empathy detection, but their performance can suffer when modalities provide conflicting cues. To understand these failures, we examine cases where unimodal and multimodal predictions diverge. Using fine-tuned models for text, audio, and video, along with a gated fusion model, we find that such disagreements often reflect underlying ambiguity, as evidenced by annotator uncertainty. Our analysis shows that dominant signals in one modality can mislead fusion when unsupported by others. We also observe that humans, like models, do not consistently benefit from multimodal input. These insights position disagreement as a useful diagnostic signal for identifying challenging examples and improving empathy system robustness.
format Preprint
id arxiv_https___arxiv_org_abs_2505_13979
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Mixed Signals: Understanding Model Disagreement in Multimodal Empathy Detection
Srikanth, Maya
Chen, Run
Hirschberg, Julia
Computation and Language
Multimodal models play a key role in empathy detection, but their performance can suffer when modalities provide conflicting cues. To understand these failures, we examine cases where unimodal and multimodal predictions diverge. Using fine-tuned models for text, audio, and video, along with a gated fusion model, we find that such disagreements often reflect underlying ambiguity, as evidenced by annotator uncertainty. Our analysis shows that dominant signals in one modality can mislead fusion when unsupported by others. We also observe that humans, like models, do not consistently benefit from multimodal input. These insights position disagreement as a useful diagnostic signal for identifying challenging examples and improving empathy system robustness.
title Mixed Signals: Understanding Model Disagreement in Multimodal Empathy Detection
topic Computation and Language
url https://arxiv.org/abs/2505.13979