Investigating Faithfulness in Large Audio Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mousavi, Pooneh, Jain, Lovenya, Ravanelli, Mirco, Subakan, Cem
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915873160167424
author Mousavi, Pooneh
Jain, Lovenya
Ravanelli, Mirco
Subakan, Cem
author_facet Mousavi, Pooneh
Jain, Lovenya
Ravanelli, Mirco
Subakan, Cem
contents Large Audio Language Models (LALMs) integrate audio encoders with pretrained Large Language Models to perform complex multimodal reasoning tasks. While these models can generate Chain-of-Thought (CoT) explanations, the faithfulness of these reasoning chains remains unclear. In this work, we propose a systematic framework to evaluate CoT faithfulness in LALMs with respect to both the input audio and the final model prediction. We define three criteria for audio faithfulness: hallucination-free, holistic, and attentive listening. We also introduce a benchmark based on both audio and CoT interventions to assess faithfulness. Experiments on Audio Flamingo 3 and Qwen2.5-Omni suggest a potential multimodal disconnect: reasoning often aligns with the final prediction but is not always strongly grounded in the audio and can be vulnerable to hallucinations or adversarial perturbations.
format Preprint
id arxiv_https___arxiv_org_abs_2509_22363
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Investigating Faithfulness in Large Audio Language Models
Mousavi, Pooneh
Jain, Lovenya
Ravanelli, Mirco
Subakan, Cem
Machine Learning
Audio and Speech Processing
Large Audio Language Models (LALMs) integrate audio encoders with pretrained Large Language Models to perform complex multimodal reasoning tasks. While these models can generate Chain-of-Thought (CoT) explanations, the faithfulness of these reasoning chains remains unclear. In this work, we propose a systematic framework to evaluate CoT faithfulness in LALMs with respect to both the input audio and the final model prediction. We define three criteria for audio faithfulness: hallucination-free, holistic, and attentive listening. We also introduce a benchmark based on both audio and CoT interventions to assess faithfulness. Experiments on Audio Flamingo 3 and Qwen2.5-Omni suggest a potential multimodal disconnect: reasoning often aligns with the final prediction but is not always strongly grounded in the audio and can be vulnerable to hallucinations or adversarial perturbations.
title Investigating Faithfulness in Large Audio Language Models
topic Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2509.22363