Evaluating Hallucinations in Audio-Visual Multimodal LLMs with Spoken Queries under Diverse Acoustic Conditions
Fuente:
arXiv
Saved in:
| Main Authors: | Park, Hansol, Ahn, Hoseong, Moon, Junwon, Lee, Yejin, Shim, Kyuhong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Whisper-CD: Accurate Long-Form Speech Recognition using Multi-Negative Contrastive Decoding
by: Ahn, Hoseong, et al.
Published: (2026)
by: Ahn, Hoseong, et al.
Published: (2026)
Correlation of Fréchet Audio Distance With Human Perception of Environmental Audio Is Embedding Dependant
by: Tailleur, Modan, et al.
Published: (2024)
by: Tailleur, Modan, et al.
Published: (2024)
Audio-Based Linguistic Feature Extraction for Enhancing Multi-lingual and Low-Resource Text-to-Speech
by: Kim, Youngjae, et al.
Published: (2024)
by: Kim, Youngjae, et al.
Published: (2024)
FlowAVSE: Efficient Audio-Visual Speech Enhancement with Conditional Flow Matching
by: Jung, Chaeyoung, et al.
Published: (2024)
by: Jung, Chaeyoung, et al.
Published: (2024)
HILCodec: High-Fidelity and Lightweight Neural Audio Codec
by: Ahn, Sunghwan, et al.
Published: (2024)
by: Ahn, Sunghwan, et al.
Published: (2024)
Voice Conversion with Diverse Intonation using Conditional Variational Auto-Encoder
by: Suh, Soobin, et al.
Published: (2025)
by: Suh, Soobin, et al.
Published: (2025)
Interpretable Audio Editing Evaluation via Chain-of-Thought Difference-Commonality Reasoning with Multimodal LLMs
by: Jia, Yuhang, et al.
Published: (2025)
by: Jia, Yuhang, et al.
Published: (2025)
LALM-as-a-Judge: Benchmarking Large Audio-Language Models for Safety Evaluation in Multi-Turn Spoken Dialogues
by: Ivry, Amir, et al.
Published: (2026)
by: Ivry, Amir, et al.
Published: (2026)
Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training
by: Udupa, Sathvik, et al.
Published: (2025)
by: Udupa, Sathvik, et al.
Published: (2025)
Neural Speech Tracking in a Virtual Acoustic Environment: Audio-Visual Benefit for Unscripted Continuous Speech
by: Daeglau, Mareike, et al.
Published: (2025)
by: Daeglau, Mareike, et al.
Published: (2025)
Generating Diverse Audio-Visual 360 Soundscapes for Sound Event Localization and Detection
by: Roman, Adrian S., et al.
Published: (2025)
by: Roman, Adrian S., et al.
Published: (2025)
From Contrast to Commonality: Audio Commonality Captioning for Enhanced Audio-Text Cross-modal Understanding in Multimodal LLMs
by: Jia, Yuhang, et al.
Published: (2025)
by: Jia, Yuhang, et al.
Published: (2025)
Attention-Based Audio Embeddings for Query-by-Example
by: Singh, Anup, et al.
Published: (2022)
by: Singh, Anup, et al.
Published: (2022)
Past, Present, and Future of Spatial Audio and Room Acoustics
by: Koyama, Shoichi, et al.
Published: (2025)
by: Koyama, Shoichi, et al.
Published: (2025)
Leveraging the Interplay Between Syntactic and Acoustic Cues for Optimizing Korean TTS Pause Formation
by: Jeon, Yejin, et al.
Published: (2024)
by: Jeon, Yejin, et al.
Published: (2024)
P2VA: Converting Persona Descriptions into Voice Attributes for Fair and Controllable Text-to-Speech
by: Lee, Yejin, et al.
Published: (2025)
by: Lee, Yejin, et al.
Published: (2025)
Structural and Statistical Audio Texture Knowledge Distillation for Acoustic Classification
by: Ritu, Jarin, et al.
Published: (2025)
by: Ritu, Jarin, et al.
Published: (2025)
CoPlay: Audio-agnostic Cognitive Scaling for Acoustic Sensing
by: Li, Yin, et al.
Published: (2024)
by: Li, Yin, et al.
Published: (2024)
Multi-Level Attention Aggregation for Language-Agnostic Speaker Replication
by: Jeon, Yejin, et al.
Published: (2024)
by: Jeon, Yejin, et al.
Published: (2024)
Multimodal Assessment of Speech Impairment in ALS Using Audio-Visual and Machine Learning Approaches
by: Pierotti, Francesco, et al.
Published: (2025)
by: Pierotti, Francesco, et al.
Published: (2025)
A Reference-free Metric for Language-Queried Audio Source Separation using Contrastive Language-Audio Pretraining
by: Xiao, Feiyang, et al.
Published: (2024)
by: Xiao, Feiyang, et al.
Published: (2024)
ALAS: Measuring Latent Speech-Text Alignment For Spoken Language Understanding In Multimodal LLMs
by: Mousavi, Pooneh, et al.
Published: (2025)
by: Mousavi, Pooneh, et al.
Published: (2025)
Leveraging Audio-Only Data for Text-Queried Target Sound Extraction
by: Saijo, Kohei, et al.
Published: (2024)
by: Saijo, Kohei, et al.
Published: (2024)
Exploring Text-Queried Sound Event Detection with Audio Source Separation
by: Yin, Han, et al.
Published: (2024)
by: Yin, Han, et al.
Published: (2024)
Evaluation of Virtual Acoustic Environments with Different Acoustic Level of Detail
by: Fichna, Stefan, et al.
Published: (2023)
by: Fichna, Stefan, et al.
Published: (2023)
Enhancing Audio Generation Diversity with Visual Information
by: Xie, Zeyu, et al.
Published: (2024)
by: Xie, Zeyu, et al.
Published: (2024)
KAD: No More FAD! An Effective and Efficient Evaluation Metric for Audio Generation
by: Chung, Yoonjin, et al.
Published: (2025)
by: Chung, Yoonjin, et al.
Published: (2025)
Leveraging Self-supervised Audio Representations for Data-Efficient Acoustic Scene Classification
by: Cai, Yiqiang, et al.
Published: (2024)
by: Cai, Yiqiang, et al.
Published: (2024)
Improving Acoustic Scene Classification in Low-Resource Conditions
by: Chen, Zhi, et al.
Published: (2024)
by: Chen, Zhi, et al.
Published: (2024)
Room Impulse Response Generation Conditioned on Acoustic Parameters
by: Arellano, Silvia, et al.
Published: (2025)
by: Arellano, Silvia, et al.
Published: (2025)
Uncovering the Visual Contribution in Audio-Visual Speech Recognition
by: Lin, Zhaofeng, et al.
Published: (2024)
by: Lin, Zhaofeng, et al.
Published: (2024)
The SMC Blind Spot: A Failure Mode Analysis of State-of-the-Art Beat Tracking
by: Ahn, Jaehoon, et al.
Published: (2026)
by: Ahn, Jaehoon, et al.
Published: (2026)
The CMU-AIST submission for the ICME 2025 Audio Encoder Challenge
by: Bharadwaj, Shikhar, et al.
Published: (2026)
by: Bharadwaj, Shikhar, et al.
Published: (2026)
MACE: Leveraging Audio for Evaluating Audio Captioning Systems
by: Dixit, Satvik, et al.
Published: (2024)
by: Dixit, Satvik, et al.
Published: (2024)
DARAS: Dynamic Audio-Room Acoustic Synthesis for Blind Room Impulse Response Estimation
by: Wang, Chunxi, et al.
Published: (2025)
by: Wang, Chunxi, et al.
Published: (2025)
Enhancing Zero-Shot Multi-Speaker TTS with Negated Speaker Representations
by: Jeon, Yejin, et al.
Published: (2024)
by: Jeon, Yejin, et al.
Published: (2024)
Audio-Conditioned Diffusion LLMs for ASR and Deliberation Processing
by: Wang, Mengqi, et al.
Published: (2025)
by: Wang, Mengqi, et al.
Published: (2025)
An Adaptive CMSA for Solving the Longest Filled Common Subsequence Problem with an Application in Audio Querying
by: Djukanovic, Marko, et al.
Published: (2025)
by: Djukanovic, Marko, et al.
Published: (2025)
OpenBEATs: A Fully Open-Source General-Purpose Audio Encoder
by: Bharadwaj, Shikhar, et al.
Published: (2025)
by: Bharadwaj, Shikhar, et al.
Published: (2025)
CONMOD: Controllable Neural Frame-based Modulation Effects
by: Lee, Gyubin, et al.
Published: (2024)
by: Lee, Gyubin, et al.
Published: (2024)
Similar Items
-
Whisper-CD: Accurate Long-Form Speech Recognition using Multi-Negative Contrastive Decoding
by: Ahn, Hoseong, et al.
Published: (2026) -
Correlation of Fréchet Audio Distance With Human Perception of Environmental Audio Is Embedding Dependant
by: Tailleur, Modan, et al.
Published: (2024) -
Audio-Based Linguistic Feature Extraction for Enhancing Multi-lingual and Low-Resource Text-to-Speech
by: Kim, Youngjae, et al.
Published: (2024) -
FlowAVSE: Efficient Audio-Visual Speech Enhancement with Conditional Flow Matching
by: Jung, Chaeyoung, et al.
Published: (2024) -
HILCodec: High-Fidelity and Lightweight Neural Audio Codec
by: Ahn, Sunghwan, et al.
Published: (2024)