Do Audio-Visual Large Language Models Really See and Hear?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Selvakumar, Ramaneswaran, Jayakumar, Kaousheik, Sakshi, S, Ghosh, Sreyan, Gao, Ruohan, Manocha, Dinesh |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
EH-MAM: Easy-to-Hard Masked Acoustic Modeling for Self-Supervised Speech Representation Learning
von: Seth, Ashish, et al.
Veröffentlicht: (2024)
von: Seth, Ashish, et al.
Veröffentlicht: (2024)
MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark
von: Sakshi, S, et al.
Veröffentlicht: (2024)
von: Sakshi, S, et al.
Veröffentlicht: (2024)
PAT: Parameter-Free Audio-Text Aligner to Boost Zero-Shot Audio Classification
von: Seth, Ashish, et al.
Veröffentlicht: (2024)
von: Seth, Ashish, et al.
Veröffentlicht: (2024)
Do Audio-Language Models Understand Linguistic Variations?
von: Selvakumar, Ramaneswaran, et al.
Veröffentlicht: (2024)
von: Selvakumar, Ramaneswaran, et al.
Veröffentlicht: (2024)
CompA: Addressing the Gap in Compositional Reasoning in Audio-Language Models
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2023)
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2023)
GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2024)
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2024)
Audio Hallucination Attacks: Probing the Reliability of Large Audio Language Models
von: Seth, Ashish, et al.
Veröffentlicht: (2026)
von: Seth, Ashish, et al.
Veröffentlicht: (2026)
RECAP: Retrieval-Augmented Audio Captioning
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2023)
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2023)
Audio Flamingo Next: Next-Generation Open Audio-Language Models for Speech, Sound, and Music
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2026)
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2026)
Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models
von: Goel, Arushi, et al.
Veröffentlicht: (2025)
von: Goel, Arushi, et al.
Veröffentlicht: (2025)
SPUR: A Plug-and-Play Framework for Integrating Spatial Audio Understanding and Reasoning into Large Audio-Language Models
von: Sakshi, S, et al.
Veröffentlicht: (2025)
von: Sakshi, S, et al.
Veröffentlicht: (2025)
ABEX: Data Augmentation for Low-Resource NLU via Expanding Abstract Descriptions
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2024)
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2024)
AV-RIR: Audio-Visual Room Impulse Response Estimation
von: Ratnarajah, Anton, et al.
Veröffentlicht: (2023)
von: Ratnarajah, Anton, et al.
Veröffentlicht: (2023)
Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2025)
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2025)
PitchBench: Measuring Pitch Hearing in Audio-Language Models
von: Dujardin, Milan Liessens, et al.
Veröffentlicht: (2026)
von: Dujardin, Milan Liessens, et al.
Veröffentlicht: (2026)
Multi-Domain Audio Question Answering Benchmark Toward Acoustic Content Reasoning
von: Yang, Chao-Han Huck, et al.
Veröffentlicht: (2025)
von: Yang, Chao-Han Huck, et al.
Veröffentlicht: (2025)
EGOILLUSION: Benchmarking Hallucinations in Egocentric Video Understanding
von: Seth, Ashish, et al.
Veröffentlicht: (2025)
von: Seth, Ashish, et al.
Veröffentlicht: (2025)
Audio-Maestro: Enhancing Large Audio-Language Models with Tool-Augmented Reasoning
von: Lee, Kuan-Yi, et al.
Veröffentlicht: (2025)
von: Lee, Kuan-Yi, et al.
Veröffentlicht: (2025)
Protecting Bystander Privacy via Selective Hearing in Audio LLMs
von: Zhan, Xiao, et al.
Veröffentlicht: (2025)
von: Zhan, Xiao, et al.
Veröffentlicht: (2025)
Do Models Hear Like Us? Probing the Representational Alignment of Audio LLMs and Naturalistic EEG
von: Yang, Haoyun, et al.
Veröffentlicht: (2026)
von: Yang, Haoyun, et al.
Veröffentlicht: (2026)
HalluAudio: A Comprehensive Benchmark for Hallucination Detection in Large Audio-Language Models
von: Zhao, Feiyu, et al.
Veröffentlicht: (2026)
von: Zhao, Feiyu, et al.
Veröffentlicht: (2026)
TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification
von: Anand, Nishit, et al.
Veröffentlicht: (2024)
von: Anand, Nishit, et al.
Veröffentlicht: (2024)
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing
von: Chen, Mingfei, et al.
Veröffentlicht: (2025)
von: Chen, Mingfei, et al.
Veröffentlicht: (2025)
Video-Robin: Autoregressive Diffusion Planning for Intent-Grounded Video-to-Music Generation
von: Lokegaonkar, Vaibhavi, et al.
Veröffentlicht: (2026)
von: Lokegaonkar, Vaibhavi, et al.
Veröffentlicht: (2026)
The World is Not Mono: Enabling Spatial Understanding in Large Audio-Language Models
von: You, Yuhuan, et al.
Veröffentlicht: (2026)
von: You, Yuhuan, et al.
Veröffentlicht: (2026)
OWL: Geometry-Aware Spatial Reasoning for Audio Large Language Models
von: Biswas, Subrata, et al.
Veröffentlicht: (2025)
von: Biswas, Subrata, et al.
Veröffentlicht: (2025)
Disentangling Reasoning in Large Audio-Language Models for Ambiguous Emotion Prediction
von: Yu, Xiaofeng, et al.
Veröffentlicht: (2026)
von: Yu, Xiaofeng, et al.
Veröffentlicht: (2026)
Are Audio-Language Models Listening? Audio-Specialist Heads for Adaptive Audio Steering
von: Glazer, Neta, et al.
Veröffentlicht: (2026)
von: Glazer, Neta, et al.
Veröffentlicht: (2026)
Towards Fine-grained Temporal Perception: Post-Training Large Audio-Language Models with Audio-Side Time Prompt
von: Shi, Yanfeng, et al.
Veröffentlicht: (2026)
von: Shi, Yanfeng, et al.
Veröffentlicht: (2026)
CALM: Class-Conditional Sparse Attention Vectors for Large Audio-Language Models
von: Mehta, Videet, et al.
Veröffentlicht: (2026)
von: Mehta, Videet, et al.
Veröffentlicht: (2026)
Eureka-Audio: Triggering Audio Intelligence in Compact Language Models
von: Zhang, Dan, et al.
Veröffentlicht: (2026)
von: Zhang, Dan, et al.
Veröffentlicht: (2026)
Building Robust and Scalable Multilingual ASR for Indian Languages
von: Gangwar, Arjun, et al.
Veröffentlicht: (2025)
von: Gangwar, Arjun, et al.
Veröffentlicht: (2025)
Temporal Contrastive Decoding: A Training-Free Method for Large Audio-Language Models
von: Li, Yanda, et al.
Veröffentlicht: (2026)
von: Li, Yanda, et al.
Veröffentlicht: (2026)
The Sonar Moment: Benchmarking Audio-Language Models in Audio Geo-Localization
von: Zhang, Ruixing, et al.
Veröffentlicht: (2026)
von: Zhang, Ruixing, et al.
Veröffentlicht: (2026)
Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models
von: Song, Zirui, et al.
Veröffentlicht: (2025)
von: Song, Zirui, et al.
Veröffentlicht: (2025)
Meerkat: Audio-Visual Large Language Model for Grounding in Space and Time
von: Chowdhury, Sanjoy, et al.
Veröffentlicht: (2024)
von: Chowdhury, Sanjoy, et al.
Veröffentlicht: (2024)
VocalParse: Towards Unified and Scalable Singing Voice Transcription with Large Audio Language Models
von: Chen, Yukun, et al.
Veröffentlicht: (2026)
von: Chen, Yukun, et al.
Veröffentlicht: (2026)
ReCLAP: Improving Zero Shot Audio Classification by Describing Sounds
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2024)
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2024)
Bridging Biological Hearing and Neuromorphic Computing: End-to-End Time-Domain Audio Signal Processing with Reservoir Computing
von: Sebastian, Rinku, et al.
Veröffentlicht: (2026)
von: Sebastian, Rinku, et al.
Veröffentlicht: (2026)
EMO-TTA: Improving Test-Time Adaptation of Audio-Language Models for Speech Emotion Recognition
von: Shi, Jiacheng, et al.
Veröffentlicht: (2025)
von: Shi, Jiacheng, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
EH-MAM: Easy-to-Hard Masked Acoustic Modeling for Self-Supervised Speech Representation Learning
von: Seth, Ashish, et al.
Veröffentlicht: (2024) -
MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark
von: Sakshi, S, et al.
Veröffentlicht: (2024) -
PAT: Parameter-Free Audio-Text Aligner to Boost Zero-Shot Audio Classification
von: Seth, Ashish, et al.
Veröffentlicht: (2024) -
Do Audio-Language Models Understand Linguistic Variations?
von: Selvakumar, Ramaneswaran, et al.
Veröffentlicht: (2024) -
CompA: Addressing the Gap in Compositional Reasoning in Audio-Language Models
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2023)