Exploring Audio Hallucination in Egocentric Video Understanding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Seth, Ashish, Mei, Xinhao, Zhao, Changsheng, Nagaraja, Varun, Chang, Ernie, Meyer, Gregory P., Lan, Gael Le, Xiong, Yunyang, Chandra, Vikas, Shi, Yangyang, Manocha, Dinesh, Cai, Zhipeng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
EgoAVU: Egocentric Audio-Visual Understanding
von: Seth, Ashish, et al.
Veröffentlicht: (2026)
von: Seth, Ashish, et al.
Veröffentlicht: (2026)
SLAP: Scalable Language-Audio Pretraining with Variable-Duration Audio and Multi-Objective Training
von: Mei, Xinhao, et al.
Veröffentlicht: (2026)
von: Mei, Xinhao, et al.
Veröffentlicht: (2026)
EGOILLUSION: Benchmarking Hallucinations in Egocentric Video Understanding
von: Seth, Ashish, et al.
Veröffentlicht: (2025)
von: Seth, Ashish, et al.
Veröffentlicht: (2025)
SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text
von: Liu, Haohe, et al.
Veröffentlicht: (2024)
von: Liu, Haohe, et al.
Veröffentlicht: (2024)
Towards a Systematic Evaluation of Hallucinations in Large-Vision Language Models
von: Seth, Ashish, et al.
Veröffentlicht: (2024)
von: Seth, Ashish, et al.
Veröffentlicht: (2024)
Audio Hallucination Attacks: Probing the Reliability of Large Audio Language Models
von: Seth, Ashish, et al.
Veröffentlicht: (2026)
von: Seth, Ashish, et al.
Veröffentlicht: (2026)
TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification
von: Anand, Nishit, et al.
Veröffentlicht: (2024)
von: Anand, Nishit, et al.
Veröffentlicht: (2024)
PAT: Parameter-Free Audio-Text Aligner to Boost Zero-Shot Audio Classification
von: Seth, Ashish, et al.
Veröffentlicht: (2024)
von: Seth, Ashish, et al.
Veröffentlicht: (2024)
GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2024)
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2024)
RPRA: Predicting an LLM-Judge for Efficient but Performant Inference
von: Ashley, Dylan R., et al.
Veröffentlicht: (2026)
von: Ashley, Dylan R., et al.
Veröffentlicht: (2026)
Do Audio-Language Models Understand Linguistic Variations?
von: Selvakumar, Ramaneswaran, et al.
Veröffentlicht: (2024)
von: Selvakumar, Ramaneswaran, et al.
Veröffentlicht: (2024)
High Fidelity Text-Guided Music Editing via Single-Stage Flow Matching
von: Lan, Gael Le, et al.
Veröffentlicht: (2024)
von: Lan, Gael Le, et al.
Veröffentlicht: (2024)
Basis Selection: Low-Rank Decomposition of Pretrained Large Language Models for Target Applications
von: Li, Yang, et al.
Veröffentlicht: (2024)
von: Li, Yang, et al.
Veröffentlicht: (2024)
MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark
von: Sakshi, S, et al.
Veröffentlicht: (2024)
von: Sakshi, S, et al.
Veröffentlicht: (2024)
Self-Vocabularizing Training for Neural Machine Translation
von: Lin, Pin-Jie, et al.
Veröffentlicht: (2025)
von: Lin, Pin-Jie, et al.
Veröffentlicht: (2025)
MobileLLM: Optimizing Sub-billion Parameter Language Models for On-Device Use Cases
von: Liu, Zechun, et al.
Veröffentlicht: (2024)
von: Liu, Zechun, et al.
Veröffentlicht: (2024)
BoMuDANet: Unsupervised Adaptation for Visual Scene Understanding in Unstructured Driving Environments
von: Kothandaraman, Divya, et al.
Veröffentlicht: (2020)
von: Kothandaraman, Divya, et al.
Veröffentlicht: (2020)
CompA: Addressing the Gap in Compositional Reasoning in Audio-Language Models
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2023)
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2023)
Agent-as-a-Judge: Evaluate Agents with Agents
von: Zhuge, Mingchen, et al.
Veröffentlicht: (2024)
von: Zhuge, Mingchen, et al.
Veröffentlicht: (2024)
Neural Computers
von: Zhuge, Mingchen, et al.
Veröffentlicht: (2026)
von: Zhuge, Mingchen, et al.
Veröffentlicht: (2026)
Breaking Down Power Barriers in On-Device Streaming ASR: Insights and Solutions
von: Li, Yang, et al.
Veröffentlicht: (2024)
von: Li, Yang, et al.
Veröffentlicht: (2024)
RECAP: Retrieval-Augmented Audio Captioning
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2023)
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2023)
EH-MAM: Easy-to-Hard Masked Acoustic Modeling for Self-Supervised Speech Representation Learning
von: Seth, Ashish, et al.
Veröffentlicht: (2024)
von: Seth, Ashish, et al.
Veröffentlicht: (2024)
MobileLLM-R1: Exploring the Limits of Sub-Billion Language Model Reasoners with Open Training Recipes
von: Zhao, Changsheng, et al.
Veröffentlicht: (2025)
von: Zhao, Changsheng, et al.
Veröffentlicht: (2025)
DepthLM: Metric Depth From Vision Language Models
von: Cai, Zhipeng, et al.
Veröffentlicht: (2025)
von: Cai, Zhipeng, et al.
Veröffentlicht: (2025)
SS-SFDA : Self-Supervised Source-Free Domain Adaptation for Road Segmentation in Hazardous Environments
von: Kothandaraman, Divya, et al.
Veröffentlicht: (2020)
von: Kothandaraman, Divya, et al.
Veröffentlicht: (2020)
Target-Aware Language Modeling via Granular Data Sampling
von: Chang, Ernie, et al.
Veröffentlicht: (2024)
von: Chang, Ernie, et al.
Veröffentlicht: (2024)
SPUR: A Plug-and-Play Framework for Integrating Spatial Audio Understanding and Reasoning into Large Audio-Language Models
von: Sakshi, S, et al.
Veröffentlicht: (2025)
von: Sakshi, S, et al.
Veröffentlicht: (2025)
Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2025)
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2025)
AutoMixer: Checkpoint Artifacts as Automatic Data Mixers
von: Chang, Ernie, et al.
Veröffentlicht: (2025)
von: Chang, Ernie, et al.
Veröffentlicht: (2025)
Scaling Parameter-Constrained Language Models with Quality Data
von: Chang, Ernie, et al.
Veröffentlicht: (2024)
von: Chang, Ernie, et al.
Veröffentlicht: (2024)
Visual Description Grounding Reduces Hallucinations and Boosts Reasoning in LVLMs
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2024)
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2024)
ReCLAP: Improving Zero Shot Audio Classification by Describing Sounds
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2024)
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2024)
MultiVox: A Benchmark for Evaluating Voice Assistants for Multimodal Interactions
von: Selvakumar, Ramaneswaran, et al.
Veröffentlicht: (2025)
von: Selvakumar, Ramaneswaran, et al.
Veröffentlicht: (2025)
LipGER: Visually-Conditioned Generative Error Correction for Robust Automatic Speech Recognition
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2024)
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2024)
EgoSocial: Benchmarking Proactive Intervention Ability of Omnimodal LLMs via Egocentric Social Interaction Perception
von: Wang, Xijun, et al.
Veröffentlicht: (2025)
von: Wang, Xijun, et al.
Veröffentlicht: (2025)
AV-RIR: Audio-Visual Room Impulse Response Estimation
von: Ratnarajah, Anton, et al.
Veröffentlicht: (2023)
von: Ratnarajah, Anton, et al.
Veröffentlicht: (2023)
VideoHallu: Evaluating and Mitigating Multi-modal Hallucinations on Synthetic Video Understanding
von: Li, Zongxia, et al.
Veröffentlicht: (2025)
von: Li, Zongxia, et al.
Veröffentlicht: (2025)
Uncovering the Representation Geometry of Minimal Cores in Overcomplete Reasoning Traces
von: Chowdhury, Sanjoy, et al.
Veröffentlicht: (2026)
von: Chowdhury, Sanjoy, et al.
Veröffentlicht: (2026)
PACE: Data-Driven Virtual Agent Interaction in Dense and Cluttered Environments
von: Mullen, James, et al.
Veröffentlicht: (2023)
von: Mullen, James, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
EgoAVU: Egocentric Audio-Visual Understanding
von: Seth, Ashish, et al.
Veröffentlicht: (2026) -
SLAP: Scalable Language-Audio Pretraining with Variable-Duration Audio and Multi-Objective Training
von: Mei, Xinhao, et al.
Veröffentlicht: (2026) -
EGOILLUSION: Benchmarking Hallucinations in Egocentric Video Understanding
von: Seth, Ashish, et al.
Veröffentlicht: (2025) -
SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text
von: Liu, Haohe, et al.
Veröffentlicht: (2024) -
Towards a Systematic Evaluation of Hallucinations in Large-Vision Language Models
von: Seth, Ashish, et al.
Veröffentlicht: (2024)