EgoAVU: Egocentric Audio-Visual Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Seth, Ashish, Mei, Xinhao, Zhao, Changsheng, Nagaraja, Varun, Chang, Ernie, Meyer, Gregory P., Lan, Gael Le, Xiong, Yunyang, Chandra, Vikas, Shi, Yangyang, Manocha, Dinesh, Cai, Zhipeng |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Exploring Audio Hallucination in Egocentric Video Understanding
by: Seth, Ashish, et al.
Published: (2026)
by: Seth, Ashish, et al.
Published: (2026)
SLAP: Scalable Language-Audio Pretraining with Variable-Duration Audio and Multi-Objective Training
by: Mei, Xinhao, et al.
Published: (2026)
by: Mei, Xinhao, et al.
Published: (2026)
SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text
by: Liu, Haohe, et al.
Published: (2024)
by: Liu, Haohe, et al.
Published: (2024)
TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification
by: Anand, Nishit, et al.
Published: (2024)
by: Anand, Nishit, et al.
Published: (2024)
PAT: Parameter-Free Audio-Text Aligner to Boost Zero-Shot Audio Classification
by: Seth, Ashish, et al.
Published: (2024)
by: Seth, Ashish, et al.
Published: (2024)
RPRA: Predicting an LLM-Judge for Efficient but Performant Inference
by: Ashley, Dylan R., et al.
Published: (2026)
by: Ashley, Dylan R., et al.
Published: (2026)
Towards a Systematic Evaluation of Hallucinations in Large-Vision Language Models
by: Seth, Ashish, et al.
Published: (2024)
by: Seth, Ashish, et al.
Published: (2024)
High Fidelity Text-Guided Music Editing via Single-Stage Flow Matching
by: Lan, Gael Le, et al.
Published: (2024)
by: Lan, Gael Le, et al.
Published: (2024)
Basis Selection: Low-Rank Decomposition of Pretrained Large Language Models for Target Applications
by: Li, Yang, et al.
Published: (2024)
by: Li, Yang, et al.
Published: (2024)
EgoSocial: Benchmarking Proactive Intervention Ability of Omnimodal LLMs via Egocentric Social Interaction Perception
by: Wang, Xijun, et al.
Published: (2025)
by: Wang, Xijun, et al.
Published: (2025)
EGOILLUSION: Benchmarking Hallucinations in Egocentric Video Understanding
by: Seth, Ashish, et al.
Published: (2025)
by: Seth, Ashish, et al.
Published: (2025)
GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities
by: Ghosh, Sreyan, et al.
Published: (2024)
by: Ghosh, Sreyan, et al.
Published: (2024)
Self-Vocabularizing Training for Neural Machine Translation
by: Lin, Pin-Jie, et al.
Published: (2025)
by: Lin, Pin-Jie, et al.
Published: (2025)
MobileLLM: Optimizing Sub-billion Parameter Language Models for On-Device Use Cases
by: Liu, Zechun, et al.
Published: (2024)
by: Liu, Zechun, et al.
Published: (2024)
EgoPrompt: Prompt Learning for Egocentric Action Recognition
by: Lyu, Huaihai, et al.
Published: (2025)
by: Lyu, Huaihai, et al.
Published: (2025)
Audio Hallucination Attacks: Probing the Reliability of Large Audio Language Models
by: Seth, Ashish, et al.
Published: (2026)
by: Seth, Ashish, et al.
Published: (2026)
EgoQR: Efficient QR Code Reading in Egocentric Settings
by: Moslehpour, Mohsen, et al.
Published: (2024)
by: Moslehpour, Mohsen, et al.
Published: (2024)
EgoAdapt: Adaptive Multisensory Distillation and Policy Learning for Efficient Egocentric Perception
by: Chowdhury, Sanjoy, et al.
Published: (2025)
by: Chowdhury, Sanjoy, et al.
Published: (2025)
CompA: Addressing the Gap in Compositional Reasoning in Audio-Language Models
by: Ghosh, Sreyan, et al.
Published: (2023)
by: Ghosh, Sreyan, et al.
Published: (2023)
Do Audio-Language Models Understand Linguistic Variations?
by: Selvakumar, Ramaneswaran, et al.
Published: (2024)
by: Selvakumar, Ramaneswaran, et al.
Published: (2024)
Agent-as-a-Judge: Evaluate Agents with Agents
by: Zhuge, Mingchen, et al.
Published: (2024)
by: Zhuge, Mingchen, et al.
Published: (2024)
Neural Computers
by: Zhuge, Mingchen, et al.
Published: (2026)
by: Zhuge, Mingchen, et al.
Published: (2026)
Breaking Down Power Barriers in On-Device Streaming ASR: Insights and Solutions
by: Li, Yang, et al.
Published: (2024)
by: Li, Yang, et al.
Published: (2024)
EgoSonics: Generating Synchronized Audio for Silent Egocentric Videos
by: Rai, Aashish, et al.
Published: (2024)
by: Rai, Aashish, et al.
Published: (2024)
MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark
by: Sakshi, S, et al.
Published: (2024)
by: Sakshi, S, et al.
Published: (2024)
RECAP: Retrieval-Augmented Audio Captioning
by: Ghosh, Sreyan, et al.
Published: (2023)
by: Ghosh, Sreyan, et al.
Published: (2023)
EH-MAM: Easy-to-Hard Masked Acoustic Modeling for Self-Supervised Speech Representation Learning
by: Seth, Ashish, et al.
Published: (2024)
by: Seth, Ashish, et al.
Published: (2024)
DepthLM: Metric Depth From Vision Language Models
by: Cai, Zhipeng, et al.
Published: (2025)
by: Cai, Zhipeng, et al.
Published: (2025)
BoMuDANet: Unsupervised Adaptation for Visual Scene Understanding in Unstructured Driving Environments
by: Kothandaraman, Divya, et al.
Published: (2020)
by: Kothandaraman, Divya, et al.
Published: (2020)
SS-SFDA : Self-Supervised Source-Free Domain Adaptation for Road Segmentation in Hazardous Environments
by: Kothandaraman, Divya, et al.
Published: (2020)
by: Kothandaraman, Divya, et al.
Published: (2020)
Target-Aware Language Modeling via Granular Data Sampling
by: Chang, Ernie, et al.
Published: (2024)
by: Chang, Ernie, et al.
Published: (2024)
AutoMixer: Checkpoint Artifacts as Automatic Data Mixers
by: Chang, Ernie, et al.
Published: (2025)
by: Chang, Ernie, et al.
Published: (2025)
Scaling Parameter-Constrained Language Models with Quality Data
by: Chang, Ernie, et al.
Published: (2024)
by: Chang, Ernie, et al.
Published: (2024)
ReCLAP: Improving Zero Shot Audio Classification by Describing Sounds
by: Ghosh, Sreyan, et al.
Published: (2024)
by: Ghosh, Sreyan, et al.
Published: (2024)
VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI
by: Cheng, Sijie, et al.
Published: (2024)
by: Cheng, Sijie, et al.
Published: (2024)
EgoDTM: Towards 3D-Aware Egocentric Video-Language Pretraining
by: Xu, Boshen, et al.
Published: (2025)
by: Xu, Boshen, et al.
Published: (2025)
MultiVox: A Benchmark for Evaluating Voice Assistants for Multimodal Interactions
by: Selvakumar, Ramaneswaran, et al.
Published: (2025)
by: Selvakumar, Ramaneswaran, et al.
Published: (2025)
LipGER: Visually-Conditioned Generative Error Correction for Robust Automatic Speech Recognition
by: Ghosh, Sreyan, et al.
Published: (2024)
by: Ghosh, Sreyan, et al.
Published: (2024)
MobileLLM-R1: Exploring the Limits of Sub-Billion Language Model Reasoners with Open Training Recipes
by: Zhao, Changsheng, et al.
Published: (2025)
by: Zhao, Changsheng, et al.
Published: (2025)
EgoGen: An Egocentric Synthetic Data Generator
by: Li, Gen, et al.
Published: (2024)
by: Li, Gen, et al.
Published: (2024)
Similar Items
-
Exploring Audio Hallucination in Egocentric Video Understanding
by: Seth, Ashish, et al.
Published: (2026) -
SLAP: Scalable Language-Audio Pretraining with Variable-Duration Audio and Multi-Objective Training
by: Mei, Xinhao, et al.
Published: (2026) -
SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text
by: Liu, Haohe, et al.
Published: (2024) -
TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification
by: Anand, Nishit, et al.
Published: (2024) -
PAT: Parameter-Free Audio-Text Aligner to Boost Zero-Shot Audio Classification
by: Seth, Ashish, et al.
Published: (2024)