Saved in:
Bibliographic Details
Main Authors: Seth, Ashish, Mei, Xinhao, Zhao, Changsheng, Nagaraja, Varun, Chang, Ernie, Meyer, Gregory P., Lan, Gael Le, Xiong, Yunyang, Chandra, Vikas, Shi, Yangyang, Manocha, Dinesh, Cai, Zhipeng
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2602.06139
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908816707158016
author Seth, Ashish
Mei, Xinhao
Zhao, Changsheng
Nagaraja, Varun
Chang, Ernie
Meyer, Gregory P.
Lan, Gael Le
Xiong, Yunyang
Chandra, Vikas
Shi, Yangyang
Manocha, Dinesh
Cai, Zhipeng
author_facet Seth, Ashish
Mei, Xinhao
Zhao, Changsheng
Nagaraja, Varun
Chang, Ernie
Meyer, Gregory P.
Lan, Gael Le
Xiong, Yunyang
Chandra, Vikas
Shi, Yangyang
Manocha, Dinesh
Cai, Zhipeng
contents Understanding egocentric videos plays a vital role for embodied intelligence. Recent multi-modal large language models (MLLMs) can accept both visual and audio inputs. However, due to the challenge of obtaining text labels with coherent joint-modality information, whether MLLMs can jointly understand both modalities in egocentric videos remains under-explored. To address this problem, we introduce EgoAVU, a scalable data engine to automatically generate egocentric audio-visual narrations, questions, and answers. EgoAVU enriches human narrations with multimodal context and generates audio-visual narrations through cross-modal correlation modeling. Token-based video filtering and modular, graph-based curation ensure both data diversity and quality. Leveraging EgoAVU, we construct EgoAVU-Instruct, a large-scale training dataset of 3M samples, and EgoAVU-Bench, a manually verified evaluation split covering diverse tasks. EgoAVU-Bench clearly reveals the limitations of existing MLLMs: they bias heavily toward visual signals, often neglecting audio cues or failing to correspond audio with the visual source. Finetuning MLLMs on EgoAVU-Instruct effectively addresses this issue, enabling up to 113% performance improvement on EgoAVU-Bench. Such benefits also transfer to other benchmarks such as EgoTempo and EgoIllusion, achieving up to 28% relative performance gain. Code will be released to the community.
format Preprint
id arxiv_https___arxiv_org_abs_2602_06139
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle EgoAVU: Egocentric Audio-Visual Understanding
Seth, Ashish
Mei, Xinhao
Zhao, Changsheng
Nagaraja, Varun
Chang, Ernie
Meyer, Gregory P.
Lan, Gael Le
Xiong, Yunyang
Chandra, Vikas
Shi, Yangyang
Manocha, Dinesh
Cai, Zhipeng
Computer Vision and Pattern Recognition
Understanding egocentric videos plays a vital role for embodied intelligence. Recent multi-modal large language models (MLLMs) can accept both visual and audio inputs. However, due to the challenge of obtaining text labels with coherent joint-modality information, whether MLLMs can jointly understand both modalities in egocentric videos remains under-explored. To address this problem, we introduce EgoAVU, a scalable data engine to automatically generate egocentric audio-visual narrations, questions, and answers. EgoAVU enriches human narrations with multimodal context and generates audio-visual narrations through cross-modal correlation modeling. Token-based video filtering and modular, graph-based curation ensure both data diversity and quality. Leveraging EgoAVU, we construct EgoAVU-Instruct, a large-scale training dataset of 3M samples, and EgoAVU-Bench, a manually verified evaluation split covering diverse tasks. EgoAVU-Bench clearly reveals the limitations of existing MLLMs: they bias heavily toward visual signals, often neglecting audio cues or failing to correspond audio with the visual source. Finetuning MLLMs on EgoAVU-Instruct effectively addresses this issue, enabling up to 113% performance improvement on EgoAVU-Bench. Such benefits also transfer to other benchmarks such as EgoTempo and EgoIllusion, achieving up to 28% relative performance gain. Code will be released to the community.
title EgoAVU: Egocentric Audio-Visual Understanding
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.06139