EAGLE: Egocentric AGgregated Language-video Engine

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bi, Jing, Tang, Yunlong, Song, Luchuan, Vosoughi, Ali, Nguyen, Nguyen, Xu, Chenliang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929516124831744
author Bi, Jing
Tang, Yunlong
Song, Luchuan
Vosoughi, Ali
Nguyen, Nguyen
Xu, Chenliang
author_facet Bi, Jing
Tang, Yunlong
Song, Luchuan
Vosoughi, Ali
Nguyen, Nguyen
Xu, Chenliang
contents The rapid evolution of egocentric video analysis brings new insights into understanding human activities and intentions from a first-person perspective. Despite this progress, the fragmentation in tasks like action recognition, procedure learning, and moment retrieval, \etc, coupled with inconsistent annotations and isolated model development, hinders a holistic interpretation of video content. In response, we introduce the EAGLE (Egocentric AGgregated Language-video Engine) model and the EAGLE-400K dataset to provide a unified framework that integrates various egocentric video understanding tasks. EAGLE-400K, the \textit{first} large-scale instruction-tuning dataset tailored for egocentric video, features 400K diverse samples to enhance a broad spectrum of tasks from activity recognition to procedure knowledge learning. Moreover, EAGLE, a strong video multimodal large language model (MLLM), is designed to effectively capture both spatial and temporal information. In addition, we propose a set of evaluation metrics designed to facilitate a thorough assessment of MLLM for egocentric video understanding. Our extensive experiments demonstrate EAGLE's superior performance over existing models, highlighting its ability to balance task-specific understanding with holistic video interpretation. With EAGLE, we aim to pave the way for research opportunities and practical applications in real-world scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2409_17523
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle EAGLE: Egocentric AGgregated Language-video Engine
Bi, Jing
Tang, Yunlong
Song, Luchuan
Vosoughi, Ali
Nguyen, Nguyen
Xu, Chenliang
Computer Vision and Pattern Recognition
Artificial Intelligence
The rapid evolution of egocentric video analysis brings new insights into understanding human activities and intentions from a first-person perspective. Despite this progress, the fragmentation in tasks like action recognition, procedure learning, and moment retrieval, \etc, coupled with inconsistent annotations and isolated model development, hinders a holistic interpretation of video content. In response, we introduce the EAGLE (Egocentric AGgregated Language-video Engine) model and the EAGLE-400K dataset to provide a unified framework that integrates various egocentric video understanding tasks. EAGLE-400K, the \textit{first} large-scale instruction-tuning dataset tailored for egocentric video, features 400K diverse samples to enhance a broad spectrum of tasks from activity recognition to procedure knowledge learning. Moreover, EAGLE, a strong video multimodal large language model (MLLM), is designed to effectively capture both spatial and temporal information. In addition, we propose a set of evaluation metrics designed to facilitate a thorough assessment of MLLM for egocentric video understanding. Our extensive experiments demonstrate EAGLE's superior performance over existing models, highlighting its ability to balance task-specific understanding with holistic video interpretation. With EAGLE, we aim to pave the way for research opportunities and practical applications in real-world scenarios.
title EAGLE: Egocentric AGgregated Language-video Engine
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2409.17523