SEAL: Semantic Attention Learning for Long Video Representation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Lan, Chen, Yujia, Tran, Du, Boddeti, Vishnu Naresh, Chu, Wen-Sheng
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915246859354112
author Wang, Lan
Chen, Yujia
Tran, Du
Boddeti, Vishnu Naresh
Chu, Wen-Sheng
author_facet Wang, Lan
Chen, Yujia
Tran, Du
Boddeti, Vishnu Naresh
Chu, Wen-Sheng
contents Long video understanding presents challenges due to the inherent high computational complexity and redundant temporal information. An effective representation for long videos must efficiently process such redundancy while preserving essential contents for downstream tasks. This paper introduces SEmantic Attention Learning (SEAL), a novel unified representation for long videos. To reduce computational complexity, long videos are decomposed into three distinct types of semantic entities: scenes, objects, and actions, allowing models to operate on a compact set of entities rather than a large number of frames or pixels. To further address redundancy, we propose an attention learning module that balances token relevance with diversity, formulated as a subset selection optimization problem. Our representation is versatile and applicable across various long video understanding tasks. Extensive experiments demonstrate that SEAL significantly outperforms state-of-the-art methods in video question answering and temporal grounding tasks across diverse benchmarks, including LVBench, MovieChat-1K, and Ego4D.
format Preprint
id arxiv_https___arxiv_org_abs_2412_01798
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SEAL: Semantic Attention Learning for Long Video Representation
Wang, Lan
Chen, Yujia
Tran, Du
Boddeti, Vishnu Naresh
Chu, Wen-Sheng
Computer Vision and Pattern Recognition
Long video understanding presents challenges due to the inherent high computational complexity and redundant temporal information. An effective representation for long videos must efficiently process such redundancy while preserving essential contents for downstream tasks. This paper introduces SEmantic Attention Learning (SEAL), a novel unified representation for long videos. To reduce computational complexity, long videos are decomposed into three distinct types of semantic entities: scenes, objects, and actions, allowing models to operate on a compact set of entities rather than a large number of frames or pixels. To further address redundancy, we propose an attention learning module that balances token relevance with diversity, formulated as a subset selection optimization problem. Our representation is versatile and applicable across various long video understanding tasks. Extensive experiments demonstrate that SEAL significantly outperforms state-of-the-art methods in video question answering and temporal grounding tasks across diverse benchmarks, including LVBench, MovieChat-1K, and Ego4D.
title SEAL: Semantic Attention Learning for Long Video Representation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.01798