Attend Before Attention: Efficient and Scalable Video Understanding via Autoregressive Gazing

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Shi, Baifeng, Fu, Stephanie, Lian, Long, Ye, Hanrong, Eigen, David, Reite, Aaron, Li, Boyi, Kautz, Jan, Han, Song, Chan, David M., Molchanov, Pavlo, Darrell, Trevor, Yin, Hongxu
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912963548413952
author Shi, Baifeng
Fu, Stephanie
Lian, Long
Ye, Hanrong
Eigen, David
Reite, Aaron
Li, Boyi
Kautz, Jan
Han, Song
Chan, David M.
Molchanov, Pavlo
Darrell, Trevor
Yin, Hongxu
author_facet Shi, Baifeng
Fu, Stephanie
Lian, Long
Ye, Hanrong
Eigen, David
Reite, Aaron
Li, Boyi
Kautz, Jan
Han, Song
Chan, David M.
Molchanov, Pavlo
Darrell, Trevor
Yin, Hongxu
contents Multi-modal large language models (MLLMs) have advanced general-purpose video understanding but struggle with long, high-resolution videos -- they process every pixel equally in their vision transformers (ViTs) or LLMs despite significant spatiotemporal redundancy. We introduce AutoGaze, a lightweight module that removes redundant patches before processed by a ViT or an MLLM. Trained with next-token prediction and reinforcement learning, AutoGaze autoregressively selects a minimal set of multi-scale patches that can reconstruct the video within a user-specified error threshold, eliminating redundancy while preserving information. Empirically, AutoGaze reduces visual tokens by 4x-100x and accelerates ViTs and MLLMs by up to 19x, enabling scaling MLLMs to 1K-frame 4K-resolution videos and achieving superior results on video benchmarks (e.g., 67.0% on VideoMME). Furthermore, we introduce HLVid: the first high-resolution, long-form video QA benchmark with 5-minute 4K-resolution videos, where an MLLM scaled with AutoGaze improves over the baseline by 10.1% and outperforms the previous best MLLM by 4.5%. Project page: https://autogaze.github.io/.
format Preprint
id arxiv_https___arxiv_org_abs_2603_12254
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Attend Before Attention: Efficient and Scalable Video Understanding via Autoregressive Gazing
Shi, Baifeng
Fu, Stephanie
Lian, Long
Ye, Hanrong
Eigen, David
Reite, Aaron
Li, Boyi
Kautz, Jan
Han, Song
Chan, David M.
Molchanov, Pavlo
Darrell, Trevor
Yin, Hongxu
Computer Vision and Pattern Recognition
Multi-modal large language models (MLLMs) have advanced general-purpose video understanding but struggle with long, high-resolution videos -- they process every pixel equally in their vision transformers (ViTs) or LLMs despite significant spatiotemporal redundancy. We introduce AutoGaze, a lightweight module that removes redundant patches before processed by a ViT or an MLLM. Trained with next-token prediction and reinforcement learning, AutoGaze autoregressively selects a minimal set of multi-scale patches that can reconstruct the video within a user-specified error threshold, eliminating redundancy while preserving information. Empirically, AutoGaze reduces visual tokens by 4x-100x and accelerates ViTs and MLLMs by up to 19x, enabling scaling MLLMs to 1K-frame 4K-resolution videos and achieving superior results on video benchmarks (e.g., 67.0% on VideoMME). Furthermore, we introduce HLVid: the first high-resolution, long-form video QA benchmark with 5-minute 4K-resolution videos, where an MLLM scaled with AutoGaze improves over the baseline by 10.1% and outperforms the previous best MLLM by 4.5%. Project page: https://autogaze.github.io/.
title Attend Before Attention: Efficient and Scalable Video Understanding via Autoregressive Gazing
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.12254