Saved in:
Bibliographic Details
Main Authors: Yamao, Sosuke, Miyahara, Natsuki, Qi, Yuankai, Takeuchi, Shun
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2603.15167
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908889269665792
author Yamao, Sosuke
Miyahara, Natsuki
Qi, Yuankai
Takeuchi, Shun
author_facet Yamao, Sosuke
Miyahara, Natsuki
Qi, Yuankai
Takeuchi, Shun
contents In the context of long-term video understanding with large multimodal models, many frameworks have been proposed. Although transformer-based visual compressors and memory-augmented approaches are often used to process long videos, they usually compress each frame independently and therefore fail to achieve strong performance on tasks that require understanding complete events, such as temporal ordering tasks in MLVU and VNBench. This motivates us to rethink the conventional one-way scheme from perception to memory, and instead establish a feedbackdriven process in which past visual contexts stored in the context memory can benefit ongoing perception. To this end, we propose Question-guided Visual Compression with Memory Feedback (QViC-MF), a framework for long-term video understanding. At its core is a Question-guided Multimodal Selective Attention (QMSA), which learns to preserve visual information related to the given question from both the current clip and the past related frames from the memory. The compressor and memory feedback work iteratively for each clip of the entire video. This simple yet effective design yields large performance gains on longterm video understanding tasks. Extensive experiments show that our method achieves significant improvement over current state-of-the-art methods by 6.1% on MLVU test, 8.3% on LVBench, 18.3% on VNBench Long, and 3.7% on VideoMME Long. The code will be released publicly.
format Preprint
id arxiv_https___arxiv_org_abs_2603_15167
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Question-guided Visual Compression with Memory Feedback for Long-Term Video Understanding
Yamao, Sosuke
Miyahara, Natsuki
Qi, Yuankai
Takeuchi, Shun
Computer Vision and Pattern Recognition
In the context of long-term video understanding with large multimodal models, many frameworks have been proposed. Although transformer-based visual compressors and memory-augmented approaches are often used to process long videos, they usually compress each frame independently and therefore fail to achieve strong performance on tasks that require understanding complete events, such as temporal ordering tasks in MLVU and VNBench. This motivates us to rethink the conventional one-way scheme from perception to memory, and instead establish a feedbackdriven process in which past visual contexts stored in the context memory can benefit ongoing perception. To this end, we propose Question-guided Visual Compression with Memory Feedback (QViC-MF), a framework for long-term video understanding. At its core is a Question-guided Multimodal Selective Attention (QMSA), which learns to preserve visual information related to the given question from both the current clip and the past related frames from the memory. The compressor and memory feedback work iteratively for each clip of the entire video. This simple yet effective design yields large performance gains on longterm video understanding tasks. Extensive experiments show that our method achieves significant improvement over current state-of-the-art methods by 6.1% on MLVU test, 8.3% on LVBench, 18.3% on VNBench Long, and 3.7% on VideoMME Long. The code will be released publicly.
title Question-guided Visual Compression with Memory Feedback for Long-Term Video Understanding
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.15167