CogStream: Context-guided Streaming Video Question Answering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Zicheng, Wang, Kangyu, Li, Shijie, Qian, Rui, Lin, Weiyao, Liu, Huabin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911340265734144
author Zhao, Zicheng
Wang, Kangyu
Li, Shijie
Qian, Rui
Lin, Weiyao
Liu, Huabin
author_facet Zhao, Zicheng
Wang, Kangyu
Li, Shijie
Qian, Rui
Lin, Weiyao
Liu, Huabin
contents Despite advancements in Video Large Language Models (Vid-LLMs) improving multimodal understanding, challenges persist in streaming video reasoning due to its reliance on contextual information. Existing paradigms feed all available historical contextual information into Vid-LLMs, resulting in a significant computational burden for visual data processing. Furthermore, the inclusion of irrelevant context distracts models from key details. This paper introduces a challenging task called Context-guided Streaming Video Reasoning (CogStream), which simulates real-world streaming video scenarios, requiring models to identify the most relevant historical contextual information to deduce answers for questions about the current stream. To support CogStream, we present a densely annotated dataset featuring extensive and hierarchical question-answer pairs, generated by a semi-automatic pipeline. Additionally, we present CogReasoner as a baseline model. It effectively tackles this task by leveraging visual stream compression and historical dialogue retrieval. Extensive experiments prove the effectiveness of this method.
format Preprint
id arxiv_https___arxiv_org_abs_2506_10516
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CogStream: Context-guided Streaming Video Question Answering
Zhao, Zicheng
Wang, Kangyu
Li, Shijie
Qian, Rui
Lin, Weiyao
Liu, Huabin
Computer Vision and Pattern Recognition
Artificial Intelligence
Despite advancements in Video Large Language Models (Vid-LLMs) improving multimodal understanding, challenges persist in streaming video reasoning due to its reliance on contextual information. Existing paradigms feed all available historical contextual information into Vid-LLMs, resulting in a significant computational burden for visual data processing. Furthermore, the inclusion of irrelevant context distracts models from key details. This paper introduces a challenging task called Context-guided Streaming Video Reasoning (CogStream), which simulates real-world streaming video scenarios, requiring models to identify the most relevant historical contextual information to deduce answers for questions about the current stream. To support CogStream, we present a densely annotated dataset featuring extensive and hierarchical question-answer pairs, generated by a semi-automatic pipeline. Additionally, we present CogReasoner as a baseline model. It effectively tackles this task by leveraging visual stream compression and historical dialogue retrieval. Extensive experiments prove the effectiveness of this method.
title CogStream: Context-guided Streaming Video Question Answering
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2506.10516