Enhancing Video Large Language Models with Structured Multi-Video Collaborative Reasoning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: He, Zhihao, He, Tianyao, Xu, Yun, Chen, Tieyuan, Liu, Huabin, Gan, Chaofan, Wu, Zuxuan, Lin, Weiyao
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909915243610112
author He, Zhihao
He, Tianyao
Xu, Yun
Chen, Tieyuan
Liu, Huabin
Gan, Chaofan
Wu, Zuxuan
Lin, Weiyao
author_facet He, Zhihao
He, Tianyao
Xu, Yun
Chen, Tieyuan
Liu, Huabin
Gan, Chaofan
Wu, Zuxuan
Lin, Weiyao
contents Despite the prosperity of the video language model, the current pursuit of comprehensive video reasoning is thwarted by the inherent spatio-temporal incompleteness within individual videos, resulting in hallucinations and inaccuracies. A promising solution is to augment the reasoning performance with multiple related videos. However, video tokens are numerous and contain redundant information, so directly feeding the relevant video data into a large language model to enhance responses could be counterproductive. To address this challenge, we propose a multi-video collaborative framework for video language models. For efficient and flexible video representation, we establish a Video Structuring Module to represent the video's knowledge as a spatio-temporal graph. Based on the structured video representation, we design the Graph Fusion Module to fuse the structured knowledge and valuable information from related videos into the augmented graph node tokens. Finally, we construct an elaborate multi-video structured prompt to integrate the graph, visual, and textual tokens as the input to the large language model. Extensive experiments substantiate the effectiveness of our framework, showcasing its potential as a promising avenue for advancing video language models. Code will be open-sourced at https://github.com/ziHoHe/SMV-CR.
format Preprint
id arxiv_https___arxiv_org_abs_2509_13161
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Enhancing Video Large Language Models with Structured Multi-Video Collaborative Reasoning
He, Zhihao
He, Tianyao
Xu, Yun
Chen, Tieyuan
Liu, Huabin
Gan, Chaofan
Wu, Zuxuan
Lin, Weiyao
Computer Vision and Pattern Recognition
Despite the prosperity of the video language model, the current pursuit of comprehensive video reasoning is thwarted by the inherent spatio-temporal incompleteness within individual videos, resulting in hallucinations and inaccuracies. A promising solution is to augment the reasoning performance with multiple related videos. However, video tokens are numerous and contain redundant information, so directly feeding the relevant video data into a large language model to enhance responses could be counterproductive. To address this challenge, we propose a multi-video collaborative framework for video language models. For efficient and flexible video representation, we establish a Video Structuring Module to represent the video's knowledge as a spatio-temporal graph. Based on the structured video representation, we design the Graph Fusion Module to fuse the structured knowledge and valuable information from related videos into the augmented graph node tokens. Finally, we construct an elaborate multi-video structured prompt to integrate the graph, visual, and textual tokens as the input to the large language model. Extensive experiments substantiate the effectiveness of our framework, showcasing its potential as a promising avenue for advancing video language models. Code will be open-sourced at https://github.com/ziHoHe/SMV-CR.
title Enhancing Video Large Language Models with Structured Multi-Video Collaborative Reasoning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.13161