VideoPro: Adaptive Program Reasoning for Long Video Understanding

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Li, Chenglin, Han, Feng, Wang, Yikun, Li, Ruilin, Dong, Shuai, Hou, Haowen, Li, Haitao, Chen, Qianglong, Tao, Feng, Tong, Jingqi, Zhang, Yin, Wang, Jiaqi
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914279255441408
author Li, Chenglin
Han, Feng
Wang, Yikun
Li, Ruilin
Dong, Shuai
Hou, Haowen
Li, Haitao
Chen, Qianglong
Tao, Feng
Tong, Jingqi
Zhang, Yin
Wang, Jiaqi
author_facet Li, Chenglin
Han, Feng
Wang, Yikun
Li, Ruilin
Dong, Shuai
Hou, Haowen
Li, Haitao
Chen, Qianglong
Tao, Feng
Tong, Jingqi
Zhang, Yin
Wang, Jiaqi
contents Large language models (LLMs) have shown promise in generating program workflows for visual tasks. However, previous approaches often rely on closed-source models, lack systematic reasoning, and struggle with long-form video question answering (videoQA). To address these challenges, we introduce the FS-VisPR framework, an adaptive visual program reasoning approach that balances fast reasoning for simple queries with slow reasoning for difficult ones. First, we design efficient visual modules (e.g., key clip retrieval and subtitle retrieval) to support long-form video tasks. Then, we construct a diverse and high-quality fast-slow reasoning dataset with a strong LLM to align open-source language models' ability to generate visual program workflows as FS-LLM. Next, we design a fast-slow reasoning framework with FS-LLM: Simple queries are directly solved by VideoLLMs, while difficult ones invoke visual program reasoning, motivated by human-like reasoning processes. During this process, low-confidence fast-thinking answers will trigger a second-stage slow-reasoning process, and a fallback mechanism to fast reasoning is activated if the program execution fails. Moreover, we improve visual programs through parameter search during both training and inference. By adjusting the parameters of the visual modules within the program, multiple variants are generated: during training, programs that yield correct answers are selected, while during inference, the program with the highest confidence result is applied. Experiments show that FS-VisPR improves both efficiency and reliability in visual program workflows. It achieves 50.4% accuracy on LVBench, surpassing GPT-4o, matching the performance of Qwen2.5VL-72B on VideoMME.
format Preprint
id arxiv_https___arxiv_org_abs_2509_17743
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VideoPro: Adaptive Program Reasoning for Long Video Understanding
Li, Chenglin
Han, Feng
Wang, Yikun
Li, Ruilin
Dong, Shuai
Hou, Haowen
Li, Haitao
Chen, Qianglong
Tao, Feng
Tong, Jingqi
Zhang, Yin
Wang, Jiaqi
Computer Vision and Pattern Recognition
Large language models (LLMs) have shown promise in generating program workflows for visual tasks. However, previous approaches often rely on closed-source models, lack systematic reasoning, and struggle with long-form video question answering (videoQA). To address these challenges, we introduce the FS-VisPR framework, an adaptive visual program reasoning approach that balances fast reasoning for simple queries with slow reasoning for difficult ones. First, we design efficient visual modules (e.g., key clip retrieval and subtitle retrieval) to support long-form video tasks. Then, we construct a diverse and high-quality fast-slow reasoning dataset with a strong LLM to align open-source language models' ability to generate visual program workflows as FS-LLM. Next, we design a fast-slow reasoning framework with FS-LLM: Simple queries are directly solved by VideoLLMs, while difficult ones invoke visual program reasoning, motivated by human-like reasoning processes. During this process, low-confidence fast-thinking answers will trigger a second-stage slow-reasoning process, and a fallback mechanism to fast reasoning is activated if the program execution fails. Moreover, we improve visual programs through parameter search during both training and inference. By adjusting the parameters of the visual modules within the program, multiple variants are generated: during training, programs that yield correct answers are selected, while during inference, the program with the highest confidence result is applied. Experiments show that FS-VisPR improves both efficiency and reliability in visual program workflows. It achieves 50.4% accuracy on LVBench, surpassing GPT-4o, matching the performance of Qwen2.5VL-72B on VideoMME.
title VideoPro: Adaptive Program Reasoning for Long Video Understanding
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.17743