Salvato in:
Dettagli Bibliografici
Autori principali: Jiang, Yifan, Wang, Yueying, Zhao, Rui, Parag, Toufiq, Chen, Zhimin, Liao, Zhenyu, Unnikrishnan, Jayakrishnan
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:https://arxiv.org/abs/2511.11113
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915944076410880
author Jiang, Yifan
Wang, Yueying
Zhao, Rui
Parag, Toufiq
Chen, Zhimin
Liao, Zhenyu
Unnikrishnan, Jayakrishnan
author_facet Jiang, Yifan
Wang, Yueying
Zhao, Rui
Parag, Toufiq
Chen, Zhimin
Liao, Zhenyu
Unnikrishnan, Jayakrishnan
contents Reinforcement fine-tuning (RFT), a two-stage framework consisting of supervised fine-tuning (SFT) and reinforcement learning (RL) has shown promising results on improving reasoning ability of large language models (LLMs). Yet extending RFT to large video language models (LVLMs) remains challenging. We propose VideoP2R, a novel process-aware video RFT framework that enhances video reasoning by modeling perception and reasoning as distinct processes. In the SFT stage, we develop a three-step pipeline to generate VideoP2R-CoT-162K, a high-quality, process-aware chain-of-thought (CoT) dataset for perception and reasoning. In the RL stage, we introduce a novel process-aware group relative policy optimization (PA-GRPO) algorithm that supplies separate rewards for perception and reasoning. Extensive experiments show that VideoP2R achieves state-of-the-art (SotA) performance on six out of seven video reasoning and understanding benchmarks. Ablation studies further confirm the effectiveness of our process-aware modeling and PA-GRPO and demonstrate that model's perception output is information-sufficient for downstream reasoning. Our project page is available at https://videop2r.github.io/videop2r/.
format Preprint
id arxiv_https___arxiv_org_abs_2511_11113
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VIDEOP2R: Video Understanding from Perception to Reasoning
Jiang, Yifan
Wang, Yueying
Zhao, Rui
Parag, Toufiq
Chen, Zhimin
Liao, Zhenyu
Unnikrishnan, Jayakrishnan
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Reinforcement fine-tuning (RFT), a two-stage framework consisting of supervised fine-tuning (SFT) and reinforcement learning (RL) has shown promising results on improving reasoning ability of large language models (LLMs). Yet extending RFT to large video language models (LVLMs) remains challenging. We propose VideoP2R, a novel process-aware video RFT framework that enhances video reasoning by modeling perception and reasoning as distinct processes. In the SFT stage, we develop a three-step pipeline to generate VideoP2R-CoT-162K, a high-quality, process-aware chain-of-thought (CoT) dataset for perception and reasoning. In the RL stage, we introduce a novel process-aware group relative policy optimization (PA-GRPO) algorithm that supplies separate rewards for perception and reasoning. Extensive experiments show that VideoP2R achieves state-of-the-art (SotA) performance on six out of seven video reasoning and understanding benchmarks. Ablation studies further confirm the effectiveness of our process-aware modeling and PA-GRPO and demonstrate that model's perception output is information-sufficient for downstream reasoning. Our project page is available at https://videop2r.github.io/videop2r/.
title VIDEOP2R: Video Understanding from Perception to Reasoning
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2511.11113