Video-CoM: Interactive Video Reasoning via Chain of Manipulations

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Rasheed, Hanoona, Zumri, Mohammed, Maaz, Muhammad, Yang, Ming-Hsuan, Khan, Fahad Shahbaz, Khan, Salman
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912736529612800
author Rasheed, Hanoona
Zumri, Mohammed
Maaz, Muhammad
Yang, Ming-Hsuan
Khan, Fahad Shahbaz
Khan, Salman
author_facet Rasheed, Hanoona
Zumri, Mohammed
Maaz, Muhammad
Yang, Ming-Hsuan
Khan, Fahad Shahbaz
Khan, Salman
contents Recent multimodal large language models (MLLMs) have advanced video understanding, yet most still "think about videos" ie once a video is encoded, reasoning unfolds entirely in text, treating visual input as a static context. This passive paradigm creates a semantic bottleneck: models cannot rewatch, refocus, or verify evidence, leading to shallow visual reasoning on tasks requiring fine grained spatio temporal understanding. In this work, we introduce Interactive Video Reasoning, a new paradigm that transforms video into an active cognitive workspace, enabling models to "think with videos". Our model, Video CoM, reasons through a Chain of Manipulations (CoM), performing iterative visual actions to gather and refine evidence. To support this behavior, we construct Video CoM Instruct, an 18K instruction tuning dataset curated for multi step manipulation reasoning. Beyond supervised learning, we further optimize the manipulation policy via reinforcement learning with reasoning aware Group Relative Policy Optimization (GRPO). Unlike prior work that relies solely on sparse answer rewards, our method introduces step level reasoning rewards, guiding the model toward grounded and consistent reasoning. Video CoM achieves strong results across nine video reasoning benchmarks, improving average performance by 3.6 percent over recent state of the art models, while training on only 25K SFT and 3K GRPO video samples, significantly fewer than comparable large scale models. Ablation studies demonstrate that reasoning aware rewards improve both accuracy and interpretability. Code: https://github.com/mbzuai-oryx/Video-CoM
format Preprint
id arxiv_https___arxiv_org_abs_2511_23477
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Video-CoM: Interactive Video Reasoning via Chain of Manipulations
Rasheed, Hanoona
Zumri, Mohammed
Maaz, Muhammad
Yang, Ming-Hsuan
Khan, Fahad Shahbaz
Khan, Salman
Computer Vision and Pattern Recognition
Recent multimodal large language models (MLLMs) have advanced video understanding, yet most still "think about videos" ie once a video is encoded, reasoning unfolds entirely in text, treating visual input as a static context. This passive paradigm creates a semantic bottleneck: models cannot rewatch, refocus, or verify evidence, leading to shallow visual reasoning on tasks requiring fine grained spatio temporal understanding. In this work, we introduce Interactive Video Reasoning, a new paradigm that transforms video into an active cognitive workspace, enabling models to "think with videos". Our model, Video CoM, reasons through a Chain of Manipulations (CoM), performing iterative visual actions to gather and refine evidence. To support this behavior, we construct Video CoM Instruct, an 18K instruction tuning dataset curated for multi step manipulation reasoning. Beyond supervised learning, we further optimize the manipulation policy via reinforcement learning with reasoning aware Group Relative Policy Optimization (GRPO). Unlike prior work that relies solely on sparse answer rewards, our method introduces step level reasoning rewards, guiding the model toward grounded and consistent reasoning. Video CoM achieves strong results across nine video reasoning benchmarks, improving average performance by 3.6 percent over recent state of the art models, while training on only 25K SFT and 3K GRPO video samples, significantly fewer than comparable large scale models. Ablation studies demonstrate that reasoning aware rewards improve both accuracy and interpretability. Code: https://github.com/mbzuai-oryx/Video-CoM
title Video-CoM: Interactive Video Reasoning via Chain of Manipulations
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.23477