Pause and Think: A Dataset and Benchmark for Video-Grounded Assistive Action Suggestion

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Singh, Shivam, Majumdar, Saptarshi, Prabhanjan, Pratik, Liu, Zicheng, Barsoum, Emad
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916070499024896
author Singh, Shivam
Majumdar, Saptarshi
Prabhanjan, Pratik
Liu, Zicheng
Barsoum, Emad
author_facet Singh, Shivam
Majumdar, Saptarshi
Prabhanjan, Pratik
Liu, Zicheng
Barsoum, Emad
contents Recent Vision-Language Models (VLMs) struggle with grounded reasoning, temporal consistency, and context aware planning in videos. We introduce pause-and-think-T, a reasoning-centric training dataset that encourages models to pause, reason over visual evidence, and produce concise, actionable responses. The dataset promotes structured reasoning prior to answer generation, guiding models toward human-like, scene-grounded assistance. We fine-tune a compact 4B-parameter model and evaluate it on our pause-and-think-B benchmark targeting contextual understanding and goal planning tasks. The model achieves 58.0% accuracy at 59x fewer parameters than Qwen3-VL-235B (58.9%), matching GPT-5.2 on scene understanding and surpassing GPT-4o. Beyond our benchmark, it also shows strong out-of-distribution performance on EgoThink and TempCompass, with substantial gains in affordance, assistance, attribution recognition, situated reasoning, and temporal order, without benchmark-specific training. Our results indicate that targeted reasoning supervision enables compact models to deliver actionable, visually grounded guidance while generalizing beyond training data, without requiring large-scale model expansion.
format Preprint
id arxiv_https___arxiv_org_abs_2606_00616
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Pause and Think: A Dataset and Benchmark for Video-Grounded Assistive Action Suggestion
Singh, Shivam
Majumdar, Saptarshi
Prabhanjan, Pratik
Liu, Zicheng
Barsoum, Emad
Computer Vision and Pattern Recognition
Artificial Intelligence
Recent Vision-Language Models (VLMs) struggle with grounded reasoning, temporal consistency, and context aware planning in videos. We introduce pause-and-think-T, a reasoning-centric training dataset that encourages models to pause, reason over visual evidence, and produce concise, actionable responses. The dataset promotes structured reasoning prior to answer generation, guiding models toward human-like, scene-grounded assistance. We fine-tune a compact 4B-parameter model and evaluate it on our pause-and-think-B benchmark targeting contextual understanding and goal planning tasks. The model achieves 58.0% accuracy at 59x fewer parameters than Qwen3-VL-235B (58.9%), matching GPT-5.2 on scene understanding and surpassing GPT-4o. Beyond our benchmark, it also shows strong out-of-distribution performance on EgoThink and TempCompass, with substantial gains in affordance, assistance, attribution recognition, situated reasoning, and temporal order, without benchmark-specific training. Our results indicate that targeted reasoning supervision enables compact models to deliver actionable, visually grounded guidance while generalizing beyond training data, without requiring large-scale model expansion.
title Pause and Think: A Dataset and Benchmark for Video-Grounded Assistive Action Suggestion
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2606.00616