Pause and Think: A Dataset and Benchmark for Video-Grounded Assistive Action Suggestion
Fuente:
arXiv
Saved in:
| Main Authors: | Singh, Shivam, Majumdar, Saptarshi, Prabhanjan, Pratik, Liu, Zicheng, Barsoum, Emad |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DUET-VLM: Dual stage Unified Efficient Token reduction for VLM Training and Inference
by: Singh, Aditya Kumar, et al.
Published: (2026)
by: Singh, Aditya Kumar, et al.
Published: (2026)
VideoSeek: Long-Horizon Video Agent with Tool-Guided Seeking
by: Lin, Jingyang, et al.
Published: (2026)
by: Lin, Jingyang, et al.
Published: (2026)
AdaptEvolve: Improving Efficiency of Evolutionary AI Agents through Adaptive Model Selection
by: Ray, Pretam, et al.
Published: (2026)
by: Ray, Pretam, et al.
Published: (2026)
Learning from Online Videos at Inference Time for Computer-Use Agents
by: Liu, Yujian, et al.
Published: (2025)
by: Liu, Yujian, et al.
Published: (2025)
XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language Models
by: Wang, Xingrui, et al.
Published: (2025)
by: Wang, Xingrui, et al.
Published: (2025)
On Occlusions in Video Action Detection: Benchmark Datasets And Training Recipes
by: Modi, Rajat, et al.
Published: (2024)
by: Modi, Rajat, et al.
Published: (2024)
Think with Grounding: Curriculum Reinforced Reasoning with Video Grounding for Long Video Understanding
by: Chen, Houlun, et al.
Published: (2026)
by: Chen, Houlun, et al.
Published: (2026)
TAG: Thinking with Action Unit Grounding for Facial Expression Recognition
by: Lin, Haobo, et al.
Published: (2026)
by: Lin, Haobo, et al.
Published: (2026)
VideoTemp-o3: Harmonizing Temporal Grounding and Video Understanding in Agentic Thinking-with-Videos
by: Liu, Wenqi, et al.
Published: (2026)
by: Liu, Wenqi, et al.
Published: (2026)
When Thinking Drifts: Evidential Grounding for Robust Video Reasoning
by: Luo, Mi, et al.
Published: (2025)
by: Luo, Mi, et al.
Published: (2025)
SpecVLM: Fast Speculative Decoding in Vision-Language Models
by: Huang, Haiduo, et al.
Published: (2025)
by: Huang, Haiduo, et al.
Published: (2025)
Datasets and Recipes for Video Temporal Grounding via Reinforcement Learning
by: Chen, Ruizhe, et al.
Published: (2025)
by: Chen, Ruizhe, et al.
Published: (2025)
VALUED -- Vision and Logical Understanding Evaluation Dataset
by: Saha, Soumadeep, et al.
Published: (2023)
by: Saha, Soumadeep, et al.
Published: (2023)
A Survey of Video Datasets for Grounded Event Understanding
by: Sanders, Kate, et al.
Published: (2024)
by: Sanders, Kate, et al.
Published: (2024)
Thinking in Streaming Video
by: Liu, Zikang, et al.
Published: (2026)
by: Liu, Zikang, et al.
Published: (2026)
SoftVQ-VAE: Efficient 1-Dimensional Continuous Tokenizer
by: Chen, Hao, et al.
Published: (2024)
by: Chen, Hao, et al.
Published: (2024)
Fast Occupancy Network
by: Lu, Mingjie, et al.
Published: (2024)
by: Lu, Mingjie, et al.
Published: (2024)
Partial Convolution Meets Visual Attention
by: Huang, Haiduo, et al.
Published: (2025)
by: Huang, Haiduo, et al.
Published: (2025)
DL-QAT: Weight-Decomposed Low-Rank Quantization-Aware Training for Large Language Models
by: Ke, Wenjin, et al.
Published: (2025)
by: Ke, Wenjin, et al.
Published: (2025)
Beyond Play and Pause: Turning GPT-4o Spatial Weakness into a Strength for In-Depth Interactive Video Learning
by: Goudarzi, Sajad, et al.
Published: (2025)
by: Goudarzi, Sajad, et al.
Published: (2025)
Geak: Introducing Triton Kernel AI Agent & Evaluation Benchmarks
by: Wang, Jianghui, et al.
Published: (2025)
by: Wang, Jianghui, et al.
Published: (2025)
ProGait: A Multi-Purpose Video Dataset and Benchmark for Transfemoral Prosthesis Users
by: Yin, Xiangyu, et al.
Published: (2025)
by: Yin, Xiangyu, et al.
Published: (2025)
Athena: Enhancing Multimodal Reasoning with Data-efficient Process Reward Models
by: Wang, Shuai, et al.
Published: (2025)
by: Wang, Shuai, et al.
Published: (2025)
VideoGEM: Training-free Action Grounding in Videos
by: Vogel, Felix, et al.
Published: (2025)
by: Vogel, Felix, et al.
Published: (2025)
BadmintonGRF: A Multimodal Dataset and Benchmark for Markerless Ground Reaction Force Estimation in Badminton
by: Niu, Kuoye, et al.
Published: (2026)
by: Niu, Kuoye, et al.
Published: (2026)
PinpointQA: A Dataset and Benchmark for Small Object-Centric Spatial Understanding in Indoor Videos
by: Zhou, Zhiyu, et al.
Published: (2026)
by: Zhou, Zhiyu, et al.
Published: (2026)
SPACT18: Spiking Human Action Recognition Benchmark Dataset with Complementary RGB and Thermal Modalities
by: Ashraf, Yasser, et al.
Published: (2025)
by: Ashraf, Yasser, et al.
Published: (2025)
Router-Suggest: Dynamic Routing for Multimodal Auto-Completion in Visually-Grounded Dialogs
by: Mishra, Sandeep, et al.
Published: (2026)
by: Mishra, Sandeep, et al.
Published: (2026)
DC-DiT: Adaptive Compute and Elastic Inference for Visual Generation via Dynamic Chunking
by: Haridas, Akash, et al.
Published: (2026)
by: Haridas, Akash, et al.
Published: (2026)
VideoExplorer: Think With Videos For Agentic Long-Video Understanding
by: Yuan, Huaying, et al.
Published: (2025)
by: Yuan, Huaying, et al.
Published: (2025)
VideoITG: Multimodal Video Understanding with Instructed Temporal Grounding
by: Wang, Shihao, et al.
Published: (2025)
by: Wang, Shihao, et al.
Published: (2025)
ACCIDENT: A Benchmark Dataset for Vehicle Accident Detection from Traffic Surveillance Videos
by: Picek, Lukas, et al.
Published: (2026)
by: Picek, Lukas, et al.
Published: (2026)
EgoEdit: Dataset, Real-Time Streaming Model, and Benchmark for Egocentric Video Editing
by: Li, Runjia, et al.
Published: (2025)
by: Li, Runjia, et al.
Published: (2025)
Commonsense Video Question Answering through Video-Grounded Entailment Tree Reasoning
by: Liu, Huabin, et al.
Published: (2025)
by: Liu, Huabin, et al.
Published: (2025)
CogStream: Context-guided Streaming Video Question Answering
by: Zhao, Zicheng, et al.
Published: (2025)
by: Zhao, Zicheng, et al.
Published: (2025)
How Confident are Video Models? Empowering Video Models to Express their Uncertainty
by: Mei, Zhiting, et al.
Published: (2025)
by: Mei, Zhiting, et al.
Published: (2025)
Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models
by: Wang, Haibo, et al.
Published: (2024)
by: Wang, Haibo, et al.
Published: (2024)
Object-Shot Enhanced Grounding Network for Egocentric Video
by: Feng, Yisen, et al.
Published: (2025)
by: Feng, Yisen, et al.
Published: (2025)
Ego2Web: A Web Agent Benchmark Grounded in Egocentric Videos
by: Yu, Shoubin, et al.
Published: (2026)
by: Yu, Shoubin, et al.
Published: (2026)
A Survey on Backbones for Deep Video Action Recognition
by: Tang, Zixuan, et al.
Published: (2024)
by: Tang, Zixuan, et al.
Published: (2024)
Similar Items
-
DUET-VLM: Dual stage Unified Efficient Token reduction for VLM Training and Inference
by: Singh, Aditya Kumar, et al.
Published: (2026) -
VideoSeek: Long-Horizon Video Agent with Tool-Guided Seeking
by: Lin, Jingyang, et al.
Published: (2026) -
AdaptEvolve: Improving Efficiency of Evolutionary AI Agents through Adaptive Model Selection
by: Ray, Pretam, et al.
Published: (2026) -
Learning from Online Videos at Inference Time for Computer-Use Agents
by: Liu, Yujian, et al.
Published: (2025) -
XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language Models
by: Wang, Xingrui, et al.
Published: (2025)