Video Models Reason Early: Exploiting Plan Commitment for Maze Solving
Fuente:
arXiv
Saved in:
| Main Authors: | Newman, Kaleb, Zhu, Tyler, Russakovsky, Olga |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Reasoning via Video: The First Evaluation of Video Models' Reasoning Abilities through Maze-Solving Tasks
by: Yang, Cheng, et al.
Published: (2025)
by: Yang, Cheng, et al.
Published: (2025)
D2D: Detector-to-Differentiable Critic for Improved Numeracy in Text-to-Image Generation
by: Yoo, Nobline, et al.
Published: (2025)
by: Yoo, Nobline, et al.
Published: (2025)
Attention IoU: Examining Biases in CelebA using Attention Maps
by: Serianni, Aaron, et al.
Published: (2025)
by: Serianni, Aaron, et al.
Published: (2025)
Unifying Specialized Visual Encoders for Video Language Models
by: Chung, Jihoon, et al.
Published: (2025)
by: Chung, Jihoon, et al.
Published: (2025)
Video Models Start to Solve Chess, Maze, Sudoku, Mental Rotation, and Raven' Matrices
by: Deng, Hokin
Published: (2025)
by: Deng, Hokin
Published: (2025)
ImageNet-OOD: Deciphering Modern Out-of-Distribution Detection Algorithms
by: Yang, William, et al.
Published: (2023)
by: Yang, William, et al.
Published: (2023)
D$^3$: Scaling Up Deepfake Detection by Learning from Discrepancy
by: Yang, Yongqi, et al.
Published: (2024)
by: Yang, Yongqi, et al.
Published: (2024)
The Silent Assistant: NoiseQuery as Implicit Guidance for Goal-Driven Image Generation
by: Wang, Ruoyu, et al.
Published: (2024)
by: Wang, Ruoyu, et al.
Published: (2024)
Vision-Language Dataset Distillation
by: Wu, Xindi, et al.
Published: (2023)
by: Wu, Xindi, et al.
Published: (2023)
Seeing Beyond the Scene: Analyzing and Mitigating Background Bias in Action Recognition
by: Zhou, Ellie, et al.
Published: (2025)
by: Zhou, Ellie, et al.
Published: (2025)
Beyond Objects: Contextual Synthetic Data Generation for Fine-Grained Classification
by: Yang, William, et al.
Published: (2025)
by: Yang, William, et al.
Published: (2025)
ConceptMix: A Compositional Image Generation Benchmark with Controllable Difficulty
by: Wu, Xindi, et al.
Published: (2024)
by: Wu, Xindi, et al.
Published: (2024)
A Sampling-Based Domain Generalization Study with Diffusion Generative Models
by: Zhu, Ye, et al.
Published: (2023)
by: Zhu, Ye, et al.
Published: (2023)
Do Pre-trained Vision-Language Models Encode Object States?
by: Newman, Kaleb, et al.
Published: (2024)
by: Newman, Kaleb, et al.
Published: (2024)
Exploiting Auxiliary Caption for Video Grounding
by: Li, Hongxiang, et al.
Published: (2023)
by: Li, Hongxiang, et al.
Published: (2023)
Visual Compositional Tuning
by: Wu, Xindi, et al.
Published: (2025)
by: Wu, Xindi, et al.
Published: (2025)
SOWing Information: Cultivating Contextual Coherence with MLLMs in Image Generation
by: Pei, Yuhan, et al.
Published: (2024)
by: Pei, Yuhan, et al.
Published: (2024)
From Pixels to BFS: High Maze Accuracy Does Not Imply Visual Planning
by: Salgado, Alberto G. Rodriguez
Published: (2026)
by: Salgado, Alberto G. Rodriguez
Published: (2026)
The Impact of Coreset Selection on Spurious Correlations and Group Robustness
by: Dharmasiri, Amaya, et al.
Published: (2025)
by: Dharmasiri, Amaya, et al.
Published: (2025)
Personalized Generative Models for Contextual Debiasing
by: Liang, Xinran, et al.
Published: (2026)
by: Liang, Xinran, et al.
Published: (2026)
EgoVITA: Learning to Plan and Verify for Egocentric Video Reasoning
by: Kulkarni, Yogesh, et al.
Published: (2025)
by: Kulkarni, Yogesh, et al.
Published: (2025)
What Makes a Maze Look Like a Maze?
by: Hsu, Joy, et al.
Published: (2024)
by: Hsu, Joy, et al.
Published: (2024)
Video Models Can Reason with Verifiable Rewards
by: Zhu, Tinghui, et al.
Published: (2026)
by: Zhu, Tinghui, et al.
Published: (2026)
Bias at the End of the Score
by: Magid, Salma Abdel, et al.
Published: (2026)
by: Magid, Salma Abdel, et al.
Published: (2026)
Analyzing the Roles of Language and Vision in Learning from Limited Data
by: Chen, Allison, et al.
Published: (2024)
by: Chen, Allison, et al.
Published: (2024)
VidMan: Exploiting Implicit Dynamics from Video Diffusion Model for Effective Robot Manipulation
by: Wen, Youpeng, et al.
Published: (2024)
by: Wen, Youpeng, et al.
Published: (2024)
Language Model Guided Interpretable Video Action Reasoning
by: Wang, Ning, et al.
Published: (2024)
by: Wang, Ning, et al.
Published: (2024)
Identifying and Solving Conditional Image Leakage in Image-to-Video Diffusion Model
by: Zhao, Min, et al.
Published: (2024)
by: Zhao, Min, et al.
Published: (2024)
MoniTor: Exploiting Large Language Models with Instruction for Online Video Anomaly Detection
by: Yang, Shengtian, et al.
Published: (2025)
by: Yang, Shengtian, et al.
Published: (2025)
Navigating the Maze of Explainable AI: A Systematic Approach to Evaluating Methods and Metrics
by: Klein, Lukas, et al.
Published: (2024)
by: Klein, Lukas, et al.
Published: (2024)
Exploiting Multimodal Spatial-temporal Patterns for Video Object Tracking
by: Hu, Xiantao, et al.
Published: (2024)
by: Hu, Xiantao, et al.
Published: (2024)
Exploiting Optical Flow Guidance for Transformer-Based Video Inpainting
by: Zhang, Kaidong, et al.
Published: (2023)
by: Zhang, Kaidong, et al.
Published: (2023)
Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times
by: Loginova, Olga, et al.
Published: (2025)
by: Loginova, Olga, et al.
Published: (2025)
V-ReasonBench: Toward Unified Reasoning Benchmark Suite for Video Generation Models
by: Luo, Yang, et al.
Published: (2025)
by: Luo, Yang, et al.
Published: (2025)
RANKVIDEO: Reasoning Reranking for Text-to-Video Retrieval
by: Skow, Tyler, et al.
Published: (2026)
by: Skow, Tyler, et al.
Published: (2026)
Solving the Clustering Reasoning Problems by Modeling a Deep-Learning-Based Probabilistic Model
by: Song, Ruizhuo, et al.
Published: (2024)
by: Song, Ruizhuo, et al.
Published: (2024)
Solving Motion Planning Tasks with a Scalable Generative Model
by: Hu, Yihan, et al.
Published: (2024)
by: Hu, Yihan, et al.
Published: (2024)
Explainable Action Form Assessment by Exploiting Multimodal Chain-of-Thoughts Reasoning
by: Qi, Mengshi, et al.
Published: (2025)
by: Qi, Mengshi, et al.
Published: (2025)
Dynamic Reflections: Probing Video Representations with Text Alignment
by: Zhu, Tyler, et al.
Published: (2025)
by: Zhu, Tyler, et al.
Published: (2025)
RiskCueBench: Benchmarking Anticipatory Reasoning from Early Risk Cues in Video-Language Models
by: Luo, Sha, et al.
Published: (2026)
by: Luo, Sha, et al.
Published: (2026)
Similar Items
-
Reasoning via Video: The First Evaluation of Video Models' Reasoning Abilities through Maze-Solving Tasks
by: Yang, Cheng, et al.
Published: (2025) -
D2D: Detector-to-Differentiable Critic for Improved Numeracy in Text-to-Image Generation
by: Yoo, Nobline, et al.
Published: (2025) -
Attention IoU: Examining Biases in CelebA using Attention Maps
by: Serianni, Aaron, et al.
Published: (2025) -
Unifying Specialized Visual Encoders for Video Language Models
by: Chung, Jihoon, et al.
Published: (2025) -
Video Models Start to Solve Chess, Maze, Sudoku, Mental Rotation, and Raven' Matrices
by: Deng, Hokin
Published: (2025)