Order from Chaos: Physical World Understanding from Glitchy Gameplay Videos
Fuente:
arXiv
Saved in:
| Main Authors: | Cao, Meng, Tang, Haoran, Zhao, Haoze, Han, Mingfei, Liu, Ruyang, Sun, Qiang, Chang, Xiaojun, Reid, Ian, Liang, Xiaodan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PhysGame: Uncovering Physical Commonsense Violations in Gameplay Videos
by: Cao, Meng, et al.
Published: (2024)
by: Cao, Meng, et al.
Published: (2024)
Ground-R1: Incentivizing Grounded Visual Reasoning via Reinforcement Learning
by: Cao, Meng, et al.
Published: (2025)
by: Cao, Meng, et al.
Published: (2025)
Video Spatial Reasoning with Object-Centric 3D Rollout
by: Tang, Haoran, et al.
Published: (2025)
by: Tang, Haoran, et al.
Published: (2025)
Flow4Agent: Long-form Video Understanding via Motion Prior from Optical Flow
by: Liu, Ruyang, et al.
Published: (2025)
by: Liu, Ruyang, et al.
Published: (2025)
MUSE: Mamba is Efficient Multi-scale Learner for Text-video Retrieval
by: Tang, Haoran, et al.
Published: (2024)
by: Tang, Haoran, et al.
Published: (2024)
Seeing through Imagination: Learning Scene Geometry via Implicit Spatial World Modeling
by: Cao, Meng, et al.
Published: (2025)
by: Cao, Meng, et al.
Published: (2025)
RAP: Efficient Text-Video Retrieval with Sparse-and-Correlated Adapter
by: Cao, Meng, et al.
Published: (2024)
by: Cao, Meng, et al.
Published: (2024)
StoryAgent: Customized Storytelling Video Generation via Multi-Agent Collaboration
by: Hu, Panwen, et al.
Published: (2024)
by: Hu, Panwen, et al.
Published: (2024)
Implicit Geometry Representations for Vision-and-Language Navigation from Web Videos
by: Han, Mingfei, et al.
Published: (2026)
by: Han, Mingfei, et al.
Published: (2026)
LongVLM: Efficient Long Video Understanding via Large Language Models
by: Weng, Yuetian, et al.
Published: (2024)
by: Weng, Yuetian, et al.
Published: (2024)
Shot2Story: A New Benchmark for Comprehensive Understanding of Multi-shot Videos
by: Han, Mingfei, et al.
Published: (2023)
by: Han, Mingfei, et al.
Published: (2023)
RoomTour3D: Geometry-Aware Video-Instruction Tuning for Embodied Navigation
by: Han, Mingfei, et al.
Published: (2024)
by: Han, Mingfei, et al.
Published: (2024)
CARE What Fails: Contrastive Anchored-REflection for Verifiable Multimodal Reasoning
by: Wang, Yongxin, et al.
Published: (2025)
by: Wang, Yongxin, et al.
Published: (2025)
Video SimpleQA: Towards Factuality Evaluation in Large Video Language Models
by: Cao, Meng, et al.
Published: (2025)
by: Cao, Meng, et al.
Published: (2025)
PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance
by: Sun, Shangkun, et al.
Published: (2024)
by: Sun, Shangkun, et al.
Published: (2024)
SpatialDreamer: Incentivizing Spatial Reasoning via Active Mental Imagery
by: Cao, Meng, et al.
Published: (2025)
by: Cao, Meng, et al.
Published: (2025)
World2Act: Latent Action Post-Training from World Model Dynamics
by: Vuong, An Dinh, et al.
Published: (2026)
by: Vuong, An Dinh, et al.
Published: (2026)
Beyond Dense Futures: World Models as Structured Planners for Robotic Manipulation
by: Jin, Minghao, et al.
Published: (2026)
by: Jin, Minghao, et al.
Published: (2026)
Progressive Online Video Understanding with Evidence-Aligned Timing and Transparent Decisions
by: Zhang, Kecheng, et al.
Published: (2026)
by: Zhang, Kecheng, et al.
Published: (2026)
Finite Automata Extraction: Low-data World Model Learning as Programs from Gameplay Video
by: Goel, Dave, et al.
Published: (2025)
by: Goel, Dave, et al.
Published: (2025)
EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation
by: Wang, Yongxin, et al.
Published: (2024)
by: Wang, Yongxin, et al.
Published: (2024)
PhyBlock: A Progressive Benchmark for Physical Understanding and Planning via 3D Block Assembly
by: Ma, Liang, et al.
Published: (2025)
by: Ma, Liang, et al.
Published: (2025)
BridgeIV: Bridging Customized Image and Video Generation through Test-Time Autoregressive Identity Propagation
by: Hu, Panwen, et al.
Published: (2025)
by: Hu, Panwen, et al.
Published: (2025)
CoNav: Collaborative Cross-Modal Reasoning for Embodied Navigation
by: Hao, Haihong, et al.
Published: (2025)
by: Hao, Haihong, et al.
Published: (2025)
JEUX Gameplay
by: Kukkonen, Karin, et al.
Published: (2025)
by: Kukkonen, Karin, et al.
Published: (2025)
Thinking with Drafts: Speculative Temporal Reasoning for Efficient Long Video Understanding
by: Hu, Pengfei, et al.
Published: (2025)
by: Hu, Pengfei, et al.
Published: (2025)
Bridging Chaos Game Representations and $k$-mer Frequencies of DNA Sequences
by: He, Haoze, et al.
Published: (2025)
by: He, Haoze, et al.
Published: (2025)
Automated Bug Frame Retrieval from Gameplay Videos Using Vision-Language Models
by: Lu, Wentao, et al.
Published: (2025)
by: Lu, Wentao, et al.
Published: (2025)
ST-LLM: Large Language Models Are Effective Temporal Learners
by: Liu, Ruyang, et al.
Published: (2024)
by: Liu, Ruyang, et al.
Published: (2024)
Gameplay Highlights Generation
by: Edithal, Vignesh, et al.
Published: (2025)
by: Edithal, Vignesh, et al.
Published: (2025)
GLaD: Geometric Latent Distillation for Vision-Language-Action Models
by: Guo, Minghao, et al.
Published: (2025)
by: Guo, Minghao, et al.
Published: (2025)
WISA: World Simulator Assistant for Physics-Aware Text-to-Video Generation
by: Wang, Jing, et al.
Published: (2025)
by: Wang, Jing, et al.
Published: (2025)
Continual LLaVA: Continual Instruction Tuning in Large Vision-Language Models
by: Cao, Meng, et al.
Published: (2024)
by: Cao, Meng, et al.
Published: (2024)
Reframing Pedagogical Practices Through the UTBO ‐ CLIL Framework in Art and Design Higher Education
by: Li Ruyang, et al.
Published: (2026)
by: Li Ruyang, et al.
Published: (2026)
The NES Video-Music Database: A Dataset of Symbolic Video Game Music Paired with Gameplay Videos
by: Cardoso, Igor, et al.
Published: (2024)
by: Cardoso, Igor, et al.
Published: (2024)
Semantic, Orthographic, and Phonological Biases in Humans' Wordle Gameplay
by: Liang, Jiadong, et al.
Published: (2024)
by: Liang, Jiadong, et al.
Published: (2024)
Volt Hockey Equipment and Gameplay
by: Avery Melam
Published: (2026)
by: Avery Melam
Published: (2026)
GameplayQA: A Benchmarking Framework for Decision-Dense POV-Synced Multi-Video Understanding of 3D Virtual Agents
by: Wang, Yunzhe, et al.
Published: (2026)
by: Wang, Yunzhe, et al.
Published: (2026)
Randomized Advantage Transformation (RAT): Computing Natural Policy Gradients via Direct Backpropagation
by: Sun, Mingfei
Published: (2026)
by: Sun, Mingfei
Published: (2026)
BeTAIL: Behavior Transformer Adversarial Imitation Learning from Human Racing Gameplay
by: Weaver, Catherine, et al.
Published: (2024)
by: Weaver, Catherine, et al.
Published: (2024)
Similar Items
-
PhysGame: Uncovering Physical Commonsense Violations in Gameplay Videos
by: Cao, Meng, et al.
Published: (2024) -
Ground-R1: Incentivizing Grounded Visual Reasoning via Reinforcement Learning
by: Cao, Meng, et al.
Published: (2025) -
Video Spatial Reasoning with Object-Centric 3D Rollout
by: Tang, Haoran, et al.
Published: (2025) -
Flow4Agent: Long-form Video Understanding via Motion Prior from Optical Flow
by: Liu, Ruyang, et al.
Published: (2025) -
MUSE: Mamba is Efficient Multi-scale Learner for Text-video Retrieval
by: Tang, Haoran, et al.
Published: (2024)