DeepPHY: Benchmarking Agentic VLMs on Physical Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Xinrun, Bu, Pi, Wang, Ye, Karlsson, Börje F., Wang, Ziming, Song, Tengtao, Zhu, Qi, Song, Jun, Ding, Zhiming, Zheng, Bo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ICPRL: Acquiring Physical Intuition from Interactive Control
by: Xu, Xinrun, et al.
Published: (2026)
by: Xu, Xinrun, et al.
Published: (2026)
A Survey on Game Playing Agents and Large Models: Methods, Applications, and Challenges
by: Xu, Xinrun, et al.
Published: (2024)
by: Xu, Xinrun, et al.
Published: (2024)
Can VLMs Play Action Role-Playing Games? Take Black Myth Wukong as a Study Case
by: Chen, Peng, et al.
Published: (2024)
by: Chen, Peng, et al.
Published: (2024)
MindRef: Mimicking Human Memory for Hierarchical Reference Retrieval with Fine-Grained Location Awareness
by: Wang, Ye, et al.
Published: (2024)
by: Wang, Ye, et al.
Published: (2024)
MLLM as Retriever: Interactively Learning Multimodal Retrieval for Embodied Agents
by: Yue, Junpeng, et al.
Published: (2024)
by: Yue, Junpeng, et al.
Published: (2024)
SWITCH: Benchmarking Modeling and Handling of Tangible Interfaces in Long-horizon Embodied Scenarios
by: Lin, Jieru, et al.
Published: (2025)
by: Lin, Jieru, et al.
Published: (2025)
Deep Pre-Alignment for VLMs
by: Yu, Tianyu, et al.
Published: (2026)
by: Yu, Tianyu, et al.
Published: (2026)
CombatVLA: An Efficient Vision-Language-Action Model for Combat Tasks in 3D Action Role-Playing Games
by: Chen, Peng, et al.
Published: (2025)
by: Chen, Peng, et al.
Published: (2025)
"See the World, Discover Knowledge": A Chinese Factuality Evaluation for Large Vision Language Models
by: Gu, Jihao, et al.
Published: (2025)
by: Gu, Jihao, et al.
Published: (2025)
VStyle: A Benchmark for Voice Style Adaptation with Spoken Instructions
by: Zhan, Jun, et al.
Published: (2025)
by: Zhan, Jun, et al.
Published: (2025)
LongDocURL: a Comprehensive Multimodal Long Document Benchmark Integrating Understanding, Reasoning, and Locating
by: Deng, Chao, et al.
Published: (2024)
by: Deng, Chao, et al.
Published: (2024)
GeoSense: Evaluating Identification and Application of Geometric Principles in Multimodal Reasoning
by: Xu, Liangyu, et al.
Published: (2025)
by: Xu, Liangyu, et al.
Published: (2025)
CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs
by: Jian, Ai, et al.
Published: (2025)
by: Jian, Ai, et al.
Published: (2025)
Taking Notes Brings Focus? Towards Multi-Turn Multimodal Dialogue Learning
by: Liu, Jiazheng, et al.
Published: (2025)
by: Liu, Jiazheng, et al.
Published: (2025)
Scaling Agentic Reinforcement Learning for Tool-Integrated Reasoning in VLMs
by: Lu, Meng, et al.
Published: (2025)
by: Lu, Meng, et al.
Published: (2025)
How Foundational Skills Influence VLM-based Embodied Agents:A Native Perspective
by: Peng, Bo, et al.
Published: (2026)
by: Peng, Bo, et al.
Published: (2026)
SecAgent: Efficient Mobile GUI Agent with Semantic Context
by: Xie, Yiping, et al.
Published: (2026)
by: Xie, Yiping, et al.
Published: (2026)
X-DiffVLA: X-Embodied Diffusion Action Heads for Vision-Language-Action Models
by: Li, Boyu, et al.
Published: (2026)
by: Li, Boyu, et al.
Published: (2026)
Token Preference Optimization with Self-Calibrated Visual-Anchored Rewards for Hallucination Mitigation
by: Gu, Jihao, et al.
Published: (2024)
by: Gu, Jihao, et al.
Published: (2024)
Synapse: Trajectory-as-Exemplar Prompting with Memory for Computer Control
by: Zheng, Longtao, et al.
Published: (2023)
by: Zheng, Longtao, et al.
Published: (2023)
Mobile-R1: Towards Interactive Capability for VLM-Based Mobile Agent via Systematic Training
by: Gu, Jihao, et al.
Published: (2025)
by: Gu, Jihao, et al.
Published: (2025)
StarDojo: Benchmarking Open-Ended Behaviors of Agentic Multimodal LLMs in Production-Living Simulations with Stardew Valley
by: Tan, Weihao, et al.
Published: (2025)
by: Tan, Weihao, et al.
Published: (2025)
ReWatch-R1: Boosting Complex Video Reasoning in Large Vision-Language Models through Agentic Data Synthesis
by: Zhang, Congzhi, et al.
Published: (2025)
by: Zhang, Congzhi, et al.
Published: (2025)
PriPHiT: Privacy-Preserving Hierarchical Training of Deep Neural Networks
by: Sepehri, Yamin, et al.
Published: (2024)
by: Sepehri, Yamin, et al.
Published: (2024)
Think Twice to See More: Iterative Visual Reasoning in Medical VLMs
by: Chen, Kaitao, et al.
Published: (2025)
by: Chen, Kaitao, et al.
Published: (2025)
VLRS-Bench: A Vision-Language Reasoning Benchmark for Remote Sensing
by: Luo, Zhiming, et al.
Published: (2026)
by: Luo, Zhiming, et al.
Published: (2026)
Being-0: A Humanoid Robotic Agent with Vision-Language Models and Modular Skills
by: Yuan, Haoqi, et al.
Published: (2025)
by: Yuan, Haoqi, et al.
Published: (2025)
正方形内接试证明(不一定为真,先存证)
by: Song, Ziming
Published: (2026)
by: Song, Ziming
Published: (2026)
GDBA Revisited: Unleashing the Power of Guided Local Search for Distributed Constraint Optimization
by: Deng, Yanchen, et al.
Published: (2025)
by: Deng, Yanchen, et al.
Published: (2025)
Resolving Latency and Inventory Risk in Market Making with Reinforcement Learning
by: Jiang, Junzhe, et al.
Published: (2025)
by: Jiang, Junzhe, et al.
Published: (2025)
PhysicsMind: Sim and Real Mechanics Benchmarking for Physical Reasoning and Prediction in Foundational VLMs and World Models
by: Mak, Chak-Wing, et al.
Published: (2026)
by: Mak, Chak-Wing, et al.
Published: (2026)
SIRI-Bench: Challenging VLMs' Spatial Intelligence through Complex Reasoning Tasks
by: Song, Zijian, et al.
Published: (2025)
by: Song, Zijian, et al.
Published: (2025)
Nondeterministic Polynomial-time Problem Challenge: An Ever-Scaling Reasoning Benchmark for LLMs
by: Yang, Chang, et al.
Published: (2025)
by: Yang, Chang, et al.
Published: (2025)
Patho-AgenticRAG: Towards Multimodal Agentic Retrieval-Augmented Generation for Pathology VLMs via Reinforcement Learning
by: Zhang, Wenchuan, et al.
Published: (2025)
by: Zhang, Wenchuan, et al.
Published: (2025)
Benchmarking VLMs' Reasoning About Persuasive Atypical Images
by: Malakouti, Sina, et al.
Published: (2024)
by: Malakouti, Sina, et al.
Published: (2024)
VRIQ: Benchmarking and Analyzing Visual-Reasoning IQ of VLMs
by: Khezresmaeilzadeh, Tina, et al.
Published: (2026)
by: Khezresmaeilzadeh, Tina, et al.
Published: (2026)
RANGER: A Monocular Zero-Shot Semantic Navigation Framework through Visual Contextual Adaptation
by: Yu, Ming-Ming, et al.
Published: (2025)
by: Yu, Ming-Ming, et al.
Published: (2025)
A Bibliography of Writings on Distance Education.
by: Holmberg, Borje
Published: (1990)
by: Holmberg, Borje
Published: (1990)
Microphytobenthic productivity in mangrove areas of different replanting regimes, Gazi Bay, Kenya.
by: Borje, Annika
Published: (2004)
by: Borje, Annika
Published: (2004)
Physically Interpretable Emulation of a Moist Convecting Atmosphere with a Recurrent Neural Network
by: Song, Qiyu, et al.
Published: (2025)
by: Song, Qiyu, et al.
Published: (2025)
Similar Items
-
ICPRL: Acquiring Physical Intuition from Interactive Control
by: Xu, Xinrun, et al.
Published: (2026) -
A Survey on Game Playing Agents and Large Models: Methods, Applications, and Challenges
by: Xu, Xinrun, et al.
Published: (2024) -
Can VLMs Play Action Role-Playing Games? Take Black Myth Wukong as a Study Case
by: Chen, Peng, et al.
Published: (2024) -
MindRef: Mimicking Human Memory for Hierarchical Reference Retrieval with Fine-Grained Location Awareness
by: Wang, Ye, et al.
Published: (2024) -
MLLM as Retriever: Interactively Learning Multimodal Retrieval for Embodied Agents
by: Yue, Junpeng, et al.
Published: (2024)