PhysBench: Benchmarking and Enhancing Vision-Language Models for Physical World Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Chow, Wei, Mao, Jiageng, Li, Boyi, Seita, Daniel, Guizilini, Vitor, Wang, Yue |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning from Massive Human Videos for Universal Humanoid Pose Control
by: Mao, Jiageng, et al.
Published: (2024)
by: Mao, Jiageng, et al.
Published: (2024)
PhysToolBench: Benchmarking Physical Tool Understanding for MLLMs
by: Zhang, Zixin, et al.
Published: (2025)
by: Zhang, Zixin, et al.
Published: (2025)
Robot Learning from Any Images
by: Zhao, Siheng, et al.
Published: (2025)
by: Zhao, Siheng, et al.
Published: (2025)
A Language Agent for Autonomous Driving
by: Mao, Jiageng, et al.
Published: (2023)
by: Mao, Jiageng, et al.
Published: (2023)
Can DeepSeek Reason Like a Surgeon? An Empirical Evaluation for Vision-Language Understanding in Robotic-Assisted Surgery
by: Ma, Boyi, et al.
Published: (2025)
by: Ma, Boyi, et al.
Published: (2025)
Robot Learning from a Physical World Model
by: Mao, Jiageng, et al.
Published: (2025)
by: Mao, Jiageng, et al.
Published: (2025)
RoboTrustBench: Benchmarking the Trustworthiness of Video World Models for Robotic Manipulation
by: Li, Huiqiong, et al.
Published: (2026)
by: Li, Huiqiong, et al.
Published: (2026)
TimeSpot: Benchmarking Geo-Temporal Understanding in Vision-Language Models in Real-World Settings
by: Wasi, Azmine Toushik, et al.
Published: (2026)
by: Wasi, Azmine Toushik, et al.
Published: (2026)
Towards Realistic Scene Generation with LiDAR Diffusion Models
by: Ran, Haoxi, et al.
Published: (2024)
by: Ran, Haoxi, et al.
Published: (2024)
EgoDyn-Bench: Evaluating Ego-Motion Understanding in Vision-Centric Foundation Models for Autonomous Driving
by: Schäfer, Finn Rasmus, et al.
Published: (2026)
by: Schäfer, Finn Rasmus, et al.
Published: (2026)
EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning
by: Chen, Yi, et al.
Published: (2023)
by: Chen, Yi, et al.
Published: (2023)
Fiducial Exoskeletons: Image-Centric Robot State Estimation
by: Smith, Cameron, et al.
Published: (2026)
by: Smith, Cameron, et al.
Published: (2026)
HRIBench: Benchmarking Vision-Language Models for Real-Time Human Perception in Human-Robot Interaction
by: Shi, Zhonghao, et al.
Published: (2025)
by: Shi, Zhonghao, et al.
Published: (2025)
ArtiBench and ArtiBrain: Benchmarking Generalizable Vision-Language Articulated Object Manipulation
by: Wu, Yuhan, et al.
Published: (2025)
by: Wu, Yuhan, et al.
Published: (2025)
Self-Supervised Geometry-Guided Initialization for Robust Monocular Visual Odometry
by: Kanai, Takayuki, et al.
Published: (2024)
by: Kanai, Takayuki, et al.
Published: (2024)
PhysVLM: Enabling Visual Language Models to Understand Robotic Physical Reachability
by: Zhou, Weijie, et al.
Published: (2025)
by: Zhou, Weijie, et al.
Published: (2025)
POINav: Benchmarking and Enhancing Final-Meters Arrival in Real-World Vision-Language Navigation
by: Gong, Ruiyan, et al.
Published: (2026)
by: Gong, Ruiyan, et al.
Published: (2026)
Virtual Community: An Open World for Humans, Robots, and Society
by: Zhou, Qinhong, et al.
Published: (2025)
by: Zhou, Qinhong, et al.
Published: (2025)
RoboDream: Compositional World Models for Scalable Robot Data Synthesis
by: Ye, Junjie, et al.
Published: (2026)
by: Ye, Junjie, et al.
Published: (2026)
ABot-PhysWorld: Interactive World Foundation Model for Robotic Manipulation with Physics Alignment
by: Chen, Yuzhi, et al.
Published: (2026)
by: Chen, Yuzhi, et al.
Published: (2026)
VLURes: Benchmarking VLM Visual and Linguistic Understanding in Low-Resource Languages
by: Atuhurra, Jesse, et al.
Published: (2025)
by: Atuhurra, Jesse, et al.
Published: (2025)
VULCA-Bench: A Multicultural Vision-Language Benchmark for Evaluating Cultural Understanding
by: Yu, Haorui, et al.
Published: (2026)
by: Yu, Haorui, et al.
Published: (2026)
ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop
by: Hong, Yining, et al.
Published: (2026)
by: Hong, Yining, et al.
Published: (2026)
DivScene: Towards Open-Vocabulary Object Navigation with Large Vision Language Models in Diverse Scenes
by: Wang, Zhaowei, et al.
Published: (2024)
by: Wang, Zhaowei, et al.
Published: (2024)
PhysBrain 1.0 Technical Report
by: Lian, Shijie, et al.
Published: (2026)
by: Lian, Shijie, et al.
Published: (2026)
D-REX: Differentiable Real-to-Sim-to-Real Engine for Learning Dexterous Grasping
by: Lou, Haozhe, et al.
Published: (2026)
by: Lou, Haozhe, et al.
Published: (2026)
FineBench: Benchmarking and Enhancing Vision-Language Models for Fine-grained Human Activity Understanding
by: Faure, Gueter Josmy, et al.
Published: (2026)
by: Faure, Gueter Josmy, et al.
Published: (2026)
Navigating Beyond Instructions: Vision-and-Language Navigation in Obstructed Environments
by: Hong, Haodong, et al.
Published: (2024)
by: Hong, Haodong, et al.
Published: (2024)
Driving Everywhere with Large Language Model Policy Adaptation
by: Li, Boyi, et al.
Published: (2024)
by: Li, Boyi, et al.
Published: (2024)
3D-VLA: A 3D Vision-Language-Action Generative World Model
by: Zhen, Haoyu, et al.
Published: (2024)
by: Zhen, Haoyu, et al.
Published: (2024)
SRPO: Self-Referential Policy Optimization for Vision-Language-Action Models
by: Fei, Senyu, et al.
Published: (2025)
by: Fei, Senyu, et al.
Published: (2025)
LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models
by: Fei, Senyu, et al.
Published: (2025)
by: Fei, Senyu, et al.
Published: (2025)
ETPNav: Evolving Topological Planning for Vision-Language Navigation in Continuous Environments
by: An, Dong, et al.
Published: (2023)
by: An, Dong, et al.
Published: (2023)
RoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for Robotics
by: Song, Chan Hee, et al.
Published: (2024)
by: Song, Chan Hee, et al.
Published: (2024)
VLN-NF: Feasibility-Aware Vision-and-Language Navigation with False-Premise Instructions
by: Su, Hung-Ting, et al.
Published: (2026)
by: Su, Hung-Ting, et al.
Published: (2026)
Mechanistic Finetuning of Vision-Language-Action Models via Few-Shot Demonstrations
by: Mitra, Chancharik, et al.
Published: (2025)
by: Mitra, Chancharik, et al.
Published: (2025)
AnchorDream: Repurposing Video Diffusion for Embodiment-Aware Robot Data Synthesis
by: Ye, Junjie, et al.
Published: (2025)
by: Ye, Junjie, et al.
Published: (2025)
Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation
by: Lu, Jinghui, et al.
Published: (2026)
by: Lu, Jinghui, et al.
Published: (2026)
End-to-End Navigation with Vision Language Models: Transforming Spatial Reasoning into Question-Answering
by: Goetting, Dylan, et al.
Published: (2024)
by: Goetting, Dylan, et al.
Published: (2024)
What Limits Vision-and-Language Navigation ?
by: Wang, Yunheng, et al.
Published: (2026)
by: Wang, Yunheng, et al.
Published: (2026)
Similar Items
-
Learning from Massive Human Videos for Universal Humanoid Pose Control
by: Mao, Jiageng, et al.
Published: (2024) -
PhysToolBench: Benchmarking Physical Tool Understanding for MLLMs
by: Zhang, Zixin, et al.
Published: (2025) -
Robot Learning from Any Images
by: Zhao, Siheng, et al.
Published: (2025) -
A Language Agent for Autonomous Driving
by: Mao, Jiageng, et al.
Published: (2023) -
Can DeepSeek Reason Like a Surgeon? An Empirical Evaluation for Vision-Language Understanding in Robotic-Assisted Surgery
by: Ma, Boyi, et al.
Published: (2025)