PhysToolBench: Benchmarking Physical Tool Understanding for MLLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Zixin, Chen, Kanghao, Lin, Xingwang, Jiang, Lutao, Zheng, Xu, Lyu, Yuanhuiyi, Guo, Litao, Li, Yinchuan, Chen, Ying-Cong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Panoramic Affordance Prediction
by: Zhang, Zixin, et al.
Published: (2026)
by: Zhang, Zixin, et al.
Published: (2026)
A4-Agent: An Agentic Framework for Zero-Shot Affordance Reasoning
by: Zhang, Zixin, et al.
Published: (2025)
by: Zhang, Zixin, et al.
Published: (2025)
T-Rex-Omni: Integrating Negative Visual Prompt in Generic Object Detection
by: Zhou, Jiazhou, et al.
Published: (2025)
by: Zhou, Jiazhou, et al.
Published: (2025)
DiMeR: Disentangled Mesh Reconstruction Model
by: Jiang, Lutao, et al.
Published: (2025)
by: Jiang, Lutao, et al.
Published: (2025)
PhysBench: Benchmarking and Enhancing Vision-Language Models for Physical World Understanding
by: Chow, Wei, et al.
Published: (2025)
by: Chow, Wei, et al.
Published: (2025)
PhysGraph: Physically-Grounded Graph-Transformer Policies for Bimanual Dexterous Hand-Tool-Object Manipulation
by: Li, Runfa Blark, et al.
Published: (2026)
by: Li, Runfa Blark, et al.
Published: (2026)
BrightDreamer: Generic 3D Gaussian Generative Framework for Fast Text-to-3D Synthesis
by: Jiang, Lutao, et al.
Published: (2024)
by: Jiang, Lutao, et al.
Published: (2024)
Chasing Day and Night: Towards Robust and Efficient All-Day Object Detection Guided by an Event Camera
by: Cao, Jiahang, et al.
Published: (2023)
by: Cao, Jiahang, et al.
Published: (2023)
PhysCodeBench: Benchmarking Physics-Aware Symbolic Simulation of 3D Scenes via Self-Corrective Multi-Agent Refinement
by: Xie, Tianyidan, et al.
Published: (2026)
by: Xie, Tianyidan, et al.
Published: (2026)
NeoPhysIx: An Ultra Fast 3D Physical Simulator as Development Tool for AI Algorithms
by: Fischer, Jörn, et al.
Published: (2024)
by: Fischer, Jörn, et al.
Published: (2024)
EIT-1M: One Million EEG-Image-Text Pairs for Human Visual-textual Recognition and More
by: Zheng, Xu, et al.
Published: (2024)
by: Zheng, Xu, et al.
Published: (2024)
ToolEENet: Tool Affordance 6D Pose Estimation
by: Wang, Yunlong, et al.
Published: (2024)
by: Wang, Yunlong, et al.
Published: (2024)
EgoIntent: An Egocentric Step-level Benchmark for Understanding What, Why, and Next
by: Pan, Ye, et al.
Published: (2026)
by: Pan, Ye, et al.
Published: (2026)
PhysBrain: Human Egocentric Data as a Bridge from Vision Language Models to Physical Intelligence
by: Lin, Xiaopeng, et al.
Published: (2025)
by: Lin, Xiaopeng, et al.
Published: (2025)
PhysVLM: Enabling Visual Language Models to Understand Robotic Physical Reachability
by: Zhou, Weijie, et al.
Published: (2025)
by: Zhou, Weijie, et al.
Published: (2025)
MAGIC++: Efficient and Resilient Modality-Agnostic Semantic Segmentation via Hierarchical Modality Selection
by: Zheng, Xu, et al.
Published: (2024)
by: Zheng, Xu, et al.
Published: (2024)
CleanUpBench: Embodied Sweeping and Grasping Benchmark
by: Li, Wenbo, et al.
Published: (2025)
by: Li, Wenbo, et al.
Published: (2025)
Hierarchical Reinforcement Learning for Articulated Tool Manipulation with Multifingered Hand
by: Xu, Wei, et al.
Published: (2025)
by: Xu, Wei, et al.
Published: (2025)
Physics-Conditioned Grasping for Stable Tool Use
by: Trupin, Noah, et al.
Published: (2025)
by: Trupin, Noah, et al.
Published: (2025)
S2R-Bench: A Sim-to-Real Evaluation Benchmark for Autonomous Driving
by: Wang, Li, et al.
Published: (2025)
by: Wang, Li, et al.
Published: (2025)
SFCo-Nav: Efficient Zero-Shot Visual Language Navigation via Collaboration of Slow LLM and Fast Attributed Graph Alignment
by: Xiong, Chaoran, et al.
Published: (2026)
by: Xiong, Chaoran, et al.
Published: (2026)
Adaptive Manipulation Potential and Haptic Estimation for Tool-Mediated Interaction
by: Yang, Lin, et al.
Published: (2026)
by: Yang, Lin, et al.
Published: (2026)
Semantic-Contact Fields for Category-Level Generalizable Tactile Tool Manipulation
by: Ma, Kevin Yuchen, et al.
Published: (2026)
by: Ma, Kevin Yuchen, et al.
Published: (2026)
ST-BiBench: Benchmarking Multi-Stream Multimodal Coordination in Bimanual Embodied Tasks for MLLMs
by: Wu, Xin, et al.
Published: (2026)
by: Wu, Xin, et al.
Published: (2026)
RoboWM-Bench: A Benchmark for Evaluating World Models in Robotic Manipulation
by: Jiang, Feng, et al.
Published: (2026)
by: Jiang, Feng, et al.
Published: (2026)
Elite-EvGS: Learning Event-based 3D Gaussian Splatting by Distilling Event-to-Video Priors
by: Zhang, Zixin, et al.
Published: (2024)
by: Zhang, Zixin, et al.
Published: (2024)
Tool-as-Interface: Learning Robot Policies from Observing Human Tool Use
by: Chen, Haonan, et al.
Published: (2025)
by: Chen, Haonan, et al.
Published: (2025)
Understanding-in-Generation: Reinforcing Generative Capability of Unified Model via Infusing Understanding into Generation
by: Lyu, Yuanhuiyi, et al.
Published: (2025)
by: Lyu, Yuanhuiyi, et al.
Published: (2025)
ABot-PhysWorld: Interactive World Foundation Model for Robotic Manipulation with Physics Alignment
by: Chen, Yuzhi, et al.
Published: (2026)
by: Chen, Yuzhi, et al.
Published: (2026)
CookBench: A Long-Horizon Embodied Planning Benchmark for Complex Cooking Scenarios
by: Cai, Muzhen, et al.
Published: (2025)
by: Cai, Muzhen, et al.
Published: (2025)
MTRDrive: Memory-Tool Synergistic Reasoning for Robust Autonomous Driving in Corner Cases
by: Luo, Ziang, et al.
Published: (2025)
by: Luo, Ziang, et al.
Published: (2025)
Affordance Agent Harness: Verification-Gated Skill Orchestration
by: Huang, Haojian, et al.
Published: (2026)
by: Huang, Haojian, et al.
Published: (2026)
RobotSmith: Generative Robotic Tool Design for Acquisition of Complex Manipulation Skills
by: Lin, Chunru, et al.
Published: (2025)
by: Lin, Chunru, et al.
Published: (2025)
RADAR: Benchmarking Vision-Language-Action Generalization via Real-World Dynamics, Spatial-Physical Intelligence, and Autonomous Evaluation
by: Chen, Yuhao, et al.
Published: (2026)
by: Chen, Yuhao, et al.
Published: (2026)
Enabling Extensible Embodied Capabilities with Tools
by: Zhou, Xueyang, et al.
Published: (2026)
by: Zhou, Xueyang, et al.
Published: (2026)
Exploring Bottlenecks in VLM-LLM Navigation: How 3D Scene Understanding Capability Impacts Zero-Shot VLN
by: Xia, Ziyi, et al.
Published: (2026)
by: Xia, Ziyi, et al.
Published: (2026)
Online Learning for Vibration Suppression in Physical Robot Interaction using Power Tools
by: Solak, Gokhan, et al.
Published: (2025)
by: Solak, Gokhan, et al.
Published: (2025)
SurgiPose: Estimating Surgical Tool Kinematics from Monocular Video for Surgical Robot Learning
by: Chen, Juo-Tung, et al.
Published: (2025)
by: Chen, Juo-Tung, et al.
Published: (2025)
Bench2Drive-Robust: Benchmarking Closed-Loop Autonomous Driving under Deployment Perturbations
by: Zhang, Zhiyuan, et al.
Published: (2026)
by: Zhang, Zhiyuan, et al.
Published: (2026)
Do Open-Loop Metrics Predict Closed-Loop Driving? A Cross-Benchmark Correlation Study of NAVSIM and Bench2Drive
by: Wang, Yiru, et al.
Published: (2026)
by: Wang, Yiru, et al.
Published: (2026)
Similar Items
-
Panoramic Affordance Prediction
by: Zhang, Zixin, et al.
Published: (2026) -
A4-Agent: An Agentic Framework for Zero-Shot Affordance Reasoning
by: Zhang, Zixin, et al.
Published: (2025) -
T-Rex-Omni: Integrating Negative Visual Prompt in Generic Object Detection
by: Zhou, Jiazhou, et al.
Published: (2025) -
DiMeR: Disentangled Mesh Reconstruction Model
by: Jiang, Lutao, et al.
Published: (2025) -
PhysBench: Benchmarking and Enhancing Vision-Language Models for Physical World Understanding
by: Chow, Wei, et al.
Published: (2025)