PhysReason: A Comprehensive Benchmark towards Physics-Based Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Xinyu, Dong, Yuxuan, Wu, Yanrui, Huang, Jiaxing, Jia, Chengyou, Fernando, Basura, Shou, Mike Zheng, Zhang, Lingling, Liu, Jun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CoFFT: Chain of Foresight-Focus Thought for Visual Language Models
by: Zhang, Xinyu, et al.
Published: (2025)
by: Zhang, Xinyu, et al.
Published: (2025)
Diagram-Driven Course Questions Generation
by: Zhang, Xinyu, et al.
Published: (2024)
by: Zhang, Xinyu, et al.
Published: (2024)
LogicGraph : Benchmarking Multi-Path Logical Reasoning via Neuro-Symbolic Generation and Verification
by: Wu, Yanrui, et al.
Published: (2026)
by: Wu, Yanrui, et al.
Published: (2026)
OptiVerse: A Comprehensive Benchmark towards Optimization Problem Solving
by: Zhang, Xinyu, et al.
Published: (2026)
by: Zhang, Xinyu, et al.
Published: (2026)
RCA: Region Conditioned Adaptation for Visual Abductive Reasoning
by: Zhang, Hao, et al.
Published: (2023)
by: Zhang, Hao, et al.
Published: (2023)
Know-Show: Benchmarking Video-Language Models on Spatio-Temporal Grounded Reasoning
by: Sugandhika, Chinthani, et al.
Published: (2025)
by: Sugandhika, Chinthani, et al.
Published: (2025)
GeoChallenge: A Multi-Answer Multiple-Choice Benchmark for Geometric Reasoning with Diagrams
by: Zhang, Yushun, et al.
Published: (2026)
by: Zhang, Yushun, et al.
Published: (2026)
PhysNote: Self-Knowledge Notes for Evolvable Physical Reasoning in Vision-Language Model
by: Zhang, Sinin, et al.
Published: (2026)
by: Zhang, Sinin, et al.
Published: (2026)
Learning to Reason Iteratively and Parallelly for Complex Visual Reasoning Scenarios
by: Jaiswal, Shantanu, et al.
Published: (2024)
by: Jaiswal, Shantanu, et al.
Published: (2024)
SeePhys: Does Seeing Help Thinking? -- Benchmarking Vision-Based Physics Reasoning
by: Xiang, Kun, et al.
Published: (2025)
by: Xiang, Kun, et al.
Published: (2025)
VProChart: Answering Chart Question through Visual Perception Alignment Agent and Programmatic Solution Reasoning
by: Huang, Muye, et al.
Published: (2024)
by: Huang, Muye, et al.
Published: (2024)
ChainReaction: Causal Chain-Guided Reasoning for Modular and Explainable Causal-Why Video Question Answering
by: Parmar, Paritosh, et al.
Published: (2025)
by: Parmar, Paritosh, et al.
Published: (2025)
Inferring Past Human Actions in Homes with Abductive Reasoning
by: Tan, Clement, et al.
Published: (2022)
by: Tan, Clement, et al.
Published: (2022)
$\textbf{AGT$^{AO}$}$: Robust and Stabilized LLM Unlearning via Adversarial Gating Training with Adaptive Orthogonality
by: Li, Pengyu, et al.
Published: (2026)
by: Li, Pengyu, et al.
Published: (2026)
Benchmarking Chinese Commonsense Reasoning of LLMs: From Chinese-Specifics to Reasoning-Memorization Correlations
by: Sun, Jiaxing, et al.
Published: (2024)
by: Sun, Jiaxing, et al.
Published: (2024)
PhysUniBench: A Multi-Modal Physics Reasoning Benchmark at Undergraduate Level
by: Wang, Lintao, et al.
Published: (2025)
by: Wang, Lintao, et al.
Published: (2025)
BilliardPhys-Bench: Benchmarking Physical Reasoning and Visual Dynamics of Multimodal LLMs
by: Wang, Ben, et al.
Published: (2026)
by: Wang, Ben, et al.
Published: (2026)
PhysInOne: Visual Physics Learning and Reasoning in One Suite
by: Zhou, Siyuan, et al.
Published: (2026)
by: Zhou, Siyuan, et al.
Published: (2026)
MME-Reasoning: A Comprehensive Benchmark for Logical Reasoning in MLLMs
by: Yuan, Jiakang, et al.
Published: (2025)
by: Yuan, Jiakang, et al.
Published: (2025)
Neuro Symbolic Knowledge Reasoning for Procedural Video Question Answering
by: Fernando, Basura, et al.
Published: (2025)
by: Fernando, Basura, et al.
Published: (2025)
Think or Not? Selective Reasoning via Reinforcement Learning for Vision-Language Models
by: Wang, Jiaqi, et al.
Published: (2025)
by: Wang, Jiaqi, et al.
Published: (2025)
FIRE: A Comprehensive Benchmark for Financial Intelligence and Reasoning Evaluation
by: Zhang, Xiyuan, et al.
Published: (2026)
by: Zhang, Xiyuan, et al.
Published: (2026)
FinanceReasoning: Benchmarking Financial Numerical Reasoning More Credible, Comprehensive and Challenging
by: Tang, Zichen, et al.
Published: (2025)
by: Tang, Zichen, et al.
Published: (2025)
AutoGPS: Automated Geometry Problem Solving via Multimodal Formalization and Deductive Reasoning
by: Ping, Bowen, et al.
Published: (2025)
by: Ping, Bowen, et al.
Published: (2025)
OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language Models
by: Jia, Mengdi, et al.
Published: (2025)
by: Jia, Mengdi, et al.
Published: (2025)
Ego-centric Predictive Model Conditioned on Hand Trajectories
by: Zhang, Binjie, et al.
Published: (2025)
by: Zhang, Binjie, et al.
Published: (2025)
GoT-CQA: Graph-of-Thought Guided Compositional Reasoning for Chart Question Answering
by: Zhang, Lingling, et al.
Published: (2024)
by: Zhang, Lingling, et al.
Published: (2024)
BaZi-Based Character Simulation Benchmark: Evaluating AI on Temporal and Persona Reasoning
by: Zheng, Siyuan, et al.
Published: (2025)
by: Zheng, Siyuan, et al.
Published: (2025)
Mitigating Easy Option Bias in Multiple-Choice Question Answering
by: Zhang, Hao, et al.
Published: (2025)
by: Zhang, Hao, et al.
Published: (2025)
EvoChart: A Benchmark and a Self-Training Approach Towards Real-World Chart Understanding
by: Huang, Muye, et al.
Published: (2024)
by: Huang, Muye, et al.
Published: (2024)
FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark
by: Fang, Rongyao, et al.
Published: (2025)
by: Fang, Rongyao, et al.
Published: (2025)
SeePhys Pro: Diagnosing Modality Transfer and Blind-Training Effects in Multimodal RLVR for Physics Reasoning
by: Xiang, Kun, et al.
Published: (2026)
by: Xiang, Kun, et al.
Published: (2026)
Predicting the Next Action by Modeling the Abstract Goal
by: Roy, Debaditya, et al.
Published: (2022)
by: Roy, Debaditya, et al.
Published: (2022)
Solving Math Word Problems via Cooperative Reasoning induced Language Models
by: Zhu, Xinyu, et al.
Published: (2022)
by: Zhu, Xinyu, et al.
Published: (2022)
UGPhysics: A Comprehensive Benchmark for Undergraduate Physics Reasoning with Large Language Models
by: Xu, Xin, et al.
Published: (2025)
by: Xu, Xin, et al.
Published: (2025)
Multimodal Reasoning for Science: Technical Report and 1st Place Solution to the ICML 2025 SeePhys Challenge
by: Liang, Hao, et al.
Published: (2025)
by: Liang, Hao, et al.
Published: (2025)
Your Reasoning Benchmark May Not Test Reasoning: Revealing Perception Bottleneck in Abstract Reasoning Benchmarks
by: Wang, Xinhe, et al.
Published: (2025)
by: Wang, Xinhe, et al.
Published: (2025)
VideoMind: A Chain-of-LoRA Agent for Temporal-Grounded Video Reasoning
by: Liu, Ye, et al.
Published: (2025)
by: Liu, Ye, et al.
Published: (2025)
macOSWorld: A Multilingual Interactive Benchmark for GUI Agents
by: Yang, Pei, et al.
Published: (2025)
by: Yang, Pei, et al.
Published: (2025)
From Recognition to Reasoning: Benchmarking and Enhancing MLLMs on Real-World Receipt Document Understanding
by: Wang, Yandi, et al.
Published: (2026)
by: Wang, Yandi, et al.
Published: (2026)
Similar Items
-
CoFFT: Chain of Foresight-Focus Thought for Visual Language Models
by: Zhang, Xinyu, et al.
Published: (2025) -
Diagram-Driven Course Questions Generation
by: Zhang, Xinyu, et al.
Published: (2024) -
LogicGraph : Benchmarking Multi-Path Logical Reasoning via Neuro-Symbolic Generation and Verification
by: Wu, Yanrui, et al.
Published: (2026) -
OptiVerse: A Comprehensive Benchmark towards Optimization Problem Solving
by: Zhang, Xinyu, et al.
Published: (2026) -
RCA: Region Conditioned Adaptation for Visual Abductive Reasoning
by: Zhang, Hao, et al.
Published: (2023)