PhysUniBench: A Multi-Modal Physics Reasoning Benchmark at Undergraduate Level
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Lintao, Su, Encheng, Liu, Jiaqi, Li, Pengze, Xiao, Jiabei, Zhang, Wenlong, Dai, Xinnan, Chen, Xi, Meng, Yuan, Bai, Lei, Ouyang, Wanli, Tang, Shixiang, Wang, Aoran, Ma, Xinzhu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SciIF: Benchmarking Scientific Instruction Following Towards Rigorous Scientific Intelligence
by: Su, Encheng, et al.
Published: (2026)
by: Su, Encheng, et al.
Published: (2026)
S3Mem: Structured Spatiotemporal Scene-Event Memory for Long-Horizon Interactive Question Answering
by: Su, Encheng, et al.
Published: (2026)
by: Su, Encheng, et al.
Published: (2026)
ARCHE: A Novel Task to Evaluate LLMs on Latent Reasoning Chain Extraction
by: Li, Pengze, et al.
Published: (2025)
by: Li, Pengze, et al.
Published: (2025)
CPRet: A Dataset, Benchmark, and Model for Retrieval in Competitive Programming
by: Deng, Han, et al.
Published: (2025)
by: Deng, Han, et al.
Published: (2025)
UniSTD: Towards Unified Spatio-Temporal Learning across Diverse Disciplines
by: Tang, Chen, et al.
Published: (2025)
by: Tang, Chen, et al.
Published: (2025)
SciReasoner: Laying the Scientific Reasoning Ground Across Disciplines
by: Wang, Yizhou, et al.
Published: (2025)
by: Wang, Yizhou, et al.
Published: (2025)
Charting Empirical Laws for LLM Fine-Tuning in Scientific Multi-Discipline Learning
by: Wang, Lintao, et al.
Published: (2026)
by: Wang, Lintao, et al.
Published: (2026)
Mimicking the Physicist's Eye:A VLM-centric Approach for Physics Formula Discovery
by: Liu, Jiaqi, et al.
Published: (2025)
by: Liu, Jiaqi, et al.
Published: (2025)
Hidden in Plain Sight: Visual-to-Symbolic Analytical Solution Inference from Field Visualizations
by: Li, Pengze, et al.
Published: (2026)
by: Li, Pengze, et al.
Published: (2026)
UniREditBench: A Unified Reasoning-based Image Editing Benchmark
by: Han, Feng, et al.
Published: (2025)
by: Han, Feng, et al.
Published: (2025)
From Sequence to Structure: Uncovering Substructure Reasoning in Transformers
by: Dai, Xinnan, et al.
Published: (2025)
by: Dai, Xinnan, et al.
Published: (2025)
Pixel2Phys: Distilling Governing Laws from Visual Dynamics
by: Li, Ruikun, et al.
Published: (2026)
by: Li, Ruikun, et al.
Published: (2026)
BilliardPhys-Bench: Benchmarking Physical Reasoning and Visual Dynamics of Multimodal LLMs
by: Wang, Ben, et al.
Published: (2026)
by: Wang, Ben, et al.
Published: (2026)
CMT: A Cascade MAR with Topology Predictor for Multimodal Conditional CAD Generation
by: Wu, Jianyu, et al.
Published: (2025)
by: Wu, Jianyu, et al.
Published: (2025)
ComfyBench: Benchmarking LLM-based Agents in ComfyUI for Autonomously Designing Collaborative AI Systems
by: Xue, Xiangyuan, et al.
Published: (2024)
by: Xue, Xiangyuan, et al.
Published: (2024)
PredBench: Benchmarking Spatio-Temporal Prediction across Diverse Disciplines
by: Wang, ZiDong, et al.
Published: (2024)
by: Wang, ZiDong, et al.
Published: (2024)
3D Object Detection from Images for Autonomous Driving: A Survey
by: Ma, Xinzhu, et al.
Published: (2022)
by: Ma, Xinzhu, et al.
Published: (2022)
ResearchBench: Benchmarking LLMs in Scientific Discovery via Inspiration-Based Task Decomposition
by: Liu, Yujie, et al.
Published: (2025)
by: Liu, Yujie, et al.
Published: (2025)
UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models
by: Xu, Xin, et al.
Published: (2025)
by: Xu, Xin, et al.
Published: (2025)
UniGenBench++: A Unified Semantic Evaluation Benchmark for Text-to-Image Generation
by: Wang, Yibin, et al.
Published: (2025)
by: Wang, Yibin, et al.
Published: (2025)
OmniMatBench: A Human-Calibrated Multimodal Reasoning Benchmark Across 19 Materials Science Subfields
by: Liu, Wanhao, et al.
Published: (2026)
by: Liu, Wanhao, et al.
Published: (2026)
Exploring Graph Learning Tasks with Pure LLMs: A Comprehensive Benchmark and Investigation
by: Wang, Yuxiang, et al.
Published: (2025)
by: Wang, Yuxiang, et al.
Published: (2025)
AdaBrain-Bench: Benchmarking Brain Foundation Models for Brain-Computer Interface Applications
by: Wu, Jiamin, et al.
Published: (2025)
by: Wu, Jiamin, et al.
Published: (2025)
Recurrent Neural Goodness-of-Fit Test for Time Series
by: Zhang, Aoran, et al.
Published: (2024)
by: Zhang, Aoran, et al.
Published: (2024)
PhysBench: Benchmarking and Enhancing Vision-Language Models for Physical World Understanding
by: Chow, Wei, et al.
Published: (2025)
by: Chow, Wei, et al.
Published: (2025)
UniGRPO: Unified Policy Optimization for Reasoning-Driven Visual Generation
by: Liu, Jie, et al.
Published: (2026)
by: Liu, Jie, et al.
Published: (2026)
Solving all laminar flows around airfoils all-at-once using a parametric neural network solver
by: Cao, Wenbo, et al.
Published: (2025)
by: Cao, Wenbo, et al.
Published: (2025)
MLLM-based Discovery of Intrinsic Coordinates and Governing Equations from High-Dimensional Data
by: Li, Ruikun, et al.
Published: (2025)
by: Li, Ruikun, et al.
Published: (2025)
FaithCoT-Bench: Benchmarking Instance-Level Faithfulness of Chain-of-Thought Reasoning
by: Shen, Xu, et al.
Published: (2025)
by: Shen, Xu, et al.
Published: (2025)
PhysToolBench: Benchmarking Physical Tool Understanding for MLLMs
by: Zhang, Zixin, et al.
Published: (2025)
by: Zhang, Zixin, et al.
Published: (2025)
Progressive Modality Cooperation for Multi-Modality Domain Adaptation
by: Zhang, Weichen, et al.
Published: (2025)
by: Zhang, Weichen, et al.
Published: (2025)
UniPAD: A Universal Pre-training Paradigm for Autonomous Driving
by: Yang, Honghui, et al.
Published: (2023)
by: Yang, Honghui, et al.
Published: (2023)
Human-Centric Foundation Models: Perception, Generation and Agentic Modeling
by: Tang, Shixiang, et al.
Published: (2025)
by: Tang, Shixiang, et al.
Published: (2025)
Learning Geometry-Guided Depth via Projective Modeling for Monocular 3D Object Detection
by: Zhang, Yinmin, et al.
Published: (2021)
by: Zhang, Yinmin, et al.
Published: (2021)
SynCast: Synergizing Contradictions in Precipitation Nowcasting via Diffusion Sequential Preference Optimization
by: Xu, Kaiyi, et al.
Published: (2025)
by: Xu, Kaiyi, et al.
Published: (2025)
Model Decides How to Tokenize: Adaptive DNA Sequence Tokenization with MxDNA
by: Qiao, Lifeng, et al.
Published: (2024)
by: Qiao, Lifeng, et al.
Published: (2024)
Understand Before You Generate: Self-Guided Training for Autoregressive Image Generation
by: Yue, Xiaoyu, et al.
Published: (2025)
by: Yue, Xiaoyu, et al.
Published: (2025)
AFBench: A Large-scale Benchmark for Airfoil Design
by: Liu, Jian, et al.
Published: (2024)
by: Liu, Jian, et al.
Published: (2024)
PhysReason: A Comprehensive Benchmark towards Physics-Based Reasoning
by: Zhang, Xinyu, et al.
Published: (2025)
by: Zhang, Xinyu, et al.
Published: (2025)
Referring Video Object Segmentation with Cross-Modality Proxy Queries
by: Sun, Baoli, et al.
Published: (2025)
by: Sun, Baoli, et al.
Published: (2025)
Similar Items
-
SciIF: Benchmarking Scientific Instruction Following Towards Rigorous Scientific Intelligence
by: Su, Encheng, et al.
Published: (2026) -
S3Mem: Structured Spatiotemporal Scene-Event Memory for Long-Horizon Interactive Question Answering
by: Su, Encheng, et al.
Published: (2026) -
ARCHE: A Novel Task to Evaluate LLMs on Latent Reasoning Chain Extraction
by: Li, Pengze, et al.
Published: (2025) -
CPRet: A Dataset, Benchmark, and Model for Retrieval in Competitive Programming
by: Deng, Han, et al.
Published: (2025) -
UniSTD: Towards Unified Spatio-Temporal Learning across Diverse Disciplines
by: Tang, Chen, et al.
Published: (2025)