FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Zeyu, Xu, Jingye, Li, Xiaogang, Xiao, Peiyao, Kong, Qinhao, Wang, Ben, Xu, Chengliang, Chen, Zichao, Zhao, Bing, Wei, Hu |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
BilliardPhys-Bench: Benchmarking Physical Reasoning and Visual Dynamics of Multimodal LLMs
by: Wang, Ben, et al.
Published: (2026)
by: Wang, Ben, et al.
Published: (2026)
SPM-Bench: Benchmarking Large Language Models for Scanning Probe Microscopy
by: Xiao, Peiyao, et al.
Published: (2026)
by: Xiao, Peiyao, et al.
Published: (2026)
CrystalXRD-Bench: Benchmarking Vision-Language Models for XRD Peak Indexing Across Diverse Crystalline Materials
by: Xu, Chengliang, et al.
Published: (2026)
by: Xu, Chengliang, et al.
Published: (2026)
MLLM-CompBench: A Comparative Reasoning Benchmark for Multimodal LLMs
by: Kil, Jihyung, et al.
Published: (2024)
by: Kil, Jihyung, et al.
Published: (2024)
Exploiting Parallelism for Fast Feynman Diagrammatics
by: Sturt, John, et al.
Published: (2024)
by: Sturt, John, et al.
Published: (2024)
EarthSpatialBench: Benchmarking Spatial Reasoning Capabilities of Multimodal LLMs on Earth Imagery
by: Xu, Zelin, et al.
Published: (2026)
by: Xu, Zelin, et al.
Published: (2026)
ORIGAMISPACE: Benchmarking Multimodal LLMs in Multi-Step Spatial Reasoning with Mathematical Constraints
by: Xu, Rui, et al.
Published: (2025)
by: Xu, Rui, et al.
Published: (2025)
IV-Bench: A Benchmark for Image-Grounded Video Perception and Reasoning in Multimodal LLMs
by: Ma, David, et al.
Published: (2025)
by: Ma, David, et al.
Published: (2025)
oMeBench: Towards Robust Benchmarking of LLMs in Organic Mechanism Elucidation and Reasoning
by: Xu, Ruiling, et al.
Published: (2025)
by: Xu, Ruiling, et al.
Published: (2025)
DPrivBench: Benchmarking LLMs' Reasoning for Differential Privacy
by: Wang, Erchi, et al.
Published: (2026)
by: Wang, Erchi, et al.
Published: (2026)
LocalBench: Benchmarking LLMs on County-Level Local Knowledge and Reasoning
by: Gao, Zihan, et al.
Published: (2025)
by: Gao, Zihan, et al.
Published: (2025)
CLM-Bench: Benchmarking and Analyzing Cross-lingual Misalignment of LLMs in Knowledge Editing
by: Hu, Yucheng, et al.
Published: (2026)
by: Hu, Yucheng, et al.
Published: (2026)
FaithCoT-Bench: Benchmarking Instance-Level Faithfulness of Chain-of-Thought Reasoning
by: Shen, Xu, et al.
Published: (2025)
by: Shen, Xu, et al.
Published: (2025)
MetaBench: A Multi-task Benchmark for Assessing LLMs in Metabolomics
by: Lu, Yuxing, et al.
Published: (2025)
by: Lu, Yuxing, et al.
Published: (2025)
AccidentBench: Benchmarking Multimodal Understanding and Reasoning in Vehicle Accidents and Beyond
by: Gu, Shangding, et al.
Published: (2025)
by: Gu, Shangding, et al.
Published: (2025)
Unmasking Reasoning Processes: A Process-aware Benchmark for Evaluating Structural Mathematical Reasoning in LLMs
by: Zheng, Xiang, et al.
Published: (2026)
by: Zheng, Xiang, et al.
Published: (2026)
HiBench: Benchmarking LLMs Capability on Hierarchical Structure Reasoning
by: Jiang, Zhuohang, et al.
Published: (2025)
by: Jiang, Zhuohang, et al.
Published: (2025)
DNR Bench: Benchmarking Over-Reasoning in Reasoning LLMs
by: Hashemi, Masoud, et al.
Published: (2025)
by: Hashemi, Masoud, et al.
Published: (2025)
Evaluating 21st-Century Competencies in Postsecondary Curricula with Large Language Models: Performance Benchmarking and Reasoning-Based Prompting Strategies
by: Xu, Zhen, et al.
Published: (2026)
by: Xu, Zhen, et al.
Published: (2026)
CSR-Bench: Benchmarking LLM Agents in Deployment of Computer Science Research Repositories
by: Xiao, Yijia, et al.
Published: (2025)
by: Xiao, Yijia, et al.
Published: (2025)
CodeSpecBench: Benchmarking LLMs for Executable Behavioral Specification Generation
by: Chen, Zaoyu, et al.
Published: (2026)
by: Chen, Zaoyu, et al.
Published: (2026)
AQA-Bench: An Interactive Benchmark for Evaluating LLMs' Sequential Reasoning Ability
by: Yang, Siwei, et al.
Published: (2024)
by: Yang, Siwei, et al.
Published: (2024)
Pervaporation recovery of aniline using the degradable material poly(butylene adipate‐terephthalate): Experiment and molecular simulation
by: Haozhe Wang, et al.
Published: (2024)
by: Haozhe Wang, et al.
Published: (2024)
CARE-Bench: A Benchmark of Diverse Client Simulations Guided by Expert Principles for Evaluating LLMs in Psychological Counseling
by: Wang, Bichen, et al.
Published: (2025)
by: Wang, Bichen, et al.
Published: (2025)
DeepVision-103K: A Visually Diverse, Broad-Coverage, and Verifiable Mathematical Dataset for Multimodal Reasoning
by: Sun, Haoxiang, et al.
Published: (2026)
by: Sun, Haoxiang, et al.
Published: (2026)
The Gordian Knot for VLMs: Diagrammatic Knot Reasoning as a Hard Benchmark
by: Liu, Hao, et al.
Published: (2026)
by: Liu, Hao, et al.
Published: (2026)
CriticBench: Benchmarking LLMs for Critique-Correct Reasoning
by: Lin, Zicheng, et al.
Published: (2024)
by: Lin, Zicheng, et al.
Published: (2024)
TopoBench: Benchmarking LLMs on Hard Topological Reasoning
by: Maniparambil, Mayug, et al.
Published: (2026)
by: Maniparambil, Mayug, et al.
Published: (2026)
MMDG-Bench: A Benchmark for Multimodal Domain Generalization
by: Zhan, Qianshan, et al.
Published: (2026)
by: Zhan, Qianshan, et al.
Published: (2026)
XFacta: Contemporary, Real-World Dataset and Evaluation for Multimodal Misinformation Detection with Multimodal LLMs
by: Xiao, Yuzhuo, et al.
Published: (2025)
by: Xiao, Yuzhuo, et al.
Published: (2025)
SciRerankBench: Benchmarking Rerankers Towards Scientific Retrieval-Augmented Generated LLMs
by: Chen, Haotian, et al.
Published: (2025)
by: Chen, Haotian, et al.
Published: (2025)
Enhancing the General Agent Capabilities of Low-Parameter LLMs through Tuning and Multi-Branch Reasoning
by: Zhou, Qinhao, et al.
Published: (2024)
by: Zhou, Qinhao, et al.
Published: (2024)
HAIBU-ReMUD: Reasoning Multimodal Ultrasound Dataset and Model Bridging to General Specific Domains
by: Wang, Shijie, et al.
Published: (2025)
by: Wang, Shijie, et al.
Published: (2025)
MCiteBench: A Multimodal Benchmark for Generating Text with Citations
by: Hu, Caiyu, et al.
Published: (2025)
by: Hu, Caiyu, et al.
Published: (2025)
SpatialBench: Benchmarking Multimodal Large Language Models for Spatial Cognition
by: Xu, Peiran, et al.
Published: (2025)
by: Xu, Peiran, et al.
Published: (2025)
PRL-Bench: A Comprehensive Benchmark Evaluating LLMs' Capabilities in Frontier Physics Research
by: Miao, Tingjia, et al.
Published: (2026)
by: Miao, Tingjia, et al.
Published: (2026)
CoS++: Towards More General and Explicit Implementations for Sampling High-Order Feynman Diagrammatic Series
by: Shi, Boyuan
Published: (2025)
by: Shi, Boyuan
Published: (2025)
ChartBench: A Benchmark for Complex Visual Reasoning in Charts
by: Xu, Zhengzhuo, et al.
Published: (2023)
by: Xu, Zhengzhuo, et al.
Published: (2023)
SpineBench: Benchmarking Multimodal LLMs for Spinal Pathology Analysis
by: Zhang, Chenghanyu, et al.
Published: (2025)
by: Zhang, Chenghanyu, et al.
Published: (2025)
Benchmarking Multimodal Mathematical Reasoning with Explicit Visual Dependency
by: Wang, Zhikai, et al.
Published: (2025)
by: Wang, Zhikai, et al.
Published: (2025)
Similar Items
-
BilliardPhys-Bench: Benchmarking Physical Reasoning and Visual Dynamics of Multimodal LLMs
by: Wang, Ben, et al.
Published: (2026) -
SPM-Bench: Benchmarking Large Language Models for Scanning Probe Microscopy
by: Xiao, Peiyao, et al.
Published: (2026) -
CrystalXRD-Bench: Benchmarking Vision-Language Models for XRD Peak Indexing Across Diverse Crystalline Materials
by: Xu, Chengliang, et al.
Published: (2026) -
MLLM-CompBench: A Comparative Reasoning Benchmark for Multimodal LLMs
by: Kil, Jihyung, et al.
Published: (2024) -
Exploiting Parallelism for Fast Feynman Diagrammatics
by: Sturt, John, et al.
Published: (2024)