SymPyBench: A Dynamic Benchmark for Scientific Reasoning with Executable Python Code
Fuente:
arXiv
Saved in:
| Main Authors: | Imani, Shima, Moon, Seungwhan, Ahmadyan, Adel, Zhang, Lu, Ahmed, Kirmani, Damavandi, Babak |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PRiSM: An Agentic Multimodal Benchmark for Scientific Reasoning via Python-Grounded Evaluation
by: Imani, Shima, et al.
Published: (2025)
by: Imani, Shima, et al.
Published: (2025)
TRACE: A Framework for Analyzing and Enhancing Stepwise Reasoning in Vision-Language Models
by: Imani, Shima, et al.
Published: (2025)
by: Imani, Shima, et al.
Published: (2025)
Doppelgänger's Watch: A Split Objective Approach to Large Language Models
by: Ghasemlou, Shervin, et al.
Published: (2024)
by: Ghasemlou, Shervin, et al.
Published: (2024)
WearVQA: A Visual Question Answering Benchmark for Wearables in Egocentric Authentic Real-world scenarios
by: Chang, Eun, et al.
Published: (2025)
by: Chang, Eun, et al.
Published: (2025)
Proactive Assistant Dialogue Generation from Streaming Egocentric Videos
by: Zhang, Yichi, et al.
Published: (2025)
by: Zhang, Yichi, et al.
Published: (2025)
CRUXEval-X: A Benchmark for Multilingual Code Reasoning, Understanding and Execution
by: Xu, Ruiyang, et al.
Published: (2024)
by: Xu, Ruiyang, et al.
Published: (2024)
AInsteinBench: Benchmarking Coding Agents on Scientific Repositories
by: Duston, Titouan, et al.
Published: (2025)
by: Duston, Titouan, et al.
Published: (2025)
ResearchEnvBench: Benchmarking Agents on Environment Synthesis for Research Code Execution
by: Wang, Yubang, et al.
Published: (2026)
by: Wang, Yubang, et al.
Published: (2026)
An Execution-Verified Multi-Language Benchmark for Code Semantic Reasoning
by: Li, Yikun, et al.
Published: (2026)
by: Li, Yikun, et al.
Published: (2026)
$A^3$-Bench: Benchmarking Memory-Driven Scientific Reasoning via Anchor and Attractor Activation
by: Zhang, Jian, et al.
Published: (2026)
by: Zhang, Jian, et al.
Published: (2026)
GeoAgentBench: A Dynamic Execution Benchmark for Tool-Augmented Agents in Spatial Analysis
by: Yu, Bo, et al.
Published: (2026)
by: Yu, Bo, et al.
Published: (2026)
DW-Bench: Benchmarking LLMs on Data Warehouse Graph Topology Reasoning
by: Ahmed, Ahmed G. A. H, et al.
Published: (2026)
by: Ahmed, Ahmed G. A. H, et al.
Published: (2026)
CRUXEval: A Benchmark for Code Reasoning, Understanding and Execution
by: Gu, Alex, et al.
Published: (2024)
by: Gu, Alex, et al.
Published: (2024)
FasterPy: An LLM-based Code Execution Efficiency Optimization Framework
by: Wu, Yue, et al.
Published: (2025)
by: Wu, Yue, et al.
Published: (2025)
Benchmarking Scientific Understanding and Reasoning for Video Generation using VideoScience-Bench
by: Hu, Lanxiang, et al.
Published: (2025)
by: Hu, Lanxiang, et al.
Published: (2025)
SymGPT: Auditing Smart Contracts via Combining Symbolic Execution with Large Language Models
by: Xia, Shihao, et al.
Published: (2025)
by: Xia, Shihao, et al.
Published: (2025)
Code2Bench: Scaling Source and Rigor for Dynamic Benchmark Construction
by: Zhang, Zhe, et al.
Published: (2025)
by: Zhang, Zhe, et al.
Published: (2025)
SciVideoBench: Benchmarking Scientific Video Reasoning in Large Multimodal Models
by: Deng, Andong, et al.
Published: (2025)
by: Deng, Andong, et al.
Published: (2025)
XDomainBench: Diagnosing Reasoning Collapse in High-Dimensional Scientific Knowledge Composition
by: Zhiren, Gong, et al.
Published: (2026)
by: Zhiren, Gong, et al.
Published: (2026)
PyGDA: A Python Library for Graph Domain Adaptation
by: Zhang, Zhen, et al.
Published: (2025)
by: Zhang, Zhen, et al.
Published: (2025)
DyPyBench: A Benchmark of Executable Python Software
by: Bouzenia, Islem, et al.
Published: (2024)
by: Bouzenia, Islem, et al.
Published: (2024)
NewtonBench: Benchmarking Generalizable Scientific Law Discovery in LLM Agents
by: Zheng, Tianshi, et al.
Published: (2025)
by: Zheng, Tianshi, et al.
Published: (2025)
PyCSP3: Modeling Combinatorial Constrained Problems in Python
by: Lecoutre, Christophe, et al.
Published: (2020)
by: Lecoutre, Christophe, et al.
Published: (2020)
OpenClawBench: Benchmarking Process-side Anomalies in Real-world Agent Execution Trajectories
by: Liu, Yibing, et al.
Published: (2026)
by: Liu, Yibing, et al.
Published: (2026)
AutoResearchBench: Benchmarking AI Agents on Complex Scientific Literature Discovery
by: Xiong, Lei, et al.
Published: (2026)
by: Xiong, Lei, et al.
Published: (2026)
RedCode: Risky Code Execution and Generation Benchmark for Code Agents
by: Guo, Chengquan, et al.
Published: (2024)
by: Guo, Chengquan, et al.
Published: (2024)
SymRTLO: Enhancing RTL Code Optimization with LLMs and Neuron-Inspired Symbolic Reasoning
by: Wang, Yiting, et al.
Published: (2025)
by: Wang, Yiting, et al.
Published: (2025)
WARC-Bench: Web Archive Based Benchmark for GUI Subtask Executions
by: Srivastava, Sanjari, et al.
Published: (2025)
by: Srivastava, Sanjari, et al.
Published: (2025)
MorphoBench: A Benchmark with Difficulty Adaptive to Model Reasoning
by: Wang, Xukai, et al.
Published: (2025)
by: Wang, Xukai, et al.
Published: (2025)
Code Execution as Grounded Supervision for LLM Reasoning
by: Jung, Dongwon, et al.
Published: (2025)
by: Jung, Dongwon, et al.
Published: (2025)
CGBench: Benchmarking Language Model Scientific Reasoning for Clinical Genetics Research
by: Queen, Owen, et al.
Published: (2025)
by: Queen, Owen, et al.
Published: (2025)
VisCoder: Fine-Tuning LLMs for Executable Python Visualization Code Generation
by: Ni, Yuansheng, et al.
Published: (2025)
by: Ni, Yuansheng, et al.
Published: (2025)
PyFCG: Fluid Construction Grammar in Python
by: Van Eecke, Paul, et al.
Published: (2025)
by: Van Eecke, Paul, et al.
Published: (2025)
BilliardPhys-Bench: Benchmarking Physical Reasoning and Visual Dynamics of Multimodal LLMs
by: Wang, Ben, et al.
Published: (2026)
by: Wang, Ben, et al.
Published: (2026)
PythonSaga: Redefining the Benchmark to Evaluate Code Generating LLMs
by: Yadav, Ankit, et al.
Published: (2024)
by: Yadav, Ankit, et al.
Published: (2024)
FeatureBench: Benchmarking Agentic Coding for Complex Feature Development
by: Zhou, Qixing, et al.
Published: (2026)
by: Zhou, Qixing, et al.
Published: (2026)
PyMilo: A Python Library for ML I/O
by: Rostami, AmirHosein, et al.
Published: (2024)
by: Rostami, AmirHosein, et al.
Published: (2024)
FEM-Bench: A Structured Scientific Reasoning Benchmark for Evaluating Code-Generating LLMs
by: Mohammadzadeh, Saeed, et al.
Published: (2025)
by: Mohammadzadeh, Saeed, et al.
Published: (2025)
VerifyBench: A Systematic Benchmark for Evaluating Reasoning Verifiers Across Domains
by: Li, Xuzhao, et al.
Published: (2025)
by: Li, Xuzhao, et al.
Published: (2025)
FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning
by: Wang, Zeyu, et al.
Published: (2026)
by: Wang, Zeyu, et al.
Published: (2026)
Similar Items
-
PRiSM: An Agentic Multimodal Benchmark for Scientific Reasoning via Python-Grounded Evaluation
by: Imani, Shima, et al.
Published: (2025) -
TRACE: A Framework for Analyzing and Enhancing Stepwise Reasoning in Vision-Language Models
by: Imani, Shima, et al.
Published: (2025) -
Doppelgänger's Watch: A Split Objective Approach to Large Language Models
by: Ghasemlou, Shervin, et al.
Published: (2024) -
WearVQA: A Visual Question Answering Benchmark for Wearables in Egocentric Authentic Real-world scenarios
by: Chang, Eun, et al.
Published: (2025) -
Proactive Assistant Dialogue Generation from Streaming Egocentric Videos
by: Zhang, Yichi, et al.
Published: (2025)