PRiSM: An Agentic Multimodal Benchmark for Scientific Reasoning via Python-Grounded Evaluation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Imani, Shima, Moon, Seungwhan, Ahmadyan, Adel, Zhang, Lu, Ahmed, Kirmani, Damavandi, Babak |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SymPyBench: A Dynamic Benchmark for Scientific Reasoning with Executable Python Code
von: Imani, Shima, et al.
Veröffentlicht: (2025)
von: Imani, Shima, et al.
Veröffentlicht: (2025)
TRACE: A Framework for Analyzing and Enhancing Stepwise Reasoning in Vision-Language Models
von: Imani, Shima, et al.
Veröffentlicht: (2025)
von: Imani, Shima, et al.
Veröffentlicht: (2025)
PRiSM: Benchmarking Phone Realization in Speech Models
von: Bharadwaj, Shikhar, et al.
Veröffentlicht: (2026)
von: Bharadwaj, Shikhar, et al.
Veröffentlicht: (2026)
Doppelgänger's Watch: A Split Objective Approach to Large Language Models
von: Ghasemlou, Shervin, et al.
Veröffentlicht: (2024)
von: Ghasemlou, Shervin, et al.
Veröffentlicht: (2024)
WearVQA: A Visual Question Answering Benchmark for Wearables in Egocentric Authentic Real-world scenarios
von: Chang, Eun, et al.
Veröffentlicht: (2025)
von: Chang, Eun, et al.
Veröffentlicht: (2025)
Proactive Assistant Dialogue Generation from Streaming Egocentric Videos
von: Zhang, Yichi, et al.
Veröffentlicht: (2025)
von: Zhang, Yichi, et al.
Veröffentlicht: (2025)
PaperArena: An Evaluation Benchmark for Tool-Augmented Agentic Reasoning on Scientific Literature
von: Wang, Daoyu, et al.
Veröffentlicht: (2025)
von: Wang, Daoyu, et al.
Veröffentlicht: (2025)
Pixel-Grounded Retrieval for Knowledgeable Large Multimodal Models
von: Kim, Jeonghwan, et al.
Veröffentlicht: (2026)
von: Kim, Jeonghwan, et al.
Veröffentlicht: (2026)
OMNIFLOW: A Physics-Grounded Multimodal Agent for Generalized Scientific Reasoning
von: Wu, Hao, et al.
Veröffentlicht: (2026)
von: Wu, Hao, et al.
Veröffentlicht: (2026)
LABSHIELD: A Multimodal Benchmark for Safety-Critical Reasoning and Planning in Scientific Laboratories
von: Sun, Qianpu, et al.
Veröffentlicht: (2026)
von: Sun, Qianpu, et al.
Veröffentlicht: (2026)
Reasoning With a Star: A Heliophysics Dataset and Benchmark for Agentic Scientific Reasoning
von: Lee, Kevin, et al.
Veröffentlicht: (2025)
von: Lee, Kevin, et al.
Veröffentlicht: (2025)
Agentic Spatio-Temporal Grounding via Collaborative Reasoning
von: Zhao, Heng, et al.
Veröffentlicht: (2026)
von: Zhao, Heng, et al.
Veröffentlicht: (2026)
Benchmarking Agentic Systems in Automated Scientific Information Extraction with ChemX
von: Vepreva, Anastasia, et al.
Veröffentlicht: (2025)
von: Vepreva, Anastasia, et al.
Veröffentlicht: (2025)
STEER: Flexible Robotic Manipulation via Dense Language Grounding
von: Smith, Laura, et al.
Veröffentlicht: (2024)
von: Smith, Laura, et al.
Veröffentlicht: (2024)
Spatial-Agent: Agentic Geo-spatial Reasoning with Scientific Core Concepts
von: Bao, Riyang, et al.
Veröffentlicht: (2026)
von: Bao, Riyang, et al.
Veröffentlicht: (2026)
MC-Search: Evaluating and Enhancing Multimodal Agentic Search with Structured Long Reasoning Chains
von: Ning, Xuying, et al.
Veröffentlicht: (2026)
von: Ning, Xuying, et al.
Veröffentlicht: (2026)
An Agentic Evaluation Framework for AI-Generated Scientific Code in PETSc
von: Zhang, Hong, et al.
Veröffentlicht: (2026)
von: Zhang, Hong, et al.
Veröffentlicht: (2026)
AMIGO: Agentic Multi-Image Grounding Oracle Benchmark
von: Wang, Min, et al.
Veröffentlicht: (2026)
von: Wang, Min, et al.
Veröffentlicht: (2026)
BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games
von: Paglieri, Davide, et al.
Veröffentlicht: (2024)
von: Paglieri, Davide, et al.
Veröffentlicht: (2024)
DeepPHY: Benchmarking Agentic VLMs on Physical Reasoning
von: Xu, Xinrun, et al.
Veröffentlicht: (2025)
von: Xu, Xinrun, et al.
Veröffentlicht: (2025)
The Path Ahead for Agentic AI: Challenges and Opportunities
von: Sibai, Nadia, et al.
Veröffentlicht: (2026)
von: Sibai, Nadia, et al.
Veröffentlicht: (2026)
SciVideoBench: Benchmarking Scientific Video Reasoning in Large Multimodal Models
von: Deng, Andong, et al.
Veröffentlicht: (2025)
von: Deng, Andong, et al.
Veröffentlicht: (2025)
Seeing and Reasoning with Confidence: Supercharging Multimodal LLMs with an Uncertainty-Aware Agentic Framework
von: Zhi, Zhuo, et al.
Veröffentlicht: (2025)
von: Zhi, Zhuo, et al.
Veröffentlicht: (2025)
Agentic Scientific Simulation: Execution-Grounded Model Construction and Reconstruction
von: Lie, Knut-Andreas, et al.
Veröffentlicht: (2026)
von: Lie, Knut-Andreas, et al.
Veröffentlicht: (2026)
CARV: A Diagnostic Benchmark for Compositional Analogical Reasoning in Multimodal LLMs
von: Du, Yongkang, et al.
Veröffentlicht: (2026)
von: Du, Yongkang, et al.
Veröffentlicht: (2026)
$A^3$-Bench: Benchmarking Memory-Driven Scientific Reasoning via Anchor and Attractor Activation
von: Zhang, Jian, et al.
Veröffentlicht: (2026)
von: Zhang, Jian, et al.
Veröffentlicht: (2026)
MolQuest: A Benchmark for Agentic Evaluation of Abductive Reasoning in Chemical Structure Elucidation
von: Han, Taolin, et al.
Veröffentlicht: (2026)
von: Han, Taolin, et al.
Veröffentlicht: (2026)
Octopus: Agentic Multimodal Reasoning with Six-Capability Orchestration
von: Guo, Yifu, et al.
Veröffentlicht: (2025)
von: Guo, Yifu, et al.
Veröffentlicht: (2025)
Grounding LLMs in Scientific Discovery via Embodied Actions
von: Zhang, Bo, et al.
Veröffentlicht: (2026)
von: Zhang, Bo, et al.
Veröffentlicht: (2026)
CARE: Towards Clinical Accountability in Multi-Modal Medical Reasoning with an Evidence-Grounded Agentic Framework
von: Du, Yuexi, et al.
Veröffentlicht: (2026)
von: Du, Yuexi, et al.
Veröffentlicht: (2026)
From Prompts to Pavement Through Time: Temporal Grounding in Agentic Scene-to-Plan Reasoning
von: Gado, Ahmed Y., et al.
Veröffentlicht: (2026)
von: Gado, Ahmed Y., et al.
Veröffentlicht: (2026)
Measuring What Matters: Benchmarking Generative, Multimodal, and Agentic AI in Healthcare
von: Desikan, Prasanna, et al.
Veröffentlicht: (2026)
von: Desikan, Prasanna, et al.
Veröffentlicht: (2026)
VTC-Bench: Evaluating Agentic Multimodal Models via Compositional Visual Tool Chaining
von: Zhu, Xuanyu, et al.
Veröffentlicht: (2026)
von: Zhu, Xuanyu, et al.
Veröffentlicht: (2026)
CGBench: Benchmarking Language Model Scientific Reasoning for Clinical Genetics Research
von: Queen, Owen, et al.
Veröffentlicht: (2025)
von: Queen, Owen, et al.
Veröffentlicht: (2025)
TimeSage-MT: A Multi-Turn Benchmark for Evaluating Agentic Time Series Reasoning
von: Kong, Yaxuan, et al.
Veröffentlicht: (2026)
von: Kong, Yaxuan, et al.
Veröffentlicht: (2026)
VisScience: An Extensive Benchmark for Evaluating K12 Educational Multi-modal Scientific Reasoning
von: Jiang, Zhihuan, et al.
Veröffentlicht: (2024)
von: Jiang, Zhihuan, et al.
Veröffentlicht: (2024)
Higher-Order Knowledge Representations for Agentic Scientific Reasoning
von: Stewart, Isabella A., et al.
Veröffentlicht: (2026)
von: Stewart, Isabella A., et al.
Veröffentlicht: (2026)
TRAJECT-Bench:A Trajectory-Aware Benchmark for Evaluating Agentic Tool Use
von: He, Pengfei, et al.
Veröffentlicht: (2025)
von: He, Pengfei, et al.
Veröffentlicht: (2025)
PythonSaga: Redefining the Benchmark to Evaluate Code Generating LLMs
von: Yadav, Ankit, et al.
Veröffentlicht: (2024)
von: Yadav, Ankit, et al.
Veröffentlicht: (2024)
AEC-Bench: A Multimodal Benchmark for Agentic Systems in Architecture, Engineering, and Construction
von: Mankodiya, Harsh, et al.
Veröffentlicht: (2026)
von: Mankodiya, Harsh, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
SymPyBench: A Dynamic Benchmark for Scientific Reasoning with Executable Python Code
von: Imani, Shima, et al.
Veröffentlicht: (2025) -
TRACE: A Framework for Analyzing and Enhancing Stepwise Reasoning in Vision-Language Models
von: Imani, Shima, et al.
Veröffentlicht: (2025) -
PRiSM: Benchmarking Phone Realization in Speech Models
von: Bharadwaj, Shikhar, et al.
Veröffentlicht: (2026) -
Doppelgänger's Watch: A Split Objective Approach to Large Language Models
von: Ghasemlou, Shervin, et al.
Veröffentlicht: (2024) -
WearVQA: A Visual Question Answering Benchmark for Wearables in Egocentric Authentic Real-world scenarios
von: Chang, Eun, et al.
Veröffentlicht: (2025)