MSQA: Benchmarking LLMs on Graduate-Level Materials Science Reasoning and Knowledge
Fuente:
arXiv
Saved in:
| Main Authors: | Cheung, Jerry Junyang, Shen, Shiyao, Zhuang, Yuchen, Li, Yinghao, Ramprasad, Rampi, Zhang, Chao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Retrieval Augmented Generation of Literature-derived Polymer Knowledge: The Example of a Biodegradable Polymer Expert System
by: Gupta, Sonakshi, et al.
Published: (2026)
by: Gupta, Sonakshi, et al.
Published: (2026)
A Simple but Effective Approach to Improve Structured Language Model Output for Information Extraction
by: Li, Yinghao, et al.
Published: (2024)
by: Li, Yinghao, et al.
Published: (2024)
TPD: Enhancing Student Language Model Reasoning via Principle Discovery and Guidance
by: Wang, Haorui, et al.
Published: (2024)
by: Wang, Haorui, et al.
Published: (2024)
From Reasoning to Generalization: Knowledge-Augmented LLMs for ARC Benchmark
by: Lei, Chao, et al.
Published: (2025)
by: Lei, Chao, et al.
Published: (2025)
LocalBench: Benchmarking LLMs on County-Level Local Knowledge and Reasoning
by: Gao, Zihan, et al.
Published: (2025)
by: Gao, Zihan, et al.
Published: (2025)
SCICONVBENCH: Benchmarking LLMs on Multi-Turn Clarification for Task Formulation in Computational Science
by: Somasekharan, Nithin, et al.
Published: (2026)
by: Somasekharan, Nithin, et al.
Published: (2026)
MatSciBench: Benchmarking the Reasoning Ability of Large Language Models in Materials Science
by: Zhang, Junkai, et al.
Published: (2025)
by: Zhang, Junkai, et al.
Published: (2025)
Enhancing Mathematical Reasoning in LLMs with Background Operators
by: Chen, Jiajun, et al.
Published: (2024)
by: Chen, Jiajun, et al.
Published: (2024)
LLM-Augmented Chemical Synthesis and Design Decision Programs
by: Wang, Haorui, et al.
Published: (2025)
by: Wang, Haorui, et al.
Published: (2025)
Evaluating Frontier LLMs on PhD-Level Mathematical Reasoning: A Benchmark on a Textbook in Theoretical Computer Science about Randomized Algorithms
by: Cao, Yang, et al.
Published: (2025)
by: Cao, Yang, et al.
Published: (2025)
Siamese Multiple Attention Temporal Convolution Networks for Human Mobility Signature Identification
by: Zheng, Zhipeng, et al.
Published: (2024)
by: Zheng, Zhipeng, et al.
Published: (2024)
MARGE: Improving Math Reasoning for LLMs with Guided Exploration
by: Gao, Jingyue, et al.
Published: (2025)
by: Gao, Jingyue, et al.
Published: (2025)
Edit Once, Update Everywhere: A Simple Framework for Cross-Lingual Knowledge Synchronization in LLMs
by: Wu, Yuchen, et al.
Published: (2025)
by: Wu, Yuchen, et al.
Published: (2025)
Reasoning about Affordances: Causal and Compositional Reasoning in LLMs
by: Gjerde, Magnus F., et al.
Published: (2025)
by: Gjerde, Magnus F., et al.
Published: (2025)
Matryoshka Pilot: Learning to Drive Black-Box LLMs with LLMs
by: Li, Changhao, et al.
Published: (2024)
by: Li, Changhao, et al.
Published: (2024)
OmniMatBench: A Human-Calibrated Multimodal Reasoning Benchmark Across 19 Materials Science Subfields
by: Liu, Wanhao, et al.
Published: (2026)
by: Liu, Wanhao, et al.
Published: (2026)
RAG in the Wild: On the (In)effectiveness of LLMs with Mixture-of-Knowledge Retrieval Augmentation
by: Xu, Ran, et al.
Published: (2025)
by: Xu, Ran, et al.
Published: (2025)
Can AI Validate Science? Benchmarking LLMs for Accurate Scientific Claim $\rightarrow$ Evidence Reasoning
by: Javaji, Shashidhar Reddy, et al.
Published: (2025)
by: Javaji, Shashidhar Reddy, et al.
Published: (2025)
CHAIRO: Contextual Hierarchical Analogical Induction and Reasoning Optimization for LLMs
by: Lu, Haotian, et al.
Published: (2026)
by: Lu, Haotian, et al.
Published: (2026)
FedVAE: Trajectory privacy preserving based on Federated Variational AutoEncoder
by: Jiang, Yuchen, et al.
Published: (2024)
by: Jiang, Yuchen, et al.
Published: (2024)
MAPS: Advancing Multi-Modal Reasoning in Expert-Level Physical Science
by: Zhu, Erle, et al.
Published: (2025)
by: Zhu, Erle, et al.
Published: (2025)
SVGenius: Benchmarking LLMs in SVG Understanding, Editing and Generation
by: Chen, Siqi, et al.
Published: (2025)
by: Chen, Siqi, et al.
Published: (2025)
TIME: A Multi-level Benchmark for Temporal Reasoning of LLMs in Real-World Scenarios
by: Wei, Shaohang, et al.
Published: (2025)
by: Wei, Shaohang, et al.
Published: (2025)
WorkForceAgent-R1: Incentivizing Reasoning Capability in LLM-based Web Agents via Reinforcement Learning
by: Zhuang, Yuchen, et al.
Published: (2025)
by: Zhuang, Yuchen, et al.
Published: (2025)
FaithCoT-Bench: Benchmarking Instance-Level Faithfulness of Chain-of-Thought Reasoning
by: Shen, Xu, et al.
Published: (2025)
by: Shen, Xu, et al.
Published: (2025)
KISS - Knowledge Infrastructure for Scientific Simulation: A Scaffolding for Agentic Earth Science
by: Li, Ziwei, et al.
Published: (2026)
by: Li, Ziwei, et al.
Published: (2026)
CorpusQA: A 10 Million Token Benchmark for Corpus-Level Analysis and Reasoning
by: Lu, Zhiyuan, et al.
Published: (2026)
by: Lu, Zhiyuan, et al.
Published: (2026)
Can Multimodal LLMs See Science Instruction? Benchmarking Pedagogical Reasoning in K-12 Classroom Videos
by: Shen, Yixuan, et al.
Published: (2026)
by: Shen, Yixuan, et al.
Published: (2026)
ABench-Physics: Benchmarking Physical Reasoning in LLMs via High-Difficulty and Dynamic Physics Problems
by: Zhang, Yiming, et al.
Published: (2025)
by: Zhang, Yiming, et al.
Published: (2025)
MathFimer: Enhancing Mathematical Reasoning by Expanding Reasoning Steps through Fill-in-the-Middle Task
by: Yan, Yuchen, et al.
Published: (2025)
by: Yan, Yuchen, et al.
Published: (2025)
Informed Routing in LLMs: Smarter Token-Level Computation for Faster Inference
by: Han, Chao, et al.
Published: (2025)
by: Han, Chao, et al.
Published: (2025)
Conflict Detection for Temporal Knowledge Graphs:A Fast Constraint Mining Algorithm and New Benchmarks
by: Chen, Jianhao, et al.
Published: (2023)
by: Chen, Jianhao, et al.
Published: (2023)
QMBench: A Research Level Benchmark for Quantum Materials Research
by: Wang, Yanzhen, et al.
Published: (2025)
by: Wang, Yanzhen, et al.
Published: (2025)
FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning
by: Wang, Zeyu, et al.
Published: (2026)
by: Wang, Zeyu, et al.
Published: (2026)
UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models
by: Xu, Xin, et al.
Published: (2025)
by: Xu, Xin, et al.
Published: (2025)
MediEval: A Unified Medical Benchmark for Patient-Contextual and Knowledge-Grounded Reasoning in LLMs
by: Qu, Zhan, et al.
Published: (2025)
by: Qu, Zhan, et al.
Published: (2025)
LLM-Assisted Op-Amp Behavioral-Level Design via Agentic Human-Mimicking Reasoning
by: Chen, Zihao, et al.
Published: (2026)
by: Chen, Zihao, et al.
Published: (2026)
Oedipus and the Sphinx: Benchmarking and Improving Visual Language Models for Complex Graphic Reasoning
by: Zhang, Jianyi, et al.
Published: (2025)
by: Zhang, Jianyi, et al.
Published: (2025)
InftyThink: Breaking the Length Limits of Long-Context Reasoning in Large Language Models
by: Yan, Yuchen, et al.
Published: (2025)
by: Yan, Yuchen, et al.
Published: (2025)
SciEval: A Benchmark for Automatic Evaluation of K-12 Science Instructional Materials
by: Li, Zhaohui, et al.
Published: (2026)
by: Li, Zhaohui, et al.
Published: (2026)
Similar Items
-
Retrieval Augmented Generation of Literature-derived Polymer Knowledge: The Example of a Biodegradable Polymer Expert System
by: Gupta, Sonakshi, et al.
Published: (2026) -
A Simple but Effective Approach to Improve Structured Language Model Output for Information Extraction
by: Li, Yinghao, et al.
Published: (2024) -
TPD: Enhancing Student Language Model Reasoning via Principle Discovery and Guidance
by: Wang, Haorui, et al.
Published: (2024) -
From Reasoning to Generalization: Knowledge-Augmented LLMs for ARC Benchmark
by: Lei, Chao, et al.
Published: (2025) -
LocalBench: Benchmarking LLMs on County-Level Local Knowledge and Reasoning
by: Gao, Zihan, et al.
Published: (2025)