OmniMatBench: A Human-Calibrated Multimodal Reasoning Benchmark Across 19 Materials Science Subfields
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Wanhao, Xie, Jiaqing, Tan, Qian, Wang, Weida, Wang, Jue, Sun, Ran, Yang, Zhuo, Ouyang, Wanli, Bai, Lei, Fu, Tianfan, Chen, Lu, Chen, Xin, Li, Yuqiang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
PolyReal: A Benchmark for Real-World Polymer Science Workflows
von: Liu, Wanhao, et al.
Veröffentlicht: (2026)
von: Liu, Wanhao, et al.
Veröffentlicht: (2026)
ChemMLLM: Chemical Multimodal Large Language Model
von: Tan, Qian, et al.
Veröffentlicht: (2025)
von: Tan, Qian, et al.
Veröffentlicht: (2025)
MatSciBench: Benchmarking the Reasoning Ability of Large Language Models in Materials Science
von: Zhang, Junkai, et al.
Veröffentlicht: (2025)
von: Zhang, Junkai, et al.
Veröffentlicht: (2025)
Chem-R: Learning to Reason as a Chemist
von: Wang, Weida, et al.
Veröffentlicht: (2025)
von: Wang, Weida, et al.
Veröffentlicht: (2025)
MolAct: An Agentic RL Framework for Molecular Editing and Property Optimization
von: Yang, Zhuo, et al.
Veröffentlicht: (2025)
von: Yang, Zhuo, et al.
Veröffentlicht: (2025)
QCBench: Evaluating Large Language Models on Domain-Specific Quantitative Chemistry
von: Xie, Jiaqing, et al.
Veröffentlicht: (2025)
von: Xie, Jiaqing, et al.
Veröffentlicht: (2025)
SkillsInjector: Dynamic Skill Context Construction for LLM Agents
von: Li, Yanchao, et al.
Veröffentlicht: (2026)
von: Li, Yanchao, et al.
Veröffentlicht: (2026)
Reasoning BO: Enhancing Bayesian Optimization with Long-Context Reasoning Power of LLMs
von: Yang, Zhuo, et al.
Veröffentlicht: (2025)
von: Yang, Zhuo, et al.
Veröffentlicht: (2025)
DeepProtein: Deep Learning Library and Benchmark for Protein Sequence Learning
von: Xie, Jiaqing, et al.
Veröffentlicht: (2024)
von: Xie, Jiaqing, et al.
Veröffentlicht: (2024)
MOOSE-Chem3: Toward Experiment-Guided Hypothesis Ranking via Simulated Experimental Feedback
von: Liu, Wanhao, et al.
Veröffentlicht: (2025)
von: Liu, Wanhao, et al.
Veröffentlicht: (2025)
OmniGenBench: A Benchmark for Omnipotent Multimodal Generation across 50+ Tasks
von: Wang, Jiayu, et al.
Veröffentlicht: (2025)
von: Wang, Jiayu, et al.
Veröffentlicht: (2025)
BabelBench: An Omni Benchmark for Code-Driven Analysis of Multimodal and Multistructured Data
von: Wang, Xuwu, et al.
Veröffentlicht: (2024)
von: Wang, Xuwu, et al.
Veröffentlicht: (2024)
MatFormBench: A Benchmarking Evaluation Framework for Target-Driven Materials Formulation
von: Wu, Linhan, et al.
Veröffentlicht: (2026)
von: Wu, Linhan, et al.
Veröffentlicht: (2026)
ComfyBench: Benchmarking LLM-based Agents in ComfyUI for Autonomously Designing Collaborative AI Systems
von: Xue, Xiangyuan, et al.
Veröffentlicht: (2024)
von: Xue, Xiangyuan, et al.
Veröffentlicht: (2024)
PhysUniBench: A Multi-Modal Physics Reasoning Benchmark at Undergraduate Level
von: Wang, Lintao, et al.
Veröffentlicht: (2025)
von: Wang, Lintao, et al.
Veröffentlicht: (2025)
OmniBrainBench: A Comprehensive Multimodal Benchmark for Brain Imaging Analysis Across Multi-stage Clinical Tasks
von: Peng, Zhihao, et al.
Veröffentlicht: (2025)
von: Peng, Zhihao, et al.
Veröffentlicht: (2025)
PredBench: Benchmarking Spatio-Temporal Prediction across Diverse Disciplines
von: Wang, ZiDong, et al.
Veröffentlicht: (2024)
von: Wang, ZiDong, et al.
Veröffentlicht: (2024)
CheMatAgent: Enhancing LLMs for Chemistry and Materials Science through Tree-Search Based Tool Learning
von: Wu, Mengsong, et al.
Veröffentlicht: (2025)
von: Wu, Mengsong, et al.
Veröffentlicht: (2025)
OmniSch: A Multimodal PCB Schematic Benchmark For Structured Diagram Visual Reasoning
von: Lu, Taiting, et al.
Veröffentlicht: (2026)
von: Lu, Taiting, et al.
Veröffentlicht: (2026)
RF-MatID: Dataset and Benchmark for Radio Frequency Material Identification
von: Chen, Xinyan, et al.
Veröffentlicht: (2026)
von: Chen, Xinyan, et al.
Veröffentlicht: (2026)
FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning
von: Wang, Zeyu, et al.
Veröffentlicht: (2026)
von: Wang, Zeyu, et al.
Veröffentlicht: (2026)
UNO-Bench: A Unified Benchmark for Exploring the Compositional Law Between Uni-modal and Omni-modal in Omni Models
von: Chen, Chen, et al.
Veröffentlicht: (2025)
von: Chen, Chen, et al.
Veröffentlicht: (2025)
MOOSE-Chem: Large Language Models for Rediscovering Unseen Chemistry Scientific Hypotheses
von: Yang, Zonglin, et al.
Veröffentlicht: (2024)
von: Yang, Zonglin, et al.
Veröffentlicht: (2024)
MatWheel: Addressing Data Scarcity in Materials Science Through Synthetic Data
von: Li, Wentao, et al.
Veröffentlicht: (2025)
von: Li, Wentao, et al.
Veröffentlicht: (2025)
ResearchBench: Benchmarking LLMs in Scientific Discovery via Inspiration-Based Task Decomposition
von: Liu, Yujie, et al.
Veröffentlicht: (2025)
von: Liu, Yujie, et al.
Veröffentlicht: (2025)
UmniBench: Unified Understand and Generation Model Oriented Omni-dimensional Benchmark
von: Liu, Kai, et al.
Veröffentlicht: (2025)
von: Liu, Kai, et al.
Veröffentlicht: (2025)
LLM4Mat-Bench: Benchmarking Large Language Models for Materials Property Prediction
von: Rubungo, Andre Niyongabo, et al.
Veröffentlicht: (2024)
von: Rubungo, Andre Niyongabo, et al.
Veröffentlicht: (2024)
MatTools: Benchmarking Large Language Models for Materials Science Tools
von: Liu, Siyu, et al.
Veröffentlicht: (2025)
von: Liu, Siyu, et al.
Veröffentlicht: (2025)
Automatic Detection of Research Values from Scientific Abstracts Across Computer Science Subfields
von: Jiang, Hang, et al.
Veröffentlicht: (2025)
von: Jiang, Hang, et al.
Veröffentlicht: (2025)
SciVideoBench: Benchmarking Scientific Video Reasoning in Large Multimodal Models
von: Deng, Andong, et al.
Veröffentlicht: (2025)
von: Deng, Andong, et al.
Veröffentlicht: (2025)
BilliardPhys-Bench: Benchmarking Physical Reasoning and Visual Dynamics of Multimodal LLMs
von: Wang, Ben, et al.
Veröffentlicht: (2026)
von: Wang, Ben, et al.
Veröffentlicht: (2026)
SeedBench: A Multi-task Benchmark for Evaluating Large Language Models in Seed Science
von: Ying, Jie, et al.
Veröffentlicht: (2025)
von: Ying, Jie, et al.
Veröffentlicht: (2025)
SpecMol: A Spectroscopy-Grounded Foundation Model for Multi-Task Molecular Learning
von: Shen, Shuaike, et al.
Veröffentlicht: (2025)
von: Shen, Shuaike, et al.
Veröffentlicht: (2025)
GlobalDentBench: A Multinational Benchmark for Evaluating LLM Clinical Reasoning in Dentistry with Expert Calibration
von: Zhao, Junjie, et al.
Veröffentlicht: (2026)
von: Zhao, Junjie, et al.
Veröffentlicht: (2026)
AdaBrain-Bench: Benchmarking Brain Foundation Models for Brain-Computer Interface Applications
von: Wu, Jiamin, et al.
Veröffentlicht: (2025)
von: Wu, Jiamin, et al.
Veröffentlicht: (2025)
Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning
von: Zhang, Di, et al.
Veröffentlicht: (2024)
von: Zhang, Di, et al.
Veröffentlicht: (2024)
SuperMat: Physically Consistent PBR Material Estimation at Interactive Rates
von: Hong, Yijia, et al.
Veröffentlicht: (2024)
von: Hong, Yijia, et al.
Veröffentlicht: (2024)
OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive Annotations
von: Ouyang, Linke, et al.
Veröffentlicht: (2024)
von: Ouyang, Linke, et al.
Veröffentlicht: (2024)
Seeing Beyond Words: MatVQA for Challenging Visual-Scientific Reasoning in Materials Science
von: Wu, Sifan, et al.
Veröffentlicht: (2025)
von: Wu, Sifan, et al.
Veröffentlicht: (2025)
AutoMat: Enabling Automated Crystal Structure Reconstruction from Microscopy via Agentic Tool Use
von: Yang, Yaotian, et al.
Veröffentlicht: (2025)
von: Yang, Yaotian, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
PolyReal: A Benchmark for Real-World Polymer Science Workflows
von: Liu, Wanhao, et al.
Veröffentlicht: (2026) -
ChemMLLM: Chemical Multimodal Large Language Model
von: Tan, Qian, et al.
Veröffentlicht: (2025) -
MatSciBench: Benchmarking the Reasoning Ability of Large Language Models in Materials Science
von: Zhang, Junkai, et al.
Veröffentlicht: (2025) -
Chem-R: Learning to Reason as a Chemist
von: Wang, Weida, et al.
Veröffentlicht: (2025) -
MolAct: An Agentic RL Framework for Molecular Editing and Property Optimization
von: Yang, Zhuo, et al.
Veröffentlicht: (2025)