ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Ziru, Chen, Shijie, Ning, Yuting, Zhang, Qianheng, Wang, Boshi, Yu, Botao, Li, Yifei, Liao, Zeyi, Wei, Chen, Lu, Zitong, Dey, Vishal, Xue, Mingyi, Baker, Frazier N., Burns, Benjamin, Adu-Ampratwum, Daniel, Huang, Xuhui, Ning, Xia, Gao, Song, Su, Yu, Sun, Huan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ChemToolAgent: The Impact of Tools on Language Agents for Chemistry Problem Solving
by: Yu, Botao, et al.
Published: (2024)
by: Yu, Botao, et al.
Published: (2024)
RLSynC: Offline-Online Reinforcement Learning for Synthon Completion
by: Baker, Frazier N., et al.
Published: (2023)
by: Baker, Frazier N., et al.
Published: (2023)
LARC: Towards Human-level Constrained Retrosynthesis Planning through an Agentic Framework
by: Baker, Frazier N., et al.
Published: (2025)
by: Baker, Frazier N., et al.
Published: (2025)
MMORF: A Multi-agent Framework for Designing Multi-objective Retrosynthesis Planning Systems
by: Baker, Frazier N., et al.
Published: (2026)
by: Baker, Frazier N., et al.
Published: (2026)
AutoSDT: Scaling Data-Driven Discovery Tasks Toward Open Co-Scientists
by: Li, Yifei, et al.
Published: (2025)
by: Li, Yifei, et al.
Published: (2025)
ARMOR: An Agentic Framework for Reaction Feasibility Prediction via Adaptive Utility-aware Multi-tool Reasoning
by: Liu, Ye, et al.
Published: (2026)
by: Liu, Ye, et al.
Published: (2026)
LlaSMol: Advancing Large Language Models for Chemistry with a Large-Scale, Comprehensive, High-Quality Instruction Tuning Dataset
by: Yu, Botao, et al.
Published: (2024)
by: Yu, Botao, et al.
Published: (2024)
GeoAnalystBench: A GeoAI benchmark for assessing large language models for spatial analysis workflow and code generation
by: Zhang, Qianheng, et al.
Published: (2025)
by: Zhang, Qianheng, et al.
Published: (2025)
log-RRIM: Yield Prediction via Local-to-global Reaction Representation Learning and Interaction Modeling
by: Hu, Xiao, et al.
Published: (2024)
by: Hu, Xiao, et al.
Published: (2024)
Generating 3D Binding Molecules Using Shape-Conditioned Diffusion Models with Guidance
by: Chen, Ziqi, et al.
Published: (2025)
by: Chen, Ziqi, et al.
Published: (2025)
DiffER: Categorical Diffusion for Chemical Retrosynthesis
by: Current, Sean, et al.
Published: (2025)
by: Current, Sean, et al.
Published: (2025)
LIDDIA: Language-based Intelligent Drug Discovery Agent
by: Averly, Reza, et al.
Published: (2025)
by: Averly, Reza, et al.
Published: (2025)
UTBoost: Rigorous Evaluation of Coding Agents on SWE-Bench
by: Yu, Boxi, et al.
Published: (2025)
by: Yu, Boxi, et al.
Published: (2025)
Enhancing Molecular Property Prediction with Auxiliary Learning and Task-Specific Adaptation
by: Dey, Vishal, et al.
Published: (2024)
by: Dey, Vishal, et al.
Published: (2024)
LUMIR: an LLM-Driven Unified Agent Framework for Multi-task Infrared Spectroscopy Reasoning
by: Xie, Zujie, et al.
Published: (2025)
by: Xie, Zujie, et al.
Published: (2025)
RedTeamCUA: Realistic Adversarial Testing of Computer-Use Agents in Hybrid Web-OS Environments
by: Liao, Zeyi, et al.
Published: (2025)
by: Liao, Zeyi, et al.
Published: (2025)
Mind2Web 2: Evaluating Agentic Search with Agent-as-a-Judge
by: Gou, Boyu, et al.
Published: (2025)
by: Gou, Boyu, et al.
Published: (2025)
Insured Agents: A Decentralized Trust Insurance Mechanism for Agentic Economy
by: Hu, Botao 'Amber', et al.
Published: (2025)
by: Hu, Botao 'Amber', et al.
Published: (2025)
CheatAgent: Attacking LLM-Empowered Recommender Systems via LLM Agent
by: Ning, Liang-bo, et al.
Published: (2025)
by: Ning, Liang-bo, et al.
Published: (2025)
Large Language Models for Controllable Multi-property Multi-objective Molecule Optimization
by: Dey, Vishal, et al.
Published: (2025)
by: Dey, Vishal, et al.
Published: (2025)
GeLLMO: Generalizing Large Language Models for Multi-property Molecule Optimization
by: Dey, Vishal, et al.
Published: (2025)
by: Dey, Vishal, et al.
Published: (2025)
Event‐Triggered Iterative Learning Control for Multi‐Agent Systems With Dos Attacks Under a Two‐Dimensional Framework
by: Qing Wang, et al.
Published: (2025)
by: Qing Wang, et al.
Published: (2025)
StutterZero and StutterFormer: End-to-End Speech Conversion for Stuttering Transcription and Correction
by: Xu, Qianheng
Published: (2025)
by: Xu, Qianheng
Published: (2025)
ScholarEval: Research Idea Evaluation Grounded in Literature
by: Moussa, Hanane Nour, et al.
Published: (2025)
by: Moussa, Hanane Nour, et al.
Published: (2025)
AstaBench: Rigorous Benchmarking of AI Agents with a Scientific Research Suite
by: Bragg, Jonathan, et al.
Published: (2025)
by: Bragg, Jonathan, et al.
Published: (2025)
AttributionBench: How Hard is Automatic Attribution Evaluation?
by: Li, Yifei, et al.
Published: (2024)
by: Li, Yifei, et al.
Published: (2024)
eCeLLM: Generalizing Large Language Models for E-commerce from Large-scale, High-quality Instruction Data
by: Peng, Bo, et al.
Published: (2024)
by: Peng, Bo, et al.
Published: (2024)
AgentBench: Evaluating LLMs as Agents
by: Liu, Xiao, et al.
Published: (2023)
by: Liu, Xiao, et al.
Published: (2023)
Prepregnancy weight loss and maternal metabolic and inflammatory biomarkers during pregnancy: An analysis of National Health and Nutrition Examination Survey
by: Yang Yu, et al.
Published: (2024)
by: Yang Yu, et al.
Published: (2024)
AutoAgent: Evolving Cognition and Elastic Memory Orchestration for Adaptive Agents
by: Wang, Xiaoxing, et al.
Published: (2026)
by: Wang, Xiaoxing, et al.
Published: (2026)
LLMSched: Uncertainty-Aware Workload Scheduling for Compound LLM Applications
by: Zhu, Botao, et al.
Published: (2025)
by: Zhu, Botao, et al.
Published: (2025)
Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent Evaluation
by: Kapoor, Sayash, et al.
Published: (2025)
by: Kapoor, Sayash, et al.
Published: (2025)
QUEST: Training Frontier Deep Research Agents with Fully Synthetic Tasks
by: Xie, Jian, et al.
Published: (2026)
by: Xie, Jian, et al.
Published: (2026)
ET-Agent: Incentivizing Effective Tool-Integrated Reasoning Agent via Behavior Calibration
by: Chen, Yifei, et al.
Published: (2026)
by: Chen, Yifei, et al.
Published: (2026)
Tap-to-Adapt: Learning User-Aligned Response Timing for Speech Agents
by: He, Zihong, et al.
Published: (2026)
by: He, Zihong, et al.
Published: (2026)
AGENTCL: Toward Rigorous Evaluation of Continual Learning in Language Agents
by: Shu, Yiheng, et al.
Published: (2026)
by: Shu, Yiheng, et al.
Published: (2026)
MIRIX: Multi-Agent Memory System for LLM-Based Agents
by: Wang, Yu, et al.
Published: (2025)
by: Wang, Yu, et al.
Published: (2025)
MobileAgentBench: An Efficient and User-Friendly Benchmark for Mobile LLM Agents
by: Wang, Luyuan, et al.
Published: (2024)
by: Wang, Luyuan, et al.
Published: (2024)
Curie: Toward Rigorous and Automated Scientific Experimentation with AI Agents
by: Kon, Patrick Tser Jern, et al.
Published: (2025)
by: Kon, Patrick Tser Jern, et al.
Published: (2025)
M3MAD-Bench: Are Multi-Agent Debates Really Effective Across Domains and Modalities?
by: Li, Ao, et al.
Published: (2026)
by: Li, Ao, et al.
Published: (2026)
Similar Items
-
ChemToolAgent: The Impact of Tools on Language Agents for Chemistry Problem Solving
by: Yu, Botao, et al.
Published: (2024) -
RLSynC: Offline-Online Reinforcement Learning for Synthon Completion
by: Baker, Frazier N., et al.
Published: (2023) -
LARC: Towards Human-level Constrained Retrosynthesis Planning through an Agentic Framework
by: Baker, Frazier N., et al.
Published: (2025) -
MMORF: A Multi-agent Framework for Designing Multi-objective Retrosynthesis Planning Systems
by: Baker, Frazier N., et al.
Published: (2026) -
AutoSDT: Scaling Data-Driven Discovery Tasks Toward Open Co-Scientists
by: Li, Yifei, et al.
Published: (2025)