CORE-Bench: Fostering the Credibility of Published Research Through a Computational Reproducibility Agent Benchmark
Fuente:
arXiv
Saved in:
| Main Authors: | Siegel, Zachary S., Kapoor, Sayash, Nagdir, Nitya, Stroebl, Benedikt, Narayanan, Arvind |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AI Agents That Matter
by: Kapoor, Sayash, et al.
Published: (2024)
by: Kapoor, Sayash, et al.
Published: (2024)
The Limits of Inference Scaling Through Resampling
by: Stroebl, Benedikt, et al.
Published: (2024)
by: Stroebl, Benedikt, et al.
Published: (2024)
WorkBench: a Benchmark Dataset for Agents in a Realistic Workplace Setting
by: Styles, Olly, et al.
Published: (2024)
by: Styles, Olly, et al.
Published: (2024)
SciReplicate-Bench: Benchmarking LLMs in Agent-driven Algorithmic Reproduction from Research Papers
by: Xiang, Yanzheng, et al.
Published: (2025)
by: Xiang, Yanzheng, et al.
Published: (2025)
Large Language Model-Based Agents for Automated Research Reproducibility: An Exploratory Study in Alzheimer's Disease
by: Dobbins, Nic, et al.
Published: (2025)
by: Dobbins, Nic, et al.
Published: (2025)
An Adversary-Resistant Multi-Agent LLM System via Credibility Scoring
by: Ebrahimi, Sana, et al.
Published: (2025)
by: Ebrahimi, Sana, et al.
Published: (2025)
MAS-Orchestra: Understanding and Improving Multi-Agent Reasoning Through Holistic Orchestration and Controlled Benchmarks
by: Ke, Zixuan, et al.
Published: (2026)
by: Ke, Zixuan, et al.
Published: (2026)
The High Cost of Incivility: Quantifying Interaction Inefficiency via Multi-Agent Monte Carlo Simulations
by: Mangold, Benedikt
Published: (2025)
by: Mangold, Benedikt
Published: (2025)
Voting or Consensus? Decision-Making in Multi-Agent Debate
by: Kaesberg, Lars Benedikt, et al.
Published: (2025)
by: Kaesberg, Lars Benedikt, et al.
Published: (2025)
AgentArch: A Comprehensive Benchmark to Evaluate Agent Architectures in Enterprise
by: Bogavelli, Tara, et al.
Published: (2025)
by: Bogavelli, Tara, et al.
Published: (2025)
Scaling Small Agents Through Strategy Auctions
by: Alazraki, Lisa, et al.
Published: (2026)
by: Alazraki, Lisa, et al.
Published: (2026)
Achieving Unanimous Consensus Through Multi-Agent Deliberation
by: Pokharel, Apurba, et al.
Published: (2025)
by: Pokharel, Apurba, et al.
Published: (2025)
WebTestBench: Evaluating Computer-Use Agents towards End-to-End Automated Web Testing
by: Kong, Fanheng, et al.
Published: (2026)
by: Kong, Fanheng, et al.
Published: (2026)
Embodied Multi-Agent Coordination by Aligning World Models Through Dialogue
by: Dongre, Vardhan, et al.
Published: (2026)
by: Dongre, Vardhan, et al.
Published: (2026)
Collab-Overcooked: Benchmarking and Evaluating Large Language Models as Collaborative Agents
by: Sun, Haochen, et al.
Published: (2025)
by: Sun, Haochen, et al.
Published: (2025)
LingxiDiagBench: A Multi-Agent Framework for Benchmarking LLMs in Chinese Psychiatric Consultation and Diagnosis
by: Xu, Shihao, et al.
Published: (2026)
by: Xu, Shihao, et al.
Published: (2026)
MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents
by: Zhu, Kunlun, et al.
Published: (2025)
by: Zhu, Kunlun, et al.
Published: (2025)
Enhancing Online Learning Efficiency Through Heterogeneous Resource Integration with a Multi-Agent RAG System
by: Srivastav, Devansh, et al.
Published: (2025)
by: Srivastav, Devansh, et al.
Published: (2025)
HSCodeComp: A Realistic and Expert-level Benchmark for Deep Search Agents in Hierarchical Rule Application
by: Yang, Yiqian, et al.
Published: (2025)
by: Yang, Yiqian, et al.
Published: (2025)
Don't Overthink It: Inter-Rollout Action Agreement as a Free Adaptive-Compute Signal for LLM Agents
by: Sethi, Khushal
Published: (2026)
by: Sethi, Khushal
Published: (2026)
Doctorina MedBench: End-to-End Evaluation of Agent-Based Medical AI
by: Kozlova, Anna, et al.
Published: (2026)
by: Kozlova, Anna, et al.
Published: (2026)
BAPPA: Benchmarking Agents, Plans, and Pipelines for Automated Text-to-SQL Generation
by: Ahmed, Fahim, et al.
Published: (2025)
by: Ahmed, Fahim, et al.
Published: (2025)
WorldView-Bench: A Benchmark for Evaluating Global Cultural Perspectives in Large Language Models
by: Mushtaq, Abdullah, et al.
Published: (2025)
by: Mushtaq, Abdullah, et al.
Published: (2025)
AssetOpsBench: Benchmarking AI Agents for Task Automation in Industrial Asset Operations and Maintenance
by: Patel, Dhaval, et al.
Published: (2025)
by: Patel, Dhaval, et al.
Published: (2025)
AgentSearchBench: A Benchmark for AI Agent Search in the Wild
by: Wu, Bin, et al.
Published: (2026)
by: Wu, Bin, et al.
Published: (2026)
BenchMARL: Benchmarking Multi-Agent Reinforcement Learning
by: Bettini, Matteo, et al.
Published: (2023)
by: Bettini, Matteo, et al.
Published: (2023)
Cognitive Insights and Stable Coalition Matching for Fostering Multi-Agent Cooperation
by: Shao, Jiaqi, et al.
Published: (2024)
by: Shao, Jiaqi, et al.
Published: (2024)
CACA Agent: Capability Collaboration based AI Agent
by: Xu, Peng, et al.
Published: (2024)
by: Xu, Peng, et al.
Published: (2024)
ResearchCodeAgent: An LLM Multi-Agent System for Automated Codification of Research Methodologies
by: Gandhi, Shubham, et al.
Published: (2025)
by: Gandhi, Shubham, et al.
Published: (2025)
Efficient Agents: Building Effective Agents While Reducing Cost
by: Wang, Ningning, et al.
Published: (2025)
by: Wang, Ningning, et al.
Published: (2025)
MedAgentBoard: Benchmarking Multi-Agent Collaboration with Conventional Methods for Diverse Medical Tasks
by: Zhu, Yinghao, et al.
Published: (2025)
by: Zhu, Yinghao, et al.
Published: (2025)
GT-HarmBench: Benchmarking AI Safety Risks Through the Lens of Game Theory
by: Cobben, Pepijn, et al.
Published: (2026)
by: Cobben, Pepijn, et al.
Published: (2026)
AblateCell: A Reproduce-then-Ablate Agent for Virtual Cell Repositories
by: Xia, Xue, et al.
Published: (2026)
by: Xia, Xue, et al.
Published: (2026)
ColorAgent: Building A Robust, Personalized, and Interactive OS Agent
by: Li, Ning, et al.
Published: (2025)
by: Li, Ning, et al.
Published: (2025)
Your Agents Are Aging Too: Agent Lifespan Engineering for Deployed Systems
by: Zhu, Jianing, et al.
Published: (2026)
by: Zhu, Jianing, et al.
Published: (2026)
Agent Primitives: Reusable Latent Building Blocks for Multi-Agent Systems
by: Jin, Haibo, et al.
Published: (2026)
by: Jin, Haibo, et al.
Published: (2026)
A Dynamic LLM-Powered Agent Network for Task-Oriented Agent Collaboration
by: Liu, Zijun, et al.
Published: (2023)
by: Liu, Zijun, et al.
Published: (2023)
SciAgent: A Unified Multi-Agent System for Generalistic Scientific Reasoning
by: Li, Xuchen, et al.
Published: (2025)
by: Li, Xuchen, et al.
Published: (2025)
AgentForesight: Online Auditing for Early Failure Prediction in Multi-Agent Systems
by: Zhang, Boxuan, et al.
Published: (2026)
by: Zhang, Boxuan, et al.
Published: (2026)
AgentSwing: Adaptive Parallel Context Management Routing for Long-Horizon Web Agents
by: Feng, Zhaopeng, et al.
Published: (2026)
by: Feng, Zhaopeng, et al.
Published: (2026)
Similar Items
-
AI Agents That Matter
by: Kapoor, Sayash, et al.
Published: (2024) -
The Limits of Inference Scaling Through Resampling
by: Stroebl, Benedikt, et al.
Published: (2024) -
WorkBench: a Benchmark Dataset for Agents in a Realistic Workplace Setting
by: Styles, Olly, et al.
Published: (2024) -
SciReplicate-Bench: Benchmarking LLMs in Agent-driven Algorithmic Reproduction from Research Papers
by: Xiang, Yanzheng, et al.
Published: (2025) -
Large Language Model-Based Agents for Automated Research Reproducibility: An Exploratory Study in Alzheimer's Disease
by: Dobbins, Nic, et al.
Published: (2025)