PaperBench: Evaluating AI's Ability to Replicate AI Research
Fuente:
arXiv
Guardado en:
| Autores principales: | Starace, Giulio, Jaffe, Oliver, Sherburn, Dane, Aung, James, Chan, Jun Shern, Maksin, Leon, Dias, Rachel, Mays, Evan, Kinsella, Benjamin, Thompson, Wyatt, Heidecke, Johannes, Glaese, Amelia, Patwardhan, Tejal |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
por: Chan, Jun Shern, et al.
Publicado: (2024)
por: Chan, Jun Shern, et al.
Publicado: (2024)
SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?
por: Miserendino, Samuel, et al.
Publicado: (2025)
por: Miserendino, Samuel, et al.
Publicado: (2025)
FrontierScience: Evaluating AI's Ability to Perform Expert-Level Scientific Tasks
por: Wang, Miles, et al.
Publicado: (2026)
por: Wang, Miles, et al.
Publicado: (2026)
ReplicationBench: Can AI Agents Replicate Astrophysics Research Papers?
por: Ye, Christine, et al.
Publicado: (2025)
por: Ye, Christine, et al.
Publicado: (2025)
Can Language Models Explain Their Own Classification Behavior?
por: Sherburn, Dane, et al.
Publicado: (2024)
por: Sherburn, Dane, et al.
Publicado: (2024)
Legal Zero-Days: A Novel Risk Vector for Advanced AI Systems
por: Sadler, Greg, et al.
Publicado: (2025)
por: Sadler, Greg, et al.
Publicado: (2025)
GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks
por: Patwardhan, Tejal, et al.
Publicado: (2025)
por: Patwardhan, Tejal, et al.
Publicado: (2025)
EVMbench: Evaluating AI Agents on Smart Contract Security
por: Wang, Justin, et al.
Publicado: (2026)
por: Wang, Justin, et al.
Publicado: (2026)
Trading Inference-Time Compute for Adversarial Robustness
por: Zaremba, Wojciech, et al.
Publicado: (2025)
por: Zaremba, Wojciech, et al.
Publicado: (2025)
The Human-AI Handshake Framework: A Bidirectional Approach to Human-AI Collaboration
por: Pyae, Aung
Publicado: (2025)
por: Pyae, Aung
Publicado: (2025)
Ethical Implications of Training Deceptive AI
por: Starace, Jason, et al.
Publicado: (2026)
por: Starace, Jason, et al.
Publicado: (2026)
What is Human-Centeredness in Human-Centered AI? Development of Human-Centeredness Framework and AI Practitioners' Perspectives
por: Pyae, Aung
Publicado: (2025)
por: Pyae, Aung
Publicado: (2025)
Python workflow for segmenting multiphase flow in porous rocks
por: Spurin, Catherine, et al.
Publicado: (2024)
por: Spurin, Catherine, et al.
Publicado: (2024)
Can LLMs Identify Critical Limitations within Scientific Research? A Systematic Evaluation on AI Research Papers
por: Xu, Zhijian, et al.
Publicado: (2025)
por: Xu, Zhijian, et al.
Publicado: (2025)
Can AI with High Reasoning Ability Replicate Human-like Decision Making in Economic Experiments?
por: Kitadai, Ayato, et al.
Publicado: (2024)
por: Kitadai, Ayato, et al.
Publicado: (2024)
Understanding Student Acceptance, Trust, and Attitudes Toward AI-Generated Images for Educational Purposes
por: Pyae, Aung
Publicado: (2024)
por: Pyae, Aung
Publicado: (2024)
AI-Generated 3D Environments as Speculative Mediators in More-Than-Human Design: An Exploratory Study
por: Pyae, Aung
Publicado: (2026)
por: Pyae, Aung
Publicado: (2026)
Persona Features Control Emergent Misalignment
por: Wang, Miles, et al.
Publicado: (2025)
por: Wang, Miles, et al.
Publicado: (2025)
Safety Assessment of Scaffolding on Construction Site using AI
por: Prabhu, Sameer, et al.
Publicado: (2025)
por: Prabhu, Sameer, et al.
Publicado: (2025)
From Prompts to Worlds: How Users Iterate, Explore, and Make Sense of AI-Generated 3D Environments
por: Pyae, Aung
Publicado: (2026)
por: Pyae, Aung
Publicado: (2026)
TPS-Bench: Evaluating AI Agents' Tool Planning \& Scheduling Abilities in Compounding Tasks
por: Xu, Hanwen, et al.
Publicado: (2025)
por: Xu, Hanwen, et al.
Publicado: (2025)
Spatial-Immune Multi-omics Refines Prognostication in Early-Stage Estrogen Receptor-Positive Breast Cancer
por: Kinsella, Zak
Publicado: (2025)
por: Kinsella, Zak
Publicado: (2025)
Nutrient recovery technology mapping datasheet - NOVAFERT project
por: Kinsella, Dónal
Publicado: (2026)
por: Kinsella, Dónal
Publicado: (2026)
Confessions of a shopaholic / Sophie Kinsella
por: Kinsella, Sophie
por: Kinsella, Sophie
ResearchGym: Evaluating Language Model Agents on Real-World AI Research
por: Garikaparthi, Aniketh, et al.
Publicado: (2026)
por: Garikaparthi, Aniketh, et al.
Publicado: (2026)
The Myo Min Aung Unified Theory (MUT) v7.34 : A Complete Theoretical Framework Integrating Spacetime Resonance, Galactic Dynamics, and Causal AI
por: Aung, Myomin
Publicado: (2026)
por: Aung, Myomin
Publicado: (2026)
Etrasimod for alopecia areata: The scenario for a less extensive and moderate form
por: M. Starace
Publicado: (2025)
por: M. Starace
Publicado: (2025)
Deep FinResearch Bench: Evaluating AI's Ability to Conduct Professional Financial Investment Research
por: Haque, Mirazul, et al.
Publicado: (2026)
por: Haque, Mirazul, et al.
Publicado: (2026)
Working with AI: Measuring the Applicability of Generative AI to Occupations
por: Tomlinson, Kiran, et al.
Publicado: (2025)
por: Tomlinson, Kiran, et al.
Publicado: (2025)
SciReplicate-Bench: Benchmarking LLMs in Agent-driven Algorithmic Reproduction from Research Papers
por: Xiang, Yanzheng, et al.
Publicado: (2025)
por: Xiang, Yanzheng, et al.
Publicado: (2025)
"I Never Had to Use the Library in High School": A Library Instruction Program for At-Risk Students
por: Fleming-May, Rachel A., et al.
Publicado: (2015)
por: Fleming-May, Rachel A., et al.
Publicado: (2015)
CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities
por: Zhu, Yuxuan, et al.
Publicado: (2025)
por: Zhu, Yuxuan, et al.
Publicado: (2025)
Farm Policies and Their Impact on Sector Composition and Risk Aversion
por: Theodoros Skevas, et al.
Publicado: (2025)
por: Theodoros Skevas, et al.
Publicado: (2025)
Improved absolute abundance estimates from spatial count data with simulation and microfossil case studies
por: Mays, Chris, et al.
Publicado: (2024)
por: Mays, Chris, et al.
Publicado: (2024)
BetterBench: Assessing AI Benchmarks, Uncovering Issues, and Establishing Best Practices
por: Reuel, Anka, et al.
Publicado: (2024)
por: Reuel, Anka, et al.
Publicado: (2024)
Deliberative Alignment: Reasoning Enables Safer Language Models
por: Guan, Melody Y., et al.
Publicado: (2024)
por: Guan, Melody Y., et al.
Publicado: (2024)
Influence of orotic acid on caprine lipid metabolism
por: Kinsella E. John
Publicado: (1969)
por: Kinsella E. John
Publicado: (1969)
Punição e Proporcionalidade: A Abordagem do Estoppel
por: N. Stephan Kinsella
Publicado: (2016)
por: N. Stephan Kinsella
Publicado: (2016)
ReplicatorBench: Benchmarking LLM Agents for Replicability in Social and Behavioral Sciences
por: Nguyen, Bang, et al.
Publicado: (2026)
por: Nguyen, Bang, et al.
Publicado: (2026)
Artificial Intelligence in Acne Diagnosis and Management: Current Applications and Future Directions
por: Shern‐Ping Choy, et al.
Publicado: (2025)
por: Shern‐Ping Choy, et al.
Publicado: (2025)
Ejemplares similares
-
MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
por: Chan, Jun Shern, et al.
Publicado: (2024) -
SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?
por: Miserendino, Samuel, et al.
Publicado: (2025) -
FrontierScience: Evaluating AI's Ability to Perform Expert-Level Scientific Tasks
por: Wang, Miles, et al.
Publicado: (2026) -
ReplicationBench: Can AI Agents Replicate Astrophysics Research Papers?
por: Ye, Christine, et al.
Publicado: (2025) -
Can Language Models Explain Their Own Classification Behavior?
por: Sherburn, Dane, et al.
Publicado: (2024)