From Reproduction to Replication: Evaluating Research Agents with Progressive Code Masking
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Kim, Gyeongwon James, Wilf, Alex, Morency, Louis-Philippe, Fried, Daniel |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Propose, Solve, Verify: Self-Play Through Formal Verification
par: Wilf, Alex, et autres
Publié: (2025)
par: Wilf, Alex, et autres
Publié: (2025)
SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents
par: Zhou, Xuhui, et autres
Publié: (2023)
par: Zhou, Xuhui, et autres
Publié: (2023)
Improving Dialogue Agents by Decomposing One Global Explicit Annotation with Local Implicit Multimodal Feedback
par: Lee, Dong Won, et autres
Publié: (2024)
par: Lee, Dong Won, et autres
Publié: (2024)
CRScore: Grounding Automated Evaluation of Code Review Comments in Code Claims and Smells
par: Naik, Atharva, et autres
Publié: (2024)
par: Naik, Atharva, et autres
Publié: (2024)
SciReplicate-Bench: Benchmarking LLMs in Agent-driven Algorithmic Reproduction from Research Papers
par: Xiang, Yanzheng, et autres
Publié: (2025)
par: Xiang, Yanzheng, et autres
Publié: (2025)
RepliBench: Evaluating the Autonomous Replication Capabilities of Language Model Agents
par: Black, Sid, et autres
Publié: (2025)
par: Black, Sid, et autres
Publié: (2025)
MultiIoT: Benchmarking Machine Learning for the Internet of Things
par: Mo, Shentong, et autres
Publié: (2023)
par: Mo, Shentong, et autres
Publié: (2023)
IoT-LM: Large Multisensory Language Models for the Internet of Things
par: Mo, Shentong, et autres
Publié: (2024)
par: Mo, Shentong, et autres
Publié: (2024)
PaperBench: Evaluating AI's Ability to Replicate AI Research
par: Starace, Giulio, et autres
Publié: (2025)
par: Starace, Giulio, et autres
Publié: (2025)
HEMM: Holistic Evaluation of Multimodal Foundation Models
par: Liang, Paul Pu, et autres
Publié: (2024)
par: Liang, Paul Pu, et autres
Publié: (2024)
Human-Agent Cooperation in Games under Incomplete Information through Natural Language Communication
par: Chen, Shenghui, et autres
Publié: (2024)
par: Chen, Shenghui, et autres
Publié: (2024)
Agent psychometrics: Task-level performance prediction in agentic coding benchmarks
par: Ge, Chris, et autres
Publié: (2026)
par: Ge, Chris, et autres
Publié: (2026)
ECCO: Can We Improve Model-Generated Code Efficiency Without Sacrificing Functional Correctness?
par: Waghjale, Siddhant, et autres
Publié: (2024)
par: Waghjale, Siddhant, et autres
Publié: (2024)
API-Assisted Code Generation for Question Answering on Varied Table Structures
par: Cao, Yihan, et autres
Publié: (2023)
par: Cao, Yihan, et autres
Publié: (2023)
Social Human Robot Embodied Conversation (SHREC) Dataset: Benchmarking Foundational Models' Social Reasoning
par: Lee, Dong Won, et autres
Publié: (2025)
par: Lee, Dong Won, et autres
Publié: (2025)
ReplicatorBench: Benchmarking LLM Agents for Replicability in Social and Behavioral Sciences
par: Nguyen, Bang, et autres
Publié: (2026)
par: Nguyen, Bang, et autres
Publié: (2026)
Empirical Evaluation of Progressive Coding for Sparse Autoencoders
par: Peter, Hans, et autres
Publié: (2025)
par: Peter, Hans, et autres
Publié: (2025)
Design Principles for Falsifiable, Replicable and Reproducible Empirical ML Research
par: Vranješ, Daniel, et autres
Publié: (2024)
par: Vranješ, Daniel, et autres
Publié: (2024)
Is the Pope Catholic? Yes, the Pope is Catholic. Generative Evaluation of Non-Literal Intent Resolution in LLMs
par: Yerukola, Akhila, et autres
Publié: (2024)
par: Yerukola, Akhila, et autres
Publié: (2024)
TroVE: Inducing Verifiable and Efficient Toolboxes for Solving Programmatic Tasks
par: Wang, Zhiruo, et autres
Publié: (2024)
par: Wang, Zhiruo, et autres
Publié: (2024)
Normative Common Ground Replication (NormCoRe): Replication-by-Translation for Studying Norms in Multi-Agent AI
par: Deck, Luca, et autres
Publié: (2026)
par: Deck, Luca, et autres
Publié: (2026)
Deep Research Bench: Evaluating AI Web Research Agents
par: FutureSearch, et autres
Publié: (2025)
par: FutureSearch, et autres
Publié: (2025)
o1-Coder: an o1 Replication for Coding
par: Zhang, Yuxiang, et autres
Publié: (2024)
par: Zhang, Yuxiang, et autres
Publié: (2024)
Tree Search for Language Model Agents
par: Koh, Jing Yu, et autres
Publié: (2024)
par: Koh, Jing Yu, et autres
Publié: (2024)
Embodied AI Agents: Modeling the World
par: Fung, Pascale, et autres
Publié: (2025)
par: Fung, Pascale, et autres
Publié: (2025)
Code Researcher: Deep Research Agent for Large Systems Code and Commit History
par: Singh, Ramneet, et autres
Publié: (2025)
par: Singh, Ramneet, et autres
Publié: (2025)
ProgAgent:A Continual RL Agent with Progress-Aware Rewards
par: Tan, Jinzhou, et autres
Publié: (2026)
par: Tan, Jinzhou, et autres
Publié: (2026)
O1 Replication Journey: A Strategic Progress Report -- Part 1
par: Qin, Yiwei, et autres
Publié: (2024)
par: Qin, Yiwei, et autres
Publié: (2024)
Dive into the Agent Matrix: A Realistic Evaluation of Self-Replication Risk in LLM Agents
par: Zhang, Boxuan, et autres
Publié: (2025)
par: Zhang, Boxuan, et autres
Publié: (2025)
Read the Paper, Write the Code: Agentic Reproduction of Social-Science Results
par: Kohler, Benjamin, et autres
Publié: (2026)
par: Kohler, Benjamin, et autres
Publié: (2026)
Evaluating Stochasticity in Deep Research Agents
par: Zhai, Haotian, et autres
Publié: (2026)
par: Zhai, Haotian, et autres
Publié: (2026)
Identifying & Interactively Refining Ambiguous User Goals for Data Visualization Code Generation
par: İnan, Mert, et autres
Publié: (2025)
par: İnan, Mert, et autres
Publié: (2025)
CodeClinic: Evaluating Automation of Coding Skills for Clinical Reasoning Agents
par: Ossowski, Timothy, et autres
Publié: (2026)
par: Ossowski, Timothy, et autres
Publié: (2026)
From Weak Cues to Real Identities: Evaluating Inference-Driven De-Anonymization in LLM Agents
par: Ko, Myeongseob, et autres
Publié: (2026)
par: Ko, Myeongseob, et autres
Publié: (2026)
From Task Executors to Research Partners: Evaluating AI Co-Pilots Through Workflow Integration in Biomedical Research
par: Weidener, Lukas, et autres
Publié: (2025)
par: Weidener, Lukas, et autres
Publié: (2025)
MaskMA: Towards Zero-Shot Multi-Agent Decision Making with Mask-Based Collaborative Learning
par: Liu, Jie, et autres
Publié: (2023)
par: Liu, Jie, et autres
Publié: (2023)
Progressive Multi-Agent Reasoning for Biological Perturbation Prediction
par: Kim, Hyomin, et autres
Publié: (2026)
par: Kim, Hyomin, et autres
Publié: (2026)
Enhancing Automated Paper Reproduction via Prompt-Free Collaborative Agents
par: Lin, Zijie, et autres
Publié: (2025)
par: Lin, Zijie, et autres
Publié: (2025)
Reflective Paper-to-Code Reproduction Enabled by Fine-Grained Verification
par: Zhou, Mingyang, et autres
Publié: (2025)
par: Zhou, Mingyang, et autres
Publié: (2025)
BiomechAgent: AI-Assisted Biomechanical Analysis Through Code-Generating Agents
par: Cotton, R. James, et autres
Publié: (2026)
par: Cotton, R. James, et autres
Publié: (2026)
Documents similaires
-
Propose, Solve, Verify: Self-Play Through Formal Verification
par: Wilf, Alex, et autres
Publié: (2025) -
SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents
par: Zhou, Xuhui, et autres
Publié: (2023) -
Improving Dialogue Agents by Decomposing One Global Explicit Annotation with Local Implicit Multimodal Feedback
par: Lee, Dong Won, et autres
Publié: (2024) -
CRScore: Grounding Automated Evaluation of Code Review Comments in Code Claims and Smells
par: Naik, Atharva, et autres
Publié: (2024) -
SciReplicate-Bench: Benchmarking LLMs in Agent-driven Algorithmic Reproduction from Research Papers
par: Xiang, Yanzheng, et autres
Publié: (2025)