RExBench: Can coding agents autonomously implement AI research extensions?
Fuente:
arXiv
Saved in:
| Main Authors: | Edwards, Nicholas, Lee, Yukyung, Mao, Yujun Audrey, Qin, Yulu, Schuster, Sebastian, Kim, Najoung |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Code Pretraining Improves Entity Tracking Abilities of Language Models
by: Kim, Najoung, et al.
Published: (2024)
by: Kim, Najoung, et al.
Published: (2024)
CheckEval: A reliable LLM-as-a-Judge framework for evaluating text generation using checklists
by: Lee, Yukyung, et al.
Published: (2024)
by: Lee, Yukyung, et al.
Published: (2024)
Ask or Assume? Uncertainty-Aware Clarification-Seeking in Coding Agents
by: Edwards, Nicholas, et al.
Published: (2026)
by: Edwards, Nicholas, et al.
Published: (2026)
Vision-and-Language Training Helps Deploy Taxonomic Knowledge but Does Not Fundamentally Alter It
by: Qin, Yulu, et al.
Published: (2025)
by: Qin, Yulu, et al.
Published: (2025)
A systematic framework for generating novel experimental hypotheses from language models
by: Misra, Kanishka, et al.
Published: (2024)
by: Misra, Kanishka, et al.
Published: (2024)
Can Structural Cues Save LLMs? Evaluating Language Models in Massive Document Streams
by: Lee, Yukyung, et al.
Published: (2026)
by: Lee, Yukyung, et al.
Published: (2026)
Do Language Models Track Entities Across State Changes?
by: Tang, Zilu, et al.
Published: (2026)
by: Tang, Zilu, et al.
Published: (2026)
Is analogy enough to draw novel adjective-noun inferences?
by: Ross, Hayley, et al.
Published: (2025)
by: Ross, Hayley, et al.
Published: (2025)
Syn-QA2: Evaluating False Assumptions in Long-tail Questions with Synthetic QA Datasets
by: Daswani, Ashwin, et al.
Published: (2024)
by: Daswani, Ashwin, et al.
Published: (2024)
Is artificial intelligence still intelligence? LLMs generalize to novel adjective-noun pairs, but don't mimic the full human distribution
by: Ross, Hayley, et al.
Published: (2024)
by: Ross, Hayley, et al.
Published: (2024)
Can LLM-Augmented autonomous agents cooperate?, An evaluation of their cooperative capabilities through Melting Pot
by: Mosquera, Manuel, et al.
Published: (2024)
by: Mosquera, Manuel, et al.
Published: (2024)
RedacBench: Can AI Erase Your Secrets?
by: Jeon, Hyunjun, et al.
Published: (2026)
by: Jeon, Hyunjun, et al.
Published: (2026)
Death of the Novel(ty): Beyond n-Gram Novelty as a Metric for Textual Creativity
by: Saakyan, Arkadiy, et al.
Published: (2025)
by: Saakyan, Arkadiy, et al.
Published: (2025)
RareBench: Can LLMs Serve as Rare Diseases Specialists?
by: Chen, Xuanzhong, et al.
Published: (2024)
by: Chen, Xuanzhong, et al.
Published: (2024)
CIRF: Tokenizing Chain-of-Thoughts into Reusable Functional Units for Efficient Latent Reasoning in Large Language Models
by: Lee, Yukyung, et al.
Published: (2026)
by: Lee, Yukyung, et al.
Published: (2026)
The why, what, and how of AI-based coding in scientific research
by: Zhuang, Tonghe, et al.
Published: (2024)
by: Zhuang, Tonghe, et al.
Published: (2024)
CHAMP: A Competition-level Dataset for Fine-Grained Analyses of LLMs' Mathematical Reasoning Capabilities
by: Mao, Yujun, et al.
Published: (2024)
by: Mao, Yujun, et al.
Published: (2024)
AuditoryBench++: Can Language Models Understand Auditory Knowledge without Hearing?
by: Ok, Hyunjong, et al.
Published: (2025)
by: Ok, Hyunjong, et al.
Published: (2025)
Personas as a Way to Model Truthfulness in Language Models
by: Joshi, Nitish, et al.
Published: (2023)
by: Joshi, Nitish, et al.
Published: (2023)
Navigating the Path of Writing: Outline-guided Text Generation with Large Language Models
by: Lee, Yukyung, et al.
Published: (2024)
by: Lee, Yukyung, et al.
Published: (2024)
ClawBench: Can AI Agents Complete Everyday Online Tasks?
by: Zhang, Yuxuan, et al.
Published: (2026)
by: Zhang, Yuxuan, et al.
Published: (2026)
REPRO-Bench: Can Agentic AI Systems Assess the Reproducibility of Social Science Research?
by: Hu, Chuxuan, et al.
Published: (2025)
by: Hu, Chuxuan, et al.
Published: (2025)
ReplicationBench: Can AI Agents Replicate Astrophysics Research Papers?
by: Ye, Christine, et al.
Published: (2025)
by: Ye, Christine, et al.
Published: (2025)
SysBench: Can Large Language Models Follow System Messages?
by: Qin, Yanzhao, et al.
Published: (2024)
by: Qin, Yanzhao, et al.
Published: (2024)
Diffusion LLMs can think EoS-by-EoS
by: Breckner, Sarah, et al.
Published: (2026)
by: Breckner, Sarah, et al.
Published: (2026)
Bench-2-CoP: Can We Trust Benchmarking for EU AI Compliance?
by: Prandi, Matteo, et al.
Published: (2025)
by: Prandi, Matteo, et al.
Published: (2025)
LLMs Still Can't Plan; Can LRMs? A Preliminary Evaluation of OpenAI's o1 on PlanBench
by: Valmeekam, Karthik, et al.
Published: (2024)
by: Valmeekam, Karthik, et al.
Published: (2024)
A systematic investigation of learnability from single child linguistic input
by: Qin, Yulu, et al.
Published: (2024)
by: Qin, Yulu, et al.
Published: (2024)
Are they lovers or friends? Evaluating LLMs' Social Reasoning in English and Korean Dialogues
by: Kim, Eunsu, et al.
Published: (2025)
by: Kim, Eunsu, et al.
Published: (2025)
HAE-RAE Bench: Evaluation of Korean Knowledge in Language Models
by: Son, Guijin, et al.
Published: (2023)
by: Son, Guijin, et al.
Published: (2023)
KVoiceBench, KOpenAudioBench, and KMMAU: Agent-Driven Korean Speech Benchmarks for Evaluating SpeechLMs
by: Kim, Haechan, et al.
Published: (2026)
by: Kim, Haechan, et al.
Published: (2026)
SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
by: Deng, Xiang, et al.
Published: (2025)
by: Deng, Xiang, et al.
Published: (2025)
FilBench: Can LLMs Understand and Generate Filipino?
by: Miranda, Lester James V., et al.
Published: (2025)
by: Miranda, Lester James V., et al.
Published: (2025)
On the Consideration of AI Openness: Can Good Intent Be Abused?
by: Kim, Yeeun, et al.
Published: (2024)
by: Kim, Yeeun, et al.
Published: (2024)
Can large language models build causal graphs?
by: Long, Stephanie, et al.
Published: (2023)
by: Long, Stephanie, et al.
Published: (2023)
CHI-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows?
by: Chen, Haolin, et al.
Published: (2026)
by: Chen, Haolin, et al.
Published: (2026)
A Behavior Tree-inspired programming language for autonomous agents
by: Biggar, Oliver, et al.
Published: (2024)
by: Biggar, Oliver, et al.
Published: (2024)
AbsenceBench: Language Models Can't Tell What's Missing
by: Fu, Harvey Yiyun, et al.
Published: (2025)
by: Fu, Harvey Yiyun, et al.
Published: (2025)
QuestBench: Can LLMs ask the right question to acquire information in reasoning tasks?
by: Li, Belinda Z., et al.
Published: (2025)
by: Li, Belinda Z., et al.
Published: (2025)
ConflictBench: Evaluating Human-AI Conflict via Interactive and Visually Grounded Environments
by: Zhao, Weixiang, et al.
Published: (2026)
by: Zhao, Weixiang, et al.
Published: (2026)
Similar Items
-
Code Pretraining Improves Entity Tracking Abilities of Language Models
by: Kim, Najoung, et al.
Published: (2024) -
CheckEval: A reliable LLM-as-a-Judge framework for evaluating text generation using checklists
by: Lee, Yukyung, et al.
Published: (2024) -
Ask or Assume? Uncertainty-Aware Clarification-Seeking in Coding Agents
by: Edwards, Nicholas, et al.
Published: (2026) -
Vision-and-Language Training Helps Deploy Taxonomic Knowledge but Does Not Fundamentally Alter It
by: Qin, Yulu, et al.
Published: (2025) -
A systematic framework for generating novel experimental hypotheses from language models
by: Misra, Kanishka, et al.
Published: (2024)