ClarEval: A Benchmark for Evaluating Clarification Skills of Code Agents under Ambiguous Instructions
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Li, Jialin, Wu, Yuan, Chang, Yi |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
ArkEval: Benchmarking and Evaluating Automated CodeRepair for ArkTS
par: Xie, Bang, et autres
Publié: (2026)
par: Xie, Bang, et autres
Publié: (2026)
HumanEvalComm: Benchmarking the Communication Competence of Code Generation for LLMs and LLM Agent
par: Wu, Jie JW, et autres
Publié: (2024)
par: Wu, Jie JW, et autres
Publié: (2024)
ComplexCodeEval: A Benchmark for Evaluating Large Code Models on More Complex Code
par: Feng, Jia, et autres
Publié: (2024)
par: Feng, Jia, et autres
Publié: (2024)
ProjectEval: A Benchmark for Programming Agents Automated Evaluation on Project-Level Code Generation
par: Liu, Kaiyuan, et autres
Publié: (2025)
par: Liu, Kaiyuan, et autres
Publié: (2025)
SolContractEval: A Benchmark for Evaluating Contract-Level Solidity Code Generation
par: Ye, Zhifan, et autres
Publié: (2025)
par: Ye, Zhifan, et autres
Publié: (2025)
FeedbackEval: A Benchmark for Evaluating Large Language Models in Feedback-Driven Code Repair Tasks
par: Dai, Dekun, et autres
Publié: (2025)
par: Dai, Dekun, et autres
Publié: (2025)
SR-Eval: Evaluating LLMs on Code Generation under Stepwise Requirement Refinement
par: Zhan, Zexun, et autres
Publié: (2025)
par: Zhan, Zexun, et autres
Publié: (2025)
ScenEval: A Benchmark for Scenario-Based Evaluation of Code Generation
par: Paul, Debalina Ghosh, et autres
Publié: (2024)
par: Paul, Debalina Ghosh, et autres
Publié: (2024)
SolEval: Benchmarking Large Language Models for Repository-level Solidity Code Generation
par: Peng, Zhiyuan, et autres
Publié: (2025)
par: Peng, Zhiyuan, et autres
Publié: (2025)
ClassEval-T: Evaluating Large Language Models in Class-Level Code Translation
par: Xue, Pengyu, et autres
Publié: (2024)
par: Xue, Pengyu, et autres
Publié: (2024)
CoderEval: A Benchmark of Pragmatic Code Generation with Generative Pre-trained Models
par: Yu, Hao, et autres
Publié: (2023)
par: Yu, Hao, et autres
Publié: (2023)
AdaptEval: A Benchmark for Evaluating Large Language Models on Code Snippet Adaptation
par: Zhang, Tanghaoran, et autres
Publié: (2026)
par: Zhang, Tanghaoran, et autres
Publié: (2026)
Claw-Eval-Live: A Live Agent Benchmark for Evolving Real-World Workflows
par: Li, Chenxin, et autres
Publié: (2026)
par: Li, Chenxin, et autres
Publié: (2026)
Can Code Language Models Learn Clarification-Seeking Behaviors?
par: Wu, Jie JW, et autres
Publié: (2025)
par: Wu, Jie JW, et autres
Publié: (2025)
Automated Repair of Ambiguous Problem Descriptions for LLM-Based Code Generation
par: Jia, Haoxiang, et autres
Publié: (2025)
par: Jia, Haoxiang, et autres
Publié: (2025)
DevEval: A Manually-Annotated Code Generation Benchmark Aligned with Real-World Code Repositories
par: Li, Jia, et autres
Publié: (2024)
par: Li, Jia, et autres
Publié: (2024)
SpecEval: Evaluating Code Comprehension in Large Language Models via Program Specifications
par: Ma, Lezhi, et autres
Publié: (2024)
par: Ma, Lezhi, et autres
Publié: (2024)
Isolating Language-Coding from Problem-Solving: Benchmarking LLMs with PseudoEval
par: Wu, Jiarong, et autres
Publié: (2025)
par: Wu, Jiarong, et autres
Publié: (2025)
DiagEval: Trajectory-Conditioned Diagnosis for Reliable Software Evaluation with GUI Agents
par: Hong, Sirui, et autres
Publié: (2026)
par: Hong, Sirui, et autres
Publié: (2026)
ScratchEval : A Multimodal Evaluation Framework for LLMs in Block-Based Programming
par: Si, Yuan, et autres
Publié: (2026)
par: Si, Yuan, et autres
Publié: (2026)
SkillSmith: Compiling Agent Skills into Boundary-Guided Runtime Interfaces
par: Xu, Duling, et autres
Publié: (2026)
par: Xu, Duling, et autres
Publié: (2026)
ReleaseEval: A Benchmark for Evaluating Language Models in Automated Release Note Generation
par: Meng, Qianru, et autres
Publié: (2025)
par: Meng, Qianru, et autres
Publié: (2025)
RedCode: Risky Code Execution and Generation Benchmark for Code Agents
par: Guo, Chengquan, et autres
Publié: (2024)
par: Guo, Chengquan, et autres
Publié: (2024)
LibEvolutionEval: A Benchmark and Study for Version-Specific Code Generation
par: Kuhar, Sachit, et autres
Publié: (2024)
par: Kuhar, Sachit, et autres
Publié: (2024)
A Benchmark for Evaluating Repository-Level Code Agents with Intermediate Reasoning on Feature Addition Task
par: Liu, Shuhan, et autres
Publié: (2026)
par: Liu, Shuhan, et autres
Publié: (2026)
StackEval: Benchmarking LLMs in Coding Assistance
par: Shah, Nidhish, et autres
Publié: (2024)
par: Shah, Nidhish, et autres
Publié: (2024)
EffiSkill: Agent Skill Based Automated Code Efficiency Optimization
par: Wang, Zimu, et autres
Publié: (2026)
par: Wang, Zimu, et autres
Publié: (2026)
ISD-Agent-Bench: A Comprehensive Benchmark for Evaluating LLM-based Instructional Design Agents
par: Jeon, YoungHoon, et autres
Publié: (2026)
par: Jeon, YoungHoon, et autres
Publié: (2026)
Scaling Coding Agents via Atomic Skills
par: Ma, Yingwei, et autres
Publié: (2026)
par: Ma, Yingwei, et autres
Publié: (2026)
Is Your AI-Generated Code Really Safe? Evaluating Large Language Models on Secure Code Generation with CodeSecEval
par: Wang, Jiexin, et autres
Publié: (2024)
par: Wang, Jiexin, et autres
Publié: (2024)
ClassEval-Pro: A Cross-Domain Benchmark for Class-Level Code Generation
par: Chen, Yeheng, et autres
Publié: (2026)
par: Chen, Yeheng, et autres
Publié: (2026)
When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions
par: Larbi, Maya, et autres
Publié: (2025)
par: Larbi, Maya, et autres
Publié: (2025)
RepoMasterEval: Evaluating Code Completion via Real-World Repositories
par: Wu, Qinyun, et autres
Publié: (2024)
par: Wu, Qinyun, et autres
Publié: (2024)
EvalSVA: Multi-Agent Evaluators for Next-Gen Software Vulnerability Assessment
par: Wen, Xin-Cheng, et autres
Publié: (2024)
par: Wen, Xin-Cheng, et autres
Publié: (2024)
CodeIF: Benchmarking the Instruction-Following Capabilities of Large Language Models for Code Generation
par: Yan, Kaiwen, et autres
Publié: (2025)
par: Yan, Kaiwen, et autres
Publié: (2025)
Automatically Benchmarking LLM Code Agents through Agent-Driven Annotation and Evaluation
par: Fu, Lingyue, et autres
Publié: (2025)
par: Fu, Lingyue, et autres
Publié: (2025)
Is Your Benchmark (Still) Useful? Dynamic Benchmarking for Code Language Models
par: Guan, Batu, et autres
Publié: (2025)
par: Guan, Batu, et autres
Publié: (2025)
A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback
par: Duan, Guoliang, et autres
Publié: (2025)
par: Duan, Guoliang, et autres
Publié: (2025)
AgentArcEval: An Architecture Evaluation Method for Foundation Model based Agents
par: Lu, Qinghua, et autres
Publié: (2025)
par: Lu, Qinghua, et autres
Publié: (2025)
CodeFuse-CommitEval: Towards Benchmarking LLM's Power on Commit Message and Code Change Inconsistency Detection
par: Zhang, Qingyu, et autres
Publié: (2025)
par: Zhang, Qingyu, et autres
Publié: (2025)
Documents similaires
-
ArkEval: Benchmarking and Evaluating Automated CodeRepair for ArkTS
par: Xie, Bang, et autres
Publié: (2026) -
HumanEvalComm: Benchmarking the Communication Competence of Code Generation for LLMs and LLM Agent
par: Wu, Jie JW, et autres
Publié: (2024) -
ComplexCodeEval: A Benchmark for Evaluating Large Code Models on More Complex Code
par: Feng, Jia, et autres
Publié: (2024) -
ProjectEval: A Benchmark for Programming Agents Automated Evaluation on Project-Level Code Generation
par: Liu, Kaiyuan, et autres
Publié: (2025) -
SolContractEval: A Benchmark for Evaluating Contract-Level Solidity Code Generation
par: Ye, Zhifan, et autres
Publié: (2025)