Agent-Diff: Benchmarking LLM Agents on Enterprise API Tasks via Code Execution with State-Diff-Based Evaluation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Pysklo, Hubert M., Zhuravel, Artem, Watson, Patrick D. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
To Diff or Not to Diff? Structure-Aware and Adaptive Output Formats for Efficient LLM-based Code Editing
von: Cheng, Wei, et al.
Veröffentlicht: (2026)
von: Cheng, Wei, et al.
Veröffentlicht: (2026)
Automatically Benchmarking LLM Code Agents through Agent-Driven Annotation and Evaluation
von: Fu, Lingyue, et al.
Veröffentlicht: (2025)
von: Fu, Lingyue, et al.
Veröffentlicht: (2025)
FireBench: Evaluating Instruction Following in Enterprise and API-Driven LLM Applications
von: Zhang, Yunfan, et al.
Veröffentlicht: (2026)
von: Zhang, Yunfan, et al.
Veröffentlicht: (2026)
Is Vibe Coding Safe? Benchmarking Vulnerability of Agent-Generated Code in Real-World Tasks
von: Zhao, Songwen, et al.
Veröffentlicht: (2025)
von: Zhao, Songwen, et al.
Veröffentlicht: (2025)
ISD-Agent-Bench: A Comprehensive Benchmark for Evaluating LLM-based Instructional Design Agents
von: Jeon, YoungHoon, et al.
Veröffentlicht: (2026)
von: Jeon, YoungHoon, et al.
Veröffentlicht: (2026)
CodeUpdateArena: Benchmarking Knowledge Editing on API Updates
von: Liu, Zeyu Leo, et al.
Veröffentlicht: (2024)
von: Liu, Zeyu Leo, et al.
Veröffentlicht: (2024)
SUPER: Evaluating Agents on Setting Up and Executing Tasks from Research Repositories
von: Bogin, Ben, et al.
Veröffentlicht: (2024)
von: Bogin, Ben, et al.
Veröffentlicht: (2024)
ProjectEval: A Benchmark for Programming Agents Automated Evaluation on Project-Level Code Generation
von: Liu, Kaiyuan, et al.
Veröffentlicht: (2025)
von: Liu, Kaiyuan, et al.
Veröffentlicht: (2025)
OmniCode: A Benchmark for Evaluating Software Engineering Agents
von: Sonwane, Atharv, et al.
Veröffentlicht: (2026)
von: Sonwane, Atharv, et al.
Veröffentlicht: (2026)
CodeBenchGen: Creating Scalable Execution-based Code Generation Benchmarks
von: Xie, Yiqing, et al.
Veröffentlicht: (2024)
von: Xie, Yiqing, et al.
Veröffentlicht: (2024)
CodeR: Issue Resolving with Multi-Agent and Task Graphs
von: Chen, Dong, et al.
Veröffentlicht: (2024)
von: Chen, Dong, et al.
Veröffentlicht: (2024)
GrowthHacker: Automated Off-Policy Evaluation Optimization Using Code-Modifying LLM Agents
von: Wu, Jie JW, et al.
Veröffentlicht: (2025)
von: Wu, Jie JW, et al.
Veröffentlicht: (2025)
SlopCodeBench: Benchmarking How Coding Agents Degrade Over Long-Horizon Iterative Tasks
von: Orlanski, Gabriel, et al.
Veröffentlicht: (2026)
von: Orlanski, Gabriel, et al.
Veröffentlicht: (2026)
Nexus: Execution-Grounded Multi-Agent Test Oracle Synthesis
von: Huang, Dong, et al.
Veröffentlicht: (2025)
von: Huang, Dong, et al.
Veröffentlicht: (2025)
CodeSpecBench: Benchmarking LLMs for Executable Behavioral Specification Generation
von: Chen, Zaoyu, et al.
Veröffentlicht: (2026)
von: Chen, Zaoyu, et al.
Veröffentlicht: (2026)
KAIJU: An Executive Kernel for Intent-Gated Execution of LLM Agents
von: Guerin, Cormac, et al.
Veröffentlicht: (2026)
von: Guerin, Cormac, et al.
Veröffentlicht: (2026)
LocAgent: Graph-Guided LLM Agents for Code Localization
von: Chen, Zhaoling, et al.
Veröffentlicht: (2025)
von: Chen, Zhaoling, et al.
Veröffentlicht: (2025)
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks
von: Jiang, Hongchao, et al.
Veröffentlicht: (2025)
von: Jiang, Hongchao, et al.
Veröffentlicht: (2025)
EffiSkill: Agent Skill Based Automated Code Efficiency Optimization
von: Wang, Zimu, et al.
Veröffentlicht: (2026)
von: Wang, Zimu, et al.
Veröffentlicht: (2026)
Terminal Agents Suffice for Enterprise Automation
von: Bechard, Patrice, et al.
Veröffentlicht: (2026)
von: Bechard, Patrice, et al.
Veröffentlicht: (2026)
StepCodeReasoner: Aligning Code Reasoning with Stepwise Execution Traces via Reinforcement Learning
von: Wang, Hao, et al.
Veröffentlicht: (2026)
von: Wang, Hao, et al.
Veröffentlicht: (2026)
Diff-XYZ: A Benchmark for Evaluating Diff Understanding
von: Glukhov, Evgeniy, et al.
Veröffentlicht: (2025)
von: Glukhov, Evgeniy, et al.
Veröffentlicht: (2025)
CodeHalu: Investigating Code Hallucinations in LLMs via Execution-based Verification
von: Tian, Yuchen, et al.
Veröffentlicht: (2024)
von: Tian, Yuchen, et al.
Veröffentlicht: (2024)
From SWE-ZERO to SWE-HERO: Execution-free to Execution-based Fine-tuning for Software Engineering Agents
von: Ludwig, Nikolai, et al.
Veröffentlicht: (2026)
von: Ludwig, Nikolai, et al.
Veröffentlicht: (2026)
CodeScope: An Execution-based Multilingual Multitask Multidimensional Benchmark for Evaluating LLMs on Code Understanding and Generation
von: Yan, Weixiang, et al.
Veröffentlicht: (2023)
von: Yan, Weixiang, et al.
Veröffentlicht: (2023)
ArkTS-CodeSearch: A Open-Source ArkTS Dataset for Code Retrieval
von: He, Yulong, et al.
Veröffentlicht: (2026)
von: He, Yulong, et al.
Veröffentlicht: (2026)
Enhancing Project-Specific Code Completion by Inferring Internal API Information
von: Deng, Le, et al.
Veröffentlicht: (2025)
von: Deng, Le, et al.
Veröffentlicht: (2025)
SWE-rebench: An Automated Pipeline for Task Collection and Decontaminated Evaluation of Software Engineering Agents
von: Badertdinov, Ibragim, et al.
Veröffentlicht: (2025)
von: Badertdinov, Ibragim, et al.
Veröffentlicht: (2025)
AgentPack: A Dataset of Code Changes, Co-Authored by Agents and Humans
von: Zi, Yangtian, et al.
Veröffentlicht: (2025)
von: Zi, Yangtian, et al.
Veröffentlicht: (2025)
Continuous Benchmark Generation for Evaluating Enterprise-scale LLM Agents
von: Saxena, Divyanshu, et al.
Veröffentlicht: (2025)
von: Saxena, Divyanshu, et al.
Veröffentlicht: (2025)
Hybrid-Gym: Training Coding Agents to Generalize Across Tasks
von: Xie, Yiqing, et al.
Veröffentlicht: (2026)
von: Xie, Yiqing, et al.
Veröffentlicht: (2026)
MERA Code: A Unified Framework for Evaluating Code Generation Across Tasks
von: Chervyakov, Artem, et al.
Veröffentlicht: (2025)
von: Chervyakov, Artem, et al.
Veröffentlicht: (2025)
Evaluating and Achieving Controllable Code Completion in Code LLM
von: Zhang, Jiajun, et al.
Veröffentlicht: (2026)
von: Zhang, Jiajun, et al.
Veröffentlicht: (2026)
BenchGuard: Who Guards the Benchmarks? Automated Auditing of LLM Agent Benchmarks
von: Tu, Xinming, et al.
Veröffentlicht: (2026)
von: Tu, Xinming, et al.
Veröffentlicht: (2026)
EvoCodeBench: An Evolving Code Generation Benchmark with Domain-Specific Evaluations
von: Li, Jia, et al.
Veröffentlicht: (2024)
von: Li, Jia, et al.
Veröffentlicht: (2024)
Comparing Developer and LLM Biases in Code Evaluation
von: Mittal, Aditya, et al.
Veröffentlicht: (2026)
von: Mittal, Aditya, et al.
Veröffentlicht: (2026)
MATCH: Task-Driven Code Evaluation through Contrastive Learning
von: Ghoummaid, Marah, et al.
Veröffentlicht: (2025)
von: Ghoummaid, Marah, et al.
Veröffentlicht: (2025)
A Conceptual Framework for API Refactoring in Enterprise Application Architectures
von: Montesi, Fabrizio, et al.
Veröffentlicht: (2024)
von: Montesi, Fabrizio, et al.
Veröffentlicht: (2024)
CodeScout: Contextual Problem Statement Enhancement for Software Agents
von: Suri, Manan, et al.
Veröffentlicht: (2026)
von: Suri, Manan, et al.
Veröffentlicht: (2026)
SWE-Pruner: Self-Adaptive Context Pruning for Coding Agents
von: Wang, Yuhang, et al.
Veröffentlicht: (2026)
von: Wang, Yuhang, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
To Diff or Not to Diff? Structure-Aware and Adaptive Output Formats for Efficient LLM-based Code Editing
von: Cheng, Wei, et al.
Veröffentlicht: (2026) -
Automatically Benchmarking LLM Code Agents through Agent-Driven Annotation and Evaluation
von: Fu, Lingyue, et al.
Veröffentlicht: (2025) -
FireBench: Evaluating Instruction Following in Enterprise and API-Driven LLM Applications
von: Zhang, Yunfan, et al.
Veröffentlicht: (2026) -
Is Vibe Coding Safe? Benchmarking Vulnerability of Agent-Generated Code in Real-World Tasks
von: Zhao, Songwen, et al.
Veröffentlicht: (2025) -
ISD-Agent-Bench: A Comprehensive Benchmark for Evaluating LLM-based Instructional Design Agents
von: Jeon, YoungHoon, et al.
Veröffentlicht: (2026)