When the Specification Emerges: Benchmarking Faithfulness Loss in Long-Horizon Coding Agents
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Yan, Lu, Chen, Xuan, Zhang, Xiangyu |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
SaaSBench: Exploring the Boundaries of Coding Agents in Long-Horizon Enterprise SaaS Engineering
par: Ren, Qingnan, et autres
Publié: (2026)
par: Ren, Qingnan, et autres
Publié: (2026)
The Conversations Beneath the Code: Triadic Data for Long-Horizon Software Engineering Agents
par: Kim, Yelin
Publié: (2026)
par: Kim, Yelin
Publié: (2026)
SlopCodeBench: Benchmarking How Coding Agents Degrade Over Long-Horizon Iterative Tasks
par: Orlanski, Gabriel, et autres
Publié: (2026)
par: Orlanski, Gabriel, et autres
Publié: (2026)
Code Review Agent Benchmark
par: Zhang, Yuntong, et autres
Publié: (2026)
par: Zhang, Yuntong, et autres
Publié: (2026)
SpecBench: Measuring Reward Hacking in Long-Horizon Coding Agents
par: Zhao, Bingchen, et autres
Publié: (2026)
par: Zhao, Bingchen, et autres
Publié: (2026)
RedCode: Risky Code Execution and Generation Benchmark for Code Agents
par: Guo, Chengquan, et autres
Publié: (2024)
par: Guo, Chengquan, et autres
Publié: (2024)
ProjDevBench: Benchmarking AI Coding Agents on End-to-End Project Development
par: Lu, Pengrui, et autres
Publié: (2026)
par: Lu, Pengrui, et autres
Publié: (2026)
ResearchEnvBench: Benchmarking Agents on Environment Synthesis for Research Code Execution
par: Wang, Yubang, et autres
Publié: (2026)
par: Wang, Yubang, et autres
Publié: (2026)
GitTaskBench: A Benchmark for Code Agents Solving Real-World Tasks Through Code Repository Leveraging
par: Ni, Ziyi, et autres
Publié: (2025)
par: Ni, Ziyi, et autres
Publié: (2025)
LibEvolutionEval: A Benchmark and Study for Version-Specific Code Generation
par: Kuhar, Sachit, et autres
Publié: (2024)
par: Kuhar, Sachit, et autres
Publié: (2024)
LoCoBench-Agent: An Interactive Benchmark for LLM Agents in Long-Context Software Engineering
par: Qiu, Jielin, et autres
Publié: (2025)
par: Qiu, Jielin, et autres
Publié: (2025)
CodeArt: Better Code Models by Attention Regularization When Symbols Are Lacking
par: Su, Zian, et autres
Publié: (2024)
par: Su, Zian, et autres
Publié: (2024)
CoRe: Benchmarking LLMs Code Reasoning Capabilities through Static Analysis Tasks
par: Xie, Danning, et autres
Publié: (2025)
par: Xie, Danning, et autres
Publié: (2025)
Lyra: A Benchmark for Turducken-Style Code Generation
par: Liang, Qingyuan, et autres
Publié: (2021)
par: Liang, Qingyuan, et autres
Publié: (2021)
A New Benchmark for the Appropriate Evaluation of RTL Code Optimization
par: Lu, Yao, et autres
Publié: (2026)
par: Lu, Yao, et autres
Publié: (2026)
DomAgent: Leveraging Knowledge Graphs and Case-Based Reasoning for Domain-Specific Code Generation
par: Wang, Shuai, et autres
Publié: (2026)
par: Wang, Shuai, et autres
Publié: (2026)
RoadmapBench: Evaluating Long-Horizon Agentic Software Development Across Version Upgrades
par: Xu, Xinbo, et autres
Publié: (2026)
par: Xu, Xinbo, et autres
Publié: (2026)
AInsteinBench: Benchmarking Coding Agents on Scientific Repositories
par: Duston, Titouan, et autres
Publié: (2025)
par: Duston, Titouan, et autres
Publié: (2025)
FreshBrew: A Benchmark for Evaluating AI Agents on Java Code Migration
par: May, Victor, et autres
Publié: (2025)
par: May, Victor, et autres
Publié: (2025)
OmniCode: A Benchmark for Evaluating Software Engineering Agents
par: Sonwane, Atharv, et autres
Publié: (2026)
par: Sonwane, Atharv, et autres
Publié: (2026)
A Benchmark for Localizing Code and Non-Code Issues in Software Projects
par: Zhang, Zejun, et autres
Publié: (2025)
par: Zhang, Zejun, et autres
Publié: (2025)
Inferring Code Correctness from Specification
par: Florian, Tambon, et autres
Publié: (2026)
par: Florian, Tambon, et autres
Publié: (2026)
Specification and Detection of LLM Code Smells
par: Mahmoudi, Brahim, et autres
Publié: (2025)
par: Mahmoudi, Brahim, et autres
Publié: (2025)
Unseen Horizons: Unveiling the Real Capability of LLM Code Generation Beyond the Familiar
par: Zhang, Yuanliang, et autres
Publié: (2024)
par: Zhang, Yuanliang, et autres
Publié: (2024)
AdaptEval: A Benchmark for Evaluating Large Language Models on Code Snippet Adaptation
par: Zhang, Tanghaoran, et autres
Publié: (2026)
par: Zhang, Tanghaoran, et autres
Publié: (2026)
Software Development Life Cycle Perspective: A Survey of Benchmarks for Code Large Language Models and Agents
par: Wang, Kaixin, et autres
Publié: (2025)
par: Wang, Kaixin, et autres
Publié: (2025)
When AI Teammates Meet Code Review: Collaboration Signals Shaping the Integration of Agent-Authored Pull Requests
par: Nachuma, Costain, et autres
Publié: (2026)
par: Nachuma, Costain, et autres
Publié: (2026)
Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks
par: Lu, Ruofan, et autres
Publié: (2025)
par: Lu, Ruofan, et autres
Publié: (2025)
Uncovering Systematic Failures of LLMs in Verifying Code Against Natural Language Specifications
par: Jin, Haolin, et autres
Publié: (2025)
par: Jin, Haolin, et autres
Publié: (2025)
ML Code Smells: From Specification to Detection
par: Mahmoudi, Brahim, et autres
Publié: (2025)
par: Mahmoudi, Brahim, et autres
Publié: (2025)
DataGovBench: Benchmarking LLM Agents for Real-World Data Governance Workflows
par: Liu, Zhou, et autres
Publié: (2025)
par: Liu, Zhou, et autres
Publié: (2025)
CodeSense: a Real-World Benchmark and Dataset for Code Semantic Reasoning
par: Roy, Monoshi Kumar, et autres
Publié: (2025)
par: Roy, Monoshi Kumar, et autres
Publié: (2025)
DOMAINEVAL: An Auto-Constructed Benchmark for Multi-Domain Code Generation
par: Zhu, Qiming, et autres
Publié: (2024)
par: Zhu, Qiming, et autres
Publié: (2024)
Semantic Similarity Loss for Neural Source Code Summarization
par: Su, Chia-Yi, et autres
Publié: (2023)
par: Su, Chia-Yi, et autres
Publié: (2023)
Deep Learning for Code Intelligence: Survey, Benchmark and Toolkit
par: Wan, Yao, et autres
Publié: (2023)
par: Wan, Yao, et autres
Publié: (2023)
Automatic Identification of Machine Learning-Specific Code Smells
par: Hamfelt, Peter, et autres
Publié: (2025)
par: Hamfelt, Peter, et autres
Publié: (2025)
Unveiling Project-Specific Bias in Neural Code Models
par: Li, Zhiming, et autres
Publié: (2022)
par: Li, Zhiming, et autres
Publié: (2022)
BabelCoder: Agentic Code Translation with Specification Alignment
par: Rabbi, Fazle, et autres
Publié: (2025)
par: Rabbi, Fazle, et autres
Publié: (2025)
SpareCodeSearch: Searching for Code Context When You Have No Spare GPU
par: Nguyen, Minh
Publié: (2025)
par: Nguyen, Minh
Publié: (2025)
SWE-AGI: Benchmarking Specification-Driven Software Construction with MoonBit in the Era of Autonomous Agents
par: Zhang, Zhirui, et autres
Publié: (2026)
par: Zhang, Zhirui, et autres
Publié: (2026)
Documents similaires
-
SaaSBench: Exploring the Boundaries of Coding Agents in Long-Horizon Enterprise SaaS Engineering
par: Ren, Qingnan, et autres
Publié: (2026) -
The Conversations Beneath the Code: Triadic Data for Long-Horizon Software Engineering Agents
par: Kim, Yelin
Publié: (2026) -
SlopCodeBench: Benchmarking How Coding Agents Degrade Over Long-Horizon Iterative Tasks
par: Orlanski, Gabriel, et autres
Publié: (2026) -
Code Review Agent Benchmark
par: Zhang, Yuntong, et autres
Publié: (2026) -
SpecBench: Measuring Reward Hacking in Long-Horizon Coding Agents
par: Zhao, Bingchen, et autres
Publié: (2026)