BUILD-AND-FIND: An Effort-Aware Protocol for Evaluating Agent-Managed Codebases
Fuente:
arXiv
Salvato in:
| Autore principale: | Lin, Jhen-Ke |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Refactoring Codebases through Library Design
di: Kovacic, Ziga, et al.
Pubblicazione: (2025)
di: Kovacic, Ziga, et al.
Pubblicazione: (2025)
SWE-CI: Evaluating Agent Capabilities in Maintaining Codebases via Continuous Integration
di: Chen, Jialong, et al.
Pubblicazione: (2026)
di: Chen, Jialong, et al.
Pubblicazione: (2026)
Meta-RAG on Large Codebases Using Code Summarization
di: Tawosi, Vali, et al.
Pubblicazione: (2025)
di: Tawosi, Vali, et al.
Pubblicazione: (2025)
Exploring the Challenges and Opportunities of AI-assisted Codebase Generation
di: Eibl, Philipp, et al.
Pubblicazione: (2025)
di: Eibl, Philipp, et al.
Pubblicazione: (2025)
FormulaCode: Evaluating Agentic Optimization on Large Codebases
di: Sehgal, Atharva, et al.
Pubblicazione: (2026)
di: Sehgal, Atharva, et al.
Pubblicazione: (2026)
AstraAI: LLMs, Retrieval, and AST-Guided Assistance for HPC Codebases
di: Natarajan, Mahesh, et al.
Pubblicazione: (2026)
di: Natarajan, Mahesh, et al.
Pubblicazione: (2026)
A Note on Code Quality Score: LLMs for Maintainable Large Codebases
di: Wong, Sherman, et al.
Pubblicazione: (2025)
di: Wong, Sherman, et al.
Pubblicazione: (2025)
Confucius Code Agent: Scalable Agent Scaffolding for Real-World Codebases
di: Wong, Sherman, et al.
Pubblicazione: (2025)
di: Wong, Sherman, et al.
Pubblicazione: (2025)
The Kitchen Loop: User-Spec-Driven Development for a Self-Evolving Codebase
di: Roy, Yannick
Pubblicazione: (2026)
di: Roy, Yannick
Pubblicazione: (2026)
Reliable Graph-RAG for Codebases: AST-Derived Graphs vs LLM-Extracted Knowledge Graphs
di: Chinthareddy, Manideep Reddy
Pubblicazione: (2026)
di: Chinthareddy, Manideep Reddy
Pubblicazione: (2026)
AI-Guided Exploration of Large-Scale Codebases
di: Alebachew, Yoseph Berhanu
Pubblicazione: (2025)
di: Alebachew, Yoseph Berhanu
Pubblicazione: (2025)
LDP: An Identity-Aware Protocol for Multi-Agent LLM Systems
di: Prakash, Sunil
Pubblicazione: (2026)
di: Prakash, Sunil
Pubblicazione: (2026)
RPG: A Repository Planning Graph for Unified and Scalable Codebase Generation
di: Luo, Jane, et al.
Pubblicazione: (2025)
di: Luo, Jane, et al.
Pubblicazione: (2025)
RFCAudit: An LLM Agent for Functional Bug Detection in Network Protocols
di: Zheng, Mingwei, et al.
Pubblicazione: (2025)
di: Zheng, Mingwei, et al.
Pubblicazione: (2025)
Agile Software Effort Estimation using Regression Techniques
di: Sima, Sisay Deresa, et al.
Pubblicazione: (2025)
di: Sima, Sisay Deresa, et al.
Pubblicazione: (2025)
AgentLens: Revealing The Lucky Pass Problem in SWE-Agent Evaluation
di: Sahoo, Priyam, et al.
Pubblicazione: (2026)
di: Sahoo, Priyam, et al.
Pubblicazione: (2026)
HAFixAgent: History-Aware Program Repair Agent
di: Shi, Yu, et al.
Pubblicazione: (2025)
di: Shi, Yu, et al.
Pubblicazione: (2025)
iPanda: An LLM-based Agent for Automated Conformance Testing of Communication Protocols
di: Sun, Xikai, et al.
Pubblicazione: (2025)
di: Sun, Xikai, et al.
Pubblicazione: (2025)
SPICE: An Automated SWE-Bench Labeling Pipeline for Issue Clarity, Test Coverage, and Effort Estimation
di: Oliva, Gustavo A., et al.
Pubblicazione: (2025)
di: Oliva, Gustavo A., et al.
Pubblicazione: (2025)
Agora: Toward Autonomous Bug Detection in Production-Level Consensus Protocols with LLM Agents
di: Liu, Xiang, et al.
Pubblicazione: (2026)
di: Liu, Xiang, et al.
Pubblicazione: (2026)
Leveraging AI for Enhanced Software Effort Estimation: A Comprehensive Study and Framework Proposal
di: Tran, Nhi, et al.
Pubblicazione: (2024)
di: Tran, Nhi, et al.
Pubblicazione: (2024)
Text Tells the Cost: Predicting and Analyzing Repayment Effort of Self-Admitted Technical Debt
di: Li, Yikun, et al.
Pubblicazione: (2023)
di: Li, Yikun, et al.
Pubblicazione: (2023)
RobuNFR: Evaluating the Robustness of Large Language Models on Non-Functional Requirements Aware Code Generation
di: Lin, Feng, et al.
Pubblicazione: (2025)
di: Lin, Feng, et al.
Pubblicazione: (2025)
OSS-UAgent: An Agent-based Usability Evaluation Framework for Open Source Software
di: Meng, Lingkai, et al.
Pubblicazione: (2025)
di: Meng, Lingkai, et al.
Pubblicazione: (2025)
Bridging Protocol and Production: Design Patterns for Deploying AI Agents with Model Context Protocol
di: Srinivasan, Vasundra
Pubblicazione: (2026)
di: Srinivasan, Vasundra
Pubblicazione: (2026)
AgentArcEval: An Architecture Evaluation Method for Foundation Model based Agents
di: Lu, Qinghua, et al.
Pubblicazione: (2025)
di: Lu, Qinghua, et al.
Pubblicazione: (2025)
Evaluating Agent-based Program Repair at Google
di: Rondon, Pat, et al.
Pubblicazione: (2025)
di: Rondon, Pat, et al.
Pubblicazione: (2025)
Efficient Failure Management for Multi-Agent Systems with Reasoning Trace Representation
di: Zhang, Lingzhe, et al.
Pubblicazione: (2026)
di: Zhang, Lingzhe, et al.
Pubblicazione: (2026)
SWE-Effi: Re-Evaluating Software AI Agent System Effectiveness Under Resource Constraints
di: Fan, Zhiyu, et al.
Pubblicazione: (2025)
di: Fan, Zhiyu, et al.
Pubblicazione: (2025)
Towards Structured, State-Aware, and Execution-Grounded Reasoning for Software Engineering Agents
di: Tse-Hsun, et al.
Pubblicazione: (2026)
di: Tse-Hsun, et al.
Pubblicazione: (2026)
EvoClaw: Evaluating AI Agents on Continuous Software Evolution
di: Deng, Gangda, et al.
Pubblicazione: (2026)
di: Deng, Gangda, et al.
Pubblicazione: (2026)
The Complexity Trap: Simple Observation Masking Is as Efficient as LLM Summarization for Agent Context Management
di: Lindenbauer, Tobias, et al.
Pubblicazione: (2025)
di: Lindenbauer, Tobias, et al.
Pubblicazione: (2025)
WebSuite: Systematically Evaluating Why Web Agents Fail
di: Li, Eric, et al.
Pubblicazione: (2024)
di: Li, Eric, et al.
Pubblicazione: (2024)
OrcaLoca: An LLM Agent Framework for Software Issue Localization
di: Yu, Zhongming, et al.
Pubblicazione: (2025)
di: Yu, Zhongming, et al.
Pubblicazione: (2025)
Vendor-Aware Industrial Agents: RAG-Enhanced LLMs for Secure On-Premise PLC Code Generation
di: Kersting, Joschka, et al.
Pubblicazione: (2025)
di: Kersting, Joschka, et al.
Pubblicazione: (2025)
From Empirical Evaluation to Context-Aware Enhancement: Repairing Regression Errors with LLMs
di: Ho, Anh, et al.
Pubblicazione: (2025)
di: Ho, Anh, et al.
Pubblicazione: (2025)
Large Language Models for Validating Network Protocol Parsers
di: Zheng, Mingwei, et al.
Pubblicazione: (2025)
di: Zheng, Mingwei, et al.
Pubblicazione: (2025)
Evaluating AGENTS.md: Are Repository-Level Context Files Helpful for Coding Agents?
di: Gloaguen, Thibaud, et al.
Pubblicazione: (2026)
di: Gloaguen, Thibaud, et al.
Pubblicazione: (2026)
DiagEval: Trajectory-Conditioned Diagnosis for Reliable Software Evaluation with GUI Agents
di: Hong, Sirui, et al.
Pubblicazione: (2026)
di: Hong, Sirui, et al.
Pubblicazione: (2026)
ToolPRMBench: Evaluating and Advancing Process Reward Models for Tool-using Agents
di: Li, Dawei, et al.
Pubblicazione: (2026)
di: Li, Dawei, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Refactoring Codebases through Library Design
di: Kovacic, Ziga, et al.
Pubblicazione: (2025) -
SWE-CI: Evaluating Agent Capabilities in Maintaining Codebases via Continuous Integration
di: Chen, Jialong, et al.
Pubblicazione: (2026) -
Meta-RAG on Large Codebases Using Code Summarization
di: Tawosi, Vali, et al.
Pubblicazione: (2025) -
Exploring the Challenges and Opportunities of AI-assisted Codebase Generation
di: Eibl, Philipp, et al.
Pubblicazione: (2025) -
FormulaCode: Evaluating Agentic Optimization on Large Codebases
di: Sehgal, Atharva, et al.
Pubblicazione: (2026)