Process-Centric Analysis of Agentic Software Systems
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Shuyang, Chen, Yang, Krishna, Rahul, Sinha, Saurabh, Ganhotra, Jatin, Jabbarvand, Reyhan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Evaluating Plan Compliance in Autonomous Programming Agents
by: Liu, Shuyang, et al.
Published: (2026)
by: Liu, Shuyang, et al.
Published: (2026)
CodeMind: Evaluating Large Language Models for Code Reasoning
by: Liu, Changshu, et al.
Published: (2024)
by: Liu, Changshu, et al.
Published: (2024)
LeTI: Learning to Generate from Textual Interactions
by: Wang, Xingyao, et al.
Published: (2023)
by: Wang, Xingyao, et al.
Published: (2023)
When Agents go Astray: Course-Correcting SWE Agents with PRMs
by: Gandhi, Shubham, et al.
Published: (2025)
by: Gandhi, Shubham, et al.
Published: (2025)
SWE-MERA: A Dynamic Benchmark for Agenticly Evaluating Large Language Models on Software Engineering Tasks
by: Adamenko, Pavel, et al.
Published: (2025)
by: Adamenko, Pavel, et al.
Published: (2025)
MemRepair: Hierarchical Memory for Agentic Repository-Level Vulnerability Repair
by: Liu, Simiao, et al.
Published: (2026)
by: Liu, Simiao, et al.
Published: (2026)
SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?
by: Yang, John, et al.
Published: (2024)
by: Yang, John, et al.
Published: (2024)
Process-Level Trajectory Evaluation for Environment Configuration in Software Engineering Agents
by: Kuang, Jiayi, et al.
Published: (2025)
by: Kuang, Jiayi, et al.
Published: (2025)
E2Edev: Benchmarking Large Language Models in End-to-End Software Development Task
by: Liu, Jingyao, et al.
Published: (2025)
by: Liu, Jingyao, et al.
Published: (2025)
A Taxonomy of Foundation Model based Systems through the Lens of Software Architecture
by: Lu, Qinghua, et al.
Published: (2023)
by: Lu, Qinghua, et al.
Published: (2023)
From Code-Centric to Intent-Centric Software Engineering: A Reflexive Thematic Analysis of Generative AI, Agentic Systems, and Engineering Accountability
by: De La Cruz, Elyson
Published: (2026)
by: De La Cruz, Elyson
Published: (2026)
Multilingual Multimodal Software Developer for Code Generation
by: Chai, Linzheng, et al.
Published: (2025)
by: Chai, Linzheng, et al.
Published: (2025)
Toward an Agentic Infused Software Ecosystem
by: Marron, Mark
Published: (2026)
by: Marron, Mark
Published: (2026)
Agents in Software Engineering: Survey, Landscape, and Vision
by: Wang, Yanlin, et al.
Published: (2024)
by: Wang, Yanlin, et al.
Published: (2024)
Vibe Coding vs. Agentic Coding: Fundamentals and Practical Implications of Agentic AI
by: Sapkota, Ranjan, et al.
Published: (2025)
by: Sapkota, Ranjan, et al.
Published: (2025)
SEW: Self-Evolving Agentic Workflows for Automated Code Generation
by: Liu, Siwei, et al.
Published: (2025)
by: Liu, Siwei, et al.
Published: (2025)
ABC-Bench: Benchmarking Agentic Backend Coding in Real-World Development
by: Yang, Jie, et al.
Published: (2026)
by: Yang, Jie, et al.
Published: (2026)
Utilizing Deep Learning to Optimize Software Development Processes
by: Li, Keqin, et al.
Published: (2024)
by: Li, Keqin, et al.
Published: (2024)
GameDevBench: Evaluating Agentic Capabilities Through Game Development
by: Chi, Wayne, et al.
Published: (2026)
by: Chi, Wayne, et al.
Published: (2026)
SWE-smith: Scaling Data for Software Engineering Agents
by: Yang, John, et al.
Published: (2025)
by: Yang, John, et al.
Published: (2025)
DevEval: Evaluating Code Generation in Practical Software Projects
by: Li, Jia, et al.
Published: (2024)
by: Li, Jia, et al.
Published: (2024)
Unifying the Perspectives of NLP and Software Engineering: A Survey on Language Models for Code
by: Zhang, Ziyin, et al.
Published: (2023)
by: Zhang, Ziyin, et al.
Published: (2023)
OmniCode: A Benchmark for Evaluating Software Engineering Agents
by: Sonwane, Atharv, et al.
Published: (2026)
by: Sonwane, Atharv, et al.
Published: (2026)
PRAXIS: Integrating Program Analysis with Observability for Root-Cause Analysis
by: Cui, Shengkun, et al.
Published: (2025)
by: Cui, Shengkun, et al.
Published: (2025)
Conjecture and Inquiry: Quantifying Software Performance Requirements via Interactive Retrieval-Augmented Preference Elicitation
by: Wang, Shihai, et al.
Published: (2026)
by: Wang, Shihai, et al.
Published: (2026)
GoNoGo: An Efficient LLM-based Multi-Agent System for Streamlining Automotive Software Release Decision-Making
by: Khoee, Arsham Gholamzadeh, et al.
Published: (2024)
by: Khoee, Arsham Gholamzadeh, et al.
Published: (2024)
FormulaCode: Evaluating Agentic Optimization on Large Codebases
by: Sehgal, Atharva, et al.
Published: (2026)
by: Sehgal, Atharva, et al.
Published: (2026)
Automated Business Process Analysis: An LLM-Based Approach to Value Assessment
by: De Michele, William, et al.
Published: (2025)
by: De Michele, William, et al.
Published: (2025)
LAW: Legal Agentic Workflows for Custody and Fund Services Contracts
by: Watson, William, et al.
Published: (2024)
by: Watson, William, et al.
Published: (2024)
CR-Bench: Evaluating the Real-World Utility of AI Code Review Agents
by: Pereira, Kristen, et al.
Published: (2026)
by: Pereira, Kristen, et al.
Published: (2026)
SyncMind: Measuring Agent Out-of-Sync Recovery in Collaborative Software Engineering
by: Guo, Xuehang, et al.
Published: (2025)
by: Guo, Xuehang, et al.
Published: (2025)
SWE-AGI: Benchmarking Specification-Driven Software Construction with MoonBit in the Era of Autonomous Agents
by: Zhang, Zhirui, et al.
Published: (2026)
by: Zhang, Zhirui, et al.
Published: (2026)
Chain-of-Programming (CoP) : Empowering Large Language Models for Geospatial Code Generation
by: Hou, Shuyang, et al.
Published: (2024)
by: Hou, Shuyang, et al.
Published: (2024)
From LLMs to LLM-based Agents for Software Engineering: A Survey of Current, Challenges and Future
by: Jin, Haolin, et al.
Published: (2024)
by: Jin, Haolin, et al.
Published: (2024)
Satori-SWE: Evolutionary Test-Time Scaling for Sample-Efficient Software Engineering
by: Zeng, Guangtao, et al.
Published: (2025)
by: Zeng, Guangtao, et al.
Published: (2025)
R2Vul: Learning to Reason about Software Vulnerabilities with Reinforcement Learning and Structured Reasoning Distillation
by: Weyssow, Martin, et al.
Published: (2025)
by: Weyssow, Martin, et al.
Published: (2025)
Introduction to Analytical Software Engineering Design Paradigm
by: Houichime, Tarik, et al.
Published: (2025)
by: Houichime, Tarik, et al.
Published: (2025)
Iterative Experience Refinement of Software-Developing Agents
by: Qian, Chen, et al.
Published: (2024)
by: Qian, Chen, et al.
Published: (2024)
Deployability-Centric Infrastructure-as-Code Generation: Fail, Learn, Refine, and Succeed through LLM-Empowered DevOps Simulation
by: Zhang, Tianyi, et al.
Published: (2025)
by: Zhang, Tianyi, et al.
Published: (2025)
Experiential Co-Learning of Software-Developing Agents
by: Qian, Chen, et al.
Published: (2023)
by: Qian, Chen, et al.
Published: (2023)
Similar Items
-
Evaluating Plan Compliance in Autonomous Programming Agents
by: Liu, Shuyang, et al.
Published: (2026) -
CodeMind: Evaluating Large Language Models for Code Reasoning
by: Liu, Changshu, et al.
Published: (2024) -
LeTI: Learning to Generate from Textual Interactions
by: Wang, Xingyao, et al.
Published: (2023) -
When Agents go Astray: Course-Correcting SWE Agents with PRMs
by: Gandhi, Shubham, et al.
Published: (2025) -
SWE-MERA: A Dynamic Benchmark for Agenticly Evaluating Large Language Models on Software Engineering Tasks
by: Adamenko, Pavel, et al.
Published: (2025)