Process-Level Trajectory Evaluation for Environment Configuration in Software Engineering Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Kuang, Jiayi, Li, Yinghui, Zhang, Xin, Li, Yangning, Yin, Di, Sun, Xing, Shen, Ying, Yu, Philip S. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
EvoConfig: Self-Evolving Multi-Agent Systems for Efficient Autonomous Environment Configuration
by: Guo, Xinshuai, et al.
Published: (2026)
by: Guo, Xinshuai, et al.
Published: (2026)
Evaluating Software Process Models for Multi-Agent Class-Level Code Generation
by: Shafin, Wasique Islam, et al.
Published: (2025)
by: Shafin, Wasique Islam, et al.
Published: (2025)
Structurally Aligned Subtask-Level Memory for Software Engineering Agents
by: Shen, Kangning, et al.
Published: (2026)
by: Shen, Kangning, et al.
Published: (2026)
From Patches to Trajectories: Privileged Process Supervision for Software-Engineering Agents
by: Ma, Murong, et al.
Published: (2026)
by: Ma, Murong, et al.
Published: (2026)
An Empirical Study of Speculative Decoding on Software Engineering Tasks
by: Li, Yijia, et al.
Published: (2026)
by: Li, Yijia, et al.
Published: (2026)
ATime-Consistent Benchmark for Repository-Level Software Engineering Evaluation
by: Xianpeng, et al.
Published: (2026)
by: Xianpeng, et al.
Published: (2026)
MEnvAgent: Scalable Polyglot Environment Construction for Verifiable Software Engineering
by: Guo, Chuanzhe, et al.
Published: (2026)
by: Guo, Chuanzhe, et al.
Published: (2026)
Pros and Cons! Evaluating ChatGPT on Software Vulnerability
by: Yin, Xin
Published: (2024)
by: Yin, Xin
Published: (2024)
Assessing and Advancing Benchmarks for Evaluating Large Language Models in Software Engineering Tasks
by: Hu, Xing, et al.
Published: (2025)
by: Hu, Xing, et al.
Published: (2025)
SEAlign: Alignment Training for Software Engineering Agent
by: Zhang, Kechi, et al.
Published: (2025)
by: Zhang, Kechi, et al.
Published: (2025)
A Comprehensive Empirical Evaluation of Agent Frameworks on Code-centric Software Engineering Tasks
by: Yin, Zhuowen, et al.
Published: (2025)
by: Yin, Zhuowen, et al.
Published: (2025)
SWE-World: Building Software Engineering Agents in Docker-Free Environments
by: Sun, Shuang, et al.
Published: (2026)
by: Sun, Shuang, et al.
Published: (2026)
Knowledge-Guided Multi-Agent Framework for Application-Level Software Code Generation
by: Xiong, Qian, et al.
Published: (2025)
by: Xiong, Qian, et al.
Published: (2025)
DiagEval: Trajectory-Conditioned Diagnosis for Reliable Software Evaluation with GUI Agents
by: Hong, Sirui, et al.
Published: (2026)
by: Hong, Sirui, et al.
Published: (2026)
Deep Learning-based Software Engineering: Progress, Challenges, and Opportunities
by: Chen, Xiangping, et al.
Published: (2024)
by: Chen, Xiangping, et al.
Published: (2024)
Training Software Engineering Agents and Verifiers with SWE-Gym
by: Pan, Jiayi, et al.
Published: (2024)
by: Pan, Jiayi, et al.
Published: (2024)
The Current Challenges of Software Engineering in the Era of Large Language Models
by: Gao, Cuiyun, et al.
Published: (2024)
by: Gao, Cuiyun, et al.
Published: (2024)
SWE-Fuse: Empowering Software Agents via Issue-free Trajectory Learning and Entropy-aware RLVR Training
by: Wen, Xin-Cheng, et al.
Published: (2026)
by: Wen, Xin-Cheng, et al.
Published: (2026)
Digital Twins for Software Engineering Processes
by: Kimmel, Robin, et al.
Published: (2025)
by: Kimmel, Robin, et al.
Published: (2025)
Multitask-based Evaluation of Open-Source LLM on Software Vulnerability
by: Yin, Xin, et al.
Published: (2024)
by: Yin, Xin, et al.
Published: (2024)
Understanding Software Engineering Agents: A Study of Thought-Action-Result Trajectories
by: Bouzenia, Islem, et al.
Published: (2025)
by: Bouzenia, Islem, et al.
Published: (2025)
A Benchmark for Evaluating Repository-Level Code Agents with Intermediate Reasoning on Feature Addition Task
by: Liu, Shuhan, et al.
Published: (2026)
by: Liu, Shuhan, et al.
Published: (2026)
Designing Software with Complex Configurations
by: Cunha, Alcino
Published: (2024)
by: Cunha, Alcino
Published: (2024)
EET: Experience-Driven Early Termination for Cost-Efficient Software Engineering Agents
by: Guo, Yaoqi, et al.
Published: (2026)
by: Guo, Yaoqi, et al.
Published: (2026)
PAGENT: Learning to Patch Software Engineering Agents
by: Xue, Haoran, et al.
Published: (2025)
by: Xue, Haoran, et al.
Published: (2025)
Unified Software Engineering Agent as AI Software Engineer
by: Applis, Leonhard, et al.
Published: (2025)
by: Applis, Leonhard, et al.
Published: (2025)
Multi-Agent Systems for Dataset Adaptation in Software Engineering: Capabilities, Limitations, and Future Directions
by: Chen, Jingyi, et al.
Published: (2025)
by: Chen, Jingyi, et al.
Published: (2025)
Guided Search Strategies in Non-Serializable Environments with Applications to Software Engineering Agents
by: Zainullina, Karina, et al.
Published: (2025)
by: Zainullina, Karina, et al.
Published: (2025)
Software Engineering Agents for Embodied Controller Generation : A Study in Minigrid Environments
by: Boulet, Timothé, et al.
Published: (2025)
by: Boulet, Timothé, et al.
Published: (2025)
Lessons from a Pioneering Software Engineering Environment: Design Principles of Software through Pictures
by: I., Anthony, et al.
Published: (2024)
by: I., Anthony, et al.
Published: (2024)
Automatic Platform Configuration and Software Integration for Software-Defined Vehicles
by: Pan, Fengjunjie, et al.
Published: (2024)
by: Pan, Fengjunjie, et al.
Published: (2024)
Architectural Support for Software Performance in Continuous Software Engineering: A Systematic Mapping Study
by: Eramo, Romina, et al.
Published: (2023)
by: Eramo, Romina, et al.
Published: (2023)
Context Engineering for AI Agents in Open-Source Software
by: Mohsenimofidi, Seyedmoein, et al.
Published: (2025)
by: Mohsenimofidi, Seyedmoein, et al.
Published: (2025)
Agent-Based Software Artifact Evaluation
by: Wu, Zhaonan, et al.
Published: (2026)
by: Wu, Zhaonan, et al.
Published: (2026)
Code Graph Model (CGM): A Graph-Integrated Large Language Model for Repository-Level Software Engineering Tasks
by: Tao, Hongyuan, et al.
Published: (2025)
by: Tao, Hongyuan, et al.
Published: (2025)
Reproducible, Explainable, and Effective Evaluations of Agentic AI for Software Engineering
by: Li, Jingyue, et al.
Published: (2026)
by: Li, Jingyue, et al.
Published: (2026)
ProSoftArena: Benchmarking Hierarchical Capabilities of Multimodal Agents in Professional Software Environments
by: Ai, Jiaxin, et al.
Published: (2025)
by: Ai, Jiaxin, et al.
Published: (2025)
What Is an App Store? The Software Engineering Perspective
by: Zhu, Wenhan, et al.
Published: (2024)
by: Zhu, Wenhan, et al.
Published: (2024)
Promptware Engineering: Software Engineering for Prompt-Enabled Systems
by: Chen, Zhenpeng, et al.
Published: (2025)
by: Chen, Zhenpeng, et al.
Published: (2025)
REAgent: Requirement-Driven LLM Agents for Software Issue Resolution
by: Kuang, Shiqi, et al.
Published: (2026)
by: Kuang, Shiqi, et al.
Published: (2026)
Similar Items
-
EvoConfig: Self-Evolving Multi-Agent Systems for Efficient Autonomous Environment Configuration
by: Guo, Xinshuai, et al.
Published: (2026) -
Evaluating Software Process Models for Multi-Agent Class-Level Code Generation
by: Shafin, Wasique Islam, et al.
Published: (2025) -
Structurally Aligned Subtask-Level Memory for Software Engineering Agents
by: Shen, Kangning, et al.
Published: (2026) -
From Patches to Trajectories: Privileged Process Supervision for Software-Engineering Agents
by: Ma, Murong, et al.
Published: (2026) -
An Empirical Study of Speculative Decoding on Software Engineering Tasks
by: Li, Yijia, et al.
Published: (2026)