Effective Strategies for Asynchronous Software Engineering Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Geng, Jiayi, Neubig, Graham |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Training Versatile Coding Agents in Synthetic Environments
by: Zhu, Yiqi, et al.
Published: (2025)
by: Zhu, Yiqi, et al.
Published: (2025)
Can Large Language Models be Trusted for Evaluation? Scalable Meta-Evaluation of LLMs as Evaluators via Agent Debate
by: Chern, Steffi, et al.
Published: (2024)
by: Chern, Steffi, et al.
Published: (2024)
Accumulating Context Changes the Beliefs of Language Models
by: Geng, Jiayi, et al.
Published: (2025)
by: Geng, Jiayi, et al.
Published: (2025)
Process-Level Trajectory Evaluation for Environment Configuration in Software Engineering Agents
by: Kuang, Jiayi, et al.
Published: (2025)
by: Kuang, Jiayi, et al.
Published: (2025)
Recursive Agent Optimization
by: Gandhi, Apurva, et al.
Published: (2026)
by: Gandhi, Apurva, et al.
Published: (2026)
Agent Data Protocol: Unifying Datasets for Diverse, Effective Fine-tuning of LLM Agents
by: Song, Yueqi, et al.
Published: (2025)
by: Song, Yueqi, et al.
Published: (2025)
Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards
by: Da, Jeff, et al.
Published: (2025)
by: Da, Jeff, et al.
Published: (2025)
OmniCode: A Benchmark for Evaluating Software Engineering Agents
by: Sonwane, Atharv, et al.
Published: (2026)
by: Sonwane, Atharv, et al.
Published: (2026)
CodeScout: An Effective Recipe for Reinforcement Learning of Code Search Agents
by: Sutawika, Lintang, et al.
Published: (2026)
by: Sutawika, Lintang, et al.
Published: (2026)
Alignment for Honesty
by: Yang, Yuqing, et al.
Published: (2023)
by: Yang, Yuqing, et al.
Published: (2023)
SELF-GUIDE: Better Task-Specific Instruction Following via Self-Synthetic Finetuning
by: Zhao, Chenyang, et al.
Published: (2024)
by: Zhao, Chenyang, et al.
Published: (2024)
What Are Tools Anyway? A Survey from the Language Model Perspective
by: Wang, Zhiruo, et al.
Published: (2024)
by: Wang, Zhiruo, et al.
Published: (2024)
Training Software Engineering Agents and Verifiers with SWE-Gym
by: Pan, Jiayi, et al.
Published: (2024)
by: Pan, Jiayi, et al.
Published: (2024)
OpenHands: An Open Platform for AI Software Developers as Generalist Agents
by: Wang, Xingyao, et al.
Published: (2024)
by: Wang, Xingyao, et al.
Published: (2024)
How Do AI Agents Do Human Work? Comparing AI and Human Workflows Across Diverse Occupations
by: Wang, Zora Zhiruo, et al.
Published: (2025)
by: Wang, Zora Zhiruo, et al.
Published: (2025)
Oolong: Evaluating Long Context Reasoning and Aggregation Capabilities
by: Bertsch, Amanda, et al.
Published: (2025)
by: Bertsch, Amanda, et al.
Published: (2025)
Agents in Software Engineering: Survey, Landscape, and Vision
by: Wang, Yanlin, et al.
Published: (2024)
by: Wang, Yanlin, et al.
Published: (2024)
Training Proactive and Personalized LLM Agents
by: Sun, Weiwei, et al.
Published: (2025)
by: Sun, Weiwei, et al.
Published: (2025)
Ambig-SWE: Interactive Agents to Overcome Underspecificity in Software Engineering
by: Vijayvargiya, Sanidhya, et al.
Published: (2025)
by: Vijayvargiya, Sanidhya, et al.
Published: (2025)
Better Instruction-Following Through Minimum Bayes Risk
by: Wu, Ian, et al.
Published: (2024)
by: Wu, Ian, et al.
Published: (2024)
Robotouille: An Asynchronous Planning Benchmark for LLM Agents
by: Gonzalez-Pumariega, Gonzalo, et al.
Published: (2025)
by: Gonzalez-Pumariega, Gonzalo, et al.
Published: (2025)
SWE-smith: Scaling Data for Software Engineering Agents
by: Yang, John, et al.
Published: (2025)
by: Yang, John, et al.
Published: (2025)
ZINA: Multimodal Fine-grained Hallucination Detection and Editing
by: Wada, Yuiga, et al.
Published: (2025)
by: Wada, Yuiga, et al.
Published: (2025)
Language Modeling with Editable External Knowledge
by: Li, Belinda Z., et al.
Published: (2024)
by: Li, Belinda Z., et al.
Published: (2024)
CowPilot: A Framework for Autonomous and Human-Agent Collaborative Web Navigation
by: Huq, Faria, et al.
Published: (2025)
by: Huq, Faria, et al.
Published: (2025)
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems
by: Geng, Jiayi, et al.
Published: (2025)
by: Geng, Jiayi, et al.
Published: (2025)
Effective Harness Engineering for Algorithm Discovery with Coding Agents
by: Ishibashi, Yoichi, et al.
Published: (2026)
by: Ishibashi, Yoichi, et al.
Published: (2026)
VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding?
by: Liu, Junpeng, et al.
Published: (2024)
by: Liu, Junpeng, et al.
Published: (2024)
SyncMind: Measuring Agent Out-of-Sync Recovery in Collaborative Software Engineering
by: Guo, Xuehang, et al.
Published: (2025)
by: Guo, Xuehang, et al.
Published: (2025)
Gym-Anything: Turn any Software into an Agent Environment
by: Aggarwal, Pranjal, et al.
Published: (2026)
by: Aggarwal, Pranjal, et al.
Published: (2026)
Real-Time Generation of Game Video Commentary with Multimodal LLMs: Pause-Aware Decoding Approaches
by: Afzal, Anum, et al.
Published: (2026)
by: Afzal, Anum, et al.
Published: (2026)
MERLIN: A Testbed for Multilingual Multimodal Entity Recognition and Linking
by: Ramamoorthy, Sathyanarayanan, et al.
Published: (2025)
by: Ramamoorthy, Sathyanarayanan, et al.
Published: (2025)
CMU's IWSLT 2024 Simultaneous Speech Translation System
by: Xu, Xi, et al.
Published: (2024)
by: Xu, Xi, et al.
Published: (2024)
Asking What Matters: Reward-Driven Clarification for Software Engineering Tasks
by: Vijayvargiya, Sanidhya, et al.
Published: (2026)
by: Vijayvargiya, Sanidhya, et al.
Published: (2026)
On Problems of Implicit Context Compression for Software Engineering Agents
by: Gelvan, Kirill, et al.
Published: (2026)
by: Gelvan, Kirill, et al.
Published: (2026)
Agentless: Demystifying LLM-based Software Engineering Agents
by: Xia, Chunqiu Steven, et al.
Published: (2024)
by: Xia, Chunqiu Steven, et al.
Published: (2024)
Asynchronous LLM Function Calling
by: Gim, In, et al.
Published: (2024)
by: Gim, In, et al.
Published: (2024)
M-Prometheus: A Suite of Open Multilingual LLM Judges
by: Pombal, José, et al.
Published: (2025)
by: Pombal, José, et al.
Published: (2025)
Overtrained Language Models Are Harder to Fine-Tune
by: Springer, Jacob Mitchell, et al.
Published: (2025)
by: Springer, Jacob Mitchell, et al.
Published: (2025)
SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents
by: Zhou, Xuhui, et al.
Published: (2023)
by: Zhou, Xuhui, et al.
Published: (2023)
Similar Items
-
Training Versatile Coding Agents in Synthetic Environments
by: Zhu, Yiqi, et al.
Published: (2025) -
Can Large Language Models be Trusted for Evaluation? Scalable Meta-Evaluation of LLMs as Evaluators via Agent Debate
by: Chern, Steffi, et al.
Published: (2024) -
Accumulating Context Changes the Beliefs of Language Models
by: Geng, Jiayi, et al.
Published: (2025) -
Process-Level Trajectory Evaluation for Environment Configuration in Software Engineering Agents
by: Kuang, Jiayi, et al.
Published: (2025) -
Recursive Agent Optimization
by: Gandhi, Apurva, et al.
Published: (2026)