Human-In-The-Loop Software Development Agents: Challenges and Future Directions

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Pasuksmit, Jirat, Takerngsaksiri, Wannita, Thongtanunam, Patanamon, Tantithamthavorn, Chakkrit, Zhang, Ruixiong, Wang, Shiyan, Jiang, Fan, Li, Jing, Cook, Evan, Chen, Kun, Wu, Ming
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911002735411200
author Pasuksmit, Jirat
Takerngsaksiri, Wannita
Thongtanunam, Patanamon
Tantithamthavorn, Chakkrit
Zhang, Ruixiong
Wang, Shiyan
Jiang, Fan
Li, Jing
Cook, Evan
Chen, Kun
Wu, Ming
author_facet Pasuksmit, Jirat
Takerngsaksiri, Wannita
Thongtanunam, Patanamon
Tantithamthavorn, Chakkrit
Zhang, Ruixiong
Wang, Shiyan
Jiang, Fan
Li, Jing
Cook, Evan
Chen, Kun
Wu, Ming
contents Multi-agent LLM-driven systems for software development are rapidly gaining traction, offering new opportunities to enhance productivity. At Atlassian, we deployed Human-in-the-Loop Software Development Agents to resolve Jira work items and evaluated the generated code quality using functional correctness testing and GPT-based similarity scoring. This paper highlights two major challenges: the high computational costs of unit testing and the variability in LLM-based evaluations. We also propose future research directions to improve evaluation frameworks for Human-In-The-Loop software development tools.
format Preprint
id arxiv_https___arxiv_org_abs_2506_11009
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Human-In-The-Loop Software Development Agents: Challenges and Future Directions
Pasuksmit, Jirat
Takerngsaksiri, Wannita
Thongtanunam, Patanamon
Tantithamthavorn, Chakkrit
Zhang, Ruixiong
Wang, Shiyan
Jiang, Fan
Li, Jing
Cook, Evan
Chen, Kun
Wu, Ming
Software Engineering
Multi-agent LLM-driven systems for software development are rapidly gaining traction, offering new opportunities to enhance productivity. At Atlassian, we deployed Human-in-the-Loop Software Development Agents to resolve Jira work items and evaluated the generated code quality using functional correctness testing and GPT-based similarity scoring. This paper highlights two major challenges: the high computational costs of unit testing and the variability in LLM-based evaluations. We also propose future research directions to improve evaluation frameworks for Human-In-The-Loop software development tools.
title Human-In-The-Loop Software Development Agents: Challenges and Future Directions
topic Software Engineering
url https://arxiv.org/abs/2506.11009