Human-In-The-Loop Software Development Agents: Challenges and Future Directions
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866911002735411200 |
|---|---|
| author | Pasuksmit, Jirat Takerngsaksiri, Wannita Thongtanunam, Patanamon Tantithamthavorn, Chakkrit Zhang, Ruixiong Wang, Shiyan Jiang, Fan Li, Jing Cook, Evan Chen, Kun Wu, Ming |
| author_facet | Pasuksmit, Jirat Takerngsaksiri, Wannita Thongtanunam, Patanamon Tantithamthavorn, Chakkrit Zhang, Ruixiong Wang, Shiyan Jiang, Fan Li, Jing Cook, Evan Chen, Kun Wu, Ming |
| contents | Multi-agent LLM-driven systems for software development are rapidly gaining traction, offering new opportunities to enhance productivity. At Atlassian, we deployed Human-in-the-Loop Software Development Agents to resolve Jira work items and evaluated the generated code quality using functional correctness testing and GPT-based similarity scoring. This paper highlights two major challenges: the high computational costs of unit testing and the variability in LLM-based evaluations. We also propose future research directions to improve evaluation frameworks for Human-In-The-Loop software development tools. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_11009 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Human-In-The-Loop Software Development Agents: Challenges and Future Directions Pasuksmit, Jirat Takerngsaksiri, Wannita Thongtanunam, Patanamon Tantithamthavorn, Chakkrit Zhang, Ruixiong Wang, Shiyan Jiang, Fan Li, Jing Cook, Evan Chen, Kun Wu, Ming Software Engineering Multi-agent LLM-driven systems for software development are rapidly gaining traction, offering new opportunities to enhance productivity. At Atlassian, we deployed Human-in-the-Loop Software Development Agents to resolve Jira work items and evaluated the generated code quality using functional correctness testing and GPT-based similarity scoring. This paper highlights two major challenges: the high computational costs of unit testing and the variability in LLM-based evaluations. We also propose future research directions to improve evaluation frameworks for Human-In-The-Loop software development tools. |
| title | Human-In-The-Loop Software Development Agents: Challenges and Future Directions |
| topic | Software Engineering |
| url | https://arxiv.org/abs/2506.11009 |