Reverse Engineering User Stories from Code using Large Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866909803874353152 |
|---|---|
| author | Ouf, Mohamed Li, Haoyu Zhang, Michael Guizani, Mariam |
| author_facet | Ouf, Mohamed Li, Haoyu Zhang, Michael Guizani, Mariam |
| contents | User stories are essential in agile development, yet often missing or outdated in legacy and poorly documented systems. We investigate whether large language models (LLMs) can automatically recover user stories directly from source code and how prompt design impacts output quality. Using 1,750 annotated C++ snippets of varying complexity, we evaluate five state-of-the-art LLMs across six prompting strategies. Results show that all models achieve, on average, an F1 score of 0.8 for code up to 200 NLOC. Our findings show that a single illustrative example enables the smallest model (8B) to match the performance of a much larger 70B model. In contrast, structured reasoning via Chain-of-Thought offers only marginal gains, primarily for larger models. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_19587 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Reverse Engineering User Stories from Code using Large Language Models Ouf, Mohamed Li, Haoyu Zhang, Michael Guizani, Mariam Software Engineering Artificial Intelligence User stories are essential in agile development, yet often missing or outdated in legacy and poorly documented systems. We investigate whether large language models (LLMs) can automatically recover user stories directly from source code and how prompt design impacts output quality. Using 1,750 annotated C++ snippets of varying complexity, we evaluate five state-of-the-art LLMs across six prompting strategies. Results show that all models achieve, on average, an F1 score of 0.8 for code up to 200 NLOC. Our findings show that a single illustrative example enables the smallest model (8B) to match the performance of a much larger 70B model. In contrast, structured reasoning via Chain-of-Thought offers only marginal gains, primarily for larger models. |
| title | Reverse Engineering User Stories from Code using Large Language Models |
| topic | Software Engineering Artificial Intelligence |
| url | https://arxiv.org/abs/2509.19587 |