From Reviews to Requirements: Can LLMs Generate Human-Like User Stories?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sakib, Shadman, Akhand, Oishy Fatema, Tasneem, Tasnia, Ahmed, Shohel
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918416911171584
author Sakib, Shadman
Akhand, Oishy Fatema
Tasneem, Tasnia
Ahmed, Shohel
author_facet Sakib, Shadman
Akhand, Oishy Fatema
Tasneem, Tasnia
Ahmed, Shohel
contents App store reviews provide a constant flow of real user feedback that can help improve software requirements. However, these reviews are often messy, informal, and difficult to analyze manually at scale. Although automated techniques exist, many do not perform well when replicated and often fail to produce clean, backlog-ready user stories for agile projects. In this study, we evaluate how well large language models (LLMs) such as GPT-3.5 Turbo, Gemini 2.0 Flash, and Mistral 7B Instruct can generate usable user stories directly from raw app reviews. Using the Mini-BAR dataset of 1,000+ health app reviews, we tested zero-shot, one-shot, and two-shot prompting methods. We evaluated the generated user stories using both human judgment (via the RUST framework) and a RoBERTa classifier fine-tuned on UStAI to assess their overall quality. Our results show that LLMs can match or even outperform humans in writing fluent, well-formatted user stories, especially when few-shot prompts are used. However, they still struggle to produce independent and unique user stories, which are essential for building a strong agile backlog. Overall, our findings show how LLMs can reliably turn unstructured app reviews into actionable software requirements, providing developers with clear guidance to turn user feedback into meaningful improvements.
format Preprint
id arxiv_https___arxiv_org_abs_2603_28163
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle From Reviews to Requirements: Can LLMs Generate Human-Like User Stories?
Sakib, Shadman
Akhand, Oishy Fatema
Tasneem, Tasnia
Ahmed, Shohel
Computation and Language
App store reviews provide a constant flow of real user feedback that can help improve software requirements. However, these reviews are often messy, informal, and difficult to analyze manually at scale. Although automated techniques exist, many do not perform well when replicated and often fail to produce clean, backlog-ready user stories for agile projects. In this study, we evaluate how well large language models (LLMs) such as GPT-3.5 Turbo, Gemini 2.0 Flash, and Mistral 7B Instruct can generate usable user stories directly from raw app reviews. Using the Mini-BAR dataset of 1,000+ health app reviews, we tested zero-shot, one-shot, and two-shot prompting methods. We evaluated the generated user stories using both human judgment (via the RUST framework) and a RoBERTa classifier fine-tuned on UStAI to assess their overall quality. Our results show that LLMs can match or even outperform humans in writing fluent, well-formatted user stories, especially when few-shot prompts are used. However, they still struggle to produce independent and unique user stories, which are essential for building a strong agile backlog. Overall, our findings show how LLMs can reliably turn unstructured app reviews into actionable software requirements, providing developers with clear guidance to turn user feedback into meaningful improvements.
title From Reviews to Requirements: Can LLMs Generate Human-Like User Stories?
topic Computation and Language
url https://arxiv.org/abs/2603.28163