From Prompt to Product: A Human-Centered Benchmark of Agentic App Generation Systems
Fuente:
arXiv
Guardado en:
| Autores principales: | Ortiz, Marcos, Hill, Justin, Overbay, Collin, Semenec, Ingrida, Sauve-Hoover, Frederic, Schwoebel, Jim, Shor, Joel |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
CentaurEval: Benchmarking Human-in-the-Loop Value in Agentic Coding
por: Luo, Hanjun, et al.
Publicado: (2025)
por: Luo, Hanjun, et al.
Publicado: (2025)
Exploring Direct Instruction and Summary-Mediated Prompting in LLM-Assisted Code Modification
por: Tang, Ningzhi, et al.
Publicado: (2025)
por: Tang, Ningzhi, et al.
Publicado: (2025)
The Best Ends by the Best Means: Ethical Concerns in App Reviews
por: Olson, Lauren, et al.
Publicado: (2024)
por: Olson, Lauren, et al.
Publicado: (2024)
Reassessing Java Code Readability Models with a Human-Centered Approach
por: Sergeyuk, Agnia, et al.
Publicado: (2024)
por: Sergeyuk, Agnia, et al.
Publicado: (2024)
Auto-Generating Personas from User Reviews in VR App Stores
por: Wang, Yi, et al.
Publicado: (2026)
por: Wang, Yi, et al.
Publicado: (2026)
MotorEase: Automated Detection of Motor Impairment Accessibility Issues in Mobile App UIs
por: Krishnavajjala, Arun, et al.
Publicado: (2024)
por: Krishnavajjala, Arun, et al.
Publicado: (2024)
Early Accessibility: Automating Alt-Text Generation for UI Icons During App Development
por: Haque, Sabrina, et al.
Publicado: (2025)
por: Haque, Sabrina, et al.
Publicado: (2025)
Inferring Alt-text For UI Icons With Large Language Models During App Development
por: Haque, Sabrina, et al.
Publicado: (2024)
por: Haque, Sabrina, et al.
Publicado: (2024)
From Correctness to Collaboration: Toward a Human-Centered Framework for Evaluating AI Agent Behavior in Software Engineering
por: Dong, Tao, et al.
Publicado: (2025)
por: Dong, Tao, et al.
Publicado: (2025)
Two Integration Pathways in Human-Centered Requirements Engineering: A Systematic Mapping Study of Structural Gaps
por: Benzarti, Imen, et al.
Publicado: (2026)
por: Benzarti, Imen, et al.
Publicado: (2026)
Examining the Use and Impact of an AI Code Assistant on Developer Productivity and Experience in the Enterprise
por: Weisz, Justin D., et al.
Publicado: (2024)
por: Weisz, Justin D., et al.
Publicado: (2024)
Age Matters: Analyzing Age-Related Discussions in App Reviews
por: Nirmania, Shashiwadana, et al.
Publicado: (2026)
por: Nirmania, Shashiwadana, et al.
Publicado: (2026)
EyeTrans: Merging Human and Machine Attention for Neural Code Summarization
por: Zhang, Yifan, et al.
Publicado: (2024)
por: Zhang, Yifan, et al.
Publicado: (2024)
Agentic Metacognition: Designing a "Self-Aware" Low-Code Agent for Failure Prediction and Human Handoff
por: Xu, Jiexi
Publicado: (2025)
por: Xu, Jiexi
Publicado: (2025)
Gendered Prompting and LLM Code Review: How Gender Cues in the Prompt Shape Code Quality and Evaluation
por: Janzen, Lynn, et al.
Publicado: (2026)
por: Janzen, Lynn, et al.
Publicado: (2026)
EyeMulator: Improving Code Language Models by Mimicking Human Visual Attention
por: Zhang, Yifan, et al.
Publicado: (2025)
por: Zhang, Yifan, et al.
Publicado: (2025)
LikeThis! Empowering App Users to Submit UI Improvement Suggestions Instead of Complaints
por: Wei, Jialiang, et al.
Publicado: (2026)
por: Wei, Jialiang, et al.
Publicado: (2026)
A Unified, Cross-Platform Framework for Automatic GUI and Plugin Generation in Structural Bioinformatics and Beyond
por: Guo, Sikao, et al.
Publicado: (2026)
por: Guo, Sikao, et al.
Publicado: (2026)
Single Conversation Methodology: A Human-Centered Protocol for AI-Assisted Software Development
por: Escobedo, Salvador D.
Publicado: (2025)
por: Escobedo, Salvador D.
Publicado: (2025)
From Junior to Senior: Allocating Agency and Navigating Professional Growth in Agentic AI-Mediated Software Engineering
por: Feng, Dana, et al.
Publicado: (2026)
por: Feng, Dana, et al.
Publicado: (2026)
Recover as It is Designed to Be: Recovering from Compatibility Mobile App Crashes by Reusing User Flows
por: Kim, Donghwi, et al.
Publicado: (2024)
por: Kim, Donghwi, et al.
Publicado: (2024)
From Prompts to Propositions: A Logic-Based Lens on Student-LLM Interactions
por: Alfageeh, Ali, et al.
Publicado: (2025)
por: Alfageeh, Ali, et al.
Publicado: (2025)
NaturalEdit: Code Modification through Direct Interaction with Adaptive Natural Language Representation
por: Tang, Ningzhi, et al.
Publicado: (2025)
por: Tang, Ningzhi, et al.
Publicado: (2025)
A Study on Developer Behaviors for Validating and Repairing LLM-Generated Code Using Eye Tracking and IDE Actions
por: Tang, Ningzhi, et al.
Publicado: (2024)
por: Tang, Ningzhi, et al.
Publicado: (2024)
Why AI Agents Still Need You: Findings from Developer-Agent Collaborations in the Wild
por: Kumar, Aayush, et al.
Publicado: (2025)
por: Kumar, Aayush, et al.
Publicado: (2025)
Human-Centered Evaluation of an LLM-Based Process Modeling Copilot: A Mixed-Methods Study with Domain Experts
por: Lauer, Chantale, et al.
Publicado: (2026)
por: Lauer, Chantale, et al.
Publicado: (2026)
Story Arena: A Multi-Agent Environment for Envisioning the Future of Software Engineering
por: Weisz, Justin D., et al.
Publicado: (2025)
por: Weisz, Justin D., et al.
Publicado: (2025)
Sources of Underproduction in Open Source Software
por: Champion, Kaylea, et al.
Publicado: (2024)
por: Champion, Kaylea, et al.
Publicado: (2024)
Prototyping with Prompts: Emerging Approaches and Challenges in Generative AI Design for Collaborative Software Teams
por: Subramonyam, Hari, et al.
Publicado: (2024)
por: Subramonyam, Hari, et al.
Publicado: (2024)
FeedAIde: Guiding App Users to Submit Rich Feedback Reports by Asking Context-Aware Follow-Up Questions
por: Pourasad, Ali Ebrahimi, et al.
Publicado: (2026)
por: Pourasad, Ali Ebrahimi, et al.
Publicado: (2026)
GazeCopilot: Evaluating Novel Gaze-Informed Prompting for AI-Supported Code Comprehension and Readability
por: Elfares, Yasmine, et al.
Publicado: (2025)
por: Elfares, Yasmine, et al.
Publicado: (2025)
From Human-to-Human to Human-to-Bot Conversations in Software Engineering
por: Khojah, Ranim, et al.
Publicado: (2024)
por: Khojah, Ranim, et al.
Publicado: (2024)
Programming by Chat: A Large-Scale Behavioral Analysis of 11,579 Real-World AI-Assisted IDE Sessions
por: Tang, Ningzhi, et al.
Publicado: (2026)
por: Tang, Ningzhi, et al.
Publicado: (2026)
app.build: A Production Framework for Scaling Agentic Prompt-to-App Generation with Environment Scaffolding
por: Kniazev, Evgenii, et al.
Publicado: (2025)
por: Kniazev, Evgenii, et al.
Publicado: (2025)
Debugging Without Error Messages: How LLM Prompting Strategy Affects Programming Error Explanation Effectiveness
por: Salmon, Audrey, et al.
Publicado: (2025)
por: Salmon, Audrey, et al.
Publicado: (2025)
Prompts Are Programs Too! Understanding How Developers Build Software Containing Prompts
por: Liang, Jenny T., et al.
Publicado: (2024)
por: Liang, Jenny T., et al.
Publicado: (2024)
Prompt-with-Me: in-IDE Structured Prompt Management for LLM-Driven Software Engineering
por: Li, Ziyou, et al.
Publicado: (2025)
por: Li, Ziyou, et al.
Publicado: (2025)
Designing Adaptive User Interfaces for mHealth Applications Targeting Chronic Disease: A User-Centered Approach
por: Wang, Wei, et al.
Publicado: (2024)
por: Wang, Wei, et al.
Publicado: (2024)
Git Takes Two: Split-View Awareness for Collaborative Learning of Distributed Workflows in Git
por: Bucher, Joel, et al.
Publicado: (2026)
por: Bucher, Joel, et al.
Publicado: (2026)
The Fast and Spurious: Developer Productivity with GenAI
por: Afroz, Sadia, et al.
Publicado: (2025)
por: Afroz, Sadia, et al.
Publicado: (2025)
Ejemplares similares
-
CentaurEval: Benchmarking Human-in-the-Loop Value in Agentic Coding
por: Luo, Hanjun, et al.
Publicado: (2025) -
Exploring Direct Instruction and Summary-Mediated Prompting in LLM-Assisted Code Modification
por: Tang, Ningzhi, et al.
Publicado: (2025) -
The Best Ends by the Best Means: Ethical Concerns in App Reviews
por: Olson, Lauren, et al.
Publicado: (2024) -
Reassessing Java Code Readability Models with a Human-Centered Approach
por: Sergeyuk, Agnia, et al.
Publicado: (2024) -
Auto-Generating Personas from User Reviews in VR App Stores
por: Wang, Yi, et al.
Publicado: (2026)