Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild
Fuente:
arXiv
Saved in:
| Main Authors: | van der Maden, Willem, Sadek, Malak, Xiao, Ziang, Mottelson, Aske, Liao, Q. Vera, Zhu, Jichen |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Rethinking Model Evaluation as Narrowing the Socio-Technical Gap
by: Liao, Q. Vera, et al.
Published: (2023)
by: Liao, Q. Vera, et al.
Published: (2023)
The Centers and Margins of Modeling Humans in Well-being Technologies: A Decentering Approach
by: Zhu, Jichen, et al.
Published: (2025)
by: Zhu, Jichen, et al.
Published: (2025)
"I Would Have Written My Code Differently'': Beginners Struggle to Understand LLM-Generated Code
by: Zi, Yangtian, et al.
Published: (2025)
by: Zi, Yangtian, et al.
Published: (2025)
Generative Echo Chamber? Effects of LLM-Powered Search Systems on Diverse Information Seeking
by: Sharma, Nikhil, et al.
Published: (2024)
by: Sharma, Nikhil, et al.
Published: (2024)
Gendered Prompting and LLM Code Review: How Gender Cues in the Prompt Shape Code Quality and Evaluation
by: Janzen, Lynn, et al.
Published: (2026)
by: Janzen, Lynn, et al.
Published: (2026)
Cultural Impact on Requirements Engineering Activities: Bangladeshi Practitioners' View
by: Muzammel, Chowdhury Shahriar, et al.
Published: (2025)
by: Muzammel, Chowdhury Shahriar, et al.
Published: (2025)
Is Seeing Believing? Evaluating Human Sensitivity to Synthetic Video
by: Wegmann, David, et al.
Published: (2026)
by: Wegmann, David, et al.
Published: (2026)
Breaking New Ground in Software Defect Prediction: Introducing Practical and Actionable Metrics with Superior Predictive Power for Enhanced Decision-Making
by: Cataño, Carlos Andrés Ramírez, et al.
Published: (2025)
by: Cataño, Carlos Andrés Ramírez, et al.
Published: (2025)
Understanding the Human-LLM Dynamic: A Literature Survey of LLM Use in Programming Tasks
by: Etsenake, Deborah, et al.
Published: (2024)
by: Etsenake, Deborah, et al.
Published: (2024)
AI for Requirements Engineering: Industry adoption and Practitioner perspectives
by: Rani, Lekshmi Murali, et al.
Published: (2025)
by: Rani, Lekshmi Murali, et al.
Published: (2025)
AI Agents for Web Testing: A Case Study in the Wild
by: Ye, Naimeng, et al.
Published: (2025)
by: Ye, Naimeng, et al.
Published: (2025)
Requirements Perception Gap across Stakeholders: A Comparative Survey of Aged Care Digital Health Software
by: Xiao, Yuqing, et al.
Published: (2026)
by: Xiao, Yuqing, et al.
Published: (2026)
Using an LLM to Help With Code Understanding
by: Nam, Daye, et al.
Published: (2023)
by: Nam, Daye, et al.
Published: (2023)
Positive AI: Key Challenges in Designing Artificial Intelligence for Wellbeing
by: van der Maden, Willem, et al.
Published: (2023)
by: van der Maden, Willem, et al.
Published: (2023)
Debugging Without Error Messages: How LLM Prompting Strategy Affects Programming Error Explanation Effectiveness
by: Salmon, Audrey, et al.
Published: (2025)
by: Salmon, Audrey, et al.
Published: (2025)
Why AI Agents Still Need You: Findings from Developer-Agent Collaborations in the Wild
by: Kumar, Aayush, et al.
Published: (2025)
by: Kumar, Aayush, et al.
Published: (2025)
PriviSense: A Frida-Based Framework for Multi-Sensor Spoofing on Android
by: Khalilov, Ibrahim, et al.
Published: (2026)
by: Khalilov, Ibrahim, et al.
Published: (2026)
Assistance or Disruption? Exploring and Evaluating the Design and Trade-offs of Proactive AI Programming Support
by: Pu, Kevin, et al.
Published: (2025)
by: Pu, Kevin, et al.
Published: (2025)
Prompts Are Programs Too! Understanding How Developers Build Software Containing Prompts
by: Liang, Jenny T., et al.
Published: (2024)
by: Liang, Jenny T., et al.
Published: (2024)
Offloading Score: Measuring AI Reliance Through Counterfactual Workflows
by: Padmakumar, Vishakh, et al.
Published: (2026)
by: Padmakumar, Vishakh, et al.
Published: (2026)
Understanding and supporting how developers prompt for LLM-powered code editing in practice
by: Nam, Daye, et al.
Published: (2025)
by: Nam, Daye, et al.
Published: (2025)
Do You Understand How I Feel?: Towards Verified Empathy in Therapy Chatbots
by: Dettori, Francesco, et al.
Published: (2026)
by: Dettori, Francesco, et al.
Published: (2026)
Bridging the Socio-Emotional Gap: The Functional Dimension of Human-AI Collaboration for Software Engineering
by: Rani, Lekshmi Murali, et al.
Published: (2026)
by: Rani, Lekshmi Murali, et al.
Published: (2026)
Mining the Gold: Student-AI Chat Logs as Rich Sources for Automated Knowledge Gap Detection
by: Fu, Quanzhi, et al.
Published: (2025)
by: Fu, Quanzhi, et al.
Published: (2025)
Two Integration Pathways in Human-Centered Requirements Engineering: A Systematic Mapping Study of Structural Gaps
by: Benzarti, Imen, et al.
Published: (2026)
by: Benzarti, Imen, et al.
Published: (2026)
A Model for Understanding and Reducing Developer Burnout
by: Trinkenreich, Bianca, et al.
Published: (2023)
by: Trinkenreich, Bianca, et al.
Published: (2023)
Bridging the Interpretation Gap in Accessibility Testing: Empathetic and Legal-Aware Bug Report Generation via Large Language Models
by: Koyama, Ryoya, et al.
Published: (2026)
by: Koyama, Ryoya, et al.
Published: (2026)
The Impact of LLM-Assistants on Software Developer Productivity: A Systematic Review and Mapping Study
by: Mohamed, Amr, et al.
Published: (2025)
by: Mohamed, Amr, et al.
Published: (2025)
Towards an Understanding of Developer Experience-Driven Transparency in Software Ecosystems
by: Zacarias, Rodrigo Oliveira, et al.
Published: (2025)
by: Zacarias, Rodrigo Oliveira, et al.
Published: (2025)
HookLens: Visual Analytics for Understanding React Hooks Structures
by: Hwang, Suyeon, et al.
Published: (2026)
by: Hwang, Suyeon, et al.
Published: (2026)
AmbiBench: Benchmarking Mobile GUI Agents Beyond One-Shot Instructions in the Wild
by: Sun, Jiazheng, et al.
Published: (2026)
by: Sun, Jiazheng, et al.
Published: (2026)
How Scientists Use Large Language Models to Program
by: O'Brien, Gabrielle
Published: (2025)
by: O'Brien, Gabrielle
Published: (2025)
Understanding the Career Mobility of Blind and Low Vision Software Professionals
by: Cha, Yoonha, et al.
Published: (2024)
by: Cha, Yoonha, et al.
Published: (2024)
Cognitive Biases in LLM-Assisted Software Development
by: Zhou, Xinyi, et al.
Published: (2026)
by: Zhou, Xinyi, et al.
Published: (2026)
The Evolution of Information Seeking in Software Development: Understanding the Role and Impact of AI Assistants
by: Haque, Ebtesam Al, et al.
Published: (2024)
by: Haque, Ebtesam Al, et al.
Published: (2024)
The Fast and Spurious: Developer Productivity with GenAI
by: Afroz, Sadia, et al.
Published: (2025)
by: Afroz, Sadia, et al.
Published: (2025)
Governance in Practice: How Open Source Projects Define and Document Roles
by: Oliveira, Pedro, et al.
Published: (2026)
by: Oliveira, Pedro, et al.
Published: (2026)
How Scientists Use Jupyter Notebooks: Goals, Quality Attributes, and Opportunities
by: Huang, Ruanqianqian, et al.
Published: (2025)
by: Huang, Ruanqianqian, et al.
Published: (2025)
How Developers Choose Debugging Strategies for Challenging Web Application Defects
by: Arab, Maryam, et al.
Published: (2025)
by: Arab, Maryam, et al.
Published: (2025)
"How do people decide?": A Model for Software Library Selection
by: Tanzil, Minaoar Hossain, et al.
Published: (2024)
by: Tanzil, Minaoar Hossain, et al.
Published: (2024)
Similar Items
-
Rethinking Model Evaluation as Narrowing the Socio-Technical Gap
by: Liao, Q. Vera, et al.
Published: (2023) -
The Centers and Margins of Modeling Humans in Well-being Technologies: A Decentering Approach
by: Zhu, Jichen, et al.
Published: (2025) -
"I Would Have Written My Code Differently'': Beginners Struggle to Understand LLM-Generated Code
by: Zi, Yangtian, et al.
Published: (2025) -
Generative Echo Chamber? Effects of LLM-Powered Search Systems on Diverse Information Seeking
by: Sharma, Nikhil, et al.
Published: (2024) -
Gendered Prompting and LLM Code Review: How Gender Cues in the Prompt Shape Code Quality and Evaluation
by: Janzen, Lynn, et al.
Published: (2026)