Towards Evaluation Guidelines for Empirical Studies involving LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wagner, Stefan, Barón, Marvin Muñoz, Falessi, Davide, Baltes, Sebastian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Guidelines for Empirical Studies in Software Engineering involving Large Language Models
von: Baltes, Sebastian, et al.
Veröffentlicht: (2025)
von: Baltes, Sebastian, et al.
Veröffentlicht: (2025)
Taming Timeout Flakiness: An Empirical Study of SAP HANA
von: Berndt, Alexander, et al.
Veröffentlicht: (2024)
von: Berndt, Alexander, et al.
Veröffentlicht: (2024)
Do Test and Environmental Complexity Increase Flakiness? An Empirical Study of SAP HANA
von: Berndt, Alexander, et al.
Veröffentlicht: (2024)
von: Berndt, Alexander, et al.
Veröffentlicht: (2024)
Characterizing Requirements Smells
von: Gentili, Emanuele, et al.
Veröffentlicht: (2024)
von: Gentili, Emanuele, et al.
Veröffentlicht: (2024)
Flaky Tests in a Large Industrial Database Management System: An Empirical Study of Fixed Issue Reports for SAP HANA
von: Berndt, Alexander, et al.
Veröffentlicht: (2026)
von: Berndt, Alexander, et al.
Veröffentlicht: (2026)
Apples, Oranges, and Software Engineering: Study Selection Challenges for Secondary Research on Latent Variables
von: Wyrich, Marvin, et al.
Veröffentlicht: (2024)
von: Wyrich, Marvin, et al.
Veröffentlicht: (2024)
The Silent Scientist: When Software Research Fails to Reach Its Audience
von: Wyrich, Marvin, et al.
Veröffentlicht: (2025)
von: Wyrich, Marvin, et al.
Veröffentlicht: (2025)
Information-Theoretic Detection of Unusual Source Code Changes
von: Torres, Adriano, et al.
Veröffentlicht: (2025)
von: Torres, Adriano, et al.
Veröffentlicht: (2025)
Anticipating Bugs: Ticket-Level Bug Prediction and Temporal Proximity Effects
von: La Prova, Daniele, et al.
Veröffentlicht: (2025)
von: La Prova, Daniele, et al.
Veröffentlicht: (2025)
Teaching Literature Reviewing for Software Engineering Research
von: Baltes, Sebastian, et al.
Veröffentlicht: (2024)
von: Baltes, Sebastian, et al.
Veröffentlicht: (2024)
UX Debt: Developers Borrow While Users Pay
von: Baltes, Sebastian, et al.
Veröffentlicht: (2021)
von: Baltes, Sebastian, et al.
Veröffentlicht: (2021)
A Multifaceted View on Discrimination in Software Development Careers
von: Chakraborty, Shalini, et al.
Veröffentlicht: (2025)
von: Chakraborty, Shalini, et al.
Veröffentlicht: (2025)
Lost in Transition: The Struggle of Women Returning to Software Engineering Research after Career Breaks
von: Chakraborty, Shalini, et al.
Veröffentlicht: (2025)
von: Chakraborty, Shalini, et al.
Veröffentlicht: (2025)
Ethics of Care for Software Engineering
von: Serebrenik, Alexander, et al.
Veröffentlicht: (2026)
von: Serebrenik, Alexander, et al.
Veröffentlicht: (2026)
Can We Classify Flaky Tests Using Only Test Code? An LLM-Based Empirical Study
von: Berndt, Alexander, et al.
Veröffentlicht: (2026)
von: Berndt, Alexander, et al.
Veröffentlicht: (2026)
AI Slop and the Software Commons
von: Baltes, Sebastian, et al.
Veröffentlicht: (2026)
von: Baltes, Sebastian, et al.
Veröffentlicht: (2026)
How Does Cognitive Capability and Personality Influence Problem Solving in Coding Interview Puzzles?
von: Hidellaarachchi, Dulaji, et al.
Veröffentlicht: (2025)
von: Hidellaarachchi, Dulaji, et al.
Veröffentlicht: (2025)
"An Endless Stream of AI Slop": The Growing Burden of AI-Assisted Software Development
von: Baltes, Sebastian, et al.
Veröffentlicht: (2026)
von: Baltes, Sebastian, et al.
Veröffentlicht: (2026)
An Extensive Comparison of Static Application Security Testing Tools
von: Esposito, Matteo, et al.
Veröffentlicht: (2024)
von: Esposito, Matteo, et al.
Veröffentlicht: (2024)
The Influence of Code Comments on the Perceived Helpfulness of Stack Overflow Posts
von: Figl, Kathrin, et al.
Veröffentlicht: (2025)
von: Figl, Kathrin, et al.
Veröffentlicht: (2025)
Context Engineering for AI Agents in Open-Source Software
von: Mohsenimofidi, Seyedmoein, et al.
Veröffentlicht: (2025)
von: Mohsenimofidi, Seyedmoein, et al.
Veröffentlicht: (2025)
Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study
von: Santana Jr, E. G., et al.
Veröffentlicht: (2025)
von: Santana Jr, E. G., et al.
Veröffentlicht: (2025)
40 Years of Designing Code Comprehension Experiments: A Systematic Mapping Study
von: Wyrich, Marvin, et al.
Veröffentlicht: (2022)
von: Wyrich, Marvin, et al.
Veröffentlicht: (2022)
Operationalizing Ethics for AI Agents: How Developers Encode Values into Repository Context Files
von: Treude, Christoph, et al.
Veröffentlicht: (2026)
von: Treude, Christoph, et al.
Veröffentlicht: (2026)
Software Engineering Podcasts: An Empirical Study of Their Potential as a Research Resource
von: Wyrich, Marvin, et al.
Veröffentlicht: (2026)
von: Wyrich, Marvin, et al.
Veröffentlicht: (2026)
Configuring Agentic AI Coding Tools: An Exploratory Study
von: Galster, Matthias, et al.
Veröffentlicht: (2026)
von: Galster, Matthias, et al.
Veröffentlicht: (2026)
On the Flakiness of LLM-Generated Tests for Industrial and Open-Source Database Management Systems
von: Berndt, Alexander, et al.
Veröffentlicht: (2026)
von: Berndt, Alexander, et al.
Veröffentlicht: (2026)
An Empirical Study of Perceptions of General LLMs and Multimodal LLMs on Hugging Face
von: Liu, Yujian, et al.
Veröffentlicht: (2026)
von: Liu, Yujian, et al.
Veröffentlicht: (2026)
Political and Ideological Pressure in Software Engineering Research: The Case of DEI Backlash
von: Hyrynsalmi, Sonja M., et al.
Veröffentlicht: (2026)
von: Hyrynsalmi, Sonja M., et al.
Veröffentlicht: (2026)
Not real or too soft? On the challenges of publishing interdisciplinary software engineering research
von: Hyrynsalmi, Sonja M., et al.
Veröffentlicht: (2025)
von: Hyrynsalmi, Sonja M., et al.
Veröffentlicht: (2025)
An Empirical Study on the Potential of LLMs in Automated Software Refactoring
von: Liu, Bo, et al.
Veröffentlicht: (2024)
von: Liu, Bo, et al.
Veröffentlicht: (2024)
An Empirical Study on the Capability of LLMs in Decomposing Bug Reports
von: Chen, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Chen, Zhiyuan, et al.
Veröffentlicht: (2025)
On the Need to Rethink Trust in AI Assistants for Software Development: A Critical Review
von: Baltes, Sebastian, et al.
Veröffentlicht: (2025)
von: Baltes, Sebastian, et al.
Veröffentlicht: (2025)
LLMs are Bug Replicators: An Empirical Study on LLMs' Capability in Completing Bug-prone Code
von: Guo, Liwei, et al.
Veröffentlicht: (2025)
von: Guo, Liwei, et al.
Veröffentlicht: (2025)
Fault Localisation and Repair for DL Systems: An Empirical Study with LLMs
von: Kim, Jinhan, et al.
Veröffentlicht: (2025)
von: Kim, Jinhan, et al.
Veröffentlicht: (2025)
Exploring the Effectiveness of LLMs in Automated Logging Generation: An Empirical Study
von: Li, Yichen, et al.
Veröffentlicht: (2023)
von: Li, Yichen, et al.
Veröffentlicht: (2023)
$Classi|Q\rangle$ Towards a Translation Framework To Bridge The Classical-Quantum Programming Gap
von: Esposito, Matteo, et al.
Veröffentlicht: (2024)
von: Esposito, Matteo, et al.
Veröffentlicht: (2024)
Policy-driven Software Bill of Materials on GitHub: An Empirical Study
von: Novikov, Oleksii, et al.
Veröffentlicht: (2025)
von: Novikov, Oleksii, et al.
Veröffentlicht: (2025)
Guidelines to Prompt Large Language Models for Code Generation: An Empirical Characterization
von: Midolo, Alessandro, et al.
Veröffentlicht: (2026)
von: Midolo, Alessandro, et al.
Veröffentlicht: (2026)
Harnessing Hype to Teach Empirical Thinking: An Experience With AI Coding Assistants
von: Wyrich, Marvin, et al.
Veröffentlicht: (2026)
von: Wyrich, Marvin, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Guidelines for Empirical Studies in Software Engineering involving Large Language Models
von: Baltes, Sebastian, et al.
Veröffentlicht: (2025) -
Taming Timeout Flakiness: An Empirical Study of SAP HANA
von: Berndt, Alexander, et al.
Veröffentlicht: (2024) -
Do Test and Environmental Complexity Increase Flakiness? An Empirical Study of SAP HANA
von: Berndt, Alexander, et al.
Veröffentlicht: (2024) -
Characterizing Requirements Smells
von: Gentili, Emanuele, et al.
Veröffentlicht: (2024) -
Flaky Tests in a Large Industrial Database Management System: An Empirical Study of Fixed Issue Reports for SAP HANA
von: Berndt, Alexander, et al.
Veröffentlicht: (2026)