On the Flakiness of LLM-Generated Tests for Industrial and Open-Source Database Management Systems
Fuente:
arXiv
Saved in:
| Main Authors: | Berndt, Alexander, Bach, Thomas, Gemulla, Rainer, Kessel, Marcus, Baltes, Sebastian |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Can We Classify Flaky Tests Using Only Test Code? An LLM-Based Empirical Study
by: Berndt, Alexander, et al.
Published: (2026)
by: Berndt, Alexander, et al.
Published: (2026)
Flaky Tests in a Large Industrial Database Management System: An Empirical Study of Fixed Issue Reports for SAP HANA
by: Berndt, Alexander, et al.
Published: (2026)
by: Berndt, Alexander, et al.
Published: (2026)
Do Test and Environmental Complexity Increase Flakiness? An Empirical Study of SAP HANA
by: Berndt, Alexander, et al.
Published: (2024)
by: Berndt, Alexander, et al.
Published: (2024)
Taming Timeout Flakiness: An Empirical Study of SAP HANA
by: Berndt, Alexander, et al.
Published: (2024)
by: Berndt, Alexander, et al.
Published: (2024)
The Vocabulary of Flaky Tests in the Context of SAP HANA
by: Berndt, Alexander, et al.
Published: (2026)
by: Berndt, Alexander, et al.
Published: (2026)
Systemic Flakiness: An Empirical Analysis of Co-Occurring Flaky Test Failures
by: Parry, Owain, et al.
Published: (2025)
by: Parry, Owain, et al.
Published: (2025)
Context Engineering for AI Agents in Open-Source Software
by: Mohsenimofidi, Seyedmoein, et al.
Published: (2025)
by: Mohsenimofidi, Seyedmoein, et al.
Published: (2025)
Using Large Language Models to Support Automation of Failure Management in CI/CD Pipelines: A Case Study in SAP HANA
by: Bui, Duong, et al.
Published: (2026)
by: Bui, Duong, et al.
Published: (2026)
Ethics of Care for Software Engineering
by: Serebrenik, Alexander, et al.
Published: (2026)
by: Serebrenik, Alexander, et al.
Published: (2026)
Treating Run-time Execution History as a First-Class Citizen: Co-Versioning Run-time Behavior alongside Code
by: Kessel, Marcus
Published: (2026)
by: Kessel, Marcus
Published: (2026)
The Effects of Computational Resources on Flaky Tests
by: Silva, Denini, et al.
Published: (2023)
by: Silva, Denini, et al.
Published: (2023)
NeuroFlake: A Neuro-Symbolic LLM Framework for Flaky Test Classification
by: Hoque, Khondaker Tasnia, et al.
Published: (2026)
by: Hoque, Khondaker Tasnia, et al.
Published: (2026)
A Dataset of Reproducible Flaky-Test Failures
by: Rafi, Suzzana, et al.
Published: (2026)
by: Rafi, Suzzana, et al.
Published: (2026)
Self-Admitted GenAI Usage in Open-Source Software
by: Xiao, Tao, et al.
Published: (2025)
by: Xiao, Tao, et al.
Published: (2025)
Information-Theoretic Detection of Unusual Source Code Changes
by: Torres, Adriano, et al.
Published: (2025)
by: Torres, Adriano, et al.
Published: (2025)
A Generic Approach to Fix Test Flakiness in Real-World Projects
by: Chen, Yang, et al.
Published: (2024)
by: Chen, Yang, et al.
Published: (2024)
Automated Test Validators for Flaky Cyber-Physical System Simulators: Approach and Evaluation
by: Jodat, Baharin A., et al.
Published: (2025)
by: Jodat, Baharin A., et al.
Published: (2025)
UX Debt: Developers Borrow While Users Pay
by: Baltes, Sebastian, et al.
Published: (2021)
by: Baltes, Sebastian, et al.
Published: (2021)
A Multifaceted View on Discrimination in Software Development Careers
by: Chakraborty, Shalini, et al.
Published: (2025)
by: Chakraborty, Shalini, et al.
Published: (2025)
Lost in Transition: The Struggle of Women Returning to Software Engineering Research after Career Breaks
by: Chakraborty, Shalini, et al.
Published: (2025)
by: Chakraborty, Shalini, et al.
Published: (2025)
Teaching Literature Reviewing for Software Engineering Research
by: Baltes, Sebastian, et al.
Published: (2024)
by: Baltes, Sebastian, et al.
Published: (2024)
A Systematic Evaluation of Environmental Flakiness in JavaScript Tests
by: Hashemi, Negar, et al.
Published: (2026)
by: Hashemi, Negar, et al.
Published: (2026)
JS-TOD: Detecting Order-Dependent Flaky Tests in Jest
by: Hashemi, Negar, et al.
Published: (2025)
by: Hashemi, Negar, et al.
Published: (2025)
Detecting and Evaluating Order-Dependent Flaky Tests in JavaScript
by: Hashemi, Negar, et al.
Published: (2025)
by: Hashemi, Negar, et al.
Published: (2025)
Detecting Flaky Tests in Quantum Software: A Dynamic Approach
by: Kim, Dongchan, et al.
Published: (2025)
by: Kim, Dongchan, et al.
Published: (2025)
Reduction of Test Re-runs by Prioritizing Potential Order Dependent Flaky Tests
by: Iqbal, Hasnain, et al.
Published: (2025)
by: Iqbal, Hasnain, et al.
Published: (2025)
A Practical Framework for Flaky Failure Triage in Distributed Database Continuous Integration
by: Zhu, Jun-Peng, et al.
Published: (2026)
by: Zhu, Jun-Peng, et al.
Published: (2026)
Cross-Project Flakiness: A Case Study of the OpenStack Ecosystem
by: Xiao, Tao, et al.
Published: (2026)
by: Xiao, Tao, et al.
Published: (2026)
WEFix: Intelligent Automatic Generation of Explicit Waits for Efficient Web End-to-End Flaky Tests
by: Liu, Xinyue, et al.
Published: (2024)
by: Liu, Xinyue, et al.
Published: (2024)
Test-driven Software Experimentation with LASSO: an LLM Prompt Benchmarking Example
by: Kessel, Marcus
Published: (2024)
by: Kessel, Marcus
Published: (2024)
Dockerfile Flakiness: Characterization and Repair
by: Shabani, Taha, et al.
Published: (2024)
by: Shabani, Taha, et al.
Published: (2024)
A Preliminary Study of Fixed Flaky Tests in Rust Projects on GitHub
by: Schroeder, Tom, et al.
Published: (2025)
by: Schroeder, Tom, et al.
Published: (2025)
230,439 Test Failures Later: An Empirical Evaluation of Flaky Failure Classifiers
by: Alshammari, Abdulrahman, et al.
Published: (2024)
by: Alshammari, Abdulrahman, et al.
Published: (2024)
AI Slop and the Software Commons
by: Baltes, Sebastian, et al.
Published: (2026)
by: Baltes, Sebastian, et al.
Published: (2026)
"An Endless Stream of AI Slop": The Growing Burden of AI-Assisted Software Development
by: Baltes, Sebastian, et al.
Published: (2026)
by: Baltes, Sebastian, et al.
Published: (2026)
How Does Cognitive Capability and Personality Influence Problem Solving in Coding Interview Puzzles?
by: Hidellaarachchi, Dulaji, et al.
Published: (2025)
by: Hidellaarachchi, Dulaji, et al.
Published: (2025)
FlakyGuard: Automatically Fixing Flaky Tests at Industry Scale
by: Li, Chengpeng, et al.
Published: (2025)
by: Li, Chengpeng, et al.
Published: (2025)
FlakyFix: Using Large Language Models for Predicting Flaky Test Fix Categories and Test Code Repair
by: Fatima, Sakina, et al.
Published: (2023)
by: Fatima, Sakina, et al.
Published: (2023)
Detecting and Mitigating Flakiness in REST API Fuzzing
by: Zhang, Man, et al.
Published: (2026)
by: Zhang, Man, et al.
Published: (2026)
Identifying Flaky Tests in Quantum Code: A Machine Learning Approach
by: Kaur, Khushdeep, et al.
Published: (2025)
by: Kaur, Khushdeep, et al.
Published: (2025)
Similar Items
-
Can We Classify Flaky Tests Using Only Test Code? An LLM-Based Empirical Study
by: Berndt, Alexander, et al.
Published: (2026) -
Flaky Tests in a Large Industrial Database Management System: An Empirical Study of Fixed Issue Reports for SAP HANA
by: Berndt, Alexander, et al.
Published: (2026) -
Do Test and Environmental Complexity Increase Flakiness? An Empirical Study of SAP HANA
by: Berndt, Alexander, et al.
Published: (2024) -
Taming Timeout Flakiness: An Empirical Study of SAP HANA
by: Berndt, Alexander, et al.
Published: (2024) -
The Vocabulary of Flaky Tests in the Context of SAP HANA
by: Berndt, Alexander, et al.
Published: (2026)