FlakyGuard: Automatically Fixing Flaky Tests at Industry Scale
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Chengpeng, Behrang, Farnaz, Shi, August, Liu, Peng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DR.FIX: Automatically Fixing Data Races at Industry Scale
by: Behrang, Farnaz, et al.
Published: (2025)
by: Behrang, Farnaz, et al.
Published: (2025)
FlakyFix: Using Large Language Models for Predicting Flaky Test Fix Categories and Test Code Repair
by: Fatima, Sakina, et al.
Published: (2023)
by: Fatima, Sakina, et al.
Published: (2023)
FlaKat: A Machine Learning-Based Categorization Framework for Flaky Tests
by: Lin, Shizhe, et al.
Published: (2024)
by: Lin, Shizhe, et al.
Published: (2024)
A Dataset of Reproducible Flaky-Test Failures
by: Rafi, Suzzana, et al.
Published: (2026)
by: Rafi, Suzzana, et al.
Published: (2026)
Identifying Flaky Tests in Quantum Code: A Machine Learning Approach
by: Kaur, Khushdeep, et al.
Published: (2025)
by: Kaur, Khushdeep, et al.
Published: (2025)
APRIL: API Synthesis with Automatic Prompt Optimization and Reinforcement Learning
by: Zhong, Hua, et al.
Published: (2025)
by: Zhong, Hua, et al.
Published: (2025)
JS-TOD: Detecting Order-Dependent Flaky Tests in Jest
by: Hashemi, Negar, et al.
Published: (2025)
by: Hashemi, Negar, et al.
Published: (2025)
Detecting and Evaluating Order-Dependent Flaky Tests in JavaScript
by: Hashemi, Negar, et al.
Published: (2025)
by: Hashemi, Negar, et al.
Published: (2025)
A Systematic Evaluation of Environmental Flakiness in JavaScript Tests
by: Hashemi, Negar, et al.
Published: (2026)
by: Hashemi, Negar, et al.
Published: (2026)
$\textbf{PLUM}$: Improving Code LMs with Execution-Guided On-Policy Preference Learning Driven By Synthetic Test Cases
by: Zhang, Dylan, et al.
Published: (2024)
by: Zhang, Dylan, et al.
Published: (2024)
Automating Detection and Root-Cause Analysis of Flaky Tests in Quantum Software
by: Sivaloganathan, Janakan, et al.
Published: (2026)
by: Sivaloganathan, Janakan, et al.
Published: (2026)
Automating Quantum Software Maintenance: Flakiness Detection and Root Cause Analysis
by: Sivaloganathan, Janakan, et al.
Published: (2024)
by: Sivaloganathan, Janakan, et al.
Published: (2024)
Systemic Flakiness: An Empirical Analysis of Co-Occurring Flaky Test Failures
by: Parry, Owain, et al.
Published: (2025)
by: Parry, Owain, et al.
Published: (2025)
AI-Assisted Fixes to Code Review Comments at Scale
by: Maddila, Chandra, et al.
Published: (2025)
by: Maddila, Chandra, et al.
Published: (2025)
Large Language Models Synergize with Automated Machine Learning
by: Xu, Jinglue, et al.
Published: (2024)
by: Xu, Jinglue, et al.
Published: (2024)
Verify Before You Fix: Agentic Execution Grounding for Trustworthy Cross-Language Code Analysis
by: Gajjar, Jugal
Published: (2026)
by: Gajjar, Jugal
Published: (2026)
A Joint Learning Model with Variational Interaction for Multilingual Program Translation
by: Du, Yali, et al.
Published: (2024)
by: Du, Yali, et al.
Published: (2024)
Raw Pointer Rewriting with LLMs for Translating C to Safer Rust
by: Gao, Yifei, et al.
Published: (2025)
by: Gao, Yifei, et al.
Published: (2025)
Can It Edit? Evaluating the Ability of Large Language Models to Follow Code Editing Instructions
by: Cassano, Federico, et al.
Published: (2023)
by: Cassano, Federico, et al.
Published: (2023)
Automatically Testing Functional Properties of Code Translation Models
by: Eniser, Hasan Ferit, et al.
Published: (2023)
by: Eniser, Hasan Ferit, et al.
Published: (2023)
Incoherence as Oracle-less Measure of Error in LLM-Based Code Generation
by: Valentin, Thomas, et al.
Published: (2025)
by: Valentin, Thomas, et al.
Published: (2025)
FormalSpecCpp: A Dataset of C++ Formal Specifications created using LLMs
by: Chakraborty, Madhurima, et al.
Published: (2025)
by: Chakraborty, Madhurima, et al.
Published: (2025)
Benchmarking LLM for Code Smells Detection: OpenAI GPT-4.0 vs DeepSeek-V3
by: Sadik, Ahmed R., et al.
Published: (2025)
by: Sadik, Ahmed R., et al.
Published: (2025)
Representing Prompting Patterns with PDL: Compliance Agent Case Study
by: Vaziri, Mandana, et al.
Published: (2025)
by: Vaziri, Mandana, et al.
Published: (2025)
ChatDBG: Augmenting Debugging with Large Language Models
by: Levin, Kyla H., et al.
Published: (2024)
by: Levin, Kyla H., et al.
Published: (2024)
Do Large Code Models Understand Programming Concepts? Counterfactual Analysis for Code Predicates
by: Hooda, Ashish, et al.
Published: (2024)
by: Hooda, Ashish, et al.
Published: (2024)
On the Effectiveness of Machine Learning-based Call Graph Pruning: An Empirical Study
by: Mir, Amir M., et al.
Published: (2024)
by: Mir, Amir M., et al.
Published: (2024)
Understanding Tool-Augmented Agents for Lean Formalization: A Factorial Analysis
by: Zhang, Ke, et al.
Published: (2026)
by: Zhang, Ke, et al.
Published: (2026)
ScenicNL: Generating Probabilistic Scenario Programs from Natural Language
by: Elmaaroufi, Karim, et al.
Published: (2024)
by: Elmaaroufi, Karim, et al.
Published: (2024)
Large Language Models for Code Summarization
by: Szalontai, Balázs, et al.
Published: (2024)
by: Szalontai, Balázs, et al.
Published: (2024)
DafnyBench: A Benchmark for Formal Software Verification
by: Loughridge, Chloe, et al.
Published: (2024)
by: Loughridge, Chloe, et al.
Published: (2024)
PerfRL: A Small Language Model Framework for Efficient Code Optimization
by: Duan, Shukai, et al.
Published: (2023)
by: Duan, Shukai, et al.
Published: (2023)
JavaBench: A Benchmark of Object-Oriented Code Generation for Evaluating Large Language Models
by: Cao, Jialun, et al.
Published: (2024)
by: Cao, Jialun, et al.
Published: (2024)
MonoCoder: Domain-Specific Code Language Model for HPC Codes and Tasks
by: Kadosh, Tal, et al.
Published: (2023)
by: Kadosh, Tal, et al.
Published: (2023)
A Multi-Expert Large Language Model Architecture for Verilog Code Generation
by: Nadimi, Bardia, et al.
Published: (2024)
by: Nadimi, Bardia, et al.
Published: (2024)
Encoding architecture algebra
by: Bersier, Stephane, et al.
Published: (2024)
by: Bersier, Stephane, et al.
Published: (2024)
A Generic Approach to Fix Test Flakiness in Real-World Projects
by: Chen, Yang, et al.
Published: (2024)
by: Chen, Yang, et al.
Published: (2024)
Is Programming by Example solved by LLMs?
by: Li, Wen-Ding, et al.
Published: (2024)
by: Li, Wen-Ding, et al.
Published: (2024)
EquiBench: Benchmarking Large Language Models' Reasoning about Program Semantics via Equivalence Checking
by: Wei, Anjiang, et al.
Published: (2025)
by: Wei, Anjiang, et al.
Published: (2025)
TokDrift: When LLM Speaks in Subwords but Code Speaks in Grammar
by: Li, Yinxi, et al.
Published: (2025)
by: Li, Yinxi, et al.
Published: (2025)
Similar Items
-
DR.FIX: Automatically Fixing Data Races at Industry Scale
by: Behrang, Farnaz, et al.
Published: (2025) -
FlakyFix: Using Large Language Models for Predicting Flaky Test Fix Categories and Test Code Repair
by: Fatima, Sakina, et al.
Published: (2023) -
FlaKat: A Machine Learning-Based Categorization Framework for Flaky Tests
by: Lin, Shizhe, et al.
Published: (2024) -
A Dataset of Reproducible Flaky-Test Failures
by: Rafi, Suzzana, et al.
Published: (2026) -
Identifying Flaky Tests in Quantum Code: A Machine Learning Approach
by: Kaur, Khushdeep, et al.
Published: (2025)