Understanding and Detecting Flaky Builds in GitHub Actions
Fuente:
arXiv
Saved in:
| Main Authors: | Ge, Wenhao, Zhang, Chen |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Experience with GitHub Copilot for Developer Productivity at Zoominfo
by: Bakal, Gal, et al.
Published: (2025)
by: Bakal, Gal, et al.
Published: (2025)
On the Mistaken Assumption of Interchangeable Deep Reinforcement Learning Implementations
by: Hundal, Rajdeep Singh, et al.
Published: (2025)
by: Hundal, Rajdeep Singh, et al.
Published: (2025)
Towards Explainable Test Case Prioritisation with Learning-to-Rank Models
by: Ramírez, Aurora, et al.
Published: (2024)
by: Ramírez, Aurora, et al.
Published: (2024)
CodeTracer: Towards Traceable Agent States
by: Li, Han, et al.
Published: (2026)
by: Li, Han, et al.
Published: (2026)
AuditRepairBench: A Paired-Execution Trace Corpus for Evaluator-Channel Ranking Instability in Agent Repair
by: Hu, Yuelin, et al.
Published: (2026)
by: Hu, Yuelin, et al.
Published: (2026)
Automating the Detection of Code Vulnerabilities by Analyzing GitHub Issues
by: Cipollone, Daniele, et al.
Published: (2025)
by: Cipollone, Daniele, et al.
Published: (2025)
GitHub Copilot and Developer Productivity: An Observational Dose-Response Analysis
by: Heilman, Alex, et al.
Published: (2026)
by: Heilman, Alex, et al.
Published: (2026)
AIRA: AI-Induced Risk Audit: A Structured Inspection Framework for AI-Generated Code
by: Parris, William M.
Published: (2026)
by: Parris, William M.
Published: (2026)
DEFault++: Automated Fault Detection, Categorization, and Diagnosis for Transformer Architectures
by: Jahan, Sigma, et al.
Published: (2026)
by: Jahan, Sigma, et al.
Published: (2026)
Mining Subscenario Refactoring Opportunities in Behaviour-Driven Software Test Suites: ML Classifiers and LLM-Judge Baselines
by: Mughal, Ali Hassaan, et al.
Published: (2026)
by: Mughal, Ali Hassaan, et al.
Published: (2026)
VulScribeR: Exploring RAG-based Vulnerability Augmentation with LLMs
by: Daneshvar, Seyed Shayan, et al.
Published: (2024)
by: Daneshvar, Seyed Shayan, et al.
Published: (2024)
CUDABeaver: Benchmarking LLM-Based Automated CUDA Debugging
by: Li, Shiyang, et al.
Published: (2026)
by: Li, Shiyang, et al.
Published: (2026)
PyPackIT: Automated Research Software Engineering for Scientific Python Applications on GitHub
by: Ariamajd, Armin, et al.
Published: (2025)
by: Ariamajd, Armin, et al.
Published: (2025)
Feedback-Normalized Developer Memory for Reinforcement-Learning Coding Agents: A Safety-Gated MCP Architecture
by: Iscan, Mehmet
Published: (2026)
by: Iscan, Mehmet
Published: (2026)
An Analysis of LLM Fine-Tuning and Few-Shot Learning for Flaky Test Detection and Classification
by: More, Riddhi, et al.
Published: (2025)
by: More, Riddhi, et al.
Published: (2025)
DRS-OSS: Practical Diff Risk Scoring with LLMs
by: Sayedsalehi, Ali, et al.
Published: (2025)
by: Sayedsalehi, Ali, et al.
Published: (2025)
Dual-Process Scaffold Reasoning for Enhancing LLM Code Debugging
by: Hsieh, Po-Chung, et al.
Published: (2025)
by: Hsieh, Po-Chung, et al.
Published: (2025)
Code Documentation and Analysis to Secure Software Development
by: Attie, Paul, et al.
Published: (2024)
by: Attie, Paul, et al.
Published: (2024)
Context Engineering for Multi-Agent LLM Code Assistants Using Elicit, NotebookLM, ChatGPT, and Claude Code
by: Haseeb, Muhammad
Published: (2025)
by: Haseeb, Muhammad
Published: (2025)
Natural Language Summarization Enables Multi-Repository Bug Localization by LLMs in Microservice Architectures
by: Oskooei, Amirkia Rafiei, et al.
Published: (2025)
by: Oskooei, Amirkia Rafiei, et al.
Published: (2025)
An LSTM-based Test Selection Method for Self-Driving Cars
by: Güllü, Ali, et al.
Published: (2025)
by: Güllü, Ali, et al.
Published: (2025)
RepoAudit: An Autonomous LLM-Agent for Repository-Level Code Auditing
by: Guo, Jinyao, et al.
Published: (2025)
by: Guo, Jinyao, et al.
Published: (2025)
ClawHub Security Signals: When VirusTotal, Static Analysis, and SkillSpector Disagree
by: Koc, Vincent, et al.
Published: (2026)
by: Koc, Vincent, et al.
Published: (2026)
SpecOps: A Fully Automated AI Agent Testing Framework in Real-World GUI Environments
by: Ahmed, Syed Yusuf, et al.
Published: (2026)
by: Ahmed, Syed Yusuf, et al.
Published: (2026)
LLMCup: Ranking-Enhanced Comment Updating with LLMs
by: Ge, Hua, et al.
Published: (2025)
by: Ge, Hua, et al.
Published: (2025)
A measurement substrate for agentic Kubernetes operations: Methodology and a case study in retrieval-compounding falsification
by: Odmark, Joshua, et al.
Published: (2026)
by: Odmark, Joshua, et al.
Published: (2026)
Runtime Execution Traces Guided Automated Program Repair with Multi-Agent Debate
by: Wu, Jiaqing, et al.
Published: (2026)
by: Wu, Jiaqing, et al.
Published: (2026)
RMCBench: Benchmarking Large Language Models' Resistance to Malicious Code
by: Chen, Jiachi, et al.
Published: (2024)
by: Chen, Jiachi, et al.
Published: (2024)
Scattered Forest Search: Smarter Code Space Exploration with LLMs
by: Light, Jonathan, et al.
Published: (2024)
by: Light, Jonathan, et al.
Published: (2024)
Highly Interactive Testing for Uninterrupted Development Flow
by: Tropin, Andrew
Published: (2025)
by: Tropin, Andrew
Published: (2025)
Generative AI and the Transformation of Software Development Practices
by: Acharya, Vivek
Published: (2025)
by: Acharya, Vivek
Published: (2025)
L2MAC: Large Language Model Automatic Computer for Extensive Code Generation
by: Holt, Samuel, et al.
Published: (2023)
by: Holt, Samuel, et al.
Published: (2023)
SLEAN: Simple Lightweight Ensemble Analysis Network for Multi-Provider LLM Coordination: Design, Implementation, and Vibe Coding Bug Investigation Case Study
by: Vargas, Matheus J. T.
Published: (2025)
by: Vargas, Matheus J. T.
Published: (2025)
MeDeT: Medical Device Digital Twins Creation with Few-shot Meta-learning
by: Sartaj, Hassan, et al.
Published: (2024)
by: Sartaj, Hassan, et al.
Published: (2024)
Demystifying the Silence of Correctness Bugs in PyTorch Compiler
by: Li, Meiziniu, et al.
Published: (2026)
by: Li, Meiziniu, et al.
Published: (2026)
Enhancing Differential Testing With LLMs For Testing Deep Learning Libraries
by: Li, Meiziniu, et al.
Published: (2024)
by: Li, Meiziniu, et al.
Published: (2024)
COMET: Coverage-guided Model Generation For Deep Learning Library Testing
by: Li, Meiziniu, et al.
Published: (2022)
by: Li, Meiziniu, et al.
Published: (2022)
Review Beats Planning: Dual-Model Interaction Patterns for Code Synthesis
by: Miller, Jan
Published: (2026)
by: Miller, Jan
Published: (2026)
Predictive Analytics for Collaborators Answers, Code Quality, and Dropout on Stack Overflow
by: Zolduoarrati, Elijah, et al.
Published: (2025)
by: Zolduoarrati, Elijah, et al.
Published: (2025)
LLMDFA: Analyzing Dataflow in Code with Large Language Models
by: Wang, Chengpeng, et al.
Published: (2024)
by: Wang, Chengpeng, et al.
Published: (2024)
Similar Items
-
Experience with GitHub Copilot for Developer Productivity at Zoominfo
by: Bakal, Gal, et al.
Published: (2025) -
On the Mistaken Assumption of Interchangeable Deep Reinforcement Learning Implementations
by: Hundal, Rajdeep Singh, et al.
Published: (2025) -
Towards Explainable Test Case Prioritisation with Learning-to-Rank Models
by: Ramírez, Aurora, et al.
Published: (2024) -
CodeTracer: Towards Traceable Agent States
by: Li, Han, et al.
Published: (2026) -
AuditRepairBench: A Paired-Execution Trace Corpus for Evaluator-Channel Ranking Instability in Agent Repair
by: Hu, Yuelin, et al.
Published: (2026)