An Analysis of LLM Fine-Tuning and Few-Shot Learning for Flaky Test Detection and Classification
Fuente:
arXiv
Salvato in:
| Autori principali: | More, Riddhi, Bradbury, Jeremy S. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Assessing Data Augmentation-Induced Bias in Training and Testing of Machine Learning Models
di: More, Riddhi, et al.
Pubblicazione: (2025)
di: More, Riddhi, et al.
Pubblicazione: (2025)
Addressing Data Leakage in HumanEval Using Combinatorial Test Design
di: Bradbury, Jeremy S., et al.
Pubblicazione: (2024)
di: Bradbury, Jeremy S., et al.
Pubblicazione: (2024)
From Untestable to Testable: Metamorphic Testing in the Age of LLMs
di: Terragni, Valerio
Pubblicazione: (2026)
di: Terragni, Valerio
Pubblicazione: (2026)
A Systematic Approach for Assessing Large Language Models' Test Case Generation Capability
di: Chang, Hung-Fu, et al.
Pubblicazione: (2025)
di: Chang, Hung-Fu, et al.
Pubblicazione: (2025)
Towards a Probabilistic Framework for Analyzing and Improving LLM-Enabled Software
di: Baldonado, Juan Manuel, et al.
Pubblicazione: (2025)
di: Baldonado, Juan Manuel, et al.
Pubblicazione: (2025)
When Retrieval Hurts Code Completion: A Diagnostic Study of Stale Repository Context
di: Weng, Haojun, et al.
Pubblicazione: (2026)
di: Weng, Haojun, et al.
Pubblicazione: (2026)
CoTran: An LLM-based Code Translator using Reinforcement Learning with Feedback from Compiler and Symbolic Execution
di: Jana, Prithwish, et al.
Pubblicazione: (2023)
di: Jana, Prithwish, et al.
Pubblicazione: (2023)
AcTracer: Active Testing of Large Language Model via Multi-Stage Sampling
di: Huang, Yuheng, et al.
Pubblicazione: (2024)
di: Huang, Yuheng, et al.
Pubblicazione: (2024)
RelRepair: Enhancing Automated Program Repair by Retrieving Relevant Code
di: Liu, Shunyu, et al.
Pubblicazione: (2025)
di: Liu, Shunyu, et al.
Pubblicazione: (2025)
AgentModernize: Preserving Business Logic in Legacy Modernization with Multi-Agent LLMs and Behavioral Specification Graphs
di: Ahmed, Sheikh Nazib, et al.
Pubblicazione: (2026)
di: Ahmed, Sheikh Nazib, et al.
Pubblicazione: (2026)
AuditRepairBench: A Paired-Execution Trace Corpus for Evaluator-Channel Ranking Instability in Agent Repair
di: Hu, Yuelin, et al.
Pubblicazione: (2026)
di: Hu, Yuelin, et al.
Pubblicazione: (2026)
Automated structural testing of LLM-based agents: methods, framework, and case studies
di: Kohl, Jens, et al.
Pubblicazione: (2026)
di: Kohl, Jens, et al.
Pubblicazione: (2026)
GALA: Multimodal Graph Alignment for Bug Localization in Automated Program Repair
di: Liu, Zhuoyao, et al.
Pubblicazione: (2026)
di: Liu, Zhuoyao, et al.
Pubblicazione: (2026)
RMCBench: Benchmarking Large Language Models' Resistance to Malicious Code
di: Chen, Jiachi, et al.
Pubblicazione: (2024)
di: Chen, Jiachi, et al.
Pubblicazione: (2024)
Bug In the Code Stack: Can LLMs Find Bugs in Large Python Code Stacks
di: Lee, Hokyung, et al.
Pubblicazione: (2024)
di: Lee, Hokyung, et al.
Pubblicazione: (2024)
Reducing Maintenance Burden in Behaviour-Driven Development: A Paraphrase-Robust Duplicate-Step Detector with a 1.1M-Step Open Benchmark
di: Mughal, Ali Hassaan, et al.
Pubblicazione: (2026)
di: Mughal, Ali Hassaan, et al.
Pubblicazione: (2026)
Retromorphic Testing with Hierarchical Verification for Hallucination Detection in RAG
di: Yu, Boxi, et al.
Pubblicazione: (2026)
di: Yu, Boxi, et al.
Pubblicazione: (2026)
LLMORPH: Automated Metamorphic Testing of Large Language Models
di: Cho, Steven, et al.
Pubblicazione: (2026)
di: Cho, Steven, et al.
Pubblicazione: (2026)
CIFE: Code Instruction-Following Evaluation
di: Gunnu, Sravani, et al.
Pubblicazione: (2025)
di: Gunnu, Sravani, et al.
Pubblicazione: (2025)
Mining Subscenario Refactoring Opportunities in Behaviour-Driven Software Test Suites: ML Classifiers and LLM-Judge Baselines
di: Mughal, Ali Hassaan, et al.
Pubblicazione: (2026)
di: Mughal, Ali Hassaan, et al.
Pubblicazione: (2026)
AgentEval: DAG-Structured Step-Level Evaluation for Agentic Workflows with Error Propagation Tracking
di: Guo, Dongxin, et al.
Pubblicazione: (2026)
di: Guo, Dongxin, et al.
Pubblicazione: (2026)
Understanding and Detecting Flaky Builds in GitHub Actions
di: Ge, Wenhao, et al.
Pubblicazione: (2026)
di: Ge, Wenhao, et al.
Pubblicazione: (2026)
SPIRA: Building an Intelligent System for Respiratory Insufficiency Detection
di: Ferreira, Renato Cordeiro, et al.
Pubblicazione: (2025)
di: Ferreira, Renato Cordeiro, et al.
Pubblicazione: (2025)
Learning Software Bug Reports: A Systematic Literature Review
di: Long, Guoming, et al.
Pubblicazione: (2025)
di: Long, Guoming, et al.
Pubblicazione: (2025)
Continuous Discovery of Vulnerabilities in LLM Serving Systems with Fuzzing
di: Zhao, Yunze, et al.
Pubblicazione: (2026)
di: Zhao, Yunze, et al.
Pubblicazione: (2026)
Making a Pipeline Production-Ready: Challenges and Lessons Learned in the Healthcare Domain
di: Lawand, Daniel Angelo Esteves, et al.
Pubblicazione: (2025)
di: Lawand, Daniel Angelo Esteves, et al.
Pubblicazione: (2025)
Enhancing Differential Testing With LLMs For Testing Deep Learning Libraries
di: Li, Meiziniu, et al.
Pubblicazione: (2024)
di: Li, Meiziniu, et al.
Pubblicazione: (2024)
Toward Architecture-Aware Evaluation Metrics for LLM Agents
di: Souza, Débora, et al.
Pubblicazione: (2026)
di: Souza, Débora, et al.
Pubblicazione: (2026)
Automated Bug Triaging using Instruction-Tuned Large Language Models
di: Kiashemshaki, Kiana, et al.
Pubblicazione: (2025)
di: Kiashemshaki, Kiana, et al.
Pubblicazione: (2025)
PROTEA: Offline Evaluation and Iterative Refinement for Multi-Agent LLM Workflows
di: Kawamura, Kazuki, et al.
Pubblicazione: (2026)
di: Kawamura, Kazuki, et al.
Pubblicazione: (2026)
Software Defined Vehicle Code Generation: A Few-Shot Prompting Approach
di: Nguyen, Quang-Dung, et al.
Pubblicazione: (2025)
di: Nguyen, Quang-Dung, et al.
Pubblicazione: (2025)
COMET: Coverage-guided Model Generation For Deep Learning Library Testing
di: Li, Meiziniu, et al.
Pubblicazione: (2022)
di: Li, Meiziniu, et al.
Pubblicazione: (2022)
RepoLaunch: Automating Build&Test Pipeline of Code Repositories on ANY Language and ANY Platform
di: Li, Kenan, et al.
Pubblicazione: (2026)
di: Li, Kenan, et al.
Pubblicazione: (2026)
SmellBench: Evaluating LLM Agents on Architectural Code Smell Repair
di: Dinu, Ion George, et al.
Pubblicazione: (2026)
di: Dinu, Ion George, et al.
Pubblicazione: (2026)
The Kieker Observability Framework Version 2
di: Yang, Shinhyung, et al.
Pubblicazione: (2025)
di: Yang, Shinhyung, et al.
Pubblicazione: (2025)
Feedback-Normalized Developer Memory for Reinforcement-Learning Coding Agents: A Safety-Gated MCP Architecture
di: Iscan, Mehmet
Pubblicazione: (2026)
di: Iscan, Mehmet
Pubblicazione: (2026)
Fine-Tuning LLMs to Analyze Multiple Dimensions of Code Review: A Maximum Entropy Regulated Long Chain-of-Thought Approach
di: Yu, Yongda, et al.
Pubblicazione: (2025)
di: Yu, Yongda, et al.
Pubblicazione: (2025)
LLMs taking shortcuts in test generation: A study with SAP HANA and LevelDB
di: Bekmyradov, Vekil, et al.
Pubblicazione: (2026)
di: Bekmyradov, Vekil, et al.
Pubblicazione: (2026)
SLEAN: Simple Lightweight Ensemble Analysis Network for Multi-Provider LLM Coordination: Design, Implementation, and Vibe Coding Bug Investigation Case Study
di: Vargas, Matheus J. T.
Pubblicazione: (2025)
di: Vargas, Matheus J. T.
Pubblicazione: (2025)
VulScribeR: Exploring RAG-based Vulnerability Augmentation with LLMs
di: Daneshvar, Seyed Shayan, et al.
Pubblicazione: (2024)
di: Daneshvar, Seyed Shayan, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Assessing Data Augmentation-Induced Bias in Training and Testing of Machine Learning Models
di: More, Riddhi, et al.
Pubblicazione: (2025) -
Addressing Data Leakage in HumanEval Using Combinatorial Test Design
di: Bradbury, Jeremy S., et al.
Pubblicazione: (2024) -
From Untestable to Testable: Metamorphic Testing in the Age of LLMs
di: Terragni, Valerio
Pubblicazione: (2026) -
A Systematic Approach for Assessing Large Language Models' Test Case Generation Capability
di: Chang, Hung-Fu, et al.
Pubblicazione: (2025) -
Towards a Probabilistic Framework for Analyzing and Improving LLM-Enabled Software
di: Baldonado, Juan Manuel, et al.
Pubblicazione: (2025)