LLM Critics Help Catch LLM Bugs
Fuente:
arXiv
Guardado en:
| Autores principales: | McAleese, Nat, Pokorny, Rai Michael, Uribe, Juan Felipe Ceron, Nitishinskaya, Evgenia, Trebacz, Maja, Leike, Jan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
BugGen: A Self-Correcting Multi-Agent LLM Pipeline for Realistic RTL Bug Synthesis
por: Jasper, Surya, et al.
Publicado: (2025)
por: Jasper, Surya, et al.
Publicado: (2025)
LLM-Powered Test Case Generation for Detecting Bugs in Plausible Programs
por: Liu, Kaibo, et al.
Publicado: (2024)
por: Liu, Kaibo, et al.
Publicado: (2024)
An Empirical Study of Bugs in Modern LLM Agent Frameworks
por: Zhu, Xinxue, et al.
Publicado: (2026)
por: Zhu, Xinxue, et al.
Publicado: (2026)
Prover-Verifier Games improve legibility of LLM outputs
por: Kirchner, Jan Hendrik, et al.
Publicado: (2024)
por: Kirchner, Jan Hendrik, et al.
Publicado: (2024)
Challenging Bug Prediction and Repair Models with Synthetic Bugs
por: Ibrahimzada, Ali Reza, et al.
Publicado: (2023)
por: Ibrahimzada, Ali Reza, et al.
Publicado: (2023)
A First Look at Bugs in LLM Inference Engines
por: Liu, Mugeng, et al.
Publicado: (2025)
por: Liu, Mugeng, et al.
Publicado: (2025)
scicode-lint: Detecting Methodology Bugs in Scientific Python Code with LLM-Generated Patterns
por: Samsonau, Sergey V.
Publicado: (2026)
por: Samsonau, Sergey V.
Publicado: (2026)
LLM-Powered Silent Bug Fuzzing in Deep Learning Libraries via Versatile and Controlled Bug Transfer
por: Zhang, Kunpeng, et al.
Publicado: (2026)
por: Zhang, Kunpeng, et al.
Publicado: (2026)
HotBugs.jar: A Benchmark of Hot Fixes for Time-Critical Bugs
por: Hanna, Carol, et al.
Publicado: (2025)
por: Hanna, Carol, et al.
Publicado: (2025)
Testing Refactoring Engine via Historical Bug Report driven LLM
por: Wang, Haibo, et al.
Publicado: (2025)
por: Wang, Haibo, et al.
Publicado: (2025)
Qualitative Evaluation of LLM-Designed GUI
por: Sawicki, Bartosz, et al.
Publicado: (2026)
por: Sawicki, Bartosz, et al.
Publicado: (2026)
AgentSZZ: Teaching the LLM Agent to Play Detective with Bug-Inducing Commits
por: Lyu, Yunbo, et al.
Publicado: (2026)
por: Lyu, Yunbo, et al.
Publicado: (2026)
SelfHeal: Empirical Fix Pattern Analysis and Bug Repair in LLM Agents
por: Islam, Niful, et al.
Publicado: (2026)
por: Islam, Niful, et al.
Publicado: (2026)
AssertFlip: Reproducing Bugs via Inversion of LLM-Generated Passing Tests
por: Khatib, Lara, et al.
Publicado: (2025)
por: Khatib, Lara, et al.
Publicado: (2025)
When Agents Fail: A Comprehensive Study of Bugs in LLM Agents with Automated Labeling
por: Islam, Niful, et al.
Publicado: (2026)
por: Islam, Niful, et al.
Publicado: (2026)
Where's the Bug? Attention Probing for Scalable Fault Localization
por: Stein, Adam, et al.
Publicado: (2025)
por: Stein, Adam, et al.
Publicado: (2025)
The Limits of Long-Context Reasoning in Automated Bug Fixing
por: Raju, Ravi, et al.
Publicado: (2026)
por: Raju, Ravi, et al.
Publicado: (2026)
Hints Help Finding and Fixing Bugs Differently in Python and Text-based Program Representations
por: Rawal, Ruchit, et al.
Publicado: (2024)
por: Rawal, Ruchit, et al.
Publicado: (2024)
Fuzzing BusyBox: Leveraging LLM and Crash Reuse for Embedded Bug Unearthing
por: Asmita, et al.
Publicado: (2024)
por: Asmita, et al.
Publicado: (2024)
PROMFUZZ: Leveraging LLM-Driven and Bug-Oriented Composite Analysis for Detecting Functional Bugs in Smart Contracts
por: Lin, Xingshuang, et al.
Publicado: (2025)
por: Lin, Xingshuang, et al.
Publicado: (2025)
Towards Enhancing the Reproducibility of Deep Learning Bugs: An Empirical Study
por: Shah, Mehil B., et al.
Publicado: (2024)
por: Shah, Mehil B., et al.
Publicado: (2024)
Minimizing False Positives in Static Bug Detection via LLM-Enhanced Path Feasibility Analysis
por: Du, Xueying, et al.
Publicado: (2025)
por: Du, Xueying, et al.
Publicado: (2025)
Empirical Research on Utilizing LLM-based Agents for Automated Bug Fixing via LangGraph
por: Wang, Jialin, et al.
Publicado: (2025)
por: Wang, Jialin, et al.
Publicado: (2025)
Enhancing LLM-Based Bug Reproduction for Android Apps via Pre-Assessment of Visual Effects
por: Xiao, Xiangyang, et al.
Publicado: (2026)
por: Xiao, Xiangyang, et al.
Publicado: (2026)
LLM-Based Detection of Tangled Code Changes for Higher-Quality Method-Level Bug Datasets
por: Opu, Md Nahidul Islam, et al.
Publicado: (2025)
por: Opu, Md Nahidul Islam, et al.
Publicado: (2025)
Does Few-Shot Learning Help LLM Performance in Code Synthesis?
por: Xu, Derek, et al.
Publicado: (2024)
por: Xu, Derek, et al.
Publicado: (2024)
An Empirical Study on LLM-based Agents for Automated Bug Fixing
por: Meng, Xiangxin, et al.
Publicado: (2024)
por: Meng, Xiangxin, et al.
Publicado: (2024)
RFCAudit: An LLM Agent for Functional Bug Detection in Network Protocols
por: Zheng, Mingwei, et al.
Publicado: (2025)
por: Zheng, Mingwei, et al.
Publicado: (2025)
Language Models are Better Bug Detector Through Code-Pair Classification
por: Alrashedy, Kamel, et al.
Publicado: (2023)
por: Alrashedy, Kamel, et al.
Publicado: (2023)
Mut4All: Fuzzing Compilers via LLM-Synthesized Mutators Learned from Bug Reports
por: Wang, Bo, et al.
Publicado: (2025)
por: Wang, Bo, et al.
Publicado: (2025)
Bug Severity Prediction in Software Projects Using Supervised Machine Learning Models
por: Nice, Nafisha Tamanna
Publicado: (2026)
por: Nice, Nafisha Tamanna
Publicado: (2026)
Combining Language and App UI Analysis for the Automated Assessment of Bug Reproduction Steps
por: Mahmud, Junayed, et al.
Publicado: (2025)
por: Mahmud, Junayed, et al.
Publicado: (2025)
When Bugs Linger: A Study of Anomalous Resolution Time Outliers and Their Themes
por: Patil, Avinash
Publicado: (2025)
por: Patil, Avinash
Publicado: (2025)
LLM-Based Design Pattern Detection
por: Schindler, Christian, et al.
Publicado: (2025)
por: Schindler, Christian, et al.
Publicado: (2025)
ReadMe.LLM: A Framework to Help LLMs Understand Your Library
por: Wijaya, Sandya, et al.
Publicado: (2025)
por: Wijaya, Sandya, et al.
Publicado: (2025)
Agora: Toward Autonomous Bug Detection in Production-Level Consensus Protocols with LLM Agents
por: Liu, Xiang, et al.
Publicado: (2026)
por: Liu, Xiang, et al.
Publicado: (2026)
HLSDebugger: Identification and Correction of Logic Bugs in HLS Code with LLM Solutions
por: Wang, Jing, et al.
Publicado: (2025)
por: Wang, Jing, et al.
Publicado: (2025)
Empirical Analysis and Detection of Hallucinations in LLM-Generated Bug Report Summaries
por: Nirujan, Hinduja, et al.
Publicado: (2026)
por: Nirujan, Hinduja, et al.
Publicado: (2026)
Can We Enhance Bug Report Quality Using LLMs?: An Empirical Study of LLM-Based Bug Report Generation
por: Acharya, Jagrit, et al.
Publicado: (2025)
por: Acharya, Jagrit, et al.
Publicado: (2025)
TreeMind: Automatically Reproducing Android Bug Reports via LLM-empowered Monte Carlo Tree Search
por: Chen, Zhengyu, et al.
Publicado: (2025)
por: Chen, Zhengyu, et al.
Publicado: (2025)
Ejemplares similares
-
BugGen: A Self-Correcting Multi-Agent LLM Pipeline for Realistic RTL Bug Synthesis
por: Jasper, Surya, et al.
Publicado: (2025) -
LLM-Powered Test Case Generation for Detecting Bugs in Plausible Programs
por: Liu, Kaibo, et al.
Publicado: (2024) -
An Empirical Study of Bugs in Modern LLM Agent Frameworks
por: Zhu, Xinxue, et al.
Publicado: (2026) -
Prover-Verifier Games improve legibility of LLM outputs
por: Kirchner, Jan Hendrik, et al.
Publicado: (2024) -
Challenging Bug Prediction and Repair Models with Synthetic Bugs
por: Ibrahimzada, Ali Reza, et al.
Publicado: (2023)