The Foundation Cracks: A Comprehensive Study on Bugs and Testing Practices in LLM Libraries
Fuente:
arXiv
Saved in:
| Main Authors: | Jiang, Weipeng, Zhang, Xiaoyu, Xie, Xiaofei, Yu, Jiongchi, Zhi, Yuhan, Ma, Shiqing, Shen, Chao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AutoEmpirical: LLM-Based Automated Research for Empirical Software Fault Analysis
by: Yu, Jiongchi, et al.
Published: (2025)
by: Yu, Jiongchi, et al.
Published: (2025)
Chimera: Harnessing Multi-Agent LLMs for Automatic Insider Threat Simulation
by: Yu, Jiongchi, et al.
Published: (2025)
by: Yu, Jiongchi, et al.
Published: (2025)
Efficient DNN-Powered Software with Fair Sparse Models
by: Gao, Xuanqi, et al.
Published: (2024)
by: Gao, Xuanqi, et al.
Published: (2024)
Weaponizing the Commons: A Taxonomy and Detection Framework of Abuse on GitHub
by: Cheng, Yuli, et al.
Published: (2026)
by: Cheng, Yuli, et al.
Published: (2026)
False Friends in the Shell: Unveiling the Emoticon Semantic Confusion in Large Language Models
by: Jiang, Weipeng, et al.
Published: (2026)
by: Jiang, Weipeng, et al.
Published: (2026)
Deep Learning Library Testing: Definition, Methods and Challenges
by: Zhang, Xiaoyu, et al.
Published: (2024)
by: Zhang, Xiaoyu, et al.
Published: (2024)
Leveraging LLM Agents for Automated Video Game Testing
by: Wang, Chengjia, et al.
Published: (2025)
by: Wang, Chengjia, et al.
Published: (2025)
CITADEL: Context Similarity Based Deep Learning Framework Bug Finding
by: Zhang, Xiaoyu, et al.
Published: (2024)
by: Zhang, Xiaoyu, et al.
Published: (2024)
ASSURE: Metamorphic Testing for AI-powered Browser Extensions
by: Gao, Xuanqi, et al.
Published: (2025)
by: Gao, Xuanqi, et al.
Published: (2025)
DREAM: Debugging and Repairing AutoML Pipelines
by: Zhang, Xiaoyu, et al.
Published: (2023)
by: Zhang, Xiaoyu, et al.
Published: (2023)
Defects4C: Benchmarking Large Language Model Repair Capability with C/C++ Bugs
by: Wang, Jian, et al.
Published: (2025)
by: Wang, Jian, et al.
Published: (2025)
An Empirical Study on LLM-based Agents for Automated Bug Fixing
by: Meng, Xiangxin, et al.
Published: (2024)
by: Meng, Xiangxin, et al.
Published: (2024)
The Invisible Hand: Unveiling Provider Bias in Large Language Models for Code Generation
by: Zhang, Xiaoyu, et al.
Published: (2025)
by: Zhang, Xiaoyu, et al.
Published: (2025)
An Empirical Study of Bugs in Modern LLM Agent Frameworks
by: Zhu, Xinxue, et al.
Published: (2026)
by: Zhu, Xinxue, et al.
Published: (2026)
Beyond Accuracy: Policy Invariance as a Reliability Test for LLM Safety Judges
by: Weng, Shihao, et al.
Published: (2026)
by: Weng, Shihao, et al.
Published: (2026)
From Effectiveness to Efficiency: Uncovering Linguistic Bias in Large Language Model-based Code Generation
by: Jiang, Weipeng, et al.
Published: (2024)
by: Jiang, Weipeng, et al.
Published: (2024)
ARGUS: Defending LLM Agents Against Context-Aware Prompt Injection
by: Weng, Shihao, et al.
Published: (2026)
by: Weng, Shihao, et al.
Published: (2026)
Go-Oracle: Automated Test Oracle for Go Concurrency Bugs
by: Tsimpourlas, Foivos, et al.
Published: (2024)
by: Tsimpourlas, Foivos, et al.
Published: (2024)
Human in the Loop for Fuzz Testing: Literature Review and the Road Ahead
by: Yu, Jiongchi, et al.
Published: (2026)
by: Yu, Jiongchi, et al.
Published: (2026)
Rethinking Technology Stack Selection with AI Coding Proficiency
by: Zhang, Xiaoyu, et al.
Published: (2025)
by: Zhang, Xiaoyu, et al.
Published: (2025)
Bias Testing and Mitigation in LLM-based Code Generation
by: Huang, Dong, et al.
Published: (2023)
by: Huang, Dong, et al.
Published: (2023)
Faster Configuration Performance Bug Testing with Neural Dual-level Prioritization
by: Ma, Youpeng, et al.
Published: (2025)
by: Ma, Youpeng, et al.
Published: (2025)
Benchmarking and Revisiting Code Generation Assessment: A Mutation-Based Approach
by: Wang, Longtian, et al.
Published: (2025)
by: Wang, Longtian, et al.
Published: (2025)
HLSDebugger: Identification and Correction of Logic Bugs in HLS Code with LLM Solutions
by: Wang, Jing, et al.
Published: (2025)
by: Wang, Jing, et al.
Published: (2025)
Can We Enhance Bug Report Quality Using LLMs?: An Empirical Study of LLM-Based Bug Report Generation
by: Acharya, Jagrit, et al.
Published: (2025)
by: Acharya, Jagrit, et al.
Published: (2025)
Rethinking the Value of Agent-Generated Tests for LLM-Based Software Engineering Agents
by: Chen, Zhi, et al.
Published: (2026)
by: Chen, Zhi, et al.
Published: (2026)
DriveTester: A Unified Platform for Simulation-Based Autonomous Driving Testing
by: Cheng, Mingfei, et al.
Published: (2024)
by: Cheng, Mingfei, et al.
Published: (2024)
Do Code Semantics Help? A Comprehensive Study on Execution Trace-Based Information for Code Large Language Models
by: Wang, Jian, et al.
Published: (2025)
by: Wang, Jian, et al.
Published: (2025)
STCLocker: Deadlock Avoidance Testing for Autonomous Driving Systems
by: Cheng, Mingfei, et al.
Published: (2025)
by: Cheng, Mingfei, et al.
Published: (2025)
Integrating Various Software Artifacts for Better LLM-based Bug Localization and Program Repair
by: Feng, Qiong, et al.
Published: (2024)
by: Feng, Qiong, et al.
Published: (2024)
Understanding the Supply Chain and Risks of Large Language Model Applications
by: Ma, Yujie, et al.
Published: (2025)
by: Ma, Yujie, et al.
Published: (2025)
Dynamic Cogeneration of Bug Reproduction Test in Agentic Program Repair
by: Cheng, Runxiang, et al.
Published: (2026)
by: Cheng, Runxiang, et al.
Published: (2026)
RFCAudit: An LLM Agent for Functional Bug Detection in Network Protocols
by: Zheng, Mingwei, et al.
Published: (2025)
by: Zheng, Mingwei, et al.
Published: (2025)
Prompt Stability in Code LLMs: Measuring Sensitivity across Emotion- and Personality-Driven Variations
by: Ma, Wei, et al.
Published: (2025)
by: Ma, Wei, et al.
Published: (2025)
A Comprehensive Study on Static Application Security Testing (SAST) Tools for Android
by: Zhu, Jingyun, et al.
Published: (2024)
by: Zhu, Jingyun, et al.
Published: (2024)
Agentic Property-Based Testing: Finding Bugs Across the Python Ecosystem
by: Maaz, Muhammad, et al.
Published: (2025)
by: Maaz, Muhammad, et al.
Published: (2025)
When Agents Fail: A Comprehensive Study of Bugs in LLM Agents with Automated Labeling
by: Islam, Niful, et al.
Published: (2026)
by: Islam, Niful, et al.
Published: (2026)
A Comprehensive Study of Governance Issues in Decentralized Finance Applications
by: Ma, Wei, et al.
Published: (2023)
by: Ma, Wei, et al.
Published: (2023)
Bug Analysis Towards Bug Resolution Time Prediction
by: Ozkan, Hasan Yagiz, et al.
Published: (2024)
by: Ozkan, Hasan Yagiz, et al.
Published: (2024)
Empirical Analysis and Detection of Hallucinations in LLM-Generated Bug Report Summaries
by: Nirujan, Hinduja, et al.
Published: (2026)
by: Nirujan, Hinduja, et al.
Published: (2026)
Similar Items
-
AutoEmpirical: LLM-Based Automated Research for Empirical Software Fault Analysis
by: Yu, Jiongchi, et al.
Published: (2025) -
Chimera: Harnessing Multi-Agent LLMs for Automatic Insider Threat Simulation
by: Yu, Jiongchi, et al.
Published: (2025) -
Efficient DNN-Powered Software with Fair Sparse Models
by: Gao, Xuanqi, et al.
Published: (2024) -
Weaponizing the Commons: A Taxonomy and Detection Framework of Abuse on GitHub
by: Cheng, Yuli, et al.
Published: (2026) -
False Friends in the Shell: Unveiling the Emoticon Semantic Confusion in Large Language Models
by: Jiang, Weipeng, et al.
Published: (2026)