Library Drift: Diagnosing and Fixing a Silent Failure Mode in Self-Evolving LLM Skill Libraries
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Xing, Cui, Yanwei, Wang, Guanghui, Li, Ziyuan, Qiu, Wei, Zhu, Bing, He, Peiyang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Ratchet: A Minimal Hygiene Recipe for Self-Evolving LLM Agents
by: Zhang, Xing, et al.
Published: (2026)
by: Zhang, Xing, et al.
Published: (2026)
Experience Compression Spectrum: Unifying Memory, Skills, and Rules in LLM Agents
by: Zhang, Xing, et al.
Published: (2026)
by: Zhang, Xing, et al.
Published: (2026)
Prompt Optimization Is a Coin Flip: Diagnosing When It Helps in Compound AI Systems
by: Zhang, Xing, et al.
Published: (2026)
by: Zhang, Xing, et al.
Published: (2026)
Guardrails Beat Guidance: A Large-Scale Study of Rules, Skills, and Persistent Configuration for Coding Agents
by: Zhang, Xing, et al.
Published: (2026)
by: Zhang, Xing, et al.
Published: (2026)
Skill Drift Is Contract Violation: Proactive Maintenance for LLM Agent Skill Libraries
by: Fan, Linfeng, et al.
Published: (2026)
by: Fan, Linfeng, et al.
Published: (2026)
SkillCraft: Can LLM Agents Learn to Use Tools Skillfully?
by: Chen, Shiqi, et al.
Published: (2026)
by: Chen, Shiqi, et al.
Published: (2026)
Drift-Bench: Diagnosing Cooperative Breakdowns in LLM Agents under Input Faults via Multi-Turn Interaction
by: Bao, Han, et al.
Published: (2026)
by: Bao, Han, et al.
Published: (2026)
Functional Consistency of LLM Code Embeddings: A Self-Evolving Data Synthesis Framework for Benchmarking
by: Li, Zhuohao, et al.
Published: (2025)
by: Li, Zhuohao, et al.
Published: (2025)
ProbeLLM: Automating Principled Diagnosis of LLM Failures
by: Huang, Yue, et al.
Published: (2026)
by: Huang, Yue, et al.
Published: (2026)
VFDelta: A Framework for Detecting Silent Vulnerability Fixes by Enhancing Code Change Learning
by: Yang, Xu, et al.
Published: (2024)
by: Yang, Xu, et al.
Published: (2024)
MultiFileTest: A Multi-File-Level LLM Unit Test Generation Benchmark and Impact of Error Fixing Mechanisms
by: Wang, Yibo, et al.
Published: (2025)
by: Wang, Yibo, et al.
Published: (2025)
Code Fingerprints: Disentangled Attribution of LLM-Generated Code
by: Guo, Jiaxun, et al.
Published: (2026)
by: Guo, Jiaxun, et al.
Published: (2026)
EPSO: A Caching-Based Efficient Superoptimizer for BPF Bytecode
by: Zhu, Qian, et al.
Published: (2025)
by: Zhu, Qian, et al.
Published: (2025)
COBOLAssist: Analyzing and Fixing Compilation Errors for LLM-Powered COBOL Code Generation
by: Dau, Anh T. V., et al.
Published: (2026)
by: Dau, Anh T. V., et al.
Published: (2026)
(Why) Is My Prompt Getting Worse? Rethinking Regression Testing for Evolving LLM APIs
by: Ma, Wanqin, et al.
Published: (2023)
by: Ma, Wanqin, et al.
Published: (2023)
MEMCoder: Multi-dimensional Evolving Memory for Private-Library-Oriented Code Generation
by: Li, Mofei, et al.
Published: (2026)
by: Li, Mofei, et al.
Published: (2026)
LLM-Powered Silent Bug Fuzzing in Deep Learning Libraries via Versatile and Controlled Bug Transfer
by: Zhang, Kunpeng, et al.
Published: (2026)
by: Zhang, Kunpeng, et al.
Published: (2026)
EvoConfig: Self-Evolving Multi-Agent Systems for Efficient Autonomous Environment Configuration
by: Guo, Xinshuai, et al.
Published: (2026)
by: Guo, Xinshuai, et al.
Published: (2026)
The Alignment Floor: How Persona Customization Breaks Safety in Weakly-Aligned LLMs
by: Zhang, Xing, et al.
Published: (2026)
by: Zhang, Xing, et al.
Published: (2026)
Evidence Absence Is Not Evidence Insufficiency: Diagnosing NEI Construction Artifacts in Fact Verification
by: Qiu, Jingxi, et al.
Published: (2026)
by: Qiu, Jingxi, et al.
Published: (2026)
Fast and Accurate Silent Vulnerability Fix Retrieval
by: Liu, Xueqing, et al.
Published: (2025)
by: Liu, Xueqing, et al.
Published: (2025)
RUM: Rule+LLM-Based Comprehensive Assessment on Testing Skills
by: Wang, Yue, et al.
Published: (2025)
by: Wang, Yue, et al.
Published: (2025)
Compositional API Recommendation for Library-Oriented Code Generation
by: Ma, Zexiong, et al.
Published: (2024)
by: Ma, Zexiong, et al.
Published: (2024)
SEVerA: Verified Synthesis of Self-Evolving Agents
by: Banerjee, Debangshu, et al.
Published: (2026)
by: Banerjee, Debangshu, et al.
Published: (2026)
ContraFix: Agentic Vulnerability Repair via Differential Runtime Evidence and Skill Reuse
by: Liu, Simiao, et al.
Published: (2026)
by: Liu, Simiao, et al.
Published: (2026)
Learning Graph-based Patch Representations for Identifying and Assessing Silent Vulnerability Fixes
by: Han, Mei, et al.
Published: (2024)
by: Han, Mei, et al.
Published: (2024)
Finding Compiler Bugs through Cross-Language Code Generator and Differential Testing
by: Feng, Qiong, et al.
Published: (2025)
by: Feng, Qiong, et al.
Published: (2025)
SEW: Self-Evolving Agentic Workflows for Automated Code Generation
by: Liu, Siwei, et al.
Published: (2025)
by: Liu, Siwei, et al.
Published: (2025)
EffiSkill: Agent Skill Based Automated Code Efficiency Optimization
by: Wang, Zimu, et al.
Published: (2026)
by: Wang, Zimu, et al.
Published: (2026)
On the Illusion of Success: An Empirical Study of Build Reruns and Silent Failures in Industrial CI
by: Aïdasso, Henri, et al.
Published: (2025)
by: Aïdasso, Henri, et al.
Published: (2025)
EigenData: A Self-Evolving Multi-Agent Platform for Function-Calling Data Synthesis, Auditing, and Repair
by: Chen, Jiaao, et al.
Published: (2026)
by: Chen, Jiaao, et al.
Published: (2026)
Automatically Benchmarking LLM Code Agents through Agent-Driven Annotation and Evaluation
by: Fu, Lingyue, et al.
Published: (2025)
by: Fu, Lingyue, et al.
Published: (2025)
BanglaForge: LLM Collaboration with Self-Refinement for Bangla Code Generation
by: Dihan, Mahir Labib, et al.
Published: (2025)
by: Dihan, Mahir Labib, et al.
Published: (2025)
ContextEcho: A Benchmark for Persona Drift in Long Agentic-Coding Sessions
by: Ding, Xianzhong, et al.
Published: (2026)
by: Ding, Xianzhong, et al.
Published: (2026)
AutoBench: Automatic Testbench Generation and Evaluation Using LLMs for HDL Design
by: Qiu, Ruidi, et al.
Published: (2024)
by: Qiu, Ruidi, et al.
Published: (2024)
Inferring Non-Failure Conditions for Declarative Programs
by: Hanus, Michael
Published: (2024)
by: Hanus, Michael
Published: (2024)
Benchmarking Failures in Tool-Augmented Language Models
by: Treviño, Eduardo, et al.
Published: (2025)
by: Treviño, Eduardo, et al.
Published: (2025)
Validated Code Translation for Projects with External Libraries
by: Zhang, Hanliang, et al.
Published: (2026)
by: Zhang, Hanliang, et al.
Published: (2026)
Coffee: Boost Your Code LLMs by Fixing Bugs with Feedback
by: Moon, Seungjun, et al.
Published: (2023)
by: Moon, Seungjun, et al.
Published: (2023)
Reasoning Runtime Behavior of a Program with LLM: How Far Are We?
by: Chen, Junkai, et al.
Published: (2024)
by: Chen, Junkai, et al.
Published: (2024)
Similar Items
-
Ratchet: A Minimal Hygiene Recipe for Self-Evolving LLM Agents
by: Zhang, Xing, et al.
Published: (2026) -
Experience Compression Spectrum: Unifying Memory, Skills, and Rules in LLM Agents
by: Zhang, Xing, et al.
Published: (2026) -
Prompt Optimization Is a Coin Flip: Diagnosing When It Helps in Compound AI Systems
by: Zhang, Xing, et al.
Published: (2026) -
Guardrails Beat Guidance: A Large-Scale Study of Rules, Skills, and Persistent Configuration for Coding Agents
by: Zhang, Xing, et al.
Published: (2026) -
Skill Drift Is Contract Violation: Proactive Maintenance for LLM Agent Skill Libraries
by: Fan, Linfeng, et al.
Published: (2026)