When LLMs Lag Behind: Knowledge Conflicts from Evolving APIs in Code Generation
Fuente:
arXiv
Guardado en:
| Autores principales: | Ashik, Ahmed Nusayer, Wang, Shaowei, Chen, Tse-Hsun, Asaduzzaman, Muhammad, Tian, Yuan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
ZS4C: Zero-Shot Synthesis of Compilable Code for Incomplete Code Snippets using LLMs
por: Kabir, Azmain, et al.
Publicado: (2024)
por: Kabir, Azmain, et al.
Publicado: (2024)
Studying the Impact of Early Test Termination Due to Assertion Failure on Code Coverage and Spectrum-based Fault Localization
por: Uddin, Md. Ashraf, et al.
Publicado: (2025)
por: Uddin, Md. Ashraf, et al.
Publicado: (2025)
CODEPROMPTZIP: Code-specific Prompt Compression for Retrieval-Augmented Generation in Coding Tasks with LMs
por: He, Pengfei, et al.
Publicado: (2025)
por: He, Pengfei, et al.
Publicado: (2025)
Static Program Slicing Using Language Models With Dataflow-Aware Pretraining and Constrained Decoding
por: He, Pengfei, et al.
Publicado: (2026)
por: He, Pengfei, et al.
Publicado: (2026)
Evaluating the Effectiveness and Efficiency of Demonstration Retrievers in RAG for Coding Tasks
por: He, Pengfei, et al.
Publicado: (2024)
por: He, Pengfei, et al.
Publicado: (2024)
SLICET5: Static Program Slicing using Language Models with Copy Mechanism and Constrained Decoding
por: He, Pengfei, et al.
Publicado: (2025)
por: He, Pengfei, et al.
Publicado: (2025)
A Survey of Code Review Benchmarks and Evaluation Practices in Pre-LLM and LLM Era
por: Khan, Taufiqul Islam, et al.
Publicado: (2026)
por: Khan, Taufiqul Islam, et al.
Publicado: (2026)
A Comprehensive Framework for Evaluating API-oriented Code Generation in Large Language Models
por: Wu, Yixi, et al.
Publicado: (2024)
por: Wu, Yixi, et al.
Publicado: (2024)
GUIWatcher: Automatically Detecting GUI Lags by Analyzing Mobile Application Screencasts
por: Liu, Wei, et al.
Publicado: (2025)
por: Liu, Wei, et al.
Publicado: (2025)
Typify: A Lightweight Usage-driven Static Analyzer for Precise Python Type Inference
por: Aman, Ali, et al.
Publicado: (2026)
por: Aman, Ali, et al.
Publicado: (2026)
Towards Better Graph Neural Network-based Fault Localization Through Enhanced Code Representation
por: Rafi, Md Nakhla, et al.
Publicado: (2024)
por: Rafi, Md Nakhla, et al.
Publicado: (2024)
Think Harder and Don't Overlook Your Options: Revisiting Issue-Commit Linking with LLM-Assisted Retrieval
por: Morgan, Cole, et al.
Publicado: (2026)
por: Morgan, Cole, et al.
Publicado: (2026)
Studying and Recommending Information Highlighting in Stack Overflow Answers
por: Ahmed, Shahla Shaan, et al.
Publicado: (2024)
por: Ahmed, Shahla Shaan, et al.
Publicado: (2024)
A Multi-Agent Approach to Fault Localization via Graph-Based Retrieval and Reflexion
por: Rafi, Md Nakhla, et al.
Publicado: (2024)
por: Rafi, Md Nakhla, et al.
Publicado: (2024)
SWE-Refactor: A Repository-Level Benchmark for Real-World LLM-Based Code Refactoring
por: Xu, Yisen, et al.
Publicado: (2026)
por: Xu, Yisen, et al.
Publicado: (2026)
RAG-Reflect: Agentic Retrieval-Augmented Generation with Reflections for Comment-Driven Code Maintenance on Stack Overflow
por: Shanto, Mehedi Hasan, et al.
Publicado: (2026)
por: Shanto, Mehedi Hasan, et al.
Publicado: (2026)
Towards Structured, State-Aware, and Execution-Grounded Reasoning for Software Engineering Agents
por: Tse-Hsun, et al.
Publicado: (2026)
por: Tse-Hsun, et al.
Publicado: (2026)
Evaluating Software Process Models for Multi-Agent Class-Level Code Generation
por: Shafin, Wasique Islam, et al.
Publicado: (2025)
por: Shafin, Wasique Islam, et al.
Publicado: (2025)
Evolving Triple Knowledge-Augmented LLMs for Code Translation in Repository Context
por: Ou, Guangsheng, et al.
Publicado: (2025)
por: Ou, Guangsheng, et al.
Publicado: (2025)
Empowering AIOps: Leveraging Large Language Models for IT Operations Management
por: Vitui, Arthur, et al.
Publicado: (2025)
por: Vitui, Arthur, et al.
Publicado: (2025)
On Rank Aggregating Test Prioritizations
por: Mondal, Shouvick, et al.
Publicado: (2024)
por: Mondal, Shouvick, et al.
Publicado: (2024)
LLMParser: An Exploratory Study on Using Large Language Models for Log Parsing
por: Ma, Zeyang, et al.
Publicado: (2024)
por: Ma, Zeyang, et al.
Publicado: (2024)
Identifying Performance-Sensitive Configurations in Software Systems through Code Analysis with LLM Agents
por: Wang, Zehao, et al.
Publicado: (2024)
por: Wang, Zehao, et al.
Publicado: (2024)
Order Matters! An Empirical Study on Large Language Models' Input Order Bias in Software Fault Localization
por: Rafi, Md Nakhla, et al.
Publicado: (2024)
por: Rafi, Md Nakhla, et al.
Publicado: (2024)
Code2API: A Tool for Generating Reusable APIs from Stack Overflow Code Snippets
por: Mai, Yubo, et al.
Publicado: (2025)
por: Mai, Yubo, et al.
Publicado: (2025)
RobuNFR: Evaluating the Robustness of Large Language Models on Non-Functional Requirements Aware Code Generation
por: Lin, Feng, et al.
Publicado: (2025)
por: Lin, Feng, et al.
Publicado: (2025)
Back to the Future! Studying Data Cleanness in Defects4J and its Impact on Fault Localization
por: Rafi, Md Nakhla, et al.
Publicado: (2023)
por: Rafi, Md Nakhla, et al.
Publicado: (2023)
An Empirical Study on the Characteristics of Database Access Bugs in Java Applications
por: Liu, Wei, et al.
Publicado: (2024)
por: Liu, Wei, et al.
Publicado: (2024)
When to Stop? Towards Efficient Code Generation in LLMs with Excess Token Prevention
por: Guo, Lianghong, et al.
Publicado: (2024)
por: Guo, Lianghong, et al.
Publicado: (2024)
HistoryFinder: Advancing Method-Level Source Code History Generation with Accurate Oracles and Enhanced Algorithm
por: Islam, Shahidul, et al.
Publicado: (2025)
por: Islam, Shahidul, et al.
Publicado: (2025)
Evidence is All We Need: Do Self-Admitted Technical Debts Impact Method-Level Maintenance?
por: Chowdhury, Shaiful, et al.
Publicado: (2024)
por: Chowdhury, Shaiful, et al.
Publicado: (2024)
CI-Repair-Bench: A Repository-Aware Benchmark for Automated Patch Validation via CI Workflows
por: Muna, Rabeya Khatun, et al.
Publicado: (2026)
por: Muna, Rabeya Khatun, et al.
Publicado: (2026)
MobileUPReg: Identifying User-Perceived Performance Regressions in Mobile OS Versions
por: Liu, Wei, et al.
Publicado: (2025)
por: Liu, Wei, et al.
Publicado: (2025)
PrediQL: Automated Testing of GraphQL APIs with LLMs
por: Liu, Shaolun, et al.
Publicado: (2025)
por: Liu, Shaolun, et al.
Publicado: (2025)
Are Benchmark Tests Strong Enough? Mutation-Guided Diagnosis and Augmentation of Regression Suites
por: Li, Chenglin, et al.
Publicado: (2026)
por: Li, Chenglin, et al.
Publicado: (2026)
From Historical Patches to Repair Plans: Outcome-Conditioned Reasoning for Repository-Level Program Repair
por: Li, Chenglin, et al.
Publicado: (2026)
por: Li, Chenglin, et al.
Publicado: (2026)
PopSweeper: Automatically Detecting and Resolving App-Blocking Pop-Ups to Assist Automated Mobile GUI Testing
por: Guo, Linqiang, et al.
Publicado: (2024)
por: Guo, Linqiang, et al.
Publicado: (2024)
SATORI: Static Test Oracle Generation for REST APIs
por: Alonso, Juan C., et al.
Publicado: (2025)
por: Alonso, Juan C., et al.
Publicado: (2025)
EMF-REST: Generation of RESTful APIs from Models
por: Ed-Douibi, Hamza, et al.
Publicado: (2015)
por: Ed-Douibi, Hamza, et al.
Publicado: (2015)
Screencast-Based Analysis of User-Perceived GUI Responsiveness
por: Liu, Wei, et al.
Publicado: (2025)
por: Liu, Wei, et al.
Publicado: (2025)
Ejemplares similares
-
ZS4C: Zero-Shot Synthesis of Compilable Code for Incomplete Code Snippets using LLMs
por: Kabir, Azmain, et al.
Publicado: (2024) -
Studying the Impact of Early Test Termination Due to Assertion Failure on Code Coverage and Spectrum-based Fault Localization
por: Uddin, Md. Ashraf, et al.
Publicado: (2025) -
CODEPROMPTZIP: Code-specific Prompt Compression for Retrieval-Augmented Generation in Coding Tasks with LMs
por: He, Pengfei, et al.
Publicado: (2025) -
Static Program Slicing Using Language Models With Dataflow-Aware Pretraining and Constrained Decoding
por: He, Pengfei, et al.
Publicado: (2026) -
Evaluating the Effectiveness and Efficiency of Demonstration Retrievers in RAG for Coding Tasks
por: He, Pengfei, et al.
Publicado: (2024)