Model See, Model Do? Exposure-Aware Evaluation of Bug-vs-Fix Preference in Code LLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | Al-Kaswan, Ali, Spiess, Claudio, Devanbu, Prem, van Deursen, Arie, Izadi, Maliheh |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Traces of Memorisation in Large Language Models for Code
di: Al-Kaswan, Ali, et al.
Pubblicazione: (2023)
di: Al-Kaswan, Ali, et al.
Pubblicazione: (2023)
Code Red! On the Harmfulness of Applying Off-the-shelf Large Language Models to Programming Tasks
di: Al-Kaswan, Ali, et al.
Pubblicazione: (2025)
di: Al-Kaswan, Ali, et al.
Pubblicazione: (2025)
Do Agents Dream of Root Shells? Partial-Credit Evaluation of LLM Agents in Capture the Flag Challenges
di: Al-Kaswan, Ali, et al.
Pubblicazione: (2026)
di: Al-Kaswan, Ali, et al.
Pubblicazione: (2026)
TreeRanker: Fast and Model-agnostic Ranking System for Code Suggestions in IDEs
di: Cipollone, Daniele, et al.
Pubblicazione: (2025)
di: Cipollone, Daniele, et al.
Pubblicazione: (2025)
An Exploratory Investigation into Code License Infringements in Large Language Model Training Datasets
di: Katzy, Jonathan, et al.
Pubblicazione: (2024)
di: Katzy, Jonathan, et al.
Pubblicazione: (2024)
A Transformer-Based Approach for Smart Invocation of Automatic Code Completion
di: de Moor, Aral, et al.
Pubblicazione: (2024)
di: de Moor, Aral, et al.
Pubblicazione: (2024)
Does In-IDE Calibration of Large Language Models work at Scale?
di: Koohestani, Roham, et al.
Pubblicazione: (2025)
di: Koohestani, Roham, et al.
Pubblicazione: (2025)
Language Models for Code Completion: A Practical Evaluation
di: Izadi, Maliheh, et al.
Pubblicazione: (2024)
di: Izadi, Maliheh, et al.
Pubblicazione: (2024)
AST-PAC: AST-guided Membership Inference for Code
di: Koohestani, Roham, et al.
Pubblicazione: (2026)
di: Koohestani, Roham, et al.
Pubblicazione: (2026)
On LLMs' Internal Representation of Code Correctness
di: Ribeiro, Francisco, et al.
Pubblicazione: (2025)
di: Ribeiro, Francisco, et al.
Pubblicazione: (2025)
How Robustly do LLMs Understand Execution Semantics?
di: Spiess, Claudio, et al.
Pubblicazione: (2026)
di: Spiess, Claudio, et al.
Pubblicazione: (2026)
Localized Calibrated Uncertainty in Code Language Models
di: Gros, David, et al.
Pubblicazione: (2025)
di: Gros, David, et al.
Pubblicazione: (2025)
How Much Do Code Language Models Remember? An Investigation on Data Extraction Attacks before and after Fine-tuning
di: Salerno, Fabio, et al.
Pubblicazione: (2025)
di: Salerno, Fabio, et al.
Pubblicazione: (2025)
A Qualitative Investigation into LLM-Generated Multilingual Code Comments and Automatic Evaluation Metrics
di: Katzy, Jonathan, et al.
Pubblicazione: (2025)
di: Katzy, Jonathan, et al.
Pubblicazione: (2025)
Investigating Autonomous Agent Contributions in the Wild: Activity Patterns and Code Change over Time
di: Popescu, Razvan Mihai, et al.
Pubblicazione: (2026)
di: Popescu, Razvan Mihai, et al.
Pubblicazione: (2026)
Evaluating Non-English Developer Support in Machine Learning for Software Engineering
di: Katzy, Jonathan, et al.
Pubblicazione: (2026)
di: Katzy, Jonathan, et al.
Pubblicazione: (2026)
Data vs. Model Machine Learning Fairness Testing: An Empirical Study
di: Shome, Arumoy, et al.
Pubblicazione: (2024)
di: Shome, Arumoy, et al.
Pubblicazione: (2024)
The Heap: A Contamination-Free Multilingual Code Dataset for Evaluating Large Language Models
di: Katzy, Jonathan, et al.
Pubblicazione: (2025)
di: Katzy, Jonathan, et al.
Pubblicazione: (2025)
Evaluating Large Language Models for Functional and Maintainable Code in Industrial Settings: A Case Study at ASML
di: Mundhra, Yash, et al.
Pubblicazione: (2025)
di: Mundhra, Yash, et al.
Pubblicazione: (2025)
Long Code Arena: a Set of Benchmarks for Long-Context Code Models
di: Bogomolov, Egor, et al.
Pubblicazione: (2024)
di: Bogomolov, Egor, et al.
Pubblicazione: (2024)
Calibration and Correctness of Language Models for Code
di: Spiess, Claudio, et al.
Pubblicazione: (2024)
di: Spiess, Claudio, et al.
Pubblicazione: (2024)
Towards Automatic Translation of Machine Learning Visual Insights to Analytical Assertions
di: Shome, Arumoy, et al.
Pubblicazione: (2024)
di: Shome, Arumoy, et al.
Pubblicazione: (2024)
Rethinking IDE Customization for Enhanced HAX: A Hyperdimensional Perspective
di: Koohestani, Roham, et al.
Pubblicazione: (2025)
di: Koohestani, Roham, et al.
Pubblicazione: (2025)
A Multi-agent Onboarding Assistant based on Large Language Models, Retrieval Augmented Generation, and Chain-of-Thought
di: Ionescu, Andrei Cristian, et al.
Pubblicazione: (2025)
di: Ionescu, Andrei Cristian, et al.
Pubblicazione: (2025)
HyperSeq: A Hyper-Adaptive Representation for Predictive Sequencing of States
di: Koohestani, Roham, et al.
Pubblicazione: (2025)
di: Koohestani, Roham, et al.
Pubblicazione: (2025)
In-IDE Human-AI Experience in the Era of Large Language Models; A Literature Review
di: Sergeyuk, Agnia, et al.
Pubblicazione: (2024)
di: Sergeyuk, Agnia, et al.
Pubblicazione: (2024)
Towards Understanding What Code Language Models Learned
di: Ahmed, Toufique, et al.
Pubblicazione: (2023)
di: Ahmed, Toufique, et al.
Pubblicazione: (2023)
McUDI: Model-Centric Unsupervised Degradation Indicator for Failure Prediction AIOps Solutions
di: Poenaru-Olaru, Lorena, et al.
Pubblicazione: (2024)
di: Poenaru-Olaru, Lorena, et al.
Pubblicazione: (2024)
Understanding Feedback Mechanisms in Machine Learning Jupyter Notebooks
di: Shome, Arumoy, et al.
Pubblicazione: (2024)
di: Shome, Arumoy, et al.
Pubblicazione: (2024)
Coffee: Boost Your Code LLMs by Fixing Bugs with Feedback
di: Moon, Seungjun, et al.
Pubblicazione: (2023)
di: Moon, Seungjun, et al.
Pubblicazione: (2023)
Leveraging Large Language Models for Enhancing the Understandability of Generated Unit Tests
di: Deljouyi, Amirhossein, et al.
Pubblicazione: (2024)
di: Deljouyi, Amirhossein, et al.
Pubblicazione: (2024)
Ecosystem of Large Language Models for Code
di: Yang, Zhou, et al.
Pubblicazione: (2024)
di: Yang, Zhou, et al.
Pubblicazione: (2024)
A Survey of Trojans in Neural Models of Source Code: Taxonomy and Techniques
di: Hussain, Aftab, et al.
Pubblicazione: (2023)
di: Hussain, Aftab, et al.
Pubblicazione: (2023)
Calibration of Large Language Models on Code Summarization
di: Virk, Yuvraj, et al.
Pubblicazione: (2024)
di: Virk, Yuvraj, et al.
Pubblicazione: (2024)
Automated Attention Pattern Discovery at Scale in Large Language Models
di: Katzy, Jonathan, et al.
Pubblicazione: (2026)
di: Katzy, Jonathan, et al.
Pubblicazione: (2026)
HAFix: History-Augmented Large Language Models for Bug Fixing
di: Shi, Yu, et al.
Pubblicazione: (2025)
di: Shi, Yu, et al.
Pubblicazione: (2025)
Benchmarking AI Models in Software Engineering: A Review, Search Tool, and Unified Approach for Elevating Benchmark Quality
di: Koohestani, Roham, et al.
Pubblicazione: (2025)
di: Koohestani, Roham, et al.
Pubblicazione: (2025)
Code4MeV2: a Research-oriented Code-completion Platform
di: Koohestani, Roham, et al.
Pubblicazione: (2025)
di: Koohestani, Roham, et al.
Pubblicazione: (2025)
Prompt-with-Me: in-IDE Structured Prompt Management for LLM-Driven Software Engineering
di: Li, Ziyou, et al.
Pubblicazione: (2025)
di: Li, Ziyou, et al.
Pubblicazione: (2025)
LLMs are Bug Replicators: An Empirical Study on LLMs' Capability in Completing Bug-prone Code
di: Guo, Liwei, et al.
Pubblicazione: (2025)
di: Guo, Liwei, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Traces of Memorisation in Large Language Models for Code
di: Al-Kaswan, Ali, et al.
Pubblicazione: (2023) -
Code Red! On the Harmfulness of Applying Off-the-shelf Large Language Models to Programming Tasks
di: Al-Kaswan, Ali, et al.
Pubblicazione: (2025) -
Do Agents Dream of Root Shells? Partial-Credit Evaluation of LLM Agents in Capture the Flag Challenges
di: Al-Kaswan, Ali, et al.
Pubblicazione: (2026) -
TreeRanker: Fast and Model-agnostic Ranking System for Code Suggestions in IDEs
di: Cipollone, Daniele, et al.
Pubblicazione: (2025) -
An Exploratory Investigation into Code License Infringements in Large Language Model Training Datasets
di: Katzy, Jonathan, et al.
Pubblicazione: (2024)