Code Red! On the Harmfulness of Applying Off-the-shelf Large Language Models to Programming Tasks
Fuente:
arXiv
Saved in:
| Main Authors: | Al-Kaswan, Ali, Deatc, Sebastian, Koç, Begüm, van Deursen, Arie, Izadi, Maliheh |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Traces of Memorisation in Large Language Models for Code
by: Al-Kaswan, Ali, et al.
Published: (2023)
by: Al-Kaswan, Ali, et al.
Published: (2023)
Model See, Model Do? Exposure-Aware Evaluation of Bug-vs-Fix Preference in Code LLMs
by: Al-Kaswan, Ali, et al.
Published: (2026)
by: Al-Kaswan, Ali, et al.
Published: (2026)
An Exploratory Investigation into Code License Infringements in Large Language Model Training Datasets
by: Katzy, Jonathan, et al.
Published: (2024)
by: Katzy, Jonathan, et al.
Published: (2024)
Do Agents Dream of Root Shells? Partial-Credit Evaluation of LLM Agents in Capture the Flag Challenges
by: Al-Kaswan, Ali, et al.
Published: (2026)
by: Al-Kaswan, Ali, et al.
Published: (2026)
TreeRanker: Fast and Model-agnostic Ranking System for Code Suggestions in IDEs
by: Cipollone, Daniele, et al.
Published: (2025)
by: Cipollone, Daniele, et al.
Published: (2025)
A Transformer-Based Approach for Smart Invocation of Automatic Code Completion
by: de Moor, Aral, et al.
Published: (2024)
by: de Moor, Aral, et al.
Published: (2024)
AST-PAC: AST-guided Membership Inference for Code
by: Koohestani, Roham, et al.
Published: (2026)
by: Koohestani, Roham, et al.
Published: (2026)
Language Models for Code Completion: A Practical Evaluation
by: Izadi, Maliheh, et al.
Published: (2024)
by: Izadi, Maliheh, et al.
Published: (2024)
Benchmarking AI Models in Software Engineering: A Review, Search Tool, and Unified Approach for Elevating Benchmark Quality
by: Koohestani, Roham, et al.
Published: (2025)
by: Koohestani, Roham, et al.
Published: (2025)
A Qualitative Investigation into LLM-Generated Multilingual Code Comments and Automatic Evaluation Metrics
by: Katzy, Jonathan, et al.
Published: (2025)
by: Katzy, Jonathan, et al.
Published: (2025)
Evaluating Large Language Models for Functional and Maintainable Code in Industrial Settings: A Case Study at ASML
by: Mundhra, Yash, et al.
Published: (2025)
by: Mundhra, Yash, et al.
Published: (2025)
Evaluating Non-English Developer Support in Machine Learning for Software Engineering
by: Katzy, Jonathan, et al.
Published: (2026)
by: Katzy, Jonathan, et al.
Published: (2026)
A Multi-agent Onboarding Assistant based on Large Language Models, Retrieval Augmented Generation, and Chain-of-Thought
by: Ionescu, Andrei Cristian, et al.
Published: (2025)
by: Ionescu, Andrei Cristian, et al.
Published: (2025)
In-IDE Human-AI Experience in the Era of Large Language Models; A Literature Review
by: Sergeyuk, Agnia, et al.
Published: (2024)
by: Sergeyuk, Agnia, et al.
Published: (2024)
Towards Automatic Translation of Machine Learning Visual Insights to Analytical Assertions
by: Shome, Arumoy, et al.
Published: (2024)
by: Shome, Arumoy, et al.
Published: (2024)
Long Code Arena: a Set of Benchmarks for Long-Context Code Models
by: Bogomolov, Egor, et al.
Published: (2024)
by: Bogomolov, Egor, et al.
Published: (2024)
Leveraging Large Language Models for Enhancing the Understandability of Generated Unit Tests
by: Deljouyi, Amirhossein, et al.
Published: (2024)
by: Deljouyi, Amirhossein, et al.
Published: (2024)
Rethinking IDE Customization for Enhanced HAX: A Hyperdimensional Perspective
by: Koohestani, Roham, et al.
Published: (2025)
by: Koohestani, Roham, et al.
Published: (2025)
Understanding Feedback Mechanisms in Machine Learning Jupyter Notebooks
by: Shome, Arumoy, et al.
Published: (2024)
by: Shome, Arumoy, et al.
Published: (2024)
HyperSeq: A Hyper-Adaptive Representation for Predictive Sequencing of States
by: Koohestani, Roham, et al.
Published: (2025)
by: Koohestani, Roham, et al.
Published: (2025)
Data vs. Model Machine Learning Fairness Testing: An Empirical Study
by: Shome, Arumoy, et al.
Published: (2024)
by: Shome, Arumoy, et al.
Published: (2024)
McUDI: Model-Centric Unsupervised Degradation Indicator for Failure Prediction AIOps Solutions
by: Poenaru-Olaru, Lorena, et al.
Published: (2024)
by: Poenaru-Olaru, Lorena, et al.
Published: (2024)
Automated Harmfulness Testing for Code Large Language Models
by: Tan, Honghao, et al.
Published: (2025)
by: Tan, Honghao, et al.
Published: (2025)
Does In-IDE Calibration of Large Language Models work at Scale?
by: Koohestani, Roham, et al.
Published: (2025)
by: Koohestani, Roham, et al.
Published: (2025)
Code4MeV2: a Research-oriented Code-completion Platform
by: Koohestani, Roham, et al.
Published: (2025)
by: Koohestani, Roham, et al.
Published: (2025)
Prompt-with-Me: in-IDE Structured Prompt Management for LLM-Driven Software Engineering
by: Li, Ziyou, et al.
Published: (2025)
by: Li, Ziyou, et al.
Published: (2025)
Is Your Anomaly Detector Ready for Change? Adapting AIOps Solutions to the Real World
by: Poenaru-Olaru, Lorena, et al.
Published: (2023)
by: Poenaru-Olaru, Lorena, et al.
Published: (2023)
The Heap: A Contamination-Free Multilingual Code Dataset for Evaluating Large Language Models
by: Katzy, Jonathan, et al.
Published: (2025)
by: Katzy, Jonathan, et al.
Published: (2025)
Developer Interaction Patterns with Proactive AI: A Five-Day Field Study
by: Kuo, Nadine, et al.
Published: (2026)
by: Kuo, Nadine, et al.
Published: (2026)
How Much Do Code Language Models Remember? An Investigation on Data Extraction Attacks before and after Fine-tuning
by: Salerno, Fabio, et al.
Published: (2025)
by: Salerno, Fabio, et al.
Published: (2025)
A Survey on Evaluating Large Language Models in Code Generation Tasks
by: Chen, Liguo, et al.
Published: (2024)
by: Chen, Liguo, et al.
Published: (2024)
TriCEGAR: A Trace-Driven Abstraction Mechanism for Agentic AI
by: Koohestani, Roham, et al.
Published: (2026)
by: Koohestani, Roham, et al.
Published: (2026)
Sustainable Machine Learning Retraining: Optimizing Energy Efficiency Without Compromising Accuracy
by: Poenaru-Olaru, Lorena, et al.
Published: (2025)
by: Poenaru-Olaru, Lorena, et al.
Published: (2025)
Human-AI Experience in Integrated Development Environments: A Systematic Literature Review
by: Sergeyuk, Agnia, et al.
Published: (2025)
by: Sergeyuk, Agnia, et al.
Published: (2025)
Instructive Code Retriever: Learn from Large Language Model's Feedback for Code Intelligence Tasks
by: Lu, Jiawei, et al.
Published: (2024)
by: Lu, Jiawei, et al.
Published: (2024)
Developer Needs and Feasible Features for AI Assistants in IDEs
by: Sergeyuk, Agnia, et al.
Published: (2024)
by: Sergeyuk, Agnia, et al.
Published: (2024)
Advancing Code Coverage: Incorporating Program Analysis with Large Language Models
by: Yang, Chen, et al.
Published: (2024)
by: Yang, Chen, et al.
Published: (2024)
SafeTune: Search-based Harmfulness Minimisation for Large Language Models
by: d'Aloisio, Giordano, et al.
Published: (2026)
by: d'Aloisio, Giordano, et al.
Published: (2026)
Prepared for the Unknown: Adapting AIOps Capacity Forecasting Models to Data Changes
by: Poenaru-Olaru, Lorena, et al.
Published: (2025)
by: Poenaru-Olaru, Lorena, et al.
Published: (2025)
Applying Bayesian Analysis Guidelines to Empirical Software Engineering Data: The Case of Programming Languages and Code Quality
by: Furia, Carlo A., et al.
Published: (2021)
by: Furia, Carlo A., et al.
Published: (2021)
Similar Items
-
Traces of Memorisation in Large Language Models for Code
by: Al-Kaswan, Ali, et al.
Published: (2023) -
Model See, Model Do? Exposure-Aware Evaluation of Bug-vs-Fix Preference in Code LLMs
by: Al-Kaswan, Ali, et al.
Published: (2026) -
An Exploratory Investigation into Code License Infringements in Large Language Model Training Datasets
by: Katzy, Jonathan, et al.
Published: (2024) -
Do Agents Dream of Root Shells? Partial-Credit Evaluation of LLM Agents in Capture the Flag Challenges
by: Al-Kaswan, Ali, et al.
Published: (2026) -
TreeRanker: Fast and Model-agnostic Ranking System for Code Suggestions in IDEs
by: Cipollone, Daniele, et al.
Published: (2025)