Do LLMs Consider Security? An Empirical Study on Responses to Programming Questions
Fuente:
arXiv
Saved in:
| Main Authors: | Sajadi, Amirali, Le, Binh, Nguyen, Anh, Damevski, Kostadin, Chatterjee, Preetha |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How Safe Are AI-Generated Patches? A Large-scale Study on Security Risks in LLM and Agentic Automated Program Repair on SWE-bench
by: Sajadi, Amirali, et al.
Published: (2025)
by: Sajadi, Amirali, et al.
Published: (2025)
Psycholinguistic Analyses in Software Engineering Text: A Systematic Literature Review
by: Sajadi, Amirali, et al.
Published: (2025)
by: Sajadi, Amirali, et al.
Published: (2025)
AXE: An Agentic eXploit Engine for Confirming Zero-Day Vulnerability Reports
by: Sajadi, Amirali, et al.
Published: (2026)
by: Sajadi, Amirali, et al.
Published: (2026)
Data Augmentation for Improving Emotion Recognition in Software Engineering Communication
by: Imran, Mia Mohammad, et al.
Published: (2022)
by: Imran, Mia Mohammad, et al.
Published: (2022)
Incivility in Open Source Projects: A Comprehensive Annotated Dataset of Locked GitHub Issue Threads
by: Ehsani, Ramtin, et al.
Published: (2024)
by: Ehsani, Ramtin, et al.
Published: (2024)
Towards Personalizing Secure Programming Education with LLM-Injected Vulnerabilities
by: Frazier, Matthew, et al.
Published: (2026)
by: Frazier, Matthew, et al.
Published: (2026)
Where Do AI Coding Agents Fail? An Empirical Study of Failed Agentic Pull Requests in GitHub
by: Ehsani, Ramtin, et al.
Published: (2026)
by: Ehsani, Ramtin, et al.
Published: (2026)
LiveCodeBench Pro: How Do Olympiad Medalists Judge LLMs in Competitive Programming?
by: Zheng, Zihan, et al.
Published: (2025)
by: Zheng, Zihan, et al.
Published: (2025)
Understanding and Predicting Derailment in Toxic Conversations on GitHub
by: Imran, Mia Mohammad, et al.
Published: (2025)
by: Imran, Mia Mohammad, et al.
Published: (2025)
Can LLMs Reason About Program Semantics? A Comprehensive Evaluation of LLMs on Formal Specification Inference
by: Le-Cong, Thanh, et al.
Published: (2025)
by: Le-Cong, Thanh, et al.
Published: (2025)
From Empirical Evaluation to Context-Aware Enhancement: Repairing Regression Errors with LLMs
by: Ho, Anh, et al.
Published: (2025)
by: Ho, Anh, et al.
Published: (2025)
Is Programming by Example solved by LLMs?
by: Li, Wen-Ding, et al.
Published: (2024)
by: Li, Wen-Ding, et al.
Published: (2024)
Toxicity Ahead: Forecasting Conversational Derailment on GitHub
by: Imran, Mia Mohammad, et al.
Published: (2025)
by: Imran, Mia Mohammad, et al.
Published: (2025)
Synergizing LLMs and Knowledge Graphs: A Novel Approach to Software Repository-Related Question Answering
by: Abedu, Samuel, et al.
Published: (2024)
by: Abedu, Samuel, et al.
Published: (2024)
A Critical Study of What Code-LLMs (Do Not) Learn
by: Anand, Abhinav, et al.
Published: (2024)
by: Anand, Abhinav, et al.
Published: (2024)
Perish or Flourish? A Holistic Evaluation of Large Language Models for Code Generation in Functional Programming
by: Lang, Nguyet-Anh H., et al.
Published: (2026)
by: Lang, Nguyet-Anh H., et al.
Published: (2026)
Do Large Code Models Understand Programming Concepts? Counterfactual Analysis for Code Predicates
by: Hooda, Ashish, et al.
Published: (2024)
by: Hooda, Ashish, et al.
Published: (2024)
On the Effectiveness of Machine Learning-based Call Graph Pruning: An Empirical Study
by: Mir, Amir M., et al.
Published: (2024)
by: Mir, Amir M., et al.
Published: (2024)
How Do Your Code LLMs Perform? Empowering Code Instruction Tuning with High-Quality Data
by: Wang, Yejie, et al.
Published: (2024)
by: Wang, Yejie, et al.
Published: (2024)
Do Code Models Suffer from the Dunning-Kruger Effect?
by: Singh, Mukul, et al.
Published: (2025)
by: Singh, Mukul, et al.
Published: (2025)
Improving Data Curation of Software Vulnerability Patches through Uncertainty Quantification
by: Chen, Hui, et al.
Published: (2024)
by: Chen, Hui, et al.
Published: (2024)
Agent Skills in the Wild: An Empirical Study of Security Vulnerabilities at Scale
by: Liu, Yi, et al.
Published: (2026)
by: Liu, Yi, et al.
Published: (2026)
SQLong: Enhanced NL2SQL for Longer Contexts with LLMs
by: Nguyen, Dai Quoc, et al.
Published: (2025)
by: Nguyen, Dai Quoc, et al.
Published: (2025)
Can LLMs Enable Verification in Mainstream Programming?
by: Shefer, Aleksandr, et al.
Published: (2025)
by: Shefer, Aleksandr, et al.
Published: (2025)
Software Mention Recognition with a Three-Stage Framework Based on BERTology Models at SOMD 2024
by: Thi, Thuy Nguyen, et al.
Published: (2024)
by: Thi, Thuy Nguyen, et al.
Published: (2024)
Can LLMs Generate Reliable Test Case Generators? A Study on Competition-Level Programming Problems
by: Cao, Yuhan, et al.
Published: (2025)
by: Cao, Yuhan, et al.
Published: (2025)
An Empirical Study on Failures in Automated Issue Solving
by: Liu, Simiao, et al.
Published: (2025)
by: Liu, Simiao, et al.
Published: (2025)
AI-Powered Commit Explorer (APCE)
by: Grees, Yousab, et al.
Published: (2025)
by: Grees, Yousab, et al.
Published: (2025)
An Empirical Study on Self-correcting Large Language Models for Data Science Code Generation
by: Quoc, Thai Tang, et al.
Published: (2024)
by: Quoc, Thai Tang, et al.
Published: (2024)
Jailbreaking ChatGPT via Prompt Engineering: An Empirical Study
by: Liu, Yi, et al.
Published: (2023)
by: Liu, Yi, et al.
Published: (2023)
LLMs Lean on Priors, Not Programming Language Semantics
by: Thimmaiah, Aditya, et al.
Published: (2025)
by: Thimmaiah, Aditya, et al.
Published: (2025)
CP-Agent: Agentic Constraint Programming
by: Szeider, Stefan
Published: (2025)
by: Szeider, Stefan
Published: (2025)
Rethinking Repetition Problems of LLMs in Code Generation
by: Dong, Yihong, et al.
Published: (2025)
by: Dong, Yihong, et al.
Published: (2025)
Structured Program Synthesis using LLMs: Results and Insights from the IPARC Challenge
by: Surana, Shraddha, et al.
Published: (2025)
by: Surana, Shraddha, et al.
Published: (2025)
AutoCode: LLMs as Problem Setters for Competitive Programming
by: Zhou, Shang, et al.
Published: (2025)
by: Zhou, Shang, et al.
Published: (2025)
Semantically Aligned Question and Code Generation for Automated Insight Generation
by: Singha, Ananya, et al.
Published: (2024)
by: Singha, Ananya, et al.
Published: (2024)
CangjieBench: Benchmarking LLMs on a Low-Resource General-Purpose Programming Language
by: Cheng, Junhang, et al.
Published: (2026)
by: Cheng, Junhang, et al.
Published: (2026)
Evaluation of LLMs on Syntax-Aware Code Fill-in-the-Middle Tasks
by: Gong, Linyuan, et al.
Published: (2024)
by: Gong, Linyuan, et al.
Published: (2024)
MCP-Solver: Integrating Language Models with Constraint Programming Systems
by: Szeider, Stefan
Published: (2024)
by: Szeider, Stefan
Published: (2024)
An Empirical Evaluation of Locally Deployed LLMs for Bug Detection in Python Code
by: Vulićević, Jelena Ilić
Published: (2026)
by: Vulićević, Jelena Ilić
Published: (2026)
Similar Items
-
How Safe Are AI-Generated Patches? A Large-scale Study on Security Risks in LLM and Agentic Automated Program Repair on SWE-bench
by: Sajadi, Amirali, et al.
Published: (2025) -
Psycholinguistic Analyses in Software Engineering Text: A Systematic Literature Review
by: Sajadi, Amirali, et al.
Published: (2025) -
AXE: An Agentic eXploit Engine for Confirming Zero-Day Vulnerability Reports
by: Sajadi, Amirali, et al.
Published: (2026) -
Data Augmentation for Improving Emotion Recognition in Software Engineering Communication
by: Imran, Mia Mohammad, et al.
Published: (2022) -
Incivility in Open Source Projects: A Comprehensive Annotated Dataset of Locked GitHub Issue Threads
by: Ehsani, Ramtin, et al.
Published: (2024)