Structured Safety Auditing for Balancing Code Correctness and Content Safety in LLM-Generated Code
Fuente:
arXiv
Saved in:
| Main Authors: | Tan, Honghao, Wang, Haibo, Tan, Shin Hwei |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LLM-Guided Issue Generation from Uncovered Code Segments
by: Pressato, Diany, et al.
Published: (2026)
by: Pressato, Diany, et al.
Published: (2026)
Automated Harmfulness Testing for Code Large Language Models
by: Tan, Honghao, et al.
Published: (2025)
by: Tan, Honghao, et al.
Published: (2025)
Testing Refactoring Engine via Historical Bug Report driven LLM
by: Wang, Haibo, et al.
Published: (2025)
by: Wang, Haibo, et al.
Published: (2025)
Investigating Code Reuse in Software Redesign: A Case Study
by: Zhang, Xiaowen, et al.
Published: (2026)
by: Zhang, Xiaowen, et al.
Published: (2026)
Ethics Testing: Proactive Identification of Generative AI System Harms
by: Tan, Shin Hwei, et al.
Published: (2026)
by: Tan, Shin Hwei, et al.
Published: (2026)
What Makes Code Generation Ethically Sourced?
by: Xu, Zhuolin, et al.
Published: (2025)
by: Xu, Zhuolin, et al.
Published: (2025)
COBOLAssist: Analyzing and Fixing Compilation Errors for LLM-Powered COBOL Code Generation
by: Dau, Anh T. V., et al.
Published: (2026)
by: Dau, Anh T. V., et al.
Published: (2026)
An Empirical Study of False Negatives and Positives of Static Code Analyzers From the Perspective of Historical Issues
by: Cui, Han, et al.
Published: (2024)
by: Cui, Han, et al.
Published: (2024)
Moving beyond Deletions: Program Simplification via Diverse Program Transformations
by: Wang, Haibo, et al.
Published: (2024)
by: Wang, Haibo, et al.
Published: (2024)
An Empirical Study of Refactoring Engine Bugs
by: Wang, Haibo, et al.
Published: (2024)
by: Wang, Haibo, et al.
Published: (2024)
Dissecting Bug Triggers and Failure Modes in Modern Agentic Frameworks: An Empirical Study
by: Zhang, Xiaowen, et al.
Published: (2026)
by: Zhang, Xiaowen, et al.
Published: (2026)
Guiding ChatGPT to Fix Web UI Tests via Explanation-Consistency Checking
by: Xu, Zhuolin, et al.
Published: (2023)
by: Xu, Zhuolin, et al.
Published: (2023)
COBOL-Coder: Domain-Adapted Large Language Models for COBOL Code Generation and Translation
by: Dau, Anh T. V., et al.
Published: (2026)
by: Dau, Anh T. V., et al.
Published: (2026)
Understanding and Detecting Annotation-Induced Faults of Static Analyzers
by: Zhang, Huaien, et al.
Published: (2024)
by: Zhang, Huaien, et al.
Published: (2024)
Aligning the Objective of LLM-based Program Repair
by: Xu, Junjielong, et al.
Published: (2024)
by: Xu, Junjielong, et al.
Published: (2024)
BloomAPR: A Bloom's Taxonomy-based Framework for Assessing the Capabilities of LLM-Powered APR Solutions
by: Ma, Yinghang, et al.
Published: (2025)
by: Ma, Yinghang, et al.
Published: (2025)
Is Vibe Coding the Future? An Empirical Assessment of LLM Generated Codes for Construction Safety
by: Uddin, S M Jamil
Published: (2026)
by: Uddin, S M Jamil
Published: (2026)
Are Benchmark Tests Strong Enough? Mutation-Guided Diagnosis and Augmentation of Regression Suites
by: Li, Chenglin, et al.
Published: (2026)
by: Li, Chenglin, et al.
Published: (2026)
From Historical Patches to Repair Plans: Outcome-Conditioned Reasoning for Repository-Level Program Repair
by: Li, Chenglin, et al.
Published: (2026)
by: Li, Chenglin, et al.
Published: (2026)
When Prompt Under-Specification Improves Code Correctness: An Exploratory Study of Prompt Wording and Structure Effects on LLM-Based Code Generation
by: AKLI, Amal, et al.
Published: (2026)
by: AKLI, Amal, et al.
Published: (2026)
SemGuard: Real-Time Semantic Evaluator for Correcting LLM-Generated Code
by: Wang, Qinglin, et al.
Published: (2025)
by: Wang, Qinglin, et al.
Published: (2025)
SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
by: Hu, Qi, et al.
Published: (2026)
by: Hu, Qi, et al.
Published: (2026)
What Breaks When LLMs Code? Characterizing Operational Safety Failures of Agentic Code Assistants
by: Hasan, Alif Al, et al.
Published: (2026)
by: Hasan, Alif Al, et al.
Published: (2026)
Usability as a Weapon: Attacking the Safety of LLM-Based Code Generation via Usability Requirements
by: Li, Yue, et al.
Published: (2026)
by: Li, Yue, et al.
Published: (2026)
Assessing Correctness in LLM-Based Code Generation via Uncertainty Estimation
by: Sharma, Arindam, et al.
Published: (2025)
by: Sharma, Arindam, et al.
Published: (2025)
Scalable Thread-Safety Analysis of Java Classes with CodeQL
by: Jåtten, Bjørnar Haugstad, et al.
Published: (2025)
by: Jåtten, Bjørnar Haugstad, et al.
Published: (2025)
Search-Induced Issues in Web-Augmented LLM Code Generation: Detecting and Repairing Error-Inducing Pages
by: Wang, Guoqing, et al.
Published: (2026)
by: Wang, Guoqing, et al.
Published: (2026)
Inducing Vulnerable Code Generation in LLM Coding Assistants
by: Zeng, Binqi, et al.
Published: (2025)
by: Zeng, Binqi, et al.
Published: (2025)
Uncovering Weaknesses in Neural Code Generation
by: Lian, Xiaoli, et al.
Published: (2024)
by: Lian, Xiaoli, et al.
Published: (2024)
Uncertainty Quantification for LLM-based Code Generation
by: Xu, Senrong, et al.
Published: (2026)
by: Xu, Senrong, et al.
Published: (2026)
CodeAttack: Revealing Safety Generalization Challenges of Large Language Models via Code Completion
by: Ren, Qibing, et al.
Published: (2024)
by: Ren, Qibing, et al.
Published: (2024)
Future of Code with Generative AI: Transparency and Safety in the Era of AI Generated Software
by: Hanson, David
Published: (2025)
by: Hanson, David
Published: (2025)
Towards Exception Safety Code Generation with Intermediate Representation Agents Framework
by: Zhang, Xuanming, et al.
Published: (2024)
by: Zhang, Xuanming, et al.
Published: (2024)
CodeCircuit: Toward Inferring LLM-Generated Code Correctness via Attribution Graphs
by: He, Yicheng, et al.
Published: (2026)
by: He, Yicheng, et al.
Published: (2026)
Predicting Developer Acceptance of AI-Generated Code Suggestions
by: Jiang, Jing, et al.
Published: (2026)
by: Jiang, Jing, et al.
Published: (2026)
Beyond Code Generation: Assessing Code LLM Maturity with Postconditions
by: He, Fusen, et al.
Published: (2024)
by: He, Fusen, et al.
Published: (2024)
IntrinTrans: LLM-based Intrinsic Code Translator for RISC-V Vector
by: Han, Liutong, et al.
Published: (2025)
by: Han, Liutong, et al.
Published: (2025)
Static Analysis as a Feedback Loop: Enhancing LLM-Generated Code Beyond Correctness
by: Blyth, Scott, et al.
Published: (2025)
by: Blyth, Scott, et al.
Published: (2025)
CodeArena: A Collective Evaluation Platform for LLM Code Generation
by: Du, Mingzhe, et al.
Published: (2025)
by: Du, Mingzhe, et al.
Published: (2025)
Beyond Functional Correctness: Exploring Hallucinations in LLM-Generated Code
by: Liu, Fang, et al.
Published: (2024)
by: Liu, Fang, et al.
Published: (2024)
Similar Items
-
LLM-Guided Issue Generation from Uncovered Code Segments
by: Pressato, Diany, et al.
Published: (2026) -
Automated Harmfulness Testing for Code Large Language Models
by: Tan, Honghao, et al.
Published: (2025) -
Testing Refactoring Engine via Historical Bug Report driven LLM
by: Wang, Haibo, et al.
Published: (2025) -
Investigating Code Reuse in Software Redesign: A Case Study
by: Zhang, Xiaowen, et al.
Published: (2026) -
Ethics Testing: Proactive Identification of Generative AI System Harms
by: Tan, Shin Hwei, et al.
Published: (2026)