Automated Harmfulness Testing for Code Large Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Tan, Honghao, Wang, Haibo, Pressato, Diany, Xu, Yisen, Tan, Shin Hwei |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
LLM-Guided Issue Generation from Uncovered Code Segments
di: Pressato, Diany, et al.
Pubblicazione: (2026)
di: Pressato, Diany, et al.
Pubblicazione: (2026)
Structured Safety Auditing for Balancing Code Correctness and Content Safety in LLM-Generated Code
di: Tan, Honghao, et al.
Pubblicazione: (2026)
di: Tan, Honghao, et al.
Pubblicazione: (2026)
Ethics Testing: Proactive Identification of Generative AI System Harms
di: Tan, Shin Hwei, et al.
Pubblicazione: (2026)
di: Tan, Shin Hwei, et al.
Pubblicazione: (2026)
Testing Refactoring Engine via Historical Bug Report driven LLM
di: Wang, Haibo, et al.
Pubblicazione: (2025)
di: Wang, Haibo, et al.
Pubblicazione: (2025)
Are Benchmark Tests Strong Enough? Mutation-Guided Diagnosis and Augmentation of Regression Suites
di: Li, Chenglin, et al.
Pubblicazione: (2026)
di: Li, Chenglin, et al.
Pubblicazione: (2026)
Exploring the Jupyter Ecosystem: An Empirical Study of Bugs and Vulnerabilities
di: Jiang, Wenyuan, et al.
Pubblicazione: (2025)
di: Jiang, Wenyuan, et al.
Pubblicazione: (2025)
Guiding ChatGPT to Fix Web UI Tests via Explanation-Consistency Checking
di: Xu, Zhuolin, et al.
Pubblicazione: (2023)
di: Xu, Zhuolin, et al.
Pubblicazione: (2023)
From Historical Patches to Repair Plans: Outcome-Conditioned Reasoning for Repository-Level Program Repair
di: Li, Chenglin, et al.
Pubblicazione: (2026)
di: Li, Chenglin, et al.
Pubblicazione: (2026)
Investigating Code Reuse in Software Redesign: A Case Study
di: Zhang, Xiaowen, et al.
Pubblicazione: (2026)
di: Zhang, Xiaowen, et al.
Pubblicazione: (2026)
An Empirical Study of Refactoring Engine Bugs
di: Wang, Haibo, et al.
Pubblicazione: (2024)
di: Wang, Haibo, et al.
Pubblicazione: (2024)
Moving beyond Deletions: Program Simplification via Diverse Program Transformations
di: Wang, Haibo, et al.
Pubblicazione: (2024)
di: Wang, Haibo, et al.
Pubblicazione: (2024)
What Makes Code Generation Ethically Sourced?
di: Xu, Zhuolin, et al.
Pubblicazione: (2025)
di: Xu, Zhuolin, et al.
Pubblicazione: (2025)
LPR: Large Language Models-Aided Program Reduction
di: Zhang, Mengxiao, et al.
Pubblicazione: (2023)
di: Zhang, Mengxiao, et al.
Pubblicazione: (2023)
COBOL-Coder: Domain-Adapted Large Language Models for COBOL Code Generation and Translation
di: Dau, Anh T. V., et al.
Pubblicazione: (2026)
di: Dau, Anh T. V., et al.
Pubblicazione: (2026)
Dissecting Bug Triggers and Failure Modes in Modern Agentic Frameworks: An Empirical Study
di: Zhang, Xiaowen, et al.
Pubblicazione: (2026)
di: Zhang, Xiaowen, et al.
Pubblicazione: (2026)
An Empirical Study of False Negatives and Positives of Static Code Analyzers From the Perspective of Historical Issues
di: Cui, Han, et al.
Pubblicazione: (2024)
di: Cui, Han, et al.
Pubblicazione: (2024)
Automatic Programming: Large Language Models and Beyond
di: Lyu, Michael R., et al.
Pubblicazione: (2024)
di: Lyu, Michael R., et al.
Pubblicazione: (2024)
Understanding and Detecting Annotation-Induced Faults of Static Analyzers
di: Zhang, Huaien, et al.
Pubblicazione: (2024)
di: Zhang, Huaien, et al.
Pubblicazione: (2024)
Aligning the Objective of LLM-based Program Repair
di: Xu, Junjielong, et al.
Pubblicazione: (2024)
di: Xu, Junjielong, et al.
Pubblicazione: (2024)
COBOLAssist: Analyzing and Fixing Compilation Errors for LLM-Powered COBOL Code Generation
di: Dau, Anh T. V., et al.
Pubblicazione: (2026)
di: Dau, Anh T. V., et al.
Pubblicazione: (2026)
REACCEPT: Automated Co-evolution of Production and Test Code Based on Dynamic Validation and Large Language Models
di: Chi, Jianlei, et al.
Pubblicazione: (2024)
di: Chi, Jianlei, et al.
Pubblicazione: (2024)
BloomAPR: A Bloom's Taxonomy-based Framework for Assessing the Capabilities of LLM-Powered APR Solutions
di: Ma, Yinghang, et al.
Pubblicazione: (2025)
di: Ma, Yinghang, et al.
Pubblicazione: (2025)
Tumbling Down the Rabbit Hole: How do Assisting Exploration Strategies Facilitate Grey-box Fuzzing?
di: Wu, Mingyuan, et al.
Pubblicazione: (2024)
di: Wu, Mingyuan, et al.
Pubblicazione: (2024)
Automated Unit Test Improvement using Large Language Models at Meta
di: Alshahwan, Nadia, et al.
Pubblicazione: (2024)
di: Alshahwan, Nadia, et al.
Pubblicazione: (2024)
Task Abstention for Large Language Models in Code Generation
di: Zhou, Yanke, et al.
Pubblicazione: (2026)
di: Zhou, Yanke, et al.
Pubblicazione: (2026)
SWE-Refactor: A Repository-Level Benchmark for Real-World LLM-Based Code Refactoring
di: Xu, Yisen, et al.
Pubblicazione: (2026)
di: Xu, Yisen, et al.
Pubblicazione: (2026)
Automated Commit Message Generation with Large Language Models: An Empirical Study and Beyond
di: Xue, Pengyu, et al.
Pubblicazione: (2024)
di: Xue, Pengyu, et al.
Pubblicazione: (2024)
Automated Classification of Human Code Review Comments with Large Language Models
di: Çağlar, Semih, et al.
Pubblicazione: (2026)
di: Çağlar, Semih, et al.
Pubblicazione: (2026)
KAT: Dependency-aware Automated API Testing with Large Language Models
di: Le, Tri, et al.
Pubblicazione: (2024)
di: Le, Tri, et al.
Pubblicazione: (2024)
Code Red! On the Harmfulness of Applying Off-the-shelf Large Language Models to Programming Tasks
di: Al-Kaswan, Ali, et al.
Pubblicazione: (2025)
di: Al-Kaswan, Ali, et al.
Pubblicazione: (2025)
Adversarial Attack Classification and Robustness Testing for Large Language Models for Code
di: Liu, Yang, et al.
Pubblicazione: (2025)
di: Liu, Yang, et al.
Pubblicazione: (2025)
SafeTune: Search-based Harmfulness Minimisation for Large Language Models
di: d'Aloisio, Giordano, et al.
Pubblicazione: (2026)
di: d'Aloisio, Giordano, et al.
Pubblicazione: (2026)
Large Language Models for Automated Web-Form-Test Generation: An Empirical Study
di: Li, Tao, et al.
Pubblicazione: (2024)
di: Li, Tao, et al.
Pubblicazione: (2024)
Automated Control Logic Test Case Generation using Large Language Models
di: Koziolek, Heiko, et al.
Pubblicazione: (2024)
di: Koziolek, Heiko, et al.
Pubblicazione: (2024)
Automated Test Transfer Across Android Apps Using Large Language Models
di: Beyzaei, Benyamin, et al.
Pubblicazione: (2024)
di: Beyzaei, Benyamin, et al.
Pubblicazione: (2024)
Greening Large Language Models of Code
di: Shi, Jieke, et al.
Pubblicazione: (2023)
di: Shi, Jieke, et al.
Pubblicazione: (2023)
Fine-Tuning and Prompt Engineering for Large Language Models-based Code Review Automation
di: Pornprasit, Chanathip, et al.
Pubblicazione: (2024)
di: Pornprasit, Chanathip, et al.
Pubblicazione: (2024)
ASTRAL: Automated Safety Testing of Large Language Models
di: Ugarte, Miriam, et al.
Pubblicazione: (2025)
di: Ugarte, Miriam, et al.
Pubblicazione: (2025)
Towards Automated Page Object Generation for Web Testing using Large Language Models
di: Karagöz, Betül, et al.
Pubblicazione: (2026)
di: Karagöz, Betül, et al.
Pubblicazione: (2026)
Enhancing Large Language Models with Retrieval Augmented Generation for Software Testing and Inspection Automation
di: Fingleton, Zoe, et al.
Pubblicazione: (2026)
di: Fingleton, Zoe, et al.
Pubblicazione: (2026)
Documenti analoghi
-
LLM-Guided Issue Generation from Uncovered Code Segments
di: Pressato, Diany, et al.
Pubblicazione: (2026) -
Structured Safety Auditing for Balancing Code Correctness and Content Safety in LLM-Generated Code
di: Tan, Honghao, et al.
Pubblicazione: (2026) -
Ethics Testing: Proactive Identification of Generative AI System Harms
di: Tan, Shin Hwei, et al.
Pubblicazione: (2026) -
Testing Refactoring Engine via Historical Bug Report driven LLM
di: Wang, Haibo, et al.
Pubblicazione: (2025) -
Are Benchmark Tests Strong Enough? Mutation-Guided Diagnosis and Augmentation of Regression Suites
di: Li, Chenglin, et al.
Pubblicazione: (2026)