MOCHA: Are Code Language Models Robust Against Multi-Turn Malicious Coding Prompts?
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911078156337152 |
|---|---|
| author | Wahed, Muntasir Zhou, Xiaona Nguyen, Kiet A. Yu, Tianjiao Diwan, Nirav Wang, Gang Hakkani-Tür, Dilek Lourentzou, Ismini |
| author_facet | Wahed, Muntasir Zhou, Xiaona Nguyen, Kiet A. Yu, Tianjiao Diwan, Nirav Wang, Gang Hakkani-Tür, Dilek Lourentzou, Ismini |
| contents | Recent advancements in Large Language Models (LLMs) have significantly enhanced their code generation capabilities. However, their robustness against adversarial misuse, particularly through multi-turn malicious coding prompts, remains underexplored. In this work, we introduce code decomposition attacks, where a malicious coding task is broken down into a series of seemingly benign subtasks across multiple conversational turns to evade safety filters. To facilitate systematic evaluation, we introduce \benchmarkname{}, a large-scale benchmark designed to evaluate the robustness of code LLMs against both single-turn and multi-turn malicious prompts. Empirical results across open- and closed-source models reveal persistent vulnerabilities, especially under multi-turn scenarios. Fine-tuning on MOCHA improves rejection rates while preserving coding ability, and importantly, enhances robustness on external adversarial datasets with up to 32.4% increase in rejection rates without any additional supervision. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2507_19598 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | MOCHA: Are Code Language Models Robust Against Multi-Turn Malicious Coding Prompts? Wahed, Muntasir Zhou, Xiaona Nguyen, Kiet A. Yu, Tianjiao Diwan, Nirav Wang, Gang Hakkani-Tür, Dilek Lourentzou, Ismini Computation and Language Artificial Intelligence Cryptography and Security Machine Learning Recent advancements in Large Language Models (LLMs) have significantly enhanced their code generation capabilities. However, their robustness against adversarial misuse, particularly through multi-turn malicious coding prompts, remains underexplored. In this work, we introduce code decomposition attacks, where a malicious coding task is broken down into a series of seemingly benign subtasks across multiple conversational turns to evade safety filters. To facilitate systematic evaluation, we introduce \benchmarkname{}, a large-scale benchmark designed to evaluate the robustness of code LLMs against both single-turn and multi-turn malicious prompts. Empirical results across open- and closed-source models reveal persistent vulnerabilities, especially under multi-turn scenarios. Fine-tuning on MOCHA improves rejection rates while preserving coding ability, and importantly, enhances robustness on external adversarial datasets with up to 32.4% increase in rejection rates without any additional supervision. |
| title | MOCHA: Are Code Language Models Robust Against Multi-Turn Malicious Coding Prompts? |
| topic | Computation and Language Artificial Intelligence Cryptography and Security Machine Learning |
| url | https://arxiv.org/abs/2507.19598 |