MGC: A Compiler Framework Exploiting Compositional Blindness in Aligned LLMs for Malware Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yan, Lu, Zhang, Zhuo, Xu, Xiangzhe, An, Shengwei, Shen, Guangyu, Xuan, Zhou, Chen, Xuan, Zhang, Xiangyu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912461642268672
author Yan, Lu
Zhang, Zhuo
Xu, Xiangzhe
An, Shengwei
Shen, Guangyu
Xuan, Zhou
Chen, Xuan
Zhang, Xiangyu
author_facet Yan, Lu
Zhang, Zhuo
Xu, Xiangzhe
An, Shengwei
Shen, Guangyu
Xuan, Zhou
Chen, Xuan
Zhang, Xiangyu
contents Large language models (LLMs) have democratized software development, reducing the expertise barrier for programming complex applications. This accessibility extends to malicious software development, raising significant security concerns. While LLM providers have implemented alignment mechanisms to prevent direct generation of overtly malicious code, these safeguards predominantly evaluate individual prompts in isolation, overlooking a critical vulnerability: malicious operations can be systematically decomposed into benign-appearing sub-tasks. In this paper, we introduce the Malware Generation Compiler (MGC), a novel framework that leverages this vulnerability through modular decomposition and alignment-evasive generation. MGC employs a specialized Malware Description Intermediate Representation (MDIR) to bridge high-level malicious intents and benign-appearing code snippets. Extensive evaluation demonstrates that our attack reliably generates functional malware across diverse task specifications and categories, outperforming jailbreaking methods by +365.79% and underground services by +78.07% in correctness on three benchmark datasets. Case studies further show that MGC can reproduce and even enhance 16 real-world malware samples. This work provides critical insights for security researchers by exposing the risks of compositional attacks against aligned AI systems. Demonstrations are available at https://sites.google.com/view/malware-generation-compiler.
format Preprint
id arxiv_https___arxiv_org_abs_2507_02057
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MGC: A Compiler Framework Exploiting Compositional Blindness in Aligned LLMs for Malware Generation
Yan, Lu
Zhang, Zhuo
Xu, Xiangzhe
An, Shengwei
Shen, Guangyu
Xuan, Zhou
Chen, Xuan
Zhang, Xiangyu
Cryptography and Security
Artificial Intelligence
Large language models (LLMs) have democratized software development, reducing the expertise barrier for programming complex applications. This accessibility extends to malicious software development, raising significant security concerns. While LLM providers have implemented alignment mechanisms to prevent direct generation of overtly malicious code, these safeguards predominantly evaluate individual prompts in isolation, overlooking a critical vulnerability: malicious operations can be systematically decomposed into benign-appearing sub-tasks. In this paper, we introduce the Malware Generation Compiler (MGC), a novel framework that leverages this vulnerability through modular decomposition and alignment-evasive generation. MGC employs a specialized Malware Description Intermediate Representation (MDIR) to bridge high-level malicious intents and benign-appearing code snippets. Extensive evaluation demonstrates that our attack reliably generates functional malware across diverse task specifications and categories, outperforming jailbreaking methods by +365.79% and underground services by +78.07% in correctness on three benchmark datasets. Case studies further show that MGC can reproduce and even enhance 16 real-world malware samples. This work provides critical insights for security researchers by exposing the risks of compositional attacks against aligned AI systems. Demonstrations are available at https://sites.google.com/view/malware-generation-compiler.
title MGC: A Compiler Framework Exploiting Compositional Blindness in Aligned LLMs for Malware Generation
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2507.02057