LLMalMorph: On The Feasibility of Generating Variant Malware using Large-Language-Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Akil, Md Ajwad, Li, Adrian Shuai, Karim, Imtiaz, Iyengar, Arun, Kundu, Ashish, Parla, Vinny, Bertino, Elisa
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908574817452032
author Akil, Md Ajwad
Li, Adrian Shuai
Karim, Imtiaz
Iyengar, Arun
Kundu, Ashish
Parla, Vinny
Bertino, Elisa
author_facet Akil, Md Ajwad
Li, Adrian Shuai
Karim, Imtiaz
Iyengar, Arun
Kundu, Ashish
Parla, Vinny
Bertino, Elisa
contents Large Language Models (LLMs) have transformed software development and automated code generation. Motivated by these advancements, this paper explores the feasibility of LLMs in modifying malware source code to generate variants. We introduce LLMalMorph, a semi-automated framework that leverages semantical and syntactical code comprehension by LLMs to generate new malware variants. LLMalMorph extracts function-level information from the malware source code and employs custom-engineered prompts coupled with strategically defined code transformations to guide the LLM in generating variants without resource-intensive fine-tuning. To evaluate LLMalMorph, we collected 10 diverse Windows malware samples of varying types, complexity and functionality and generated 618 variants. Our experiments demonstrate that LLMalMorph variants can effectively evade antivirus engines, achieving typical detection rate reductions of 10-15% across multiple complex samples. Furthermore, without explicitly targeting learning-based detectors, LLMalMorph attained attack success rates of up to 91% against a Machine Learning (ML) based malware detector. We also discuss the limitations of current LLM capabilities in generating malware variants from source code and assess where this emerging technology stands in the broader context of malware variant generation.
format Preprint
id arxiv_https___arxiv_org_abs_2507_09411
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LLMalMorph: On The Feasibility of Generating Variant Malware using Large-Language-Models
Akil, Md Ajwad
Li, Adrian Shuai
Karim, Imtiaz
Iyengar, Arun
Kundu, Ashish
Parla, Vinny
Bertino, Elisa
Cryptography and Security
Large Language Models (LLMs) have transformed software development and automated code generation. Motivated by these advancements, this paper explores the feasibility of LLMs in modifying malware source code to generate variants. We introduce LLMalMorph, a semi-automated framework that leverages semantical and syntactical code comprehension by LLMs to generate new malware variants. LLMalMorph extracts function-level information from the malware source code and employs custom-engineered prompts coupled with strategically defined code transformations to guide the LLM in generating variants without resource-intensive fine-tuning. To evaluate LLMalMorph, we collected 10 diverse Windows malware samples of varying types, complexity and functionality and generated 618 variants. Our experiments demonstrate that LLMalMorph variants can effectively evade antivirus engines, achieving typical detection rate reductions of 10-15% across multiple complex samples. Furthermore, without explicitly targeting learning-based detectors, LLMalMorph attained attack success rates of up to 91% against a Machine Learning (ML) based malware detector. We also discuss the limitations of current LLM capabilities in generating malware variants from source code and assess where this emerging technology stands in the broader context of malware variant generation.
title LLMalMorph: On The Feasibility of Generating Variant Malware using Large-Language-Models
topic Cryptography and Security
url https://arxiv.org/abs/2507.09411