On Benchmarking Code LLMs for Android Malware Analysis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: He, Yiling, She, Hongyu, Qian, Xingzhi, Zheng, Xinran, Chen, Zhuo, Qin, Zhan, Cavallaro, Lorenzo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913805319012352
author He, Yiling
She, Hongyu
Qian, Xingzhi
Zheng, Xinran
Chen, Zhuo
Qin, Zhan
Cavallaro, Lorenzo
author_facet He, Yiling
She, Hongyu
Qian, Xingzhi
Zheng, Xinran
Chen, Zhuo
Qin, Zhan
Cavallaro, Lorenzo
contents Large Language Models (LLMs) have demonstrated strong capabilities in various code intelligence tasks. However, their effectiveness for Android malware analysis remains underexplored. Decompiled Android malware code presents unique challenges for analysis, due to the malicious logic being buried within a large number of functions and the frequent lack of meaningful function names. This paper presents CAMA, a benchmarking framework designed to systematically evaluate the effectiveness of Code LLMs in Android malware analysis. CAMA specifies structured model outputs to support key malware analysis tasks, including malicious function identification and malware purpose summarization. Built on these, it integrates three domain-specific evaluation metrics (consistency, fidelity, and semantic relevance), enabling rigorous stability and effectiveness assessment and cross-model comparison. We construct a benchmark dataset of 118 Android malware samples from 13 families collected in recent years, encompassing over 7.5 million distinct functions, and use CAMA to evaluate four popular open-source Code LLMs. Our experiments provide insights into how Code LLMs interpret decompiled code and quantify the sensitivity to function renaming, highlighting both their potential and current limitations in malware analysis.
format Preprint
id arxiv_https___arxiv_org_abs_2504_00694
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle On Benchmarking Code LLMs for Android Malware Analysis
He, Yiling
She, Hongyu
Qian, Xingzhi
Zheng, Xinran
Chen, Zhuo
Qin, Zhan
Cavallaro, Lorenzo
Cryptography and Security
Machine Learning
Large Language Models (LLMs) have demonstrated strong capabilities in various code intelligence tasks. However, their effectiveness for Android malware analysis remains underexplored. Decompiled Android malware code presents unique challenges for analysis, due to the malicious logic being buried within a large number of functions and the frequent lack of meaningful function names. This paper presents CAMA, a benchmarking framework designed to systematically evaluate the effectiveness of Code LLMs in Android malware analysis. CAMA specifies structured model outputs to support key malware analysis tasks, including malicious function identification and malware purpose summarization. Built on these, it integrates three domain-specific evaluation metrics (consistency, fidelity, and semantic relevance), enabling rigorous stability and effectiveness assessment and cross-model comparison. We construct a benchmark dataset of 118 Android malware samples from 13 families collected in recent years, encompassing over 7.5 million distinct functions, and use CAMA to evaluate four popular open-source Code LLMs. Our experiments provide insights into how Code LLMs interpret decompiled code and quantify the sensitivity to function renaming, highlighting both their potential and current limitations in malware analysis.
title On Benchmarking Code LLMs for Android Malware Analysis
topic Cryptography and Security
Machine Learning
url https://arxiv.org/abs/2504.00694