Beyond Classification: Evaluating LLMs for Fine-Grained Automatic Malware Behavior Auditing
Fuente:
arXiv
Saved in:
| Main Authors: | Zheng, Xinran, Qian, Xingzhi, He, Yiling, Yang, Shuo, Cavallaro, Lorenzo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LAMD: Context-driven Android Malware Detection and Classification with LLMs
by: Qian, Xingzhi, et al.
Published: (2025)
by: Qian, Xingzhi, et al.
Published: (2025)
Defending against Adversarial Malware Attacks on ML-based Android Malware Detection Systems
by: He, Ping, et al.
Published: (2025)
by: He, Ping, et al.
Published: (2025)
On the Security Risks of ML-based Malware Detection Systems: A Survey
by: He, Ping, et al.
Published: (2025)
by: He, Ping, et al.
Published: (2025)
Veritas: A Semantically Grounded Agentic Framework for Memory Corruption Vulnerability Detection in Binaries
by: Zheng, Xinran, et al.
Published: (2026)
by: Zheng, Xinran, et al.
Published: (2026)
On Benchmarking Code LLMs for Android Malware Analysis
by: He, Yiling, et al.
Published: (2025)
by: He, Yiling, et al.
Published: (2025)
Rethinking and Exploring String-Based Malware Family Classification in the Era of LLMs and RAG
by: Chen, Yufan, et al.
Published: (2025)
by: Chen, Yufan, et al.
Published: (2025)
TEMPLATEFUZZ: Fine-Grained Chat Template Fuzzing for Jailbreaking and Red Teaming LLMs
by: Shen, Qingchao, et al.
Published: (2026)
by: Shen, Qingchao, et al.
Published: (2026)
MASKDROID: Robust Android Malware Detection with Masked Graph Representations
by: Zheng, Jingnan, et al.
Published: (2024)
by: Zheng, Jingnan, et al.
Published: (2024)
Mind the Gap: Evaluating LLMs for High-Level Malicious Package Detection vs. Fine-Grained Indicator Identification
by: Ryan, Ahmed, et al.
Published: (2026)
by: Ryan, Ahmed, et al.
Published: (2026)
Scam2Prompt: A Scalable Framework for Auditing Malicious Scam Endpoints in Production LLMs
by: Chen, Zhiyang, et al.
Published: (2025)
by: Chen, Zhiyang, et al.
Published: (2025)
Detecting Android Malware by Visualizing App Behaviors from Multiple Complementary Views
by: Meng, Zhaoyi, et al.
Published: (2024)
by: Meng, Zhaoyi, et al.
Published: (2024)
Chimera: Harnessing Multi-Agent LLMs for Automatic Insider Threat Simulation
by: Yu, Jiongchi, et al.
Published: (2025)
by: Yu, Jiongchi, et al.
Published: (2025)
FCGHunter: Towards Evaluating Robustness of Graph-Based Android Malware Detection
by: Song, Shiwen, et al.
Published: (2025)
by: Song, Shiwen, et al.
Published: (2025)
Identifying Adversary Tactics and Techniques in Malware Binaries with an LLM Agent
by: Xuan, Zhou, et al.
Published: (2026)
by: Xuan, Zhou, et al.
Published: (2026)
Learning Temporal Invariance in Android Malware Detectors
by: Zheng, Xinran, et al.
Published: (2025)
by: Zheng, Xinran, et al.
Published: (2025)
MALSIGHT: Exploring Malicious Source Code and Benign Pseudocode for Iterative Binary Malware Summarization
by: Lu, Haolang, et al.
Published: (2024)
by: Lu, Haolang, et al.
Published: (2024)
DetectBERT: Towards Full App-Level Representation Learning to Detect Android Malware
by: Sun, Tiezhu, et al.
Published: (2024)
by: Sun, Tiezhu, et al.
Published: (2024)
Knowdit: Agentic Smart Contract Vulnerability Detection with Auditing Knowledge Summarization
by: Kong, Ziqiao, et al.
Published: (2026)
by: Kong, Ziqiao, et al.
Published: (2026)
Learning to Focus: Context Extraction for Efficient Code Vulnerability Detection with Language Models
by: Zheng, Xinran, et al.
Published: (2025)
by: Zheng, Xinran, et al.
Published: (2025)
Six Million (Suspected) Fake Stars in GitHub: A Growing Spiral of Popularity Contests, Spams, and Malware
by: He, Hao, et al.
Published: (2024)
by: He, Hao, et al.
Published: (2024)
BIDO: An Out-Of-Distribution Resistant Image-based Malware Detector
by: Wang, Wei, et al.
Published: (2025)
by: Wang, Wei, et al.
Published: (2025)
SmartOracle: Generating Smart Contract Oracle via Fine-Grained Invariant Detection
by: Su, Jianzhong, et al.
Published: (2024)
by: Su, Jianzhong, et al.
Published: (2024)
AEGIS: From Clues to Verdicts -- Graph-Guided Deep Vulnerability Reasoning via Dialectics and Meta-Auditing
by: Fang, Sen, et al.
Published: (2026)
by: Fang, Sen, et al.
Published: (2026)
Do Fine-Tuned LLMs Understand Vulnerabilities? An Investigation into the Semantic Trap
by: Huang, Feiyang, et al.
Published: (2026)
by: Huang, Feiyang, et al.
Published: (2026)
Automatic Red Teaming LLM-based Agents with Model Context Protocol Tools
by: He, Ping, et al.
Published: (2025)
by: He, Ping, et al.
Published: (2025)
LLM4Vuln: A Unified Evaluation Framework for Decoupling and Enhancing LLMs' Vulnerability Reasoning
by: Sun, Yuqiang, et al.
Published: (2024)
by: Sun, Yuqiang, et al.
Published: (2024)
Static Semantics Reconstruction for Enhancing JavaScript-WebAssembly Multilingual Malware Detection
by: Xia, Yifan, et al.
Published: (2023)
by: Xia, Yifan, et al.
Published: (2023)
Automatically Generating Rules of Malicious Software Packages via Large Language Model
by: Zhang, XiangRui, et al.
Published: (2025)
by: Zhang, XiangRui, et al.
Published: (2025)
Evaluating LLMs for One-Shot Patching of Real and Artificial Vulnerabilities
by: Garg, Aayush, et al.
Published: (2025)
by: Garg, Aayush, et al.
Published: (2025)
A Study of Malware Prevention in Linux Distributions
by: Vu, Duc-Ly, et al.
Published: (2024)
by: Vu, Duc-Ly, et al.
Published: (2024)
RiskTagger: An LLM-based Agent for Automatic Annotation of Web3 Crypto Money Laundering Behaviors
by: Lin, Dan, et al.
Published: (2025)
by: Lin, Dan, et al.
Published: (2025)
Accelerating Automatic Program Repair with Dual Retrieval-Augmented Fine-Tuning and Patch Generation on Large Language Models
by: Guo, Hanyang, et al.
Published: (2025)
by: Guo, Hanyang, et al.
Published: (2025)
Towards Secure Logging: Characterizing and Benchmarking Logging Code Security Issues with LLMs
by: Yuan, He Yang, et al.
Published: (2026)
by: Yuan, He Yang, et al.
Published: (2026)
An Adversarial Robust Behavior Sequence Anomaly Detection Approach Based on Critical Behavior Unit Learning
by: Zhan, Dongyang, et al.
Published: (2025)
by: Zhan, Dongyang, et al.
Published: (2025)
From LLMs to Agents: A Comparative Evaluation of LLMs and LLM-based Agents in Security Patch Detection
by: Han, Junxiao, et al.
Published: (2025)
by: Han, Junxiao, et al.
Published: (2025)
Decaf: Improving Neural Decompilation with Automatic Feedback and Search
by: Shypula, Alexander, et al.
Published: (2026)
by: Shypula, Alexander, et al.
Published: (2026)
MARD: A Multi-Agent Framework for Robust Android Malware Detection
by: Zeng, Xueying, et al.
Published: (2026)
by: Zeng, Xueying, et al.
Published: (2026)
A Time Series Analysis of Malware Uploads to Programming Language Ecosystems
by: Ruohonen, Jukka, et al.
Published: (2025)
by: Ruohonen, Jukka, et al.
Published: (2025)
Track and Trace: Automatically Uncovering Cross-chain Transactions in the Multi-blockchain Ecosystems
by: Lin, Dan, et al.
Published: (2025)
by: Lin, Dan, et al.
Published: (2025)
SafeToolBench: Pioneering a Prospective Benchmark to Evaluating Tool Utilization Safety in LLMs
by: Xia, Hongfei, et al.
Published: (2025)
by: Xia, Hongfei, et al.
Published: (2025)
Similar Items
-
LAMD: Context-driven Android Malware Detection and Classification with LLMs
by: Qian, Xingzhi, et al.
Published: (2025) -
Defending against Adversarial Malware Attacks on ML-based Android Malware Detection Systems
by: He, Ping, et al.
Published: (2025) -
On the Security Risks of ML-based Malware Detection Systems: A Survey
by: He, Ping, et al.
Published: (2025) -
Veritas: A Semantically Grounded Agentic Framework for Memory Corruption Vulnerability Detection in Binaries
by: Zheng, Xinran, et al.
Published: (2026) -
On Benchmarking Code LLMs for Android Malware Analysis
by: He, Yiling, et al.
Published: (2025)