Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Hao, Li, Hanchen, Mang, Qiuyang, Cheung, Alvin, Sen, Koushik, Song, Dawn |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SCGAgent: Recreating the Benefits of Reasoning Models for Secure Code Generation with Agentic Workflows
by: Saul, Rebecca, et al.
Published: (2025)
by: Saul, Rebecca, et al.
Published: (2025)
Explainable Android Malware Detection and Malicious Code Localization Using Graph Attention
by: Ipek, Merve Cigdem, et al.
Published: (2025)
by: Ipek, Merve Cigdem, et al.
Published: (2025)
The Blockchain Imitation Game
by: Qin, Kaihua, et al.
Published: (2023)
by: Qin, Kaihua, et al.
Published: (2023)
Multi-label Classification for Android Malware Based on Active Learning
by: Qiao, Qijing, et al.
Published: (2024)
by: Qiao, Qijing, et al.
Published: (2024)
FCGHunter: Towards Evaluating Robustness of Graph-Based Android Malware Detection
by: Song, Shiwen, et al.
Published: (2025)
by: Song, Shiwen, et al.
Published: (2025)
Do Android App Developers Accurately Report Collection of Privacy-Related Data?
by: Khedkar, Mugdha, et al.
Published: (2024)
by: Khedkar, Mugdha, et al.
Published: (2024)
Benchmarking Android Malware Detection: Traditional vs. Deep Learning Models
by: Liu, Guojun, et al.
Published: (2025)
by: Liu, Guojun, et al.
Published: (2025)
VeilAudit: Breaking the Deadlock Between Privacy and Accountability Across Blockchains
by: Qiao, Minhao, et al.
Published: (2025)
by: Qiao, Minhao, et al.
Published: (2025)
Breaking the Code: Security Assessment of AI Code Agents Through Systematic Jailbreaking Attacks
by: Saha, Shoumik, et al.
Published: (2025)
by: Saha, Shoumik, et al.
Published: (2025)
SC-Bench: A Large-Scale Dataset for Smart Contract Auditing
by: Xia, Shihao, et al.
Published: (2024)
by: Xia, Shihao, et al.
Published: (2024)
ThreatIntel-Andro: Expert-Verified Benchmarking for Robust Android Malware Research
by: Bai, Hongpeng, et al.
Published: (2025)
by: Bai, Hongpeng, et al.
Published: (2025)
SpoofTrackBench: Interpretable AI for Spoof-Aware UAV Tracking and Benchmarking
by: Le, Van, et al.
Published: (2025)
by: Le, Van, et al.
Published: (2025)
On Benchmarking Code LLMs for Android Malware Analysis
by: He, Yiling, et al.
Published: (2025)
by: He, Yiling, et al.
Published: (2025)
AudAgent: Automated Auditing of Privacy Policy Compliance in AI Agents
by: Zheng, Ye, et al.
Published: (2025)
by: Zheng, Ye, et al.
Published: (2025)
Cybersecurity AI Benchmark (CAIBench): A Meta-Benchmark for Evaluating Cybersecurity AI Agents
by: Sanz-Gómez, María, et al.
Published: (2025)
by: Sanz-Gómez, María, et al.
Published: (2025)
MalTool: Malicious Tool Attacks on LLM Agents
by: Hu, Yuepeng, et al.
Published: (2026)
by: Hu, Yuepeng, et al.
Published: (2026)
Progent: Securing AI Agents with Privilege Control
by: Shi, Tianneng, et al.
Published: (2025)
by: Shi, Tianneng, et al.
Published: (2025)
LogJack: Indirect Prompt Injection Through Cloud Logs Against LLM Debugging Agents
by: Shah, Harsh
Published: (2026)
by: Shah, Harsh
Published: (2026)
Are You Getting What You Pay For? Auditing Model Substitution in LLM APIs
by: Cai, Will, et al.
Published: (2025)
by: Cai, Will, et al.
Published: (2025)
AEGIS: No Tool Call Left Unchecked -- A Pre-Execution Firewall and Audit Layer for AI Agents
by: Yuan, Aojie, et al.
Published: (2026)
by: Yuan, Aojie, et al.
Published: (2026)
How Far are App Secrets from Being Stolen? A Case Study on Android
by: Wei, Lili, et al.
Published: (2025)
by: Wei, Lili, et al.
Published: (2025)
RoboJailBench: Benchmarking Adversarial Attacks and Defenses in Embodied Robotic Agents
by: Yeke, Doguhuan, et al.
Published: (2026)
by: Yeke, Doguhuan, et al.
Published: (2026)
GhostEI-Bench: Do Mobile Agents Resilience to Environmental Injection in Dynamic On-Device Environments?
by: Chen, Chiyu, et al.
Published: (2025)
by: Chen, Chiyu, et al.
Published: (2025)
ShieldShare: Building a VPN-backed Android Hotspot for Secure Internet Sharing with Per-User Traffic Accounting
by: Edorh, Carlos Semeho, et al.
Published: (2026)
by: Edorh, Carlos Semeho, et al.
Published: (2026)
Dockerized Android: a container-based platform to build mobile Android scenarios for Cyber Ranges
by: Capone, Daniele, et al.
Published: (2022)
by: Capone, Daniele, et al.
Published: (2022)
AutoPenBench: Benchmarking Generative Agents for Penetration Testing
by: Gioacchini, Luca, et al.
Published: (2024)
by: Gioacchini, Luca, et al.
Published: (2024)
Securing LLM Agents Need Intent-to-Execution Integrity
by: Qu, Wenjie, et al.
Published: (2026)
by: Qu, Wenjie, et al.
Published: (2026)
Enhancing Android Malware Detection: The Influence of ChatGPT on Decision-centric Task
by: Li, Yao, et al.
Published: (2024)
by: Li, Yao, et al.
Published: (2024)
Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents
by: Zhang, Hanrong, et al.
Published: (2024)
by: Zhang, Hanrong, et al.
Published: (2024)
Advancing Android Privacy Assessments with Automation
by: Khedkar, Mugdha, et al.
Published: (2024)
by: Khedkar, Mugdha, et al.
Published: (2024)
HarmfulSkillBench: How Do Harmful Skills Weaponize Your Agents?
by: Jiang, Yukun, et al.
Published: (2026)
by: Jiang, Yukun, et al.
Published: (2026)
ExploitBench: A Capability Ladder Benchmark for LLM Cybersecurity Agents
by: Lee, Seunghyun, et al.
Published: (2026)
by: Lee, Seunghyun, et al.
Published: (2026)
SLEIGHT-Bench: A Benchmark of Evasion Attacks Against Agent Monitors
by: Najt, Elle, et al.
Published: (2026)
by: Najt, Elle, et al.
Published: (2026)
LAMDA: A Longitudinal Android Malware Benchmark for Concept Drift Analysis
by: Haque, Md Ahsanul, et al.
Published: (2025)
by: Haque, Md Ahsanul, et al.
Published: (2025)
Who Tests the Testers? Systematic Enumeration and Coverage Audit of LLM Agent Tool Call Safety
by: Chen, Xuan, et al.
Published: (2026)
by: Chen, Xuan, et al.
Published: (2026)
Auditing M-LLMs for Privacy Risks: A Synthetic Benchmark and Evaluation Framework
by: Li, Junhao, et al.
Published: (2025)
by: Li, Junhao, et al.
Published: (2025)
You Told Me to Do It: Measuring Instructional Text-induced Private Data Leakage in LLM Agents
by: Kao, Ching-Yu, et al.
Published: (2026)
by: Kao, Ching-Yu, et al.
Published: (2026)
Do Agents Dream of Root Shells? Partial-Credit Evaluation of LLM Agents in Capture the Flag Challenges
by: Al-Kaswan, Ali, et al.
Published: (2026)
by: Al-Kaswan, Ali, et al.
Published: (2026)
Breaking Bad: Interpretability-Based Safety Audits of State-of-the-Art LLMs
by: Agarwal, Krishiv, et al.
Published: (2026)
by: Agarwal, Krishiv, et al.
Published: (2026)
MCPSecBench: A Systematic Security Benchmark and Playground for Testing Model Context Protocols
by: Yang, Yixuan, et al.
Published: (2025)
by: Yang, Yixuan, et al.
Published: (2025)
Similar Items
-
SCGAgent: Recreating the Benefits of Reasoning Models for Secure Code Generation with Agentic Workflows
by: Saul, Rebecca, et al.
Published: (2025) -
Explainable Android Malware Detection and Malicious Code Localization Using Graph Attention
by: Ipek, Merve Cigdem, et al.
Published: (2025) -
The Blockchain Imitation Game
by: Qin, Kaihua, et al.
Published: (2023) -
Multi-label Classification for Android Malware Based on Active Learning
by: Qiao, Qijing, et al.
Published: (2024) -
FCGHunter: Towards Evaluating Robustness of Graph-Based Android Malware Detection
by: Song, Shiwen, et al.
Published: (2025)