Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack
Fuente:
arXiv
Guardado en:
| Autores principales: | Wang, Hao, Li, Hanchen, Mang, Qiuyang, Cheung, Alvin, Sen, Koushik, Song, Dawn |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
SCGAgent: Recreating the Benefits of Reasoning Models for Secure Code Generation with Agentic Workflows
por: Saul, Rebecca, et al.
Publicado: (2025)
por: Saul, Rebecca, et al.
Publicado: (2025)
Explainable Android Malware Detection and Malicious Code Localization Using Graph Attention
por: Ipek, Merve Cigdem, et al.
Publicado: (2025)
por: Ipek, Merve Cigdem, et al.
Publicado: (2025)
The Blockchain Imitation Game
por: Qin, Kaihua, et al.
Publicado: (2023)
por: Qin, Kaihua, et al.
Publicado: (2023)
Multi-label Classification for Android Malware Based on Active Learning
por: Qiao, Qijing, et al.
Publicado: (2024)
por: Qiao, Qijing, et al.
Publicado: (2024)
FCGHunter: Towards Evaluating Robustness of Graph-Based Android Malware Detection
por: Song, Shiwen, et al.
Publicado: (2025)
por: Song, Shiwen, et al.
Publicado: (2025)
Do Android App Developers Accurately Report Collection of Privacy-Related Data?
por: Khedkar, Mugdha, et al.
Publicado: (2024)
por: Khedkar, Mugdha, et al.
Publicado: (2024)
Benchmarking Android Malware Detection: Traditional vs. Deep Learning Models
por: Liu, Guojun, et al.
Publicado: (2025)
por: Liu, Guojun, et al.
Publicado: (2025)
VeilAudit: Breaking the Deadlock Between Privacy and Accountability Across Blockchains
por: Qiao, Minhao, et al.
Publicado: (2025)
por: Qiao, Minhao, et al.
Publicado: (2025)
Breaking the Code: Security Assessment of AI Code Agents Through Systematic Jailbreaking Attacks
por: Saha, Shoumik, et al.
Publicado: (2025)
por: Saha, Shoumik, et al.
Publicado: (2025)
SC-Bench: A Large-Scale Dataset for Smart Contract Auditing
por: Xia, Shihao, et al.
Publicado: (2024)
por: Xia, Shihao, et al.
Publicado: (2024)
ThreatIntel-Andro: Expert-Verified Benchmarking for Robust Android Malware Research
por: Bai, Hongpeng, et al.
Publicado: (2025)
por: Bai, Hongpeng, et al.
Publicado: (2025)
SpoofTrackBench: Interpretable AI for Spoof-Aware UAV Tracking and Benchmarking
por: Le, Van, et al.
Publicado: (2025)
por: Le, Van, et al.
Publicado: (2025)
On Benchmarking Code LLMs for Android Malware Analysis
por: He, Yiling, et al.
Publicado: (2025)
por: He, Yiling, et al.
Publicado: (2025)
AudAgent: Automated Auditing of Privacy Policy Compliance in AI Agents
por: Zheng, Ye, et al.
Publicado: (2025)
por: Zheng, Ye, et al.
Publicado: (2025)
Cybersecurity AI Benchmark (CAIBench): A Meta-Benchmark for Evaluating Cybersecurity AI Agents
por: Sanz-Gómez, María, et al.
Publicado: (2025)
por: Sanz-Gómez, María, et al.
Publicado: (2025)
MalTool: Malicious Tool Attacks on LLM Agents
por: Hu, Yuepeng, et al.
Publicado: (2026)
por: Hu, Yuepeng, et al.
Publicado: (2026)
Progent: Securing AI Agents with Privilege Control
por: Shi, Tianneng, et al.
Publicado: (2025)
por: Shi, Tianneng, et al.
Publicado: (2025)
LogJack: Indirect Prompt Injection Through Cloud Logs Against LLM Debugging Agents
por: Shah, Harsh
Publicado: (2026)
por: Shah, Harsh
Publicado: (2026)
Are You Getting What You Pay For? Auditing Model Substitution in LLM APIs
por: Cai, Will, et al.
Publicado: (2025)
por: Cai, Will, et al.
Publicado: (2025)
AEGIS: No Tool Call Left Unchecked -- A Pre-Execution Firewall and Audit Layer for AI Agents
por: Yuan, Aojie, et al.
Publicado: (2026)
por: Yuan, Aojie, et al.
Publicado: (2026)
How Far are App Secrets from Being Stolen? A Case Study on Android
por: Wei, Lili, et al.
Publicado: (2025)
por: Wei, Lili, et al.
Publicado: (2025)
RoboJailBench: Benchmarking Adversarial Attacks and Defenses in Embodied Robotic Agents
por: Yeke, Doguhuan, et al.
Publicado: (2026)
por: Yeke, Doguhuan, et al.
Publicado: (2026)
GhostEI-Bench: Do Mobile Agents Resilience to Environmental Injection in Dynamic On-Device Environments?
por: Chen, Chiyu, et al.
Publicado: (2025)
por: Chen, Chiyu, et al.
Publicado: (2025)
ShieldShare: Building a VPN-backed Android Hotspot for Secure Internet Sharing with Per-User Traffic Accounting
por: Edorh, Carlos Semeho, et al.
Publicado: (2026)
por: Edorh, Carlos Semeho, et al.
Publicado: (2026)
Dockerized Android: a container-based platform to build mobile Android scenarios for Cyber Ranges
por: Capone, Daniele, et al.
Publicado: (2022)
por: Capone, Daniele, et al.
Publicado: (2022)
AutoPenBench: Benchmarking Generative Agents for Penetration Testing
por: Gioacchini, Luca, et al.
Publicado: (2024)
por: Gioacchini, Luca, et al.
Publicado: (2024)
Securing LLM Agents Need Intent-to-Execution Integrity
por: Qu, Wenjie, et al.
Publicado: (2026)
por: Qu, Wenjie, et al.
Publicado: (2026)
Enhancing Android Malware Detection: The Influence of ChatGPT on Decision-centric Task
por: Li, Yao, et al.
Publicado: (2024)
por: Li, Yao, et al.
Publicado: (2024)
Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents
por: Zhang, Hanrong, et al.
Publicado: (2024)
por: Zhang, Hanrong, et al.
Publicado: (2024)
Advancing Android Privacy Assessments with Automation
por: Khedkar, Mugdha, et al.
Publicado: (2024)
por: Khedkar, Mugdha, et al.
Publicado: (2024)
HarmfulSkillBench: How Do Harmful Skills Weaponize Your Agents?
por: Jiang, Yukun, et al.
Publicado: (2026)
por: Jiang, Yukun, et al.
Publicado: (2026)
ExploitBench: A Capability Ladder Benchmark for LLM Cybersecurity Agents
por: Lee, Seunghyun, et al.
Publicado: (2026)
por: Lee, Seunghyun, et al.
Publicado: (2026)
SLEIGHT-Bench: A Benchmark of Evasion Attacks Against Agent Monitors
por: Najt, Elle, et al.
Publicado: (2026)
por: Najt, Elle, et al.
Publicado: (2026)
LAMDA: A Longitudinal Android Malware Benchmark for Concept Drift Analysis
por: Haque, Md Ahsanul, et al.
Publicado: (2025)
por: Haque, Md Ahsanul, et al.
Publicado: (2025)
Who Tests the Testers? Systematic Enumeration and Coverage Audit of LLM Agent Tool Call Safety
por: Chen, Xuan, et al.
Publicado: (2026)
por: Chen, Xuan, et al.
Publicado: (2026)
Auditing M-LLMs for Privacy Risks: A Synthetic Benchmark and Evaluation Framework
por: Li, Junhao, et al.
Publicado: (2025)
por: Li, Junhao, et al.
Publicado: (2025)
You Told Me to Do It: Measuring Instructional Text-induced Private Data Leakage in LLM Agents
por: Kao, Ching-Yu, et al.
Publicado: (2026)
por: Kao, Ching-Yu, et al.
Publicado: (2026)
Do Agents Dream of Root Shells? Partial-Credit Evaluation of LLM Agents in Capture the Flag Challenges
por: Al-Kaswan, Ali, et al.
Publicado: (2026)
por: Al-Kaswan, Ali, et al.
Publicado: (2026)
Breaking Bad: Interpretability-Based Safety Audits of State-of-the-Art LLMs
por: Agarwal, Krishiv, et al.
Publicado: (2026)
por: Agarwal, Krishiv, et al.
Publicado: (2026)
MCPSecBench: A Systematic Security Benchmark and Playground for Testing Model Context Protocols
por: Yang, Yixuan, et al.
Publicado: (2025)
por: Yang, Yixuan, et al.
Publicado: (2025)
Ejemplares similares
-
SCGAgent: Recreating the Benefits of Reasoning Models for Secure Code Generation with Agentic Workflows
por: Saul, Rebecca, et al.
Publicado: (2025) -
Explainable Android Malware Detection and Malicious Code Localization Using Graph Attention
por: Ipek, Merve Cigdem, et al.
Publicado: (2025) -
The Blockchain Imitation Game
por: Qin, Kaihua, et al.
Publicado: (2023) -
Multi-label Classification for Android Malware Based on Active Learning
por: Qiao, Qijing, et al.
Publicado: (2024) -
FCGHunter: Towards Evaluating Robustness of Graph-Based Android Malware Detection
por: Song, Shiwen, et al.
Publicado: (2025)