Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Hao, Li, Hanchen, Mang, Qiuyang, Cheung, Alvin, Sen, Koushik, Song, Dawn |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SCGAgent: Recreating the Benefits of Reasoning Models for Secure Code Generation with Agentic Workflows
von: Saul, Rebecca, et al.
Veröffentlicht: (2025)
von: Saul, Rebecca, et al.
Veröffentlicht: (2025)
Explainable Android Malware Detection and Malicious Code Localization Using Graph Attention
von: Ipek, Merve Cigdem, et al.
Veröffentlicht: (2025)
von: Ipek, Merve Cigdem, et al.
Veröffentlicht: (2025)
The Blockchain Imitation Game
von: Qin, Kaihua, et al.
Veröffentlicht: (2023)
von: Qin, Kaihua, et al.
Veröffentlicht: (2023)
Multi-label Classification for Android Malware Based on Active Learning
von: Qiao, Qijing, et al.
Veröffentlicht: (2024)
von: Qiao, Qijing, et al.
Veröffentlicht: (2024)
FCGHunter: Towards Evaluating Robustness of Graph-Based Android Malware Detection
von: Song, Shiwen, et al.
Veröffentlicht: (2025)
von: Song, Shiwen, et al.
Veröffentlicht: (2025)
Do Android App Developers Accurately Report Collection of Privacy-Related Data?
von: Khedkar, Mugdha, et al.
Veröffentlicht: (2024)
von: Khedkar, Mugdha, et al.
Veröffentlicht: (2024)
Benchmarking Android Malware Detection: Traditional vs. Deep Learning Models
von: Liu, Guojun, et al.
Veröffentlicht: (2025)
von: Liu, Guojun, et al.
Veröffentlicht: (2025)
VeilAudit: Breaking the Deadlock Between Privacy and Accountability Across Blockchains
von: Qiao, Minhao, et al.
Veröffentlicht: (2025)
von: Qiao, Minhao, et al.
Veröffentlicht: (2025)
Breaking the Code: Security Assessment of AI Code Agents Through Systematic Jailbreaking Attacks
von: Saha, Shoumik, et al.
Veröffentlicht: (2025)
von: Saha, Shoumik, et al.
Veröffentlicht: (2025)
SC-Bench: A Large-Scale Dataset for Smart Contract Auditing
von: Xia, Shihao, et al.
Veröffentlicht: (2024)
von: Xia, Shihao, et al.
Veröffentlicht: (2024)
ThreatIntel-Andro: Expert-Verified Benchmarking for Robust Android Malware Research
von: Bai, Hongpeng, et al.
Veröffentlicht: (2025)
von: Bai, Hongpeng, et al.
Veröffentlicht: (2025)
SpoofTrackBench: Interpretable AI for Spoof-Aware UAV Tracking and Benchmarking
von: Le, Van, et al.
Veröffentlicht: (2025)
von: Le, Van, et al.
Veröffentlicht: (2025)
On Benchmarking Code LLMs for Android Malware Analysis
von: He, Yiling, et al.
Veröffentlicht: (2025)
von: He, Yiling, et al.
Veröffentlicht: (2025)
AudAgent: Automated Auditing of Privacy Policy Compliance in AI Agents
von: Zheng, Ye, et al.
Veröffentlicht: (2025)
von: Zheng, Ye, et al.
Veröffentlicht: (2025)
Cybersecurity AI Benchmark (CAIBench): A Meta-Benchmark for Evaluating Cybersecurity AI Agents
von: Sanz-Gómez, María, et al.
Veröffentlicht: (2025)
von: Sanz-Gómez, María, et al.
Veröffentlicht: (2025)
MalTool: Malicious Tool Attacks on LLM Agents
von: Hu, Yuepeng, et al.
Veröffentlicht: (2026)
von: Hu, Yuepeng, et al.
Veröffentlicht: (2026)
Progent: Securing AI Agents with Privilege Control
von: Shi, Tianneng, et al.
Veröffentlicht: (2025)
von: Shi, Tianneng, et al.
Veröffentlicht: (2025)
LogJack: Indirect Prompt Injection Through Cloud Logs Against LLM Debugging Agents
von: Shah, Harsh
Veröffentlicht: (2026)
von: Shah, Harsh
Veröffentlicht: (2026)
Are You Getting What You Pay For? Auditing Model Substitution in LLM APIs
von: Cai, Will, et al.
Veröffentlicht: (2025)
von: Cai, Will, et al.
Veröffentlicht: (2025)
AEGIS: No Tool Call Left Unchecked -- A Pre-Execution Firewall and Audit Layer for AI Agents
von: Yuan, Aojie, et al.
Veröffentlicht: (2026)
von: Yuan, Aojie, et al.
Veröffentlicht: (2026)
How Far are App Secrets from Being Stolen? A Case Study on Android
von: Wei, Lili, et al.
Veröffentlicht: (2025)
von: Wei, Lili, et al.
Veröffentlicht: (2025)
RoboJailBench: Benchmarking Adversarial Attacks and Defenses in Embodied Robotic Agents
von: Yeke, Doguhuan, et al.
Veröffentlicht: (2026)
von: Yeke, Doguhuan, et al.
Veröffentlicht: (2026)
GhostEI-Bench: Do Mobile Agents Resilience to Environmental Injection in Dynamic On-Device Environments?
von: Chen, Chiyu, et al.
Veröffentlicht: (2025)
von: Chen, Chiyu, et al.
Veröffentlicht: (2025)
ShieldShare: Building a VPN-backed Android Hotspot for Secure Internet Sharing with Per-User Traffic Accounting
von: Edorh, Carlos Semeho, et al.
Veröffentlicht: (2026)
von: Edorh, Carlos Semeho, et al.
Veröffentlicht: (2026)
Dockerized Android: a container-based platform to build mobile Android scenarios for Cyber Ranges
von: Capone, Daniele, et al.
Veröffentlicht: (2022)
von: Capone, Daniele, et al.
Veröffentlicht: (2022)
AutoPenBench: Benchmarking Generative Agents for Penetration Testing
von: Gioacchini, Luca, et al.
Veröffentlicht: (2024)
von: Gioacchini, Luca, et al.
Veröffentlicht: (2024)
Securing LLM Agents Need Intent-to-Execution Integrity
von: Qu, Wenjie, et al.
Veröffentlicht: (2026)
von: Qu, Wenjie, et al.
Veröffentlicht: (2026)
Enhancing Android Malware Detection: The Influence of ChatGPT on Decision-centric Task
von: Li, Yao, et al.
Veröffentlicht: (2024)
von: Li, Yao, et al.
Veröffentlicht: (2024)
Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents
von: Zhang, Hanrong, et al.
Veröffentlicht: (2024)
von: Zhang, Hanrong, et al.
Veröffentlicht: (2024)
Advancing Android Privacy Assessments with Automation
von: Khedkar, Mugdha, et al.
Veröffentlicht: (2024)
von: Khedkar, Mugdha, et al.
Veröffentlicht: (2024)
HarmfulSkillBench: How Do Harmful Skills Weaponize Your Agents?
von: Jiang, Yukun, et al.
Veröffentlicht: (2026)
von: Jiang, Yukun, et al.
Veröffentlicht: (2026)
ExploitBench: A Capability Ladder Benchmark for LLM Cybersecurity Agents
von: Lee, Seunghyun, et al.
Veröffentlicht: (2026)
von: Lee, Seunghyun, et al.
Veröffentlicht: (2026)
SLEIGHT-Bench: A Benchmark of Evasion Attacks Against Agent Monitors
von: Najt, Elle, et al.
Veröffentlicht: (2026)
von: Najt, Elle, et al.
Veröffentlicht: (2026)
LAMDA: A Longitudinal Android Malware Benchmark for Concept Drift Analysis
von: Haque, Md Ahsanul, et al.
Veröffentlicht: (2025)
von: Haque, Md Ahsanul, et al.
Veröffentlicht: (2025)
Who Tests the Testers? Systematic Enumeration and Coverage Audit of LLM Agent Tool Call Safety
von: Chen, Xuan, et al.
Veröffentlicht: (2026)
von: Chen, Xuan, et al.
Veröffentlicht: (2026)
Auditing M-LLMs for Privacy Risks: A Synthetic Benchmark and Evaluation Framework
von: Li, Junhao, et al.
Veröffentlicht: (2025)
von: Li, Junhao, et al.
Veröffentlicht: (2025)
You Told Me to Do It: Measuring Instructional Text-induced Private Data Leakage in LLM Agents
von: Kao, Ching-Yu, et al.
Veröffentlicht: (2026)
von: Kao, Ching-Yu, et al.
Veröffentlicht: (2026)
Do Agents Dream of Root Shells? Partial-Credit Evaluation of LLM Agents in Capture the Flag Challenges
von: Al-Kaswan, Ali, et al.
Veröffentlicht: (2026)
von: Al-Kaswan, Ali, et al.
Veröffentlicht: (2026)
Breaking Bad: Interpretability-Based Safety Audits of State-of-the-Art LLMs
von: Agarwal, Krishiv, et al.
Veröffentlicht: (2026)
von: Agarwal, Krishiv, et al.
Veröffentlicht: (2026)
MCPSecBench: A Systematic Security Benchmark and Playground for Testing Model Context Protocols
von: Yang, Yixuan, et al.
Veröffentlicht: (2025)
von: Yang, Yixuan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
SCGAgent: Recreating the Benefits of Reasoning Models for Secure Code Generation with Agentic Workflows
von: Saul, Rebecca, et al.
Veröffentlicht: (2025) -
Explainable Android Malware Detection and Malicious Code Localization Using Graph Attention
von: Ipek, Merve Cigdem, et al.
Veröffentlicht: (2025) -
The Blockchain Imitation Game
von: Qin, Kaihua, et al.
Veröffentlicht: (2023) -
Multi-label Classification for Android Malware Based on Active Learning
von: Qiao, Qijing, et al.
Veröffentlicht: (2024) -
FCGHunter: Towards Evaluating Robustness of Graph-Based Android Malware Detection
von: Song, Shiwen, et al.
Veröffentlicht: (2025)