ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Zhun, Schiller, Nico, Li, Hongwei, Narayana, Srijiith Sesha, Nasr, Milad, Carlini, Nicholas, Qi, Xiangyu, Wallace, Eric, Bursztein, Elie, Invernizzi, Luca, Thomas, Kurt, Shoshitaishvili, Yan, Guo, Wenbo, He, Jingxuan, Holz, Thorsten, Song, Dawn |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Evaluating the Robustness of a Production Malware Detection System to Transferable Adversarial Attacks
von: Nasr, Milad, et al.
Veröffentlicht: (2025)
von: Nasr, Milad, et al.
Veröffentlicht: (2025)
Remote Timing Attacks on Efficient Language Model Inference
von: Carlini, Nicholas, et al.
Veröffentlicht: (2024)
von: Carlini, Nicholas, et al.
Veröffentlicht: (2024)
CyberGym: Evaluating AI Agents' Real-World Cybersecurity Capabilities at Scale
von: Wang, Zhun, et al.
Veröffentlicht: (2025)
von: Wang, Zhun, et al.
Veröffentlicht: (2025)
Progent: Securing AI Agents with Privilege Control
von: Shi, Tianneng, et al.
Veröffentlicht: (2025)
von: Shi, Tianneng, et al.
Veröffentlicht: (2025)
Generalized Power Attacks against Crypto Hardware using Long-Range Deep Learning
von: Bursztein, Elie, et al.
Veröffentlicht: (2023)
von: Bursztein, Elie, et al.
Veröffentlicht: (2023)
VulnLLM-R: Specialized Reasoning LLM with Agent Scaffold for Vulnerability Detection
von: Nie, Yuzhou, et al.
Veröffentlicht: (2025)
von: Nie, Yuzhou, et al.
Veröffentlicht: (2025)
Query-Based Adversarial Prompt Generation
von: Hayase, Jonathan, et al.
Veröffentlicht: (2024)
von: Hayase, Jonathan, et al.
Veröffentlicht: (2024)
AutoAdvExBench: Benchmarking autonomous exploitation of adversarial example defenses
von: Carlini, Nicholas, et al.
Veröffentlicht: (2025)
von: Carlini, Nicholas, et al.
Veröffentlicht: (2025)
Frontier AI's Impact on the Cybersecurity Landscape
von: Potter, Yujin, et al.
Veröffentlicht: (2025)
von: Potter, Yujin, et al.
Veröffentlicht: (2025)
Magika: AI-Powered Content-Type Detection
von: Fratantonio, Yanick, et al.
Veröffentlicht: (2024)
von: Fratantonio, Yanick, et al.
Veröffentlicht: (2024)
RETVec: Resilient and Efficient Text Vectorizer
von: Bursztein, Elie, et al.
Veröffentlicht: (2023)
von: Bursztein, Elie, et al.
Veröffentlicht: (2023)
Avoiding Generative Model Writer's Block With Embedding Nudging
von: Zand, Ali, et al.
Veröffentlicht: (2024)
von: Zand, Ali, et al.
Veröffentlicht: (2024)
Video-Native Koopman Operator Learning for Active Wildfire Risk Mapping
von: Achyuta, Sesha
Veröffentlicht: (2025)
von: Achyuta, Sesha
Veröffentlicht: (2025)
VictualMark™: Building a Verified, Physics-Backed Food Economy on Top of DataCulture, 512TD, BRIXBOX, and NutriBox
von: Achyuta, Sesha
Veröffentlicht: (2025)
von: Achyuta, Sesha
Veröffentlicht: (2025)
Privacy Side Channels in Machine Learning Systems
von: Debenedetti, Edoardo, et al.
Veröffentlicht: (2023)
von: Debenedetti, Edoardo, et al.
Veröffentlicht: (2023)
Gym-Anything: Turn any Software into an Agent Environment
von: Aggarwal, Pranjal, et al.
Veröffentlicht: (2026)
von: Aggarwal, Pranjal, et al.
Veröffentlicht: (2026)
Pattern Enhanced Multi-Turn Jailbreaking: Exploiting Structural Vulnerabilities in Large Language Models
von: Nihal, Ragib Amin, et al.
Veröffentlicht: (2025)
von: Nihal, Ragib Amin, et al.
Veröffentlicht: (2025)
AbideGym: Turning Static RL Worlds into Adaptive Challenges
von: Aryan, Abi, et al.
Veröffentlicht: (2025)
von: Aryan, Abi, et al.
Veröffentlicht: (2025)
DevOps-Gym: Benchmarking AI Agents in Software DevOps Cycle
von: Tang, Yuheng, et al.
Veröffentlicht: (2026)
von: Tang, Yuheng, et al.
Veröffentlicht: (2026)
A Framework for Formalizing LLM Agent Security
von: Siu, Vincent, et al.
Veröffentlicht: (2026)
von: Siu, Vincent, et al.
Veröffentlicht: (2026)
TOWARDS GOOD GOVERNANCE: E-GOVERNANCE AS A CATALYST FOR SUSTAINABLE DEVELOPMENT
von: A Sesha Reddy
Veröffentlicht: (2023)
von: A Sesha Reddy
Veröffentlicht: (2023)
DarthShader: Fuzzing WebGPU Shader Translators & Compilers
von: Bernhard, Lukas, et al.
Veröffentlicht: (2024)
von: Bernhard, Lukas, et al.
Veröffentlicht: (2024)
No Peer, no Cry: Network Application Fuzzing via Fault Injection
von: Bars, Nils, et al.
Veröffentlicht: (2024)
von: Bars, Nils, et al.
Veröffentlicht: (2024)
An Explorative Study of Pig Butchering Scams
von: Acharya, Bhupendra, et al.
Veröffentlicht: (2024)
von: Acharya, Bhupendra, et al.
Veröffentlicht: (2024)
Bad Neighbors: On Understanding VPN Provider Networks
von: Rytilahti, Teemu, et al.
Veröffentlicht: (2024)
von: Rytilahti, Teemu, et al.
Veröffentlicht: (2024)
DROIDCCT: Cryptographic Compliance Test via Trillion-Scale Measurement
von: Moghimi, Daniel, et al.
Veröffentlicht: (2026)
von: Moghimi, Daniel, et al.
Veröffentlicht: (2026)
Profiling Resilient to Change in Probe Position
von: Bursztein, Elie, et al.
Veröffentlicht: (2026)
von: Bursztein, Elie, et al.
Veröffentlicht: (2026)
OpenSage: Self-programming Agent Generation Engine
von: Li, Hongwei, et al.
Veröffentlicht: (2026)
von: Li, Hongwei, et al.
Veröffentlicht: (2026)
ImpossibleBench: Measuring LLMs' Propensity of Exploiting Test Cases
von: Zhong, Ziqian, et al.
Veröffentlicht: (2025)
von: Zhong, Ziqian, et al.
Veröffentlicht: (2025)
Kelvin and Easterly Wave Interactions and Their Modulation by Diurnal Cycle over Eastern Atlantic and Tropical Africa
von: Mantripragada, Rama Sesha Sridhar
Veröffentlicht: (2021)
von: Mantripragada, Rama Sesha Sridhar
Veröffentlicht: (2021)
Anota: Identifying Business Logic Vulnerabilities via Annotation-Based Sanitization
von: Wang, Meng, et al.
Veröffentlicht: (2025)
von: Wang, Meng, et al.
Veröffentlicht: (2025)
The Attack and Defense Landscape of Agentic AI: A Comprehensive Survey
von: Kim, Juhee, et al.
Veröffentlicht: (2026)
von: Kim, Juhee, et al.
Veröffentlicht: (2026)
LeakAgent: RL-based Red-teaming Agent for LLM Privacy Leakage
von: Nie, Yuzhou, et al.
Veröffentlicht: (2024)
von: Nie, Yuzhou, et al.
Veröffentlicht: (2024)
Exploiting LLM Quantization
von: Egashira, Kazuki, et al.
Veröffentlicht: (2024)
von: Egashira, Kazuki, et al.
Veröffentlicht: (2024)
PoC-Gym: Towards More Reliable LLM-Assisted Proof-of-Concept Exploit Generation
von: Gezgin, Derin, et al.
Veröffentlicht: (2026)
von: Gezgin, Derin, et al.
Veröffentlicht: (2026)
Isabel Sarli camaleónica. Hibridaciones entre gender y genre en las coproducciones latinoamericanas de la dupla con Armando Bó
von: Agostina Invernizzi
Veröffentlicht: (2025)
von: Agostina Invernizzi
Veröffentlicht: (2025)
NANOTECNOLOGÍA EN LOS MEDIOS: ¿QUÉ INFORMACIÓN LLEGA AL PÚBLICO?
von: Noela Invernizzi
Veröffentlicht: (2009)
von: Noela Invernizzi
Veröffentlicht: (2009)
PRESENTACIÓN DE LOS LIBROS SINFONÍA DE LA ARMONÍA DE LAS REVELACIONES CELESTIALES DE HILDEGARD DE BINGEN Y THE VOICE OF SILENCE
von: Lucía Invernizzi
Veröffentlicht: (2005)
von: Lucía Invernizzi
Veröffentlicht: (2005)
El despegue de las nanotecnologías
von: Noela Invernizzi
Veröffentlicht: (2005)
von: Noela Invernizzi
Veröffentlicht: (2005)
Niños y adolescentes trabajadores en las calles de Lima: vida cotidiana y estrategias familiares de supervivencia
von: Antonella Invernizzi
Veröffentlicht: (2014)
von: Antonella Invernizzi
Veröffentlicht: (2014)
Ähnliche Einträge
-
Evaluating the Robustness of a Production Malware Detection System to Transferable Adversarial Attacks
von: Nasr, Milad, et al.
Veröffentlicht: (2025) -
Remote Timing Attacks on Efficient Language Model Inference
von: Carlini, Nicholas, et al.
Veröffentlicht: (2024) -
CyberGym: Evaluating AI Agents' Real-World Cybersecurity Capabilities at Scale
von: Wang, Zhun, et al.
Veröffentlicht: (2025) -
Progent: Securing AI Agents with Privilege Control
von: Shi, Tianneng, et al.
Veröffentlicht: (2025) -
Generalized Power Attacks against Crypto Hardware using Long-Range Deep Learning
von: Bursztein, Elie, et al.
Veröffentlicht: (2023)