CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhu, Yuxuan, Kellermann, Antony, Bowman, Dylan, Li, Philip, Gupta, Akul, Danda, Adarsh, Fang, Richard, Jensen, Conner, Ihli, Eric, Benn, Jason, Geronimo, Jet, Dhir, Avi, Rao, Sudhit, Yu, Kaicheng, Stone, Twm, Kang, Daniel |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Teams of LLM Agents can Exploit Zero-Day Vulnerabilities
von: Zhu, Yuxuan, et al.
Veröffentlicht: (2024)
von: Zhu, Yuxuan, et al.
Veröffentlicht: (2024)
N-Agent Ad Hoc Teamwork
von: Wang, Caroline, et al.
Veröffentlicht: (2024)
von: Wang, Caroline, et al.
Veröffentlicht: (2024)
ROTATE: Regret-driven Open-ended Training for Ad Hoc Teamwork
von: Wang, Caroline, et al.
Veröffentlicht: (2025)
von: Wang, Caroline, et al.
Veröffentlicht: (2025)
Vision-based Situational Graphs Exploiting Fiducial Markers for the Integration of Semantic Entities
von: Tourani, Ali, et al.
Veröffentlicht: (2023)
von: Tourani, Ali, et al.
Veröffentlicht: (2023)
SAGE: A Strategy-Aware Graph-Enhanced Generation Framework For Online Counseling
von: Aharon, Eliya Naomi, et al.
Veröffentlicht: (2026)
von: Aharon, Eliya Naomi, et al.
Veröffentlicht: (2026)
DM$^2$: Decentralized Multi-Agent Reinforcement Learning for Distribution Matching
von: Wang, Caroline, et al.
Veröffentlicht: (2022)
von: Wang, Caroline, et al.
Veröffentlicht: (2022)
Towards a Robust Soft Baby Robot With Rich Interaction Ability for Advanced Machine Learning Algorithms
von: Alhakami, Mohannad, et al.
Veröffentlicht: (2024)
von: Alhakami, Mohannad, et al.
Veröffentlicht: (2024)
Evaluating the Systematic Reasoning Abilities of Large Language Models through Graph Coloring
von: Heyman, Alex, et al.
Veröffentlicht: (2025)
von: Heyman, Alex, et al.
Veröffentlicht: (2025)
SALLIE: Safeguarding Against Latent Language & Image Exploits
von: Azov, Guy, et al.
Veröffentlicht: (2026)
von: Azov, Guy, et al.
Veröffentlicht: (2026)
5G Traffic Prediction with Time Series Analysis
von: Nayak, Nikhil, et al.
Veröffentlicht: (2021)
von: Nayak, Nikhil, et al.
Veröffentlicht: (2021)
Building Minimal and Reusable Causal State Abstractions for Reinforcement Learning
von: Wang, Zizhao, et al.
Veröffentlicht: (2024)
von: Wang, Zizhao, et al.
Veröffentlicht: (2024)
Evaluating and Learning Robust Bandit Policies Under Uncertain Causal Mechanisms
von: Avery, Katherine, et al.
Veröffentlicht: (2025)
von: Avery, Katherine, et al.
Veröffentlicht: (2025)
D-Shape: Demonstration-Shaped Reinforcement Learning via Goal Conditioning
von: Wang, Caroline, et al.
Veröffentlicht: (2022)
von: Wang, Caroline, et al.
Veröffentlicht: (2022)
VulScribeR: Exploring RAG-based Vulnerability Augmentation with LLMs
von: Daneshvar, Seyed Shayan, et al.
Veröffentlicht: (2024)
von: Daneshvar, Seyed Shayan, et al.
Veröffentlicht: (2024)
Ultra-Reduced-Impact-Encased-Logging (URIEL): propose a new method for selective sustainable logging and post-harvest silvicultural treatment in tropical forest using airborne robotics systems
von: Albiero, Daniel, et al.
Veröffentlicht: (2026)
von: Albiero, Daniel, et al.
Veröffentlicht: (2026)
An Explainable Collaborative Dialogue System using a Theory of Mind
von: Cohen, Philip R., et al.
Veröffentlicht: (2023)
von: Cohen, Philip R., et al.
Veröffentlicht: (2023)
AI and Machine Learning Approaches for Predicting Nanoparticles Toxicity The Critical Role of Physiochemical Properties
von: Yousaf, Iqra
Veröffentlicht: (2024)
von: Yousaf, Iqra
Veröffentlicht: (2024)
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
von: Raoufi, Behnam, et al.
Veröffentlicht: (2025)
von: Raoufi, Behnam, et al.
Veröffentlicht: (2025)
Social Learning through Interactions with Other Agents: A Survey
von: Hillier, Dylan, et al.
Veröffentlicht: (2024)
von: Hillier, Dylan, et al.
Veröffentlicht: (2024)
vS-Graphs: Tightly Coupling Visual SLAM and 3D Scene Graphs Exploiting Hierarchical Scene Understanding
von: Tourani, Ali, et al.
Veröffentlicht: (2025)
von: Tourani, Ali, et al.
Veröffentlicht: (2025)
Deployment-Time Reliability of Learned Robot Policies
von: Agia, Christopher
Veröffentlicht: (2026)
von: Agia, Christopher
Veröffentlicht: (2026)
Gyan: An Explainable Neuro-Symbolic Language Model
von: Srinivasan, Venkat, et al.
Veröffentlicht: (2026)
von: Srinivasan, Venkat, et al.
Veröffentlicht: (2026)
GraphWalk: Enabling Reasoning in Large Language Models through Tool-Based Graph Navigation
von: Ghandi, Taraneh, et al.
Veröffentlicht: (2026)
von: Ghandi, Taraneh, et al.
Veröffentlicht: (2026)
elsciRL: Integrating Language Solutions into Reinforcement Learning Problem Settings
von: Osborne, Philip, et al.
Veröffentlicht: (2025)
von: Osborne, Philip, et al.
Veröffentlicht: (2025)
Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents
von: Tang, Wenjie, et al.
Veröffentlicht: (2026)
von: Tang, Wenjie, et al.
Veröffentlicht: (2026)
A Domain-Agnostic Neurosymbolic Approach for Big Social Data Analysis: Evaluating Mental Health Sentiment on Social Media during COVID-19
von: Khandelwal, Vedant, et al.
Veröffentlicht: (2024)
von: Khandelwal, Vedant, et al.
Veröffentlicht: (2024)
Automated Circuit Interpretation via Probe Prompting
von: Birardi, Giuseppe
Veröffentlicht: (2025)
von: Birardi, Giuseppe
Veröffentlicht: (2025)
Quantum Abduction: A New Paradigm for Reasoning under Uncertainty
von: Pareschi, Remo
Veröffentlicht: (2025)
von: Pareschi, Remo
Veröffentlicht: (2025)
CCVA-FL: Cross-Client Variations Adaptive Federated Learning for Medical Imaging
von: Gupta, Sunny, et al.
Veröffentlicht: (2024)
von: Gupta, Sunny, et al.
Veröffentlicht: (2024)
Taming the Tail: Leveraging Asymmetric Loss and Pade Approximation to Overcome Medical Image Long-Tailed Class Imbalance
von: Kashyap, Pankhi, et al.
Veröffentlicht: (2024)
von: Kashyap, Pankhi, et al.
Veröffentlicht: (2024)
IMUVIE: Pickup Timeline Action Localization via Motion Movies
von: Clapham, John, et al.
Veröffentlicht: (2024)
von: Clapham, John, et al.
Veröffentlicht: (2024)
REMIND: Input Loss Landscapes Reveal Residual Memorization in Post-Unlearning LLMs
von: Cohen, Liran, et al.
Veröffentlicht: (2025)
von: Cohen, Liran, et al.
Veröffentlicht: (2025)
Can AI Assist in Olympiad Coding
von: Ren, Samuel
Veröffentlicht: (2025)
von: Ren, Samuel
Veröffentlicht: (2025)
Multi-Agent Design Assistant for the Simulation of Inertial Fusion Energy
von: Shachar, Meir H., et al.
Veröffentlicht: (2025)
von: Shachar, Meir H., et al.
Veröffentlicht: (2025)
Federated Learning for Anomaly Detection in Energy Consumption Data: Assessing the Vulnerability to Adversarial Attacks
von: Telila, Yohannis Kifle, et al.
Veröffentlicht: (2025)
von: Telila, Yohannis Kifle, et al.
Veröffentlicht: (2025)
In Context Learning with Vision Transformers: Case Study
von: Zhao, Antony, et al.
Veröffentlicht: (2025)
von: Zhao, Antony, et al.
Veröffentlicht: (2025)
From Demonstrations to Safe Deployment: Path-Consistent Safety Filtering for Diffusion Policies
von: Römer, Ralf, et al.
Veröffentlicht: (2025)
von: Römer, Ralf, et al.
Veröffentlicht: (2025)
OpenMap: Instruction Grounding via Open-Vocabulary Visual-Language Mapping
von: Li, Danyang, et al.
Veröffentlicht: (2025)
von: Li, Danyang, et al.
Veröffentlicht: (2025)
Dream to Fly: Model-Based Reinforcement Learning for Vision-Based Drone Flight
von: Romero, Angel, et al.
Veröffentlicht: (2025)
von: Romero, Angel, et al.
Veröffentlicht: (2025)
FSFM: A Biologically-Inspired Framework for Selective Forgetting of Agent Memory
von: Gu, Yingjie, et al.
Veröffentlicht: (2026)
von: Gu, Yingjie, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Teams of LLM Agents can Exploit Zero-Day Vulnerabilities
von: Zhu, Yuxuan, et al.
Veröffentlicht: (2024) -
N-Agent Ad Hoc Teamwork
von: Wang, Caroline, et al.
Veröffentlicht: (2024) -
ROTATE: Regret-driven Open-ended Training for Ad Hoc Teamwork
von: Wang, Caroline, et al.
Veröffentlicht: (2025) -
Vision-based Situational Graphs Exploiting Fiducial Markers for the Integration of Semantic Entities
von: Tourani, Ali, et al.
Veröffentlicht: (2023) -
SAGE: A Strategy-Aware Graph-Enhanced Generation Framework For Online Counseling
von: Aharon, Eliya Naomi, et al.
Veröffentlicht: (2026)