Teams of LLM Agents can Exploit Zero-Day Vulnerabilities
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhu, Yuxuan, Kellermann, Antony, Gupta, Akul, Li, Philip, Fang, Richard, Bindu, Rohan, Kang, Daniel |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Predicting Known Vulnerabilities from Attack Descriptions Using Sentence Transformers
di: Othman, Refat
Pubblicazione: (2026)
di: Othman, Refat
Pubblicazione: (2026)
CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities
di: Zhu, Yuxuan, et al.
Pubblicazione: (2025)
di: Zhu, Yuxuan, et al.
Pubblicazione: (2025)
The Automation Advantage in AI Red Teaming
di: Mulla, Rob, et al.
Pubblicazione: (2025)
di: Mulla, Rob, et al.
Pubblicazione: (2025)
AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models
di: Dawson, Ads, et al.
Pubblicazione: (2025)
di: Dawson, Ads, et al.
Pubblicazione: (2025)
DecisionBench: A Benchmark for Emergent Delegation in Long-Horizon Agentic Workflows
di: Gao, Yuxuan, et al.
Pubblicazione: (2026)
di: Gao, Yuxuan, et al.
Pubblicazione: (2026)
Benefits and Limitations of Communication in Multi-Agent Reasoning
di: Rizvi-Martel, Michael, et al.
Pubblicazione: (2025)
di: Rizvi-Martel, Michael, et al.
Pubblicazione: (2025)
Hidden in Memory: Sleeper Memory Poisoning in LLM Agents
di: Pulipaka, Sidharth, et al.
Pubblicazione: (2026)
di: Pulipaka, Sidharth, et al.
Pubblicazione: (2026)
The Power of Stories: Narrative Priming Shapes How LLM Agents Collaborate and Compete
di: Großmann, Gerrit, et al.
Pubblicazione: (2025)
di: Großmann, Gerrit, et al.
Pubblicazione: (2025)
Refusal Evaluation in Coding LLMs and Code Agents: A Systematic Review of Thirteen Malicious-Code Prompt Corpora (2023-2025)
di: Young, Richard J., et al.
Pubblicazione: (2026)
di: Young, Richard J., et al.
Pubblicazione: (2026)
AgentSentry: Mitigating Indirect Prompt Injection in LLM Agents via Temporal Causal Diagnostics and Context Purification
di: Zhang, Tian, et al.
Pubblicazione: (2026)
di: Zhang, Tian, et al.
Pubblicazione: (2026)
Self-Emotion Blended Dialogue Generation in Social Simulation Agents
di: Zhang, Qiang, et al.
Pubblicazione: (2024)
di: Zhang, Qiang, et al.
Pubblicazione: (2024)
PAVE: A Cognitive Architecture for Legitimate Violation in Generative Agent Societies
di: Yehia, Ahmad, et al.
Pubblicazione: (2026)
di: Yehia, Ahmad, et al.
Pubblicazione: (2026)
Exploring Design of Multi-Agent LLM Dialogues for Research Ideation
di: Ueda, Keisuke, et al.
Pubblicazione: (2025)
di: Ueda, Keisuke, et al.
Pubblicazione: (2025)
Multi-Agent Systems Powered by Large Language Models: Applications in Swarm Intelligence
di: Jimenez-Romero, Cristian, et al.
Pubblicazione: (2025)
di: Jimenez-Romero, Cristian, et al.
Pubblicazione: (2025)
PestMA: LLM-based Multi-Agent System for Informed Pest Management
di: Shi, Hongrui, et al.
Pubblicazione: (2025)
di: Shi, Hongrui, et al.
Pubblicazione: (2025)
Quantifying Return on Security Controls in LLM Systems
di: Moulton, Richard Helder, et al.
Pubblicazione: (2025)
di: Moulton, Richard Helder, et al.
Pubblicazione: (2025)
Review of Case-Based Reasoning for LLM Agents: Theoretical Foundations, Architectural Components, and Cognitive Integration
di: Hatalis, Kostas, et al.
Pubblicazione: (2025)
di: Hatalis, Kostas, et al.
Pubblicazione: (2025)
Literature Review Of Multi-Agent Debate For Problem-Solving
di: Tillmann, Arne
Pubblicazione: (2025)
di: Tillmann, Arne
Pubblicazione: (2025)
MEMSAD: Gradient-Coupled Anomaly Detection for Memory Poisoning in Retrieval-Augmented Agents
di: Gowda, Ishrith
Pubblicazione: (2026)
di: Gowda, Ishrith
Pubblicazione: (2026)
Building Browser Agents: Architecture, Security, and Practical Solutions
di: Vardanyan, Aram
Pubblicazione: (2025)
di: Vardanyan, Aram
Pubblicazione: (2025)
Talk is Cheap, Communication is Hard: Dynamic Grounding Failures and Repair in Multi-Agent Negotiation
di: Yao, Yiheng, et al.
Pubblicazione: (2026)
di: Yao, Yiheng, et al.
Pubblicazione: (2026)
ReDel: A Toolkit for LLM-Powered Recursive Multi-Agent Systems
di: Zhu, Andrew, et al.
Pubblicazione: (2024)
di: Zhu, Andrew, et al.
Pubblicazione: (2024)
Governance Architecture for Autonomous Agent Systems: Threats, Framework, and Engineering Practice
di: Ge, Yuxu
Pubblicazione: (2026)
di: Ge, Yuxu
Pubblicazione: (2026)
Solving Zebra Puzzles Using Constraint-Guided Multi-Agent Systems
di: Berman, Shmuel, et al.
Pubblicazione: (2024)
di: Berman, Shmuel, et al.
Pubblicazione: (2024)
Detecting Prompt Injection Attacks Against Application Using Classifiers
di: Shaheer, Safwan, et al.
Pubblicazione: (2025)
di: Shaheer, Safwan, et al.
Pubblicazione: (2025)
Beyond the Benchmark: Innovative Defenses Against Prompt Injection Attacks
di: Shaheer, Safwan, et al.
Pubblicazione: (2025)
di: Shaheer, Safwan, et al.
Pubblicazione: (2025)
CRAwDAD: Causal Reasoning Augmentation with Dual-Agent Debate
di: Vamosi, Finn G., et al.
Pubblicazione: (2025)
di: Vamosi, Finn G., et al.
Pubblicazione: (2025)
Agent WARPP: Workflow Adherence via Runtime Parallel Personalization
di: Mazzolenis, Maria Emilia, et al.
Pubblicazione: (2025)
di: Mazzolenis, Maria Emilia, et al.
Pubblicazione: (2025)
LLM Agents can Autonomously Exploit One-day Vulnerabilities
di: Fang, Richard, et al.
Pubblicazione: (2024)
di: Fang, Richard, et al.
Pubblicazione: (2024)
SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use
di: Patel, Hitesh Laxmichand, et al.
Pubblicazione: (2025)
di: Patel, Hitesh Laxmichand, et al.
Pubblicazione: (2025)
GSAR: Typed Grounding for Hallucination Detection and Recovery in Multi-Agent LLMs
di: Kamelhar, Federico A.
Pubblicazione: (2026)
di: Kamelhar, Federico A.
Pubblicazione: (2026)
A Knowledge-Based Language Model: Deducing Grammatical Knowledge in a Multi-Agent Language Acquisition Simulation
di: Shakouri, David Ph., et al.
Pubblicazione: (2025)
di: Shakouri, David Ph., et al.
Pubblicazione: (2025)
Continuous Discovery of Vulnerabilities in LLM Serving Systems with Fuzzing
di: Zhao, Yunze, et al.
Pubblicazione: (2026)
di: Zhao, Yunze, et al.
Pubblicazione: (2026)
TRIZ Agents: A Multi-Agent LLM Approach for TRIZ-Based Innovation
di: Szczepanik, Kamil, et al.
Pubblicazione: (2025)
di: Szczepanik, Kamil, et al.
Pubblicazione: (2025)
RPRA: Predicting an LLM-Judge for Efficient but Performant Inference
di: Ashley, Dylan R., et al.
Pubblicazione: (2026)
di: Ashley, Dylan R., et al.
Pubblicazione: (2026)
DURA-CPS: A Multi-Role Orchestrator for Dependability Assurance in LLM-Enabled Cyber-Physical Systems
di: Srinivasan, Trisanth, et al.
Pubblicazione: (2025)
di: Srinivasan, Trisanth, et al.
Pubblicazione: (2025)
ORACLE-SWE: Quantifying the Contribution of Oracle Information Signals on SWE Agents
di: Li, Kenan, et al.
Pubblicazione: (2026)
di: Li, Kenan, et al.
Pubblicazione: (2026)
SBASH: a Framework for Designing and Evaluating RAG vs. Prompt-Tuned LLM Honeypots
di: Adebimpe, Adetayo, et al.
Pubblicazione: (2025)
di: Adebimpe, Adetayo, et al.
Pubblicazione: (2025)
AegisShield: Democratizing Cyber Threat Modeling with Generative AI
di: Grofsky, Matthew
Pubblicazione: (2025)
di: Grofsky, Matthew
Pubblicazione: (2025)
Evaluating the Reliability of Digital Forensic Evidence Discovered by Large Language Model: A Case Study
di: Khatiwala, Jeel Piyushkumar, et al.
Pubblicazione: (2026)
di: Khatiwala, Jeel Piyushkumar, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Predicting Known Vulnerabilities from Attack Descriptions Using Sentence Transformers
di: Othman, Refat
Pubblicazione: (2026) -
CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities
di: Zhu, Yuxuan, et al.
Pubblicazione: (2025) -
The Automation Advantage in AI Red Teaming
di: Mulla, Rob, et al.
Pubblicazione: (2025) -
AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models
di: Dawson, Ads, et al.
Pubblicazione: (2025) -
DecisionBench: A Benchmark for Emergent Delegation in Long-Horizon Agentic Workflows
di: Gao, Yuxuan, et al.
Pubblicazione: (2026)