PentestJudge: Judging Agent Behavior Against Operational Requirements
Fuente:
arXiv
Guardado en:
| Autores principales: | Caldwell, Shane, Harley, Max, Kouremetis, Michael, Abruzzo, Vincent, Pearce, Will |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
AutoPentester: An LLM Agent-based Framework for Automated Pentesting
por: Ginige, Yasod, et al.
Publicado: (2025)
por: Ginige, Yasod, et al.
Publicado: (2025)
AutoPentest: Enhancing Vulnerability Management With Autonomous LLM Agents
por: Henke, Julius
Publicado: (2025)
por: Henke, Julius
Publicado: (2025)
From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World
por: Conde, Pedro, et al.
Publicado: (2026)
por: Conde, Pedro, et al.
Publicado: (2026)
OCCULT: Evaluating Large Language Models for Offensive Cyber Operation Capabilities
por: Kouremetis, Michael, et al.
Publicado: (2025)
por: Kouremetis, Michael, et al.
Publicado: (2025)
ARACNE: An LLM-Based Autonomous Shell Pentesting Agent
por: Nieponice, Tomas, et al.
Publicado: (2025)
por: Nieponice, Tomas, et al.
Publicado: (2025)
PentestMCP: A Toolkit for Agentic Penetration Testing
por: Ezetta, Zachary, et al.
Publicado: (2025)
por: Ezetta, Zachary, et al.
Publicado: (2025)
BadJudge: Backdoor Vulnerabilities of LLM-as-a-Judge
por: Tong, Terry, et al.
Publicado: (2025)
por: Tong, Terry, et al.
Publicado: (2025)
A Preliminary Study on Using Large Language Models in Software Pentesting
por: Shashwat, Kumar, et al.
Publicado: (2024)
por: Shashwat, Kumar, et al.
Publicado: (2024)
Security in LLM-as-a-Judge: A Comprehensive SoK
por: Masoud, Aiman Al, et al.
Publicado: (2026)
por: Masoud, Aiman Al, et al.
Publicado: (2026)
Optimization-based Prompt Injection Attack to LLM-as-a-Judge
por: Shi, Jiawen, et al.
Publicado: (2024)
por: Shi, Jiawen, et al.
Publicado: (2024)
Towards Effective Offensive Security LLM Agents: Hyperparameter Tuning, LLM as a Judge, and a Lightweight CTF Benchmark
por: Shao, Minghao, et al.
Publicado: (2025)
por: Shao, Minghao, et al.
Publicado: (2025)
Generative Artificial Intelligence-Supported Pentesting: A Comparison between Claude Opus, GPT-4, and Copilot
por: Martínez, Antonio López, et al.
Publicado: (2025)
por: Martínez, Antonio López, et al.
Publicado: (2025)
Hacking, The Lazy Way: LLM Augmented Pentesting
por: Goyal, Dhruva, et al.
Publicado: (2024)
por: Goyal, Dhruva, et al.
Publicado: (2024)
PentestEval: Benchmarking LLM-based Penetration Testing with Modular and Stage-Level Design
por: Yang, Ruozhao, et al.
Publicado: (2025)
por: Yang, Ruozhao, et al.
Publicado: (2025)
Rubrics as an Attack Surface: Stealthy Preference Drift in LLM Judges
por: Ding, Ruomeng, et al.
Publicado: (2026)
por: Ding, Ruomeng, et al.
Publicado: (2026)
No Attack Required: Semantic Fuzzing for Specification Violations in Agent Skills
por: Li, Ying, et al.
Publicado: (2026)
por: Li, Ying, et al.
Publicado: (2026)
Turning Bias into Bugs: Bandit-Guided Style Manipulation Attacks on LLM Judges
por: Yang, Xianglin, et al.
Publicado: (2026)
por: Yang, Xianglin, et al.
Publicado: (2026)
AttackPilot: Autonomous Inference Attacks Against ML Services With LLM-Based Agents
por: Wu, Yixin, et al.
Publicado: (2025)
por: Wu, Yixin, et al.
Publicado: (2025)
Securing AI Agents Against Prompt Injection Attacks
por: Ramakrishnan, Badrinath, et al.
Publicado: (2025)
por: Ramakrishnan, Badrinath, et al.
Publicado: (2025)
ADVERSA: Measuring Multi-Turn Guardrail Degradation and Judge Reliability in Large Language Models
por: Owiredu-Ashley, Harry
Publicado: (2026)
por: Owiredu-Ashley, Harry
Publicado: (2026)
AgentMark: Utility-Preserving Behavioral Watermarking for Agents
por: Huang, Kaibo, et al.
Publicado: (2026)
por: Huang, Kaibo, et al.
Publicado: (2026)
Sequential Behavioral Watermarking for LLM Agents
por: An, Hyeseon, et al.
Publicado: (2026)
por: An, Hyeseon, et al.
Publicado: (2026)
RTBAS: Defending LLM Agents Against Prompt Injection and Privacy Leakage
por: Zhong, Peter Yong, et al.
Publicado: (2025)
por: Zhong, Peter Yong, et al.
Publicado: (2025)
To Protect the LLM Agent Against the Prompt Injection Attack with Polymorphic Prompt
por: Wang, Zhilong, et al.
Publicado: (2025)
por: Wang, Zhilong, et al.
Publicado: (2025)
WASP: Benchmarking Web Agent Security Against Prompt Injection Attacks
por: Evtimov, Ivan, et al.
Publicado: (2025)
por: Evtimov, Ivan, et al.
Publicado: (2025)
SLEIGHT-Bench: A Benchmark of Evasion Attacks Against Agent Monitors
por: Najt, Elle, et al.
Publicado: (2026)
por: Najt, Elle, et al.
Publicado: (2026)
WARD: Adversarially Robust Defense of Web Agents Against Prompt Injections
por: Cao, Tri, et al.
Publicado: (2026)
por: Cao, Tri, et al.
Publicado: (2026)
Behavior-Aware and Generalizable Defense Against Black-Box Adversarial Attacks for ML-Based IDS
por: Ennaji, Sabrine, et al.
Publicado: (2025)
por: Ennaji, Sabrine, et al.
Publicado: (2025)
Agent Operating Systems (AOS): Integrating Agentic Control Planes into, and Beyond, Traditional Operating Systems
por: Sharma, Ankur, et al.
Publicado: (2026)
por: Sharma, Ankur, et al.
Publicado: (2026)
MELON: Provable Defense Against Indirect Prompt Injection Attacks in AI Agents
por: Zhu, Kaijie, et al.
Publicado: (2025)
por: Zhu, Kaijie, et al.
Publicado: (2025)
LivePI: More Realistic Benchmarking of Agents Against Indirect Prompt Injection
por: Zhao, Lei, et al.
Publicado: (2026)
por: Zhao, Lei, et al.
Publicado: (2026)
Temporal UI State Inconsistency in Desktop GUI Agents: Formalizing and Defending Against TOCTOU Attacks on Computer-Use Agents
por: Xu, Wenpeng
Publicado: (2026)
por: Xu, Wenpeng
Publicado: (2026)
Agent Privilege Separation in OpenClaw: A Structural Defense Against Prompt Injection
por: Cheng, Darren, et al.
Publicado: (2026)
por: Cheng, Darren, et al.
Publicado: (2026)
Evaluating Jailbreaking Vulnerabilities in LLMs Deployed as Assistants for Smart Grid Operations: A Benchmark Against NERC Standards
por: Hammadia, Taha, et al.
Publicado: (2026)
por: Hammadia, Taha, et al.
Publicado: (2026)
MCP Security Bench (MSB): Benchmarking Attacks Against Model Context Protocol in LLM Agents
por: Zhang, Dongsen, et al.
Publicado: (2025)
por: Zhang, Dongsen, et al.
Publicado: (2025)
CoopGuard: Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Round Attacks
por: Li, Siyuan, et al.
Publicado: (2026)
por: Li, Siyuan, et al.
Publicado: (2026)
VIGIL: Defending LLM Agents Against Tool Stream Injection via Verify-Before-Commit
por: Lin, Junda, et al.
Publicado: (2026)
por: Lin, Junda, et al.
Publicado: (2026)
MUZZLE: Adaptive Agentic Red-Teaming of Web Agents Against Indirect Prompt Injection Attacks
por: Syros, Georgios, et al.
Publicado: (2026)
por: Syros, Georgios, et al.
Publicado: (2026)
Architecting Secure AI Agents: Perspectives on System-Level Defenses Against Indirect Prompt Injection Attacks
por: Xiang, Chong, et al.
Publicado: (2026)
por: Xiang, Chong, et al.
Publicado: (2026)
PentestAgent: Incorporating LLM Agents to Automated Penetration Testing
por: Shen, Xiangmin, et al.
Publicado: (2024)
por: Shen, Xiangmin, et al.
Publicado: (2024)
Ejemplares similares
-
AutoPentester: An LLM Agent-based Framework for Automated Pentesting
por: Ginige, Yasod, et al.
Publicado: (2025) -
AutoPentest: Enhancing Vulnerability Management With Autonomous LLM Agents
por: Henke, Julius
Publicado: (2025) -
From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World
por: Conde, Pedro, et al.
Publicado: (2026) -
OCCULT: Evaluating Large Language Models for Offensive Cyber Operation Capabilities
por: Kouremetis, Michael, et al.
Publicado: (2025) -
ARACNE: An LLM-Based Autonomous Shell Pentesting Agent
por: Nieponice, Tomas, et al.
Publicado: (2025)