Contextual Agent Security: A Policy for Every Purpose
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tsai, Lillian, Bagdasarian, Eugene |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AirGapAgent: Protecting Privacy-Conscious Conversational Agents
von: Bagdasarian, Eugene, et al.
Veröffentlicht: (2024)
von: Bagdasarian, Eugene, et al.
Veröffentlicht: (2024)
AI Agents May Always Fall for Prompt Injections
von: Abdelnabi, Sahar, et al.
Veröffentlicht: (2026)
von: Abdelnabi, Sahar, et al.
Veröffentlicht: (2026)
ContinuousBench: Can Differentially Private Synthetic Text Improve Capabilities?
von: Liu, Peihan, et al.
Veröffentlicht: (2026)
von: Liu, Peihan, et al.
Veröffentlicht: (2026)
Chain of Attack: a Semantic-Driven Contextual Multi-Turn attacker for LLM
von: Yang, Xikang, et al.
Veröffentlicht: (2024)
von: Yang, Xikang, et al.
Veröffentlicht: (2024)
Throttling Web Agents Using Reasoning Gates
von: Kumar, Abhinav, et al.
Veröffentlicht: (2025)
von: Kumar, Abhinav, et al.
Veröffentlicht: (2025)
Every Character Counts: From Vulnerability to Defense in Phishing Detection
von: Chiper, Maria, et al.
Veröffentlicht: (2025)
von: Chiper, Maria, et al.
Veröffentlicht: (2025)
Policy-Invisible Violations in LLM-Based Agents
von: Wu, Jie, et al.
Veröffentlicht: (2026)
von: Wu, Jie, et al.
Veröffentlicht: (2026)
SecureNet: A Comparative Study of DeBERTa and Large Language Models for Phishing Detection
von: Mahendru, Sakshi, et al.
Veröffentlicht: (2024)
von: Mahendru, Sakshi, et al.
Veröffentlicht: (2024)
Securing Large Language Models (LLMs) from Prompt Injection Attacks
von: Suri, Omar Farooq Khan, et al.
Veröffentlicht: (2025)
von: Suri, Omar Farooq Khan, et al.
Veröffentlicht: (2025)
Graphene: Infrastructure Security Posture Analysis with AI-generated Attack Graphs
von: Jin, Xin, et al.
Veröffentlicht: (2023)
von: Jin, Xin, et al.
Veröffentlicht: (2023)
Transferable Embedding Inversion Attack: Uncovering Privacy Risks in Text Embeddings without Model Queries
von: Huang, Yu-Hsiang, et al.
Veröffentlicht: (2024)
von: Huang, Yu-Hsiang, et al.
Veröffentlicht: (2024)
Robust and Secure Code Watermarking for Large Language Models via ML/Crypto Codesign
von: Zhang, Ruisi, et al.
Veröffentlicht: (2025)
von: Zhang, Ruisi, et al.
Veröffentlicht: (2025)
Deep Active Learning with Crowdsourcing Data for Privacy Policy Classification
von: Qiu, Wenjun, et al.
Veröffentlicht: (2020)
von: Qiu, Wenjun, et al.
Veröffentlicht: (2020)
LiSA: Lifelong Safety Adaptation via Conservative Policy Induction
von: Kim, Minbeom, et al.
Veröffentlicht: (2026)
von: Kim, Minbeom, et al.
Veröffentlicht: (2026)
Backdooring Bias ($B^2$) into Stable Diffusion Models
von: Naseh, Ali, et al.
Veröffentlicht: (2024)
von: Naseh, Ali, et al.
Veröffentlicht: (2024)
Cyber-Zero: Training Cybersecurity Agents without Runtime
von: Zhuo, Terry Yue, et al.
Veröffentlicht: (2025)
von: Zhuo, Terry Yue, et al.
Veröffentlicht: (2025)
Assessing Deanonymization Risks with Stylometry-Assisted LLM Agent
von: Zhang, Boyang, et al.
Veröffentlicht: (2026)
von: Zhang, Boyang, et al.
Veröffentlicht: (2026)
AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks
von: Zeng, Yifan, et al.
Veröffentlicht: (2024)
von: Zeng, Yifan, et al.
Veröffentlicht: (2024)
When Routine Chats Turn Toxic: Unintended Long-Term State Poisoning in Personalized Agents
von: Xu, Xiaoyu, et al.
Veröffentlicht: (2026)
von: Xu, Xiaoyu, et al.
Veröffentlicht: (2026)
OverThink: Slowdown Attacks on Reasoning LLMs
von: Kumar, Abhinav, et al.
Veröffentlicht: (2025)
von: Kumar, Abhinav, et al.
Veröffentlicht: (2025)
SecureCode: A Production-Grade Multi-Turn Dataset for Training Security-Aware Code Generation Models
von: Thornton, Scott
Veröffentlicht: (2025)
von: Thornton, Scott
Veröffentlicht: (2025)
SecureBreak -- A dataset towards safe and secure models
von: Arazzi, Marco, et al.
Veröffentlicht: (2026)
von: Arazzi, Marco, et al.
Veröffentlicht: (2026)
Prompted Contextual Vectors for Spear-Phishing Detection
von: Nahmias, Daniel, et al.
Veröffentlicht: (2024)
von: Nahmias, Daniel, et al.
Veröffentlicht: (2024)
Terrarium: Revisiting the Blackboard for Multi-Agent Safety, Privacy, and Security Studies
von: Nakamura, Mason, et al.
Veröffentlicht: (2025)
von: Nakamura, Mason, et al.
Veröffentlicht: (2025)
Gandalf the Red: Adaptive Security for LLMs
von: Pfister, Niklas, et al.
Veröffentlicht: (2025)
von: Pfister, Niklas, et al.
Veröffentlicht: (2025)
Self-interpreting Adversarial Images
von: Zhang, Tingwei, et al.
Veröffentlicht: (2024)
von: Zhang, Tingwei, et al.
Veröffentlicht: (2024)
SecEncoder: Logs are All You Need in Security
von: Bulut, Muhammed Fatih, et al.
Veröffentlicht: (2024)
von: Bulut, Muhammed Fatih, et al.
Veröffentlicht: (2024)
BadAgent: Inserting and Activating Backdoor Attacks in LLM Agents
von: Wang, Yifei, et al.
Veröffentlicht: (2024)
von: Wang, Yifei, et al.
Veröffentlicht: (2024)
GCG Attack On A Diffusion LLM
von: Neyroud, Ruben, et al.
Veröffentlicht: (2025)
von: Neyroud, Ruben, et al.
Veröffentlicht: (2025)
A Watermark for Large Language Models
von: Kirchenbauer, John, et al.
Veröffentlicht: (2023)
von: Kirchenbauer, John, et al.
Veröffentlicht: (2023)
A Transfer Attack to Image Watermarks
von: Hu, Yuepeng, et al.
Veröffentlicht: (2024)
von: Hu, Yuepeng, et al.
Veröffentlicht: (2024)
A StrongREJECT for Empty Jailbreaks
von: Souly, Alexandra, et al.
Veröffentlicht: (2024)
von: Souly, Alexandra, et al.
Veröffentlicht: (2024)
A Systematic Review of Federated Generative Models
von: Gargary, Ashkan Vedadi, et al.
Veröffentlicht: (2024)
von: Gargary, Ashkan Vedadi, et al.
Veröffentlicht: (2024)
A Watermark for Black-Box Language Models
von: Bahri, Dara, et al.
Veröffentlicht: (2024)
von: Bahri, Dara, et al.
Veröffentlicht: (2024)
Watermark under Fire: A Robustness Evaluation of LLM Watermarking
von: Liang, Jiacheng, et al.
Veröffentlicht: (2024)
von: Liang, Jiacheng, et al.
Veröffentlicht: (2024)
Exploring Vulnerabilities and Protections in Large Language Models: A Survey
von: Liu, Frank Weizhen, et al.
Veröffentlicht: (2024)
von: Liu, Frank Weizhen, et al.
Veröffentlicht: (2024)
Logit-Gap Steering: A Forward-Pass Diagnostic for Alignment Robustness
von: Li, Tung-Ling, et al.
Veröffentlicht: (2025)
von: Li, Tung-Ling, et al.
Veröffentlicht: (2025)
Confidence Elicitation: A New Attack Vector for Large Language Models
von: Formento, Brian, et al.
Veröffentlicht: (2025)
von: Formento, Brian, et al.
Veröffentlicht: (2025)
TOSSS: a CVE-based Software Security Benchmark for Large Language Models
von: Damie, Marc, et al.
Veröffentlicht: (2026)
von: Damie, Marc, et al.
Veröffentlicht: (2026)
Securing Multi-turn Conversational Language Models From Distributed Backdoor Triggers
von: Tong, Terry, et al.
Veröffentlicht: (2024)
von: Tong, Terry, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
AirGapAgent: Protecting Privacy-Conscious Conversational Agents
von: Bagdasarian, Eugene, et al.
Veröffentlicht: (2024) -
AI Agents May Always Fall for Prompt Injections
von: Abdelnabi, Sahar, et al.
Veröffentlicht: (2026) -
ContinuousBench: Can Differentially Private Synthetic Text Improve Capabilities?
von: Liu, Peihan, et al.
Veröffentlicht: (2026) -
Chain of Attack: a Semantic-Driven Contextual Multi-Turn attacker for LLM
von: Yang, Xikang, et al.
Veröffentlicht: (2024) -
Throttling Web Agents Using Reasoning Gates
von: Kumar, Abhinav, et al.
Veröffentlicht: (2025)