ConVerse: Benchmarking Contextual Safety in Agent-to-Agent Conversations
Fuente:
arXiv
Saved in:
| Main Authors: | Gomaa, Amr, Salem, Ahmed, Abdelnabi, Sahar |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AI Agents May Always Fall for Prompt Injections
by: Abdelnabi, Sahar, et al.
Published: (2026)
by: Abdelnabi, Sahar, et al.
Published: (2026)
Get my drift? Catching LLM Task Drift with Activation Deltas
by: Abdelnabi, Sahar, et al.
Published: (2024)
by: Abdelnabi, Sahar, et al.
Published: (2024)
Firewalls to Secure Dynamic LLM Agentic Networks
by: Abdelnabi, Sahar, et al.
Published: (2025)
by: Abdelnabi, Sahar, et al.
Published: (2025)
The Hawthorne Effect in Reasoning Models: Evaluating and Steering Test Awareness
by: Abdelnabi, Sahar, et al.
Published: (2025)
by: Abdelnabi, Sahar, et al.
Published: (2025)
CI-Work: Benchmarking Contextual Integrity in Enterprise LLM Agents
by: Fu, Wenjie, et al.
Published: (2026)
by: Fu, Wenjie, et al.
Published: (2026)
Can LLMs Infer Conversational Agent Users' Personality Traits from Chat History?
by: Cögendez, Derya, et al.
Published: (2026)
by: Cögendez, Derya, et al.
Published: (2026)
Terrarium: Revisiting the Blackboard for Multi-Agent Safety, Privacy, and Security Studies
by: Nakamura, Mason, et al.
Published: (2025)
by: Nakamura, Mason, et al.
Published: (2025)
Cooperation, Competition, and Maliciousness: LLM-Stakeholders Interactive Negotiation
by: Abdelnabi, Sahar, et al.
Published: (2023)
by: Abdelnabi, Sahar, et al.
Published: (2023)
Protecting Users From Themselves: Safeguarding Contextual Privacy in Interactions with Conversational Agents
by: Ngong, Ivoline, et al.
Published: (2025)
by: Ngong, Ivoline, et al.
Published: (2025)
SafeCOMM: A Study on Safety Degradation in Fine-Tuned Telecom Large Language Models
by: Djuhera, Aladin, et al.
Published: (2025)
by: Djuhera, Aladin, et al.
Published: (2025)
Nuclear Deployed: Analyzing Catastrophic Risks in Decision-making of Autonomous LLM Agents
by: Xu, Rongwu, et al.
Published: (2025)
by: Xu, Rongwu, et al.
Published: (2025)
How Well Can LLM Agents Simulate End-User Security and Privacy Attitudes and Behaviors?
by: Li, Yuxuan, et al.
Published: (2026)
by: Li, Yuxuan, et al.
Published: (2026)
Contextualized Privacy Defense for LLM Agents
by: Wen, Yule, et al.
Published: (2026)
by: Wen, Yule, et al.
Published: (2026)
InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents
by: Zhan, Qiusi, et al.
Published: (2024)
by: Zhan, Qiusi, et al.
Published: (2024)
Phare: A Safety Probe for Large Language Models
by: Jeune, Pierre Le, et al.
Published: (2025)
by: Jeune, Pierre Le, et al.
Published: (2025)
Expected Harm: Rethinking Safety Evaluation of (Mis)Aligned LLMs
by: Chen, Yen-Shan, et al.
Published: (2026)
by: Chen, Yen-Shan, et al.
Published: (2026)
ContextualJailbreak: Evolutionary Red-Teaming via Simulated Conversational Priming
by: Béjar, Mario Rodríguez, et al.
Published: (2026)
by: Béjar, Mario Rodríguez, et al.
Published: (2026)
Medical Malice: A Dataset for Context-Aware Safety in Healthcare LLMs
by: D'addario, Andrew Maranhão Ventura
Published: (2025)
by: D'addario, Andrew Maranhão Ventura
Published: (2025)
The Earth is Flat because...: Investigating LLMs' Belief towards Misinformation via Persuasive Conversation
by: Xu, Rongwu, et al.
Published: (2023)
by: Xu, Rongwu, et al.
Published: (2023)
Stateless Yet Not Forgetful: Implicit Memory as a Hidden Channel in LLMs
by: Salem, Ahmed, et al.
Published: (2026)
by: Salem, Ahmed, et al.
Published: (2026)
AirGapAgent: Protecting Privacy-Conscious Conversational Agents
by: Bagdasarian, Eugene, et al.
Published: (2024)
by: Bagdasarian, Eugene, et al.
Published: (2024)
Personalized Attacks of Social Engineering in Multi-turn Conversations: LLM Agents for Simulation and Detection
by: Kumarage, Tharindu, et al.
Published: (2025)
by: Kumarage, Tharindu, et al.
Published: (2025)
ROK-FORTRESS: Measuring the Effect of Geopolitical Transcreation for National Security and Public Safety
by: Lee, Michael S., et al.
Published: (2026)
by: Lee, Michael S., et al.
Published: (2026)
LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments
by: Zhang, Chiyu, et al.
Published: (2026)
by: Zhang, Chiyu, et al.
Published: (2026)
Contextual Agent Security: A Policy for Every Purpose
by: Tsai, Lillian, et al.
Published: (2025)
by: Tsai, Lillian, et al.
Published: (2025)
Beyond Jailbreaking: Auditing Contextual Privacy in LLM Agents
by: Das, Saswat, et al.
Published: (2025)
by: Das, Saswat, et al.
Published: (2025)
How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States
by: Zhou, Zhenhong, et al.
Published: (2024)
by: Zhou, Zhenhong, et al.
Published: (2024)
Beyond Context: Large Language Models' Failure to Grasp Users' Intent
by: Hussain, Ahmed M., et al.
Published: (2025)
by: Hussain, Ahmed M., et al.
Published: (2025)
On the Suitability of LLM-Driven Agents for Dark Pattern Audits
by: Sun, Chen, et al.
Published: (2026)
by: Sun, Chen, et al.
Published: (2026)
Safety and Security Analysis of Large Language Models: Benchmarking Risk Profile and Harm Potential
by: Akiri, Charankumar, et al.
Published: (2025)
by: Akiri, Charankumar, et al.
Published: (2025)
Let's Measure the Elephant in the Room: Facilitating Personalized Automated Analysis of Privacy Policies at Scale
by: Zhao, Rui, et al.
Published: (2025)
by: Zhao, Rui, et al.
Published: (2025)
Steering the CensorShip: Uncovering Representation Vectors for LLM "Thought" Control
by: Cyberey, Hannah, et al.
Published: (2025)
by: Cyberey, Hannah, et al.
Published: (2025)
An Investigation into Misuse of Java Security APIs by Large Language Models
by: Mousavi, Zahra, et al.
Published: (2024)
by: Mousavi, Zahra, et al.
Published: (2024)
SecureForge: Finding and Preventing Vulnerabilities in LLM-Generated Code via Prompt Optimization
by: Liu, Houjun, et al.
Published: (2026)
by: Liu, Houjun, et al.
Published: (2026)
Domain-Independent Deception: A New Taxonomy and Linguistic Analysis
by: Verma, Rakesh M., et al.
Published: (2024)
by: Verma, Rakesh M., et al.
Published: (2024)
Data Defenses Against Large Language Models
by: Agnew, William, et al.
Published: (2024)
by: Agnew, William, et al.
Published: (2024)
How Susceptible are Large Language Models to Ideological Manipulation?
by: Chen, Kai, et al.
Published: (2024)
by: Chen, Kai, et al.
Published: (2024)
Infrastructure for Valuable, Tradable, and Verifiable Agent Memory
by: Li, Mengyuan, et al.
Published: (2026)
by: Li, Mengyuan, et al.
Published: (2026)
Jailbreak Distillation: Renewable Safety Benchmarking
by: Zhang, Jingyu, et al.
Published: (2025)
by: Zhang, Jingyu, et al.
Published: (2025)
Just Do It!? Computer-Use Agents Exhibit Blind Goal-Directedness
by: Shayegani, Erfan, et al.
Published: (2025)
by: Shayegani, Erfan, et al.
Published: (2025)
Similar Items
-
AI Agents May Always Fall for Prompt Injections
by: Abdelnabi, Sahar, et al.
Published: (2026) -
Get my drift? Catching LLM Task Drift with Activation Deltas
by: Abdelnabi, Sahar, et al.
Published: (2024) -
Firewalls to Secure Dynamic LLM Agentic Networks
by: Abdelnabi, Sahar, et al.
Published: (2025) -
The Hawthorne Effect in Reasoning Models: Evaluating and Steering Test Awareness
by: Abdelnabi, Sahar, et al.
Published: (2025) -
CI-Work: Benchmarking Contextual Integrity in Enterprise LLM Agents
by: Fu, Wenjie, et al.
Published: (2026)