MURMUR: Using cross-user chatter to break collaborative language agents in groups
Fuente:
arXiv
Saved in:
| Main Authors: | Patlan, Atharv Singh, Sheng, Peiyao, Hebbar, S. Ashwin, Mittal, Prateek, Viswanath, Pramod |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Context manipulation attacks : Web agents are susceptible to corrupted memory
by: Patlan, Atharv Singh, et al.
Published: (2025)
by: Patlan, Atharv Singh, et al.
Published: (2025)
Real AI Agents with Fake Memories: Fatal Context Manipulation Attacks on Web3 Agents
by: Patlan, Atharv Singh, et al.
Published: (2025)
by: Patlan, Atharv Singh, et al.
Published: (2025)
Proof of Diligence: Cryptoeconomic Security for Rollups
by: Sheng, Peiyao, et al.
Published: (2024)
by: Sheng, Peiyao, et al.
Published: (2024)
Unconditionally Safe Light Client
by: Moshrefi, Niusha, et al.
Published: (2024)
by: Moshrefi, Niusha, et al.
Published: (2024)
Undetectable Backdoors in Model Parameters: Hiding Sparse Secrets in High Dimensions
by: Choudhary, Sarthak, et al.
Published: (2026)
by: Choudhary, Sarthak, et al.
Published: (2026)
Teach LLMs to Phish: Stealing Private Information from Language Models
by: Panda, Ashwinee, et al.
Published: (2024)
by: Panda, Ashwinee, et al.
Published: (2024)
Scalable Fingerprinting of Large Language Models
by: Nasery, Anshul, et al.
Published: (2025)
by: Nasery, Anshul, et al.
Published: (2025)
CFT-Forensics: High-Performance Byzantine Accountability for Crash Fault Tolerant Protocols
by: Tang, Weizhao, et al.
Published: (2023)
by: Tang, Weizhao, et al.
Published: (2023)
Optimizing watermarks for large language models
by: Wouters, Bram
Published: (2023)
by: Wouters, Bram
Published: (2023)
Are Robust LLM Fingerprints Adversarially Robust?
by: Nasery, Anshul, et al.
Published: (2025)
by: Nasery, Anshul, et al.
Published: (2025)
Private Fine-tuning of Large Language Models with Zeroth-order Optimization
by: Tang, Xinyu, et al.
Published: (2024)
by: Tang, Xinyu, et al.
Published: (2024)
sudo rm -rf agentic_security
by: Lee, Sejin, et al.
Published: (2025)
by: Lee, Sejin, et al.
Published: (2025)
ReliabilityRAG: Effective and Provably Robust Defense for RAG-based Web-Search
by: Shen, Zeyu, et al.
Published: (2025)
by: Shen, Zeyu, et al.
Published: (2025)
Instructional Segment Embedding: Improving LLM Safety with Instruction Hierarchy
by: Wu, Tong, et al.
Published: (2024)
by: Wu, Tong, et al.
Published: (2024)
Sandwich attack: Multi-language Mixture Adaptive Attack on LLMs
by: Upadhayay, Bibek, et al.
Published: (2024)
by: Upadhayay, Bibek, et al.
Published: (2024)
StealthRL: Reinforcement Learning Paraphrase Attacks for Multi-Detector Evasion of AI-Text Detectors
by: Ranganath, Suraj, et al.
Published: (2026)
by: Ranganath, Suraj, et al.
Published: (2026)
PRIVATEEDIT: A Privacy-Preserving Pipeline for Face-Centric Generative Image Editing
by: Tamboli, Dipesh, et al.
Published: (2026)
by: Tamboli, Dipesh, et al.
Published: (2026)
Certifiably Robust RAG against Retrieval Corruption
by: Xiang, Chong, et al.
Published: (2024)
by: Xiang, Chong, et al.
Published: (2024)
ExCyTIn-Bench: Evaluating LLM agents on Cyber Threat Investigation
by: Wu, Yiran, et al.
Published: (2025)
by: Wu, Yiran, et al.
Published: (2025)
CVE-LLM : Automatic vulnerability evaluation in medical device industry using large language models
by: Ghosh, Rikhiya, et al.
Published: (2024)
by: Ghosh, Rikhiya, et al.
Published: (2024)
Training AI to be Loyal
by: Oh, Sewoong, et al.
Published: (2025)
by: Oh, Sewoong, et al.
Published: (2025)
A New Linear Scaling Rule for Private Adaptive Hyperparameter Optimization
by: Panda, Ashwinee, et al.
Published: (2022)
by: Panda, Ashwinee, et al.
Published: (2022)
Safety Alignment Should Be Made More Than Just a Few Tokens Deep
by: Qi, Xiangyu, et al.
Published: (2024)
by: Qi, Xiangyu, et al.
Published: (2024)
AgentWall: A Runtime Safety Layer for Local AI Agents
by: Aravind, Ashwin
Published: (2026)
by: Aravind, Ashwin
Published: (2026)
AgentSOC: A Multi-Layer Agentic AI Framework for Security Operations Automation
by: Roy, Joyjit, et al.
Published: (2026)
by: Roy, Joyjit, et al.
Published: (2026)
Recursive language models for jailbreak detection: a procedural defense for tool-augmented agents
by: Shavit, Doron
Published: (2026)
by: Shavit, Doron
Published: (2026)
Subtoxic Questions: Dive Into Attitude Change of LLM's Response in Jailbreak Attempts
by: Zhang, Tianyu, et al.
Published: (2024)
by: Zhang, Tianyu, et al.
Published: (2024)
Privacy Auditing of Large Language Models
by: Panda, Ashwinee, et al.
Published: (2025)
by: Panda, Ashwinee, et al.
Published: (2025)
PersonaMark: Personalized LLM watermarking for model protection and user attribution
by: Zhang, Yuehan, et al.
Published: (2024)
by: Zhang, Yuehan, et al.
Published: (2024)
The Ethics of Interaction: Mitigating Security Threats in LLMs
by: Kumar, Ashutosh, et al.
Published: (2024)
by: Kumar, Ashutosh, et al.
Published: (2024)
The Fire Thief Is Also the Keeper: Balancing Usability and Privacy in Prompts
by: Shen, Zhili, et al.
Published: (2024)
by: Shen, Zhili, et al.
Published: (2024)
OML: A Primitive for Reconciling Open Access with Owner Control in AI Model Distribution
by: Cheng, Zerui, et al.
Published: (2024)
by: Cheng, Zerui, et al.
Published: (2024)
Byzantine-Robust Federated Learning: An Overview With Focus on Developing Sybil-based Attacks to Backdoor Augmented Secure Aggregation Protocols
by: Deshmukh, Atharv
Published: (2024)
by: Deshmukh, Atharv
Published: (2024)
Prompt Leakage effect and defense strategies for multi-turn LLM interactions
by: Agarwal, Divyansh, et al.
Published: (2024)
by: Agarwal, Divyansh, et al.
Published: (2024)
On Evaluating the Durability of Safeguards for Open-Weight LLMs
by: Qi, Xiangyu, et al.
Published: (2024)
by: Qi, Xiangyu, et al.
Published: (2024)
Attestable Audits: Verifiable AI Safety Benchmarks Using Trusted Execution Environments
by: Schnabl, Christoph, et al.
Published: (2025)
by: Schnabl, Christoph, et al.
Published: (2025)
GPT-4 Jailbreaks Itself with Near-Perfect Success Using Self-Explanation
by: Ramesh, Govind, et al.
Published: (2024)
by: Ramesh, Govind, et al.
Published: (2024)
from Benign import Toxic: Jailbreaking the Language Model via Adversarial Metaphors
by: Yan, Yu, et al.
Published: (2025)
by: Yan, Yu, et al.
Published: (2025)
Waterfall: Framework for Robust and Scalable Text Watermarking and Provenance for LLMs
by: Lau, Gregory Kang Ruey, et al.
Published: (2024)
by: Lau, Gregory Kang Ruey, et al.
Published: (2024)
Adversarial Attacks Against Automated Fact-Checking: A Survey
by: Liu, Fanzhen, et al.
Published: (2025)
by: Liu, Fanzhen, et al.
Published: (2025)
Similar Items
-
Context manipulation attacks : Web agents are susceptible to corrupted memory
by: Patlan, Atharv Singh, et al.
Published: (2025) -
Real AI Agents with Fake Memories: Fatal Context Manipulation Attacks on Web3 Agents
by: Patlan, Atharv Singh, et al.
Published: (2025) -
Proof of Diligence: Cryptoeconomic Security for Rollups
by: Sheng, Peiyao, et al.
Published: (2024) -
Unconditionally Safe Light Client
by: Moshrefi, Niusha, et al.
Published: (2024) -
Undetectable Backdoors in Model Parameters: Hiding Sparse Secrets in High Dimensions
by: Choudhary, Sarthak, et al.
Published: (2026)