HARP: Measuring Harm Amplification in Multi-Agent LLM Systems
Fuente:
arXiv
Saved in:
| Main Authors: | Rahman, Md Hafizur, Haider, Zafaryab, Mahfuz, Tanzim, Chakraborty, Prabuddha |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
X-DFS: Explainable Artificial Intelligence Guided Design-for-Security Solution Space Exploration
by: Mahfuz, Tanzim, et al.
Published: (2024)
by: Mahfuz, Tanzim, et al.
Published: (2024)
How a Bit Becomes a Story: Semantic Steering via Differentiable Fault Injection
by: Haider, Zafaryab, et al.
Published: (2025)
by: Haider, Zafaryab, et al.
Published: (2025)
LeMo-NADe: Multi-Parameter Neural Architecture Discovery with LLMs
by: Rahman, Md Hafizur, et al.
Published: (2024)
by: Rahman, Md Hafizur, et al.
Published: (2024)
Semantic Intent Fragmentation: A Single-Shot Compositional Attack on Multi-Agent AI Pipelines
by: Ahad, Tanzim, et al.
Published: (2026)
by: Ahad, Tanzim, et al.
Published: (2026)
The Art of the Jailbreak: Formulating Jailbreak Attacks for LLM Security Beyond Binary Scoring
by: Hossain, Ismail, et al.
Published: (2026)
by: Hossain, Ismail, et al.
Published: (2026)
SALTY: Explainable Artificial Intelligence Guided Structural Analysis for Hardware Trojan Detection
by: Mahfuz, Tanzim, et al.
Published: (2025)
by: Mahfuz, Tanzim, et al.
Published: (2025)
POLARIS: Explainable Artificial Intelligence for Mitigating Power Side-Channel Leakage
by: Mahfuz, Tanzim, et al.
Published: (2025)
by: Mahfuz, Tanzim, et al.
Published: (2025)
Trackly: A Unified SaaS Platform for User Behavior Analytics and Real Time Rule Based Anomaly Detection
by: Haque, Md Zahurul, et al.
Published: (2026)
by: Haque, Md Zahurul, et al.
Published: (2026)
Federated Learning in Healthcare: Model Misconducts, Security, Challenges, Applications, and Future Research Directions -- A Systematic Review
by: Ali, Md Shahin, et al.
Published: (2024)
by: Ali, Md Shahin, et al.
Published: (2024)
False Data Injection Attack Detection in Edge-based Smart Metering Networks with Federated Learning
by: Uddin, Md Raihan, et al.
Published: (2024)
by: Uddin, Md Raihan, et al.
Published: (2024)
Mind the Gap: Missing Cyber Threat Coverage in NIDS Datasets for the Energy Sector
by: Tory, Adrita Rahman, et al.
Published: (2025)
by: Tory, Adrita Rahman, et al.
Published: (2025)
Blockchain-Enabled Explainable AI for Trusted Healthcare Systems
by: Mohsin, Md Talha
Published: (2025)
by: Mohsin, Md Talha
Published: (2025)
Enhancing Cloud Network Resilience via a Robust LLM-Empowered Multi-Agent Reinforcement Learning Framework
by: Peng, Yixiao, et al.
Published: (2026)
by: Peng, Yixiao, et al.
Published: (2026)
Reliable Weak-to-Strong Monitoring of LLM Agents
by: Kale, Neil, et al.
Published: (2025)
by: Kale, Neil, et al.
Published: (2025)
LeakAgent: RL-based Red-teaming Agent for LLM Privacy Leakage
by: Nie, Yuzhou, et al.
Published: (2024)
by: Nie, Yuzhou, et al.
Published: (2024)
A Generative Approach to LLM Harmfulness Mitigation with Red Flag Tokens
by: Dobre, David, et al.
Published: (2025)
by: Dobre, David, et al.
Published: (2025)
Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey
by: Huang, Tiansheng, et al.
Published: (2024)
by: Huang, Tiansheng, et al.
Published: (2024)
Feature-Aware Anisotropic Local Differential Privacy for Utility-Preserving Graph Representation Learning in Metal Additive Manufacturing
by: Islam, MD Shafikul, et al.
Published: (2026)
by: Islam, MD Shafikul, et al.
Published: (2026)
SUB-PLAY: Adversarial Policies against Partially Observed Multi-Agent Reinforcement Learning Systems
by: Ma, Oubo, et al.
Published: (2024)
by: Ma, Oubo, et al.
Published: (2024)
The Autonomy Tax: Defense Training Breaks LLM Agents
by: Li, Shawn, et al.
Published: (2026)
by: Li, Shawn, et al.
Published: (2026)
SHADE-Arena: Evaluating Sabotage and Monitoring in LLM Agents
by: Kutasov, Jonathan, et al.
Published: (2025)
by: Kutasov, Jonathan, et al.
Published: (2025)
Amplification Effects in Test-Time Reinforcement Learning: Safety and Reasoning Vulnerabilities
by: Khattar, Vanshaj, et al.
Published: (2026)
by: Khattar, Vanshaj, et al.
Published: (2026)
Pharmacist: Safety Alignment Data Curation for Large Language Models against Harmful Fine-tuning
by: Liu, Guozhi, et al.
Published: (2025)
by: Liu, Guozhi, et al.
Published: (2025)
BLAST: A Stealthy Backdoor Leverage Attack against Cooperative Multi-Agent Deep Reinforcement Learning based Systems
by: Fang, Jing, et al.
Published: (2025)
by: Fang, Jing, et al.
Published: (2025)
SPARD: Defending Harmful Fine-Tuning Attack via Safety Projection with Relevance-Diversity Data Selection
by: Chen, Shuhao, et al.
Published: (2026)
by: Chen, Shuhao, et al.
Published: (2026)
MemoryGraft: Persistent Compromise of LLM Agents via Poisoned Experience Retrieval
by: Srivastava, Saksham Sahai, et al.
Published: (2025)
by: Srivastava, Saksham Sahai, et al.
Published: (2025)
Attention Eclipse: Manipulating Attention to Bypass LLM Safety-Alignment
by: Zaree, Pedram, et al.
Published: (2025)
by: Zaree, Pedram, et al.
Published: (2025)
BadImplant: Injection-based Multi-Targeted Graph Backdoor Attack
by: Khan, Md Nabi Newaz, et al.
Published: (2026)
by: Khan, Md Nabi Newaz, et al.
Published: (2026)
Personalized Federated Learning Techniques: Empirical Analysis
by: Khan, Azal Ahmad, et al.
Published: (2024)
by: Khan, Azal Ahmad, et al.
Published: (2024)
Adversarial Reinforcement Learning for Offensive and Defensive Agents in a Simulated Zero-Sum Network Environment
by: Shahid, Abrar, et al.
Published: (2025)
by: Shahid, Abrar, et al.
Published: (2025)
Training RL Agents for Multi-Objective Network Defense Tasks
by: Molina-Markham, Andres, et al.
Published: (2025)
by: Molina-Markham, Andres, et al.
Published: (2025)
OEP: Poisoning Self-Evolving LLM Agents via Locally Correct but Non-Transferable Experiences
by: Wang, Kaixiang, et al.
Published: (2026)
by: Wang, Kaixiang, et al.
Published: (2026)
Fall into a Pit, Gain in a Wit: Cognitive-Guided Harmful Meme Detection via Misjudgment Risk Pattern Retrieval
by: Wang, Wenshuo, et al.
Published: (2025)
by: Wang, Wenshuo, et al.
Published: (2025)
A Novel Ensemble Learning Approach for Enhanced IoT Attack Detection: Redefining Security Paradigms in Connected Systems
by: Abdeljaber, Hikmat A. M., et al.
Published: (2025)
by: Abdeljaber, Hikmat A. M., et al.
Published: (2025)
UDora: A Unified Red Teaming Framework against LLM Agents by Dynamically Hijacking Their Own Reasoning
by: Zhang, Jiawei, et al.
Published: (2025)
by: Zhang, Jiawei, et al.
Published: (2025)
FHE-Agent: Automating CKKS Configuration for Practical Encrypted Inference via an LLM-Guided Agentic Framework
by: Xu, Nuo, et al.
Published: (2025)
by: Xu, Nuo, et al.
Published: (2025)
Fuzz-Testing Meets LLM-Based Agents: An Automated and Efficient Framework for Jailbreaking Text-To-Image Generation Models
by: Dong, Yingkai, et al.
Published: (2024)
by: Dong, Yingkai, et al.
Published: (2024)
Enhancing IoT Cyber Attack Detection in the Presence of Highly Imbalanced Data
by: Haque, Md. Ehsanul, et al.
Published: (2025)
by: Haque, Md. Ehsanul, et al.
Published: (2025)
Odysseus: Jailbreaking Commercial Multimodal LLM-integrated Systems via Dual Steganography
by: Li, Songze, et al.
Published: (2025)
by: Li, Songze, et al.
Published: (2025)
A Multi-Dimensional Quality Scoring Framework for Decentralized LLM Inference with Proof of Quality
by: Tian, Arther, et al.
Published: (2026)
by: Tian, Arther, et al.
Published: (2026)
Similar Items
-
X-DFS: Explainable Artificial Intelligence Guided Design-for-Security Solution Space Exploration
by: Mahfuz, Tanzim, et al.
Published: (2024) -
How a Bit Becomes a Story: Semantic Steering via Differentiable Fault Injection
by: Haider, Zafaryab, et al.
Published: (2025) -
LeMo-NADe: Multi-Parameter Neural Architecture Discovery with LLMs
by: Rahman, Md Hafizur, et al.
Published: (2024) -
Semantic Intent Fragmentation: A Single-Shot Compositional Attack on Multi-Agent AI Pipelines
by: Ahad, Tanzim, et al.
Published: (2026) -
The Art of the Jailbreak: Formulating Jailbreak Attacks for LLM Security Beyond Binary Scoring
by: Hossain, Ismail, et al.
Published: (2026)