Emerging Safety Attack and Defense in Federated Instruction Tuning of Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Ye, Rui, Chai, Jingyi, Liu, Xiangrui, Yang, Yaodong, Wang, Yanfeng, Chen, Siheng |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Leveraging Unstructured Text Data for Federated Instruction Tuning of Large Language Models
by: Ye, Rui, et al.
Published: (2024)
by: Ye, Rui, et al.
Published: (2024)
PsySafe: A Comprehensive Framework for Psychological-based Attack, Defense, and Evaluation of Multi-agent System Safety
by: Zhang, Zaibin, et al.
Published: (2024)
by: Zhang, Zaibin, et al.
Published: (2024)
Memory Poisoning Attack and Defense on Memory Based LLM-Agents
by: Sunil, Balachandra Devarangadi, et al.
Published: (2026)
by: Sunil, Balachandra Devarangadi, et al.
Published: (2026)
Secure Forgetting: A Framework for Privacy-Driven Unlearning in Large Language Model (LLM)-Based Agents
by: Ye, Dayong, et al.
Published: (2026)
by: Ye, Dayong, et al.
Published: (2026)
SkillSafetyBench: Evaluating Agent Safety under Skill-Facing Attack Surfaces
by: Jin, Chang, et al.
Published: (2026)
by: Jin, Chang, et al.
Published: (2026)
When Embedding-Based Defenses Fail: Rethinking Safety in LLM-Based Multi-Agent Systems
by: Zhang, Lingxi, et al.
Published: (2026)
by: Zhang, Lingxi, et al.
Published: (2026)
FedLLM-Bench: Realistic Benchmarks for Federated Learning of Large Language Models
by: Ye, Rui, et al.
Published: (2024)
by: Ye, Rui, et al.
Published: (2024)
Quantitative Resilience Modeling for Autonomous Cyber Defense
by: Cadet, Xavier, et al.
Published: (2025)
by: Cadet, Xavier, et al.
Published: (2025)
EncGPT: A Multi-Agent Workflow for Dynamic Encryption Algorithms
by: Li, Donghe, et al.
Published: (2025)
by: Li, Donghe, et al.
Published: (2025)
Beyond Input Guardrails: Reconstructing Cross-Agent Semantic Flows for Execution-Aware Attack Detection
by: Wei, Yangyang, et al.
Published: (2026)
by: Wei, Yangyang, et al.
Published: (2026)
CEE: An Inference-Time Jailbreak Defense for Embodied Intelligence via Subspace Concept Rotation
by: Yang, Jirui, et al.
Published: (2025)
by: Yang, Jirui, et al.
Published: (2025)
X-Teaming: Multi-Turn Jailbreaks and Defenses with Adaptive Multi-Agents
by: Rahman, Salman, et al.
Published: (2025)
by: Rahman, Salman, et al.
Published: (2025)
Multi-Agent Actor-Critics in Autonomous Cyber Defense
by: Wang, Mingjun, et al.
Published: (2024)
by: Wang, Mingjun, et al.
Published: (2024)
Hierarchical Multi-agent Reinforcement Learning for Cyber Network Defense
by: Singh, Aditya Vikram, et al.
Published: (2024)
by: Singh, Aditya Vikram, et al.
Published: (2024)
OpenFedLLM: Training Large Language Models on Decentralized Private Data via Federated Learning
by: Ye, Rui, et al.
Published: (2024)
by: Ye, Rui, et al.
Published: (2024)
DrunkAgent: Stealthy Memory Corruption in LLM-Powered Recommender Agents
by: Yang, Shiyi, et al.
Published: (2025)
by: Yang, Shiyi, et al.
Published: (2025)
MAS-Shield: A Defense Framework for Secure and Efficient LLM MAS
by: Wang, Kaixiang, et al.
Published: (2025)
by: Wang, Kaixiang, et al.
Published: (2025)
AutoRISE: Agent-Driven Strategy Evolution for Red-Teaming Large Language Models
by: Gautam, Tanmay, et al.
Published: (2026)
by: Gautam, Tanmay, et al.
Published: (2026)
Explainable Autonomous Cyber Defense using Adversarial Multi-Agent Reinforcement Learning
by: Zhang, Yiyao, et al.
Published: (2026)
by: Zhang, Yiyao, et al.
Published: (2026)
PILLAR: an AI-Powered Privacy Threat Modeling Tool
by: Mollaeefar, Majid, et al.
Published: (2024)
by: Mollaeefar, Majid, et al.
Published: (2024)
SV-LLM: An Agentic Approach for SoC Security Verification using Large Language Models
by: Saha, Dipayan, et al.
Published: (2025)
by: Saha, Dipayan, et al.
Published: (2025)
Breaking the Secret: Economic Interventions for Combating Collusion in Embodied Multi-Agent Systems
by: Liu, Qi, et al.
Published: (2026)
by: Liu, Qi, et al.
Published: (2026)
Aegis: Towards Governance, Integrity, and Security of AI Voice Agents
by: Li, Xiang, et al.
Published: (2026)
by: Li, Xiang, et al.
Published: (2026)
Enhancing the Robustness of QMIX against State-adversarial Attacks
by: Guo, Weiran, et al.
Published: (2023)
by: Guo, Weiran, et al.
Published: (2023)
Formalizing the Safety, Security, and Functional Properties of Agentic AI Systems
by: Allegrini, Edoardo, et al.
Published: (2025)
by: Allegrini, Edoardo, et al.
Published: (2025)
Web Fraud Attacks Against LLM-Driven Multi-Agent Systems
by: Kong, Dezhang, et al.
Published: (2025)
by: Kong, Dezhang, et al.
Published: (2025)
Attacks and Mitigations for Distributed Governance of Agentic AI under Byzantine Adversaries
by: Laws, Matthew D., et al.
Published: (2026)
by: Laws, Matthew D., et al.
Published: (2026)
Effective Red-Teaming of Policy-Adherent Agents
by: Nakash, Itay, et al.
Published: (2025)
by: Nakash, Itay, et al.
Published: (2025)
A Neuro-Symbolic Multi-Agent Approach to Legal-Cybersecurity Knowledge Integration
by: Bonfanti, Chiara, et al.
Published: (2025)
by: Bonfanti, Chiara, et al.
Published: (2025)
ScamAgents: How AI Agents Can Simulate Human-Level Scam Calls
by: Badhe, Sanket
Published: (2025)
by: Badhe, Sanket
Published: (2025)
Voting-Bloc Entropy: A New Metric for DAO Decentralization
by: Fábrega, Andrés, et al.
Published: (2025)
by: Fábrega, Andrés, et al.
Published: (2025)
Trustworthy Decentralized Autonomous Machines: A New Paradigm in Automation Economy
by: Castillo, Fernando, et al.
Published: (2025)
by: Castillo, Fernando, et al.
Published: (2025)
Decentralized Multi-Agent System with Trust-Aware Communication
by: Ding, Yepeng, et al.
Published: (2025)
by: Ding, Yepeng, et al.
Published: (2025)
AI Agents with Decentralized Identifiers and Verifiable Credentials
by: Garzon, Sandro Rodriguez, et al.
Published: (2025)
by: Garzon, Sandro Rodriguez, et al.
Published: (2025)
Incentive Mechanism Design for Privacy-Preserving Decentralized Blockchain Relayers
by: Jebari, Boutaina, et al.
Published: (2026)
by: Jebari, Boutaina, et al.
Published: (2026)
BMC4TimeSec: Verification Of Timed Security Protocols
by: Zbrzezny, Agnieszka M.
Published: (2026)
by: Zbrzezny, Agnieszka M.
Published: (2026)
Towards Transparent and Incentive-Compatible Collaboration in Decentralized LLM Multi-Agent Systems: A Blockchain-Driven Approach
by: Qi, Minfeng, et al.
Published: (2025)
by: Qi, Minfeng, et al.
Published: (2025)
SoK: Security of Autonomous LLM Agents in Agentic Commerce
by: Mao, Qian'ang, et al.
Published: (2026)
by: Mao, Qian'ang, et al.
Published: (2026)
The Best-Laid SCHEMEs: Coordinated Sabotage and Monitoring in Multi-Agent Systems
by: Radev, Nikolay, et al.
Published: (2026)
by: Radev, Nikolay, et al.
Published: (2026)
ClawCoin: An Agentic AI-Native Cryptocurrency for Decentralized Agent Economies
by: Li, Shaoyu, et al.
Published: (2026)
by: Li, Shaoyu, et al.
Published: (2026)
Similar Items
-
Leveraging Unstructured Text Data for Federated Instruction Tuning of Large Language Models
by: Ye, Rui, et al.
Published: (2024) -
PsySafe: A Comprehensive Framework for Psychological-based Attack, Defense, and Evaluation of Multi-agent System Safety
by: Zhang, Zaibin, et al.
Published: (2024) -
Memory Poisoning Attack and Defense on Memory Based LLM-Agents
by: Sunil, Balachandra Devarangadi, et al.
Published: (2026) -
Secure Forgetting: A Framework for Privacy-Driven Unlearning in Large Language Model (LLM)-Based Agents
by: Ye, Dayong, et al.
Published: (2026) -
SkillSafetyBench: Evaluating Agent Safety under Skill-Facing Attack Surfaces
by: Jin, Chang, et al.
Published: (2026)