CEE: An Inference-Time Jailbreak Defense for Embodied Intelligence via Subspace Concept Rotation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yang, Jirui, Lin, Zheyu, Lu, Zhihui, Wang, Yinggui, Wang, Lei, Wei, Tao, Duan, Qiang, Du, Xin, Yang, Shuhan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Breaking the Secret: Economic Interventions for Combating Collusion in Embodied Multi-Agent Systems
von: Liu, Qi, et al.
Veröffentlicht: (2026)
von: Liu, Qi, et al.
Veröffentlicht: (2026)
Memory Poisoning Attack and Defense on Memory Based LLM-Agents
von: Sunil, Balachandra Devarangadi, et al.
Veröffentlicht: (2026)
von: Sunil, Balachandra Devarangadi, et al.
Veröffentlicht: (2026)
OrchJail: Jailbreaking Tool-Calling Text-to-Image Agents by Orchestration-Guided Fuzzing
von: Chen, Jianming, et al.
Veröffentlicht: (2026)
von: Chen, Jianming, et al.
Veröffentlicht: (2026)
Prefix Probing: Lightweight Harmful Content Detection for Large Language Models
von: Yang, Jirui, et al.
Veröffentlicht: (2025)
von: Yang, Jirui, et al.
Veröffentlicht: (2025)
Multi-Agent Actor-Critics in Autonomous Cyber Defense
von: Wang, Mingjun, et al.
Veröffentlicht: (2024)
von: Wang, Mingjun, et al.
Veröffentlicht: (2024)
MAS-Shield: A Defense Framework for Secure and Efficient LLM MAS
von: Wang, Kaixiang, et al.
Veröffentlicht: (2025)
von: Wang, Kaixiang, et al.
Veröffentlicht: (2025)
X-Teaming: Multi-Turn Jailbreaks and Defenses with Adaptive Multi-Agents
von: Rahman, Salman, et al.
Veröffentlicht: (2025)
von: Rahman, Salman, et al.
Veröffentlicht: (2025)
EncGPT: A Multi-Agent Workflow for Dynamic Encryption Algorithms
von: Li, Donghe, et al.
Veröffentlicht: (2025)
von: Li, Donghe, et al.
Veröffentlicht: (2025)
Quantitative Resilience Modeling for Autonomous Cyber Defense
von: Cadet, Xavier, et al.
Veröffentlicht: (2025)
von: Cadet, Xavier, et al.
Veröffentlicht: (2025)
Emerging Safety Attack and Defense in Federated Instruction Tuning of Large Language Models
von: Ye, Rui, et al.
Veröffentlicht: (2024)
von: Ye, Rui, et al.
Veröffentlicht: (2024)
Hierarchical Multi-agent Reinforcement Learning for Cyber Network Defense
von: Singh, Aditya Vikram, et al.
Veröffentlicht: (2024)
von: Singh, Aditya Vikram, et al.
Veröffentlicht: (2024)
Explainable Autonomous Cyber Defense using Adversarial Multi-Agent Reinforcement Learning
von: Zhang, Yiyao, et al.
Veröffentlicht: (2026)
von: Zhang, Yiyao, et al.
Veröffentlicht: (2026)
When Embedding-Based Defenses Fail: Rethinking Safety in LLM-Based Multi-Agent Systems
von: Zhang, Lingxi, et al.
Veröffentlicht: (2026)
von: Zhang, Lingxi, et al.
Veröffentlicht: (2026)
SoK: Security of Autonomous LLM Agents in Agentic Commerce
von: Mao, Qian'ang, et al.
Veröffentlicht: (2026)
von: Mao, Qian'ang, et al.
Veröffentlicht: (2026)
Secure Forgetting: A Framework for Privacy-Driven Unlearning in Large Language Model (LLM)-Based Agents
von: Ye, Dayong, et al.
Veröffentlicht: (2026)
von: Ye, Dayong, et al.
Veröffentlicht: (2026)
G-Safeguard: A Topology-Guided Security Lens and Treatment on LLM-based Multi-agent Systems
von: Wang, Shilong, et al.
Veröffentlicht: (2025)
von: Wang, Shilong, et al.
Veröffentlicht: (2025)
Out of Sight, Not Out of Mind: Unveiling Latent Attack in Latent-based Multi-Agent Systems
von: Wang, Chenxi, et al.
Veröffentlicht: (2026)
von: Wang, Chenxi, et al.
Veröffentlicht: (2026)
Agent Name Service (ANS): A Proof-of-Concept Trust Layer for Secure AI Agent Discovery, Identity, and Governance in Kubernetes
von: Mittal, Akshay, et al.
Veröffentlicht: (2026)
von: Mittal, Akshay, et al.
Veröffentlicht: (2026)
Voting-Bloc Entropy: A New Metric for DAO Decentralization
von: Fábrega, Andrés, et al.
Veröffentlicht: (2025)
von: Fábrega, Andrés, et al.
Veröffentlicht: (2025)
Trustworthy Decentralized Autonomous Machines: A New Paradigm in Automation Economy
von: Castillo, Fernando, et al.
Veröffentlicht: (2025)
von: Castillo, Fernando, et al.
Veröffentlicht: (2025)
Decentralized Multi-Agent System with Trust-Aware Communication
von: Ding, Yepeng, et al.
Veröffentlicht: (2025)
von: Ding, Yepeng, et al.
Veröffentlicht: (2025)
AI Agents with Decentralized Identifiers and Verifiable Credentials
von: Garzon, Sandro Rodriguez, et al.
Veröffentlicht: (2025)
von: Garzon, Sandro Rodriguez, et al.
Veröffentlicht: (2025)
Incentive Mechanism Design for Privacy-Preserving Decentralized Blockchain Relayers
von: Jebari, Boutaina, et al.
Veröffentlicht: (2026)
von: Jebari, Boutaina, et al.
Veröffentlicht: (2026)
BMC4TimeSec: Verification Of Timed Security Protocols
von: Zbrzezny, Agnieszka M.
Veröffentlicht: (2026)
von: Zbrzezny, Agnieszka M.
Veröffentlicht: (2026)
Towards Transparent and Incentive-Compatible Collaboration in Decentralized LLM Multi-Agent Systems: A Blockchain-Driven Approach
von: Qi, Minfeng, et al.
Veröffentlicht: (2025)
von: Qi, Minfeng, et al.
Veröffentlicht: (2025)
The Best-Laid SCHEMEs: Coordinated Sabotage and Monitoring in Multi-Agent Systems
von: Radev, Nikolay, et al.
Veröffentlicht: (2026)
von: Radev, Nikolay, et al.
Veröffentlicht: (2026)
ClawCoin: An Agentic AI-Native Cryptocurrency for Decentralized Agent Economies
von: Li, Shaoyu, et al.
Veröffentlicht: (2026)
von: Li, Shaoyu, et al.
Veröffentlicht: (2026)
PILLAR: an AI-Powered Privacy Threat Modeling Tool
von: Mollaeefar, Majid, et al.
Veröffentlicht: (2024)
von: Mollaeefar, Majid, et al.
Veröffentlicht: (2024)
Beyond Input Guardrails: Reconstructing Cross-Agent Semantic Flows for Execution-Aware Attack Detection
von: Wei, Yangyang, et al.
Veröffentlicht: (2026)
von: Wei, Yangyang, et al.
Veröffentlicht: (2026)
CoTDeceptor:Adversarial Code Obfuscation Against CoT-Enhanced LLM Code Agents
von: Li, Haoyang, et al.
Veröffentlicht: (2025)
von: Li, Haoyang, et al.
Veröffentlicht: (2025)
Aegis: Towards Governance, Integrity, and Security of AI Voice Agents
von: Li, Xiang, et al.
Veröffentlicht: (2026)
von: Li, Xiang, et al.
Veröffentlicht: (2026)
PsySafe: A Comprehensive Framework for Psychological-based Attack, Defense, and Evaluation of Multi-agent System Safety
von: Zhang, Zaibin, et al.
Veröffentlicht: (2024)
von: Zhang, Zaibin, et al.
Veröffentlicht: (2024)
A Vision for Access Control in LLM-based Agent Systems
von: Li, Xinfeng, et al.
Veröffentlicht: (2025)
von: Li, Xinfeng, et al.
Veröffentlicht: (2025)
UIFV: Data Reconstruction Attack in Vertical Federated Learning
von: Yang, Jirui, et al.
Veröffentlicht: (2024)
von: Yang, Jirui, et al.
Veröffentlicht: (2024)
Adversary-Augmented Simulation for Fairness Evaluation and Defense in Hyperledger Fabric
von: Mahe, Erwan, et al.
Veröffentlicht: (2025)
von: Mahe, Erwan, et al.
Veröffentlicht: (2025)
Enhancing the Robustness of QMIX against State-adversarial Attacks
von: Guo, Weiran, et al.
Veröffentlicht: (2023)
von: Guo, Weiran, et al.
Veröffentlicht: (2023)
Locally Differentially Private Distributed Online Learning with Guaranteed Optimality
von: Chen, Ziqin, et al.
Veröffentlicht: (2023)
von: Chen, Ziqin, et al.
Veröffentlicht: (2023)
SafeSteer: Adaptive Subspace Steering for Efficient Jailbreak Defense in Vision-Language Models
von: Zeng, Xiyu, et al.
Veröffentlicht: (2025)
von: Zeng, Xiyu, et al.
Veröffentlicht: (2025)
Differentially Private Distributed Inference
von: Papachristou, Marios, et al.
Veröffentlicht: (2024)
von: Papachristou, Marios, et al.
Veröffentlicht: (2024)
Scalable Multi-Agent Reinforcement Learning for Residential Load Scheduling under Data Governance
von: Qin, Zhaoming, et al.
Veröffentlicht: (2021)
von: Qin, Zhaoming, et al.
Veröffentlicht: (2021)
Ähnliche Einträge
-
Breaking the Secret: Economic Interventions for Combating Collusion in Embodied Multi-Agent Systems
von: Liu, Qi, et al.
Veröffentlicht: (2026) -
Memory Poisoning Attack and Defense on Memory Based LLM-Agents
von: Sunil, Balachandra Devarangadi, et al.
Veröffentlicht: (2026) -
OrchJail: Jailbreaking Tool-Calling Text-to-Image Agents by Orchestration-Guided Fuzzing
von: Chen, Jianming, et al.
Veröffentlicht: (2026) -
Prefix Probing: Lightweight Harmful Content Detection for Large Language Models
von: Yang, Jirui, et al.
Veröffentlicht: (2025) -
Multi-Agent Actor-Critics in Autonomous Cyber Defense
von: Wang, Mingjun, et al.
Veröffentlicht: (2024)