Plant, Persist, Trigger: Sleeper Attack on Large Language Model Agents
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Yongxiang, Li, Moxin, Ma, Zhixin, Zhu, Fengbin, Liu, Dongrui, Wang, Wenjie, Feng, Fuli |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
TAT-LLM: A Specialized Language Model for Discrete Reasoning over Tabular and Textual Data
di: Zhu, Fengbin, et al.
Pubblicazione: (2024)
di: Zhu, Fengbin, et al.
Pubblicazione: (2024)
TROJail: Trajectory-Level Optimization for Multi-Turn Large Language Model Jailbreaks with Process Rewards
di: Xiong, Xiqiao, et al.
Pubblicazione: (2025)
di: Xiong, Xiqiao, et al.
Pubblicazione: (2025)
Evaluating Mathematical Reasoning of Large Language Models: A Focus on Error Identification and Correction
di: Li, Xiaoyuan, et al.
Pubblicazione: (2024)
di: Li, Xiaoyuan, et al.
Pubblicazione: (2024)
Robust Prompt Optimization for Large Language Models Against Distribution Shifts
di: Li, Moxin, et al.
Pubblicazione: (2023)
di: Li, Moxin, et al.
Pubblicazione: (2023)
Doc2SoarGraph: Discrete Reasoning over Visually-Rich Table-Text Documents via Semantic-Oriented Hierarchical Graphs
di: Zhu, Fengbin, et al.
Pubblicazione: (2023)
di: Zhu, Fengbin, et al.
Pubblicazione: (2023)
One Adapts to Any: Meta Reward Modeling for Personalized LLM Alignment
di: Cai, Hongru, et al.
Pubblicazione: (2026)
di: Cai, Hongru, et al.
Pubblicazione: (2026)
Large Language Models Empowered Personalized Web Agents
di: Cai, Hongru, et al.
Pubblicazione: (2024)
di: Cai, Hongru, et al.
Pubblicazione: (2024)
Think Twice Before Trusting: Self-Detection for Large Language Models through Comprehensive Answer Reflection
di: Li, Moxin, et al.
Pubblicazione: (2024)
di: Li, Moxin, et al.
Pubblicazione: (2024)
CAPSUL: A Comprehensive Human Protein Benchmark for Subcellular Localization
di: Hu, Yicheng, et al.
Pubblicazione: (2026)
di: Hu, Yicheng, et al.
Pubblicazione: (2026)
MMDocBench: Benchmarking Large Vision-Language Models for Fine-Grained Visual Document Understanding
di: Zhu, Fengbin, et al.
Pubblicazione: (2024)
di: Zhu, Fengbin, et al.
Pubblicazione: (2024)
MURE: Hierarchical Multi-Resolution Encoding via Vision-Language Models for Visual Document Retrieval
di: Zhu, Fengbin, et al.
Pubblicazione: (2026)
di: Zhu, Fengbin, et al.
Pubblicazione: (2026)
Towards Temporal-Aware Multi-Modal Retrieval Augmented Generation in Finance
di: Zhu, Fengbin, et al.
Pubblicazione: (2025)
di: Zhu, Fengbin, et al.
Pubblicazione: (2025)
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
di: Hubinger, Evan, et al.
Pubblicazione: (2024)
di: Hubinger, Evan, et al.
Pubblicazione: (2024)
DynaTrust: Defending Multi-Agent Systems Against Sleeper Agents via Dynamic Trust Graphs
di: Li, Yu, et al.
Pubblicazione: (2026)
di: Li, Yu, et al.
Pubblicazione: (2026)
When Emotion Becomes Trigger: Emotion-style dynamic Backdoor Attack Parasitising Large Language Models
di: Liu, Ziyu, et al.
Pubblicazione: (2026)
di: Liu, Ziyu, et al.
Pubblicazione: (2026)
Detecting Sleeper Agents in Large Language Models via Semantic Drift Analysis
di: Zanbaghi, Shahin, et al.
Pubblicazione: (2025)
di: Zanbaghi, Shahin, et al.
Pubblicazione: (2025)
Can Large Language Models Derive New Knowledge? A Dynamic Benchmark for Biological Knowledge Discovery
di: Yang, Chaoqun, et al.
Pubblicazione: (2026)
di: Yang, Chaoqun, et al.
Pubblicazione: (2026)
Toward Personalized LLM-Powered Agents: Foundations, Evaluation, and Future Directions
di: Xu, Yue, et al.
Pubblicazione: (2026)
di: Xu, Yue, et al.
Pubblicazione: (2026)
BadThink: Triggered Overthinking Attacks on Chain-of-Thought Reasoning in Large Language Models
di: Liu, Shuaitong, et al.
Pubblicazione: (2025)
di: Liu, Shuaitong, et al.
Pubblicazione: (2025)
System Prompt Poisoning: Persistent Attacks on Large Language Models Beyond User Injection
di: Li, Zongze, et al.
Pubblicazione: (2025)
di: Li, Zongze, et al.
Pubblicazione: (2025)
Invisible to Humans, Triggered by Agents: Stealthy Jailbreak Attacks on Mobile Vision-Language Agents
di: Ding, Renhua, et al.
Pubblicazione: (2025)
di: Ding, Renhua, et al.
Pubblicazione: (2025)
LPS-Bench: Benchmarking Safety Awareness of Computer-Use Agents in Long-Horizon Planning under Benign and Adversarial Scenarios
di: Chen, Tianyu, et al.
Pubblicazione: (2026)
di: Chen, Tianyu, et al.
Pubblicazione: (2026)
Personalized Image Generation with Large Multimodal Models
di: Xu, Yiyan, et al.
Pubblicazione: (2024)
di: Xu, Yiyan, et al.
Pubblicazione: (2024)
Adaptive Tool Use in Large Language Models with Meta-Cognition Trigger
di: Li, Wenjun, et al.
Pubblicazione: (2025)
di: Li, Wenjun, et al.
Pubblicazione: (2025)
The Tug of War Within: Mitigating the Fairness-Privacy Conflicts in Large Language Models
di: Qian, Chen, et al.
Pubblicazione: (2024)
di: Qian, Chen, et al.
Pubblicazione: (2024)
Single-Node Trigger Backdoor Attacks in Graph-Based Recommendation Systems
di: Li, Runze, et al.
Pubblicazione: (2025)
di: Li, Runze, et al.
Pubblicazione: (2025)
S$^4$ST: A Strong, Self-transferable, faSt, and Simple Scale Transformation for Transferable Targeted Attack
di: Liu, Yongxiang, et al.
Pubblicazione: (2024)
di: Liu, Yongxiang, et al.
Pubblicazione: (2024)
CORBA: Contagious Recursive Blocking Attacks on Multi-Agent Systems Based on Large Language Models
di: Zhou, Zhenhong, et al.
Pubblicazione: (2025)
di: Zhou, Zhenhong, et al.
Pubblicazione: (2025)
Evolving Jailbreaks: Automated Multi-Objective Long-Tail Attacks on Large Language Models
di: Hong, Wenjing, et al.
Pubblicazione: (2026)
di: Hong, Wenjing, et al.
Pubblicazione: (2026)
Medical Reasoning with Large Language Models: A Survey and MR-Bench
di: Ren, Xiaohan, et al.
Pubblicazione: (2026)
di: Ren, Xiaohan, et al.
Pubblicazione: (2026)
Omni-R1: Towards the Unified Generative Paradigm for Multimodal Reasoning
di: Cheng, Dongjie, et al.
Pubblicazione: (2026)
di: Cheng, Dongjie, et al.
Pubblicazione: (2026)
COMAP: Co-Evolving World Models and Agent Policies for LLM Agents
di: Liu, Youwei, et al.
Pubblicazione: (2026)
di: Liu, Youwei, et al.
Pubblicazione: (2026)
PEPA: a Persistently Autonomous Embodied Agent with Personalities
di: Liu, Kaige, et al.
Pubblicazione: (2026)
di: Liu, Kaige, et al.
Pubblicazione: (2026)
Opening the Black Box: A Survey on the Mechanisms of Multi-Step Reasoning in Large Language Models
di: Pan, Liangming, et al.
Pubblicazione: (2026)
di: Pan, Liangming, et al.
Pubblicazione: (2026)
The Butterfly Effect of Model Editing: Few Edits Can Trigger Large Language Models Collapse
di: Yang, Wanli, et al.
Pubblicazione: (2024)
di: Yang, Wanli, et al.
Pubblicazione: (2024)
$\textit{LinkPrompt}$: Natural and Universal Adversarial Attacks on Prompt-based Language Models
di: Xu, Yue, et al.
Pubblicazione: (2024)
di: Xu, Yue, et al.
Pubblicazione: (2024)
AppAgent-Pro: A Proactive GUI Agent System for Multidomain Information Integration and User Assistance
di: Zhao, Yuyang, et al.
Pubblicazione: (2025)
di: Zhao, Yuyang, et al.
Pubblicazione: (2025)
A Survey on Model Compression for Large Language Models
di: Zhu, Xunyu, et al.
Pubblicazione: (2023)
di: Zhu, Xunyu, et al.
Pubblicazione: (2023)
Scale over Preference: The Impact of AI-Generated Content on Online Content Ecology
di: Shi, Tianhao, et al.
Pubblicazione: (2026)
di: Shi, Tianhao, et al.
Pubblicazione: (2026)
BackdoorLLM: A Comprehensive Benchmark for Backdoor Attacks and Defenses on Large Language Models
di: Li, Yige, et al.
Pubblicazione: (2024)
di: Li, Yige, et al.
Pubblicazione: (2024)
Documenti analoghi
-
TAT-LLM: A Specialized Language Model for Discrete Reasoning over Tabular and Textual Data
di: Zhu, Fengbin, et al.
Pubblicazione: (2024) -
TROJail: Trajectory-Level Optimization for Multi-Turn Large Language Model Jailbreaks with Process Rewards
di: Xiong, Xiqiao, et al.
Pubblicazione: (2025) -
Evaluating Mathematical Reasoning of Large Language Models: A Focus on Error Identification and Correction
di: Li, Xiaoyuan, et al.
Pubblicazione: (2024) -
Robust Prompt Optimization for Large Language Models Against Distribution Shifts
di: Li, Moxin, et al.
Pubblicazione: (2023) -
Doc2SoarGraph: Discrete Reasoning over Visually-Rich Table-Text Documents via Semantic-Oriented Hierarchical Graphs
di: Zhu, Fengbin, et al.
Pubblicazione: (2023)