Monitoring Decomposition Attacks in LLMs with Lightweight Sequential Monitors
Fuente:
arXiv
Saved in:
| Main Authors: | Yueh-Han, Chen, Joshi, Nitish, Chen, Yulin, Andriushchenko, Maksym, Angell, Rico, He, He |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
by: Andriushchenko, Maksym, et al.
Published: (2024)
by: Andriushchenko, Maksym, et al.
Published: (2024)
Adaptive Attacks on Trusted Monitors Subvert AI Control Protocols
by: Terekhov, Mikhail, et al.
Published: (2025)
by: Terekhov, Mikhail, et al.
Published: (2025)
Claudini: Autoresearch Discovers State-of-the-Art Adversarial Attack Algorithms for LLMs
by: Panfilov, Alexander, et al.
Published: (2026)
by: Panfilov, Alexander, et al.
Published: (2026)
Preventing Non-intrusive Load Monitoring Privacy Invasion: A Precise Adversarial Attack Scheme for Networked Smart Meters
by: He, Jialing, et al.
Published: (2024)
by: He, Jialing, et al.
Published: (2024)
SLEIGHT-Bench: A Benchmark of Evasion Attacks Against Agent Monitors
by: Najt, Elle, et al.
Published: (2026)
by: Najt, Elle, et al.
Published: (2026)
Is Monitoring Enough? Strategic Agent Selection For Stealthy Attack in Multi-Agent Discussions
by: Xiang, Qiuchi, et al.
Published: (2026)
by: Xiang, Qiuchi, et al.
Published: (2026)
Strategic Dishonesty Can Undermine AI Safety Evaluations of Frontier LLMs
by: Panfilov, Alexander, et al.
Published: (2025)
by: Panfilov, Alexander, et al.
Published: (2025)
Capability-Based Scaling Trends for LLM-Based Red-Teaming
by: Panfilov, Alexander, et al.
Published: (2025)
by: Panfilov, Alexander, et al.
Published: (2025)
FlipAttack: Jailbreak LLMs via Flipping
by: Liu, Yue, et al.
Published: (2024)
by: Liu, Yue, et al.
Published: (2024)
MonitoringBench: Semi-Automated Red-Teaming for Agent Monitoring
by: Jotautaitė, Monika, et al.
Published: (2026)
by: Jotautaitė, Monika, et al.
Published: (2026)
LLMs Can Covertly Sandbag on Capability Evaluations Against Chain-of-Thought Monitoring
by: Li, Chloe, et al.
Published: (2025)
by: Li, Chloe, et al.
Published: (2025)
Competition Report: Finding Universal Jailbreak Backdoors in Aligned LLMs
by: Rando, Javier, et al.
Published: (2024)
by: Rando, Javier, et al.
Published: (2024)
FedSecurity: Benchmarking Attacks and Defenses in Federated Learning and Federated LLMs
by: Han, Shanshan, et al.
Published: (2023)
by: Han, Shanshan, et al.
Published: (2023)
CoT-Guard: Small Models for Strong Monitoring
by: Diwan, Nirav, et al.
Published: (2026)
by: Diwan, Nirav, et al.
Published: (2026)
AttackSeqBench: Benchmarking the Capabilities of LLMs for Attack Sequences Understanding
by: Ma, Haokai, et al.
Published: (2025)
by: Ma, Haokai, et al.
Published: (2025)
Bidirectional Intention Inference Enhances LLMs' Defense Against Multi-Turn Jailbreak Attacks
by: Tong, Haibo, et al.
Published: (2025)
by: Tong, Haibo, et al.
Published: (2025)
Persistent Backdoor Attacks under Continual Fine-Tuning of LLMs
by: Cui, Jing, et al.
Published: (2025)
by: Cui, Jing, et al.
Published: (2025)
Invisible Backdoor Attack Through Singular Value Decomposition
by: Chen, Wenmin, et al.
Published: (2024)
by: Chen, Wenmin, et al.
Published: (2024)
ShadowCode: Towards (Automatic) External Prompt Injection Attack against Code LLMs
by: Yang, Yuchen, et al.
Published: (2024)
by: Yang, Yuchen, et al.
Published: (2024)
Reasoning-Style Poisoning of LLM Agents via Stealthy Style Transfer: Process-Level Attacks and Runtime Monitoring in RSV Space
by: Zhou, Xingfu, et al.
Published: (2025)
by: Zhou, Xingfu, et al.
Published: (2025)
When Efficiency Backfires: Cascading LLMs Trigger Cascade Failure under Adversarial Attack
by: Sun, Zehan, et al.
Published: (2026)
by: Sun, Zehan, et al.
Published: (2026)
CL-Attack: Textual Backdoor Attacks via Cross-Lingual Triggers
by: Zheng, Jingyi, et al.
Published: (2024)
by: Zheng, Jingyi, et al.
Published: (2024)
DITTO: A Spoofing Attack Framework on Watermarked LLMs via Knowledge Distillation
by: An, Hyeseon, et al.
Published: (2025)
by: An, Hyeseon, et al.
Published: (2025)
ADAM: A Systematic Data Extraction Attack on Agent Memory via Adaptive Querying
by: Lyu, Xingyu, et al.
Published: (2026)
by: Lyu, Xingyu, et al.
Published: (2026)
Backdoor Attack against One-Class Sequential Anomaly Detection Models
by: Cheng, He, et al.
Published: (2024)
by: Cheng, He, et al.
Published: (2024)
PrivacyXray: Detecting Privacy Breaches in LLMs through Semantic Consistency and Probability Certainty
by: He, Jinwen, et al.
Published: (2025)
by: He, Jinwen, et al.
Published: (2025)
Ensemble Privacy Defense for Knowledge-Intensive LLMs against Membership Inference Attacks
by: Fu, Haowei, et al.
Published: (2025)
by: Fu, Haowei, et al.
Published: (2025)
Factor(U,T): Controlling Untrusted AI by Monitoring their Plans
by: Lip, Edward Lue Chee, et al.
Published: (2025)
by: Lip, Edward Lue Chee, et al.
Published: (2025)
CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations
by: Li, Xiaohu, et al.
Published: (2025)
by: Li, Xiaohu, et al.
Published: (2025)
Retrieval-Confused Generation is a Good Defender for Privacy Violation Attack of Large Language Models
by: Peng, Wanli, et al.
Published: (2025)
by: Peng, Wanli, et al.
Published: (2025)
Explainer-guided Targeted Adversarial Attacks against Binary Code Similarity Detection Models
by: Chen, Mingjie, et al.
Published: (2025)
by: Chen, Mingjie, et al.
Published: (2025)
PASA: A Principled Embedding-Space Watermarking Approach for LLM-Generated Text under Semantic-Invariant Attacks
by: Ai, Zhenxin, et al.
Published: (2026)
by: Ai, Zhenxin, et al.
Published: (2026)
LLM Agent Honeypot: Monitoring AI Hacking Agents in the Wild
by: Reworr, et al.
Published: (2024)
by: Reworr, et al.
Published: (2024)
Stop Tracking Me! Proactive Defense Against Attribute Inference Attack in LLMs
by: Yan, Dong, et al.
Published: (2026)
by: Yan, Dong, et al.
Published: (2026)
SeqAR: Jailbreak LLMs with Sequential Auto-Generated Characters
by: Yang, Yan, et al.
Published: (2024)
by: Yang, Yan, et al.
Published: (2024)
PromptKeeper: Safeguarding System Prompts for LLMs
by: Jiang, Zhifeng, et al.
Published: (2024)
by: Jiang, Zhifeng, et al.
Published: (2024)
Attention Tracker: Detecting Prompt Injection Attacks in LLMs
by: Hung, Kuo-Han, et al.
Published: (2024)
by: Hung, Kuo-Han, et al.
Published: (2024)
AdapTools: Adaptive Tool-based Indirect Prompt Injection Attacks on Agentic LLMs
by: Wang, Che, et al.
Published: (2026)
by: Wang, Che, et al.
Published: (2026)
ELBA-Bench: An Efficient Learning Backdoor Attacks Benchmark for Large Language Models
by: Liu, Xuxu, et al.
Published: (2025)
by: Liu, Xuxu, et al.
Published: (2025)
Scam Shield: Multi-Model Voting and Fine-Tuned LLMs Against Adversarial Attacks
by: Chang, Chen-Wei, et al.
Published: (2025)
by: Chang, Chen-Wei, et al.
Published: (2025)
Similar Items
-
Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
by: Andriushchenko, Maksym, et al.
Published: (2024) -
Adaptive Attacks on Trusted Monitors Subvert AI Control Protocols
by: Terekhov, Mikhail, et al.
Published: (2025) -
Claudini: Autoresearch Discovers State-of-the-Art Adversarial Attack Algorithms for LLMs
by: Panfilov, Alexander, et al.
Published: (2026) -
Preventing Non-intrusive Load Monitoring Privacy Invasion: A Precise Adversarial Attack Scheme for Networked Smart Meters
by: He, Jialing, et al.
Published: (2024) -
SLEIGHT-Bench: A Benchmark of Evasion Attacks Against Agent Monitors
by: Najt, Elle, et al.
Published: (2026)