Obfuscated Activations Bypass LLM Latent-Space Defenses
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bailey, Luke, Serrano, Alex, Sheshadri, Abhay, Seleznyov, Mikhail, Taylor, Jordan, Jenner, Erik, Hilton, Jacob, Casper, Stephen, Guestrin, Carlos, Emmons, Scott |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Neural Chameleons: Language Models Can Learn to Hide Their Thoughts from Unseen Activation Monitors
von: McGuinness, Max, et al.
Veröffentlicht: (2025)
von: McGuinness, Max, et al.
Veröffentlicht: (2025)
RL-Obfuscation: Can Language Models Learn to Evade Latent-Space Monitors?
von: Gupta, Rohan, et al.
Veröffentlicht: (2025)
von: Gupta, Rohan, et al.
Veröffentlicht: (2025)
Image Hijacks: Adversarial Images can Control Generative Models at Runtime
von: Bailey, Luke, et al.
Veröffentlicht: (2023)
von: Bailey, Luke, et al.
Veröffentlicht: (2023)
Practical Principles for AI Cost and Compute Accounting
von: Casper, Stephen, et al.
Veröffentlicht: (2025)
von: Casper, Stephen, et al.
Veröffentlicht: (2025)
When Your AIs Deceive You: Challenges of Partial Observability in Reinforcement Learning from Human Feedback
von: Lang, Leon, et al.
Veröffentlicht: (2024)
von: Lang, Leon, et al.
Veröffentlicht: (2024)
Evidence of Learned Look-Ahead in a Chess-Playing Neural Network
von: Jenner, Erik, et al.
Veröffentlicht: (2024)
von: Jenner, Erik, et al.
Veröffentlicht: (2024)
Can Reasoning Models Obfuscate Reasoning? Stress-Testing Chain-of-Thought Monitorability
von: Zolkowski, Artur, et al.
Veröffentlicht: (2025)
von: Zolkowski, Artur, et al.
Veröffentlicht: (2025)
Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs
von: Sheshadri, Abhay, et al.
Veröffentlicht: (2024)
von: Sheshadri, Abhay, et al.
Veröffentlicht: (2024)
Frontier Models Can Take Actions at Low Probabilities
von: Serrano, Alex, et al.
Veröffentlicht: (2026)
von: Serrano, Alex, et al.
Veröffentlicht: (2026)
Output Supervision Can Obfuscate the Chain of Thought
von: Drori, Jacob, et al.
Veröffentlicht: (2025)
von: Drori, Jacob, et al.
Veröffentlicht: (2025)
When Chain of Thought is Necessary, Language Models Struggle to Evade Monitors
von: Emmons, Scott, et al.
Veröffentlicht: (2025)
von: Emmons, Scott, et al.
Veröffentlicht: (2025)
Shield Bash: Abusing Defensive Coherence State Retrieval to Break Timing Obfuscation
von: Ramkrishnan, Kartik, et al.
Veröffentlicht: (2025)
von: Ramkrishnan, Kartik, et al.
Veröffentlicht: (2025)
Why Do Some Language Models Fake Alignment While Others Don't?
von: Sheshadri, Abhay, et al.
Veröffentlicht: (2025)
von: Sheshadri, Abhay, et al.
Veröffentlicht: (2025)
Breaking the Chain: A Causal Analysis of LLM Faithfulness to Intermediate Structures
von: Somov, Oleg, et al.
Veröffentlicht: (2026)
von: Somov, Oleg, et al.
Veröffentlicht: (2026)
Uncovering Latent Human Wellbeing in Language Model Embeddings
von: Freire, Pedro, et al.
Veröffentlicht: (2024)
von: Freire, Pedro, et al.
Veröffentlicht: (2024)
On the Geometric Limits of Transformer Defenses against Obfuscation Attacks: Latent Embedding Collapse & Performance Robustness Gap
von: Mashaido, Becky, et al.
Veröffentlicht: (2026)
von: Mashaido, Becky, et al.
Veröffentlicht: (2026)
A Mechanistic Analysis of a Transformer Trained on a Symbolic Multi-Step Reasoning Task
von: Brinkmann, Jannik, et al.
Veröffentlicht: (2024)
von: Brinkmann, Jannik, et al.
Veröffentlicht: (2024)
ViSTa Dataset: Do vision-language models understand sequential tasks?
von: Wybitul, Evžen, et al.
Veröffentlicht: (2024)
von: Wybitul, Evžen, et al.
Veröffentlicht: (2024)
On the Transit Obfuscation Problem
von: Takahashi, Hideaki, et al.
Veröffentlicht: (2024)
von: Takahashi, Hideaki, et al.
Veröffentlicht: (2024)
Mechanistic Unlearning: Robust Knowledge Unlearning and Editing via Mechanistic Localization
von: Guo, Phillip, et al.
Veröffentlicht: (2024)
von: Guo, Phillip, et al.
Veröffentlicht: (2024)
When Punctuation Matters: A Large-Scale Comparison of Prompt Robustness Methods for LLMs
von: Seleznyov, Mikhail, et al.
Veröffentlicht: (2025)
von: Seleznyov, Mikhail, et al.
Veröffentlicht: (2025)
Welcome First--Books Later; The Service Center Branch, Richmond Public Library, December 1967 - June 1971.
von: Emmons, Karen
Veröffentlicht: (1971)
von: Emmons, Karen
Veröffentlicht: (1971)
She's Practiced What She Teaches
von: Emmons, Julia
Veröffentlicht: (1976)
von: Emmons, Julia
Veröffentlicht: (1976)
A Call to Excellence and Innovation: A Survey of the East Saint Louis, Illinois, Public Library.
von: Jordan, Casper L.
Veröffentlicht: (1972)
von: Jordan, Casper L.
Veröffentlicht: (1972)
Library Service to Black Americans
von: Jordan, Casper Leroy
Veröffentlicht: (1971)
von: Jordan, Casper Leroy
Veröffentlicht: (1971)
xCOMET-lite: Bridging the Gap Between Efficiency and Quality in Learned MT Evaluation Metrics
von: Larionov, Daniil, et al.
Veröffentlicht: (2024)
von: Larionov, Daniil, et al.
Veröffentlicht: (2024)
ESLD (External Surrogate Latent Defense): A Latent-Space Architecture for Faster, Stronger Prompt-Injection Defense
von: Narendra, Yash
Veröffentlicht: (2026)
von: Narendra, Yash
Veröffentlicht: (2026)
ActivationReasoning: Logical Reasoning in Latent Activation Spaces
von: Helff, Lukas, et al.
Veröffentlicht: (2025)
von: Helff, Lukas, et al.
Veröffentlicht: (2025)
Stealthy Poisoning Attacks Bypass Defenses in Regression Settings
von: Carnerero-Cano, Javier, et al.
Veröffentlicht: (2026)
von: Carnerero-Cano, Javier, et al.
Veröffentlicht: (2026)
Bypassing DARCY Defense: Indistinguishable Universal Adversarial Triggers
von: Peng, Zuquan, et al.
Veröffentlicht: (2024)
von: Peng, Zuquan, et al.
Veröffentlicht: (2024)
Diffusion On Syntax Trees For Program Synthesis
von: Kapur, Shreyas, et al.
Veröffentlicht: (2024)
von: Kapur, Shreyas, et al.
Veröffentlicht: (2024)
Advancing Jailbreak Strategies: A Hybrid Approach to Exploiting LLM Vulnerabilities and Bypassing Modern Defenses
von: Ahmed, Mohamed, et al.
Veröffentlicht: (2025)
von: Ahmed, Mohamed, et al.
Veröffentlicht: (2025)
Black Academic Libraries: An Inventory.
von: Jordan, Casper LeRoy
Veröffentlicht: (1970)
von: Jordan, Casper LeRoy
Veröffentlicht: (1970)
Defending Against Unforeseen Failure Modes with Latent Adversarial Training
von: Casper, Stephen, et al.
Veröffentlicht: (2024)
von: Casper, Stephen, et al.
Veröffentlicht: (2024)
Defense-to-Attack: Bypassing Weak Defenses Enables Stronger Jailbreaks in Vision-Language Models
von: Zhao, Yunhan, et al.
Veröffentlicht: (2025)
von: Zhao, Yunhan, et al.
Veröffentlicht: (2025)
CLOAK: Contrastive Guidance for Latent Diffusion-Based Data Obfuscation
von: Yang, Xin, et al.
Veröffentlicht: (2025)
von: Yang, Xin, et al.
Veröffentlicht: (2025)
Operational Latent Spaces
von: Hawley, Scott H., et al.
Veröffentlicht: (2024)
von: Hawley, Scott H., et al.
Veröffentlicht: (2024)
BadActs: A Universal Backdoor Defense in the Activation Space
von: Yi, Biao, et al.
Veröffentlicht: (2024)
von: Yi, Biao, et al.
Veröffentlicht: (2024)
Evaluation of Prompt Injection Defenses in Large Language Models
von: Deep, Priyal, et al.
Veröffentlicht: (2026)
von: Deep, Priyal, et al.
Veröffentlicht: (2026)
Evolutionary Search for Automated Design of Uncertainty Quantification Methods
von: Seleznyov, Mikhail, et al.
Veröffentlicht: (2026)
von: Seleznyov, Mikhail, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Neural Chameleons: Language Models Can Learn to Hide Their Thoughts from Unseen Activation Monitors
von: McGuinness, Max, et al.
Veröffentlicht: (2025) -
RL-Obfuscation: Can Language Models Learn to Evade Latent-Space Monitors?
von: Gupta, Rohan, et al.
Veröffentlicht: (2025) -
Image Hijacks: Adversarial Images can Control Generative Models at Runtime
von: Bailey, Luke, et al.
Veröffentlicht: (2023) -
Practical Principles for AI Cost and Compute Accounting
von: Casper, Stephen, et al.
Veröffentlicht: (2025) -
When Your AIs Deceive You: Challenges of Partial Observability in Reinforcement Learning from Human Feedback
von: Lang, Leon, et al.
Veröffentlicht: (2024)