RL-Obfuscation: Can Language Models Learn to Evade Latent-Space Monitors?
Fuente:
arXiv
Saved in:
| Main Authors: | Gupta, Rohan, Jenner, Erik |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Obfuscated Activations Bypass LLM Latent-Space Defenses
by: Bailey, Luke, et al.
Published: (2024)
by: Bailey, Luke, et al.
Published: (2024)
Frontier Models Can Take Actions at Low Probabilities
by: Serrano, Alex, et al.
Published: (2026)
by: Serrano, Alex, et al.
Published: (2026)
Evading Deep Learning-Based Malware Detectors via Obfuscation: A Deep Reinforcement Learning Approach
by: Etter, Brian, et al.
Published: (2024)
by: Etter, Brian, et al.
Published: (2024)
Can Reasoning Models Obfuscate Reasoning? Stress-Testing Chain-of-Thought Monitorability
by: Zolkowski, Artur, et al.
Published: (2025)
by: Zolkowski, Artur, et al.
Published: (2025)
SALSA-RL: Stability Analysis in the Latent Space of Actions for Reinforcement Learning
by: Li, Xuyang, et al.
Published: (2025)
by: Li, Xuyang, et al.
Published: (2025)
Training on Documents About Monitoring Leads to CoT Obfuscation
by: Haskins, Reilly, et al.
Published: (2026)
by: Haskins, Reilly, et al.
Published: (2026)
Prompt Obfuscation for Large Language Models
by: Pape, David, et al.
Published: (2024)
by: Pape, David, et al.
Published: (2024)
RL-Guided Data Selection for Language Model Finetuning
by: Jha, Animesh, et al.
Published: (2025)
by: Jha, Animesh, et al.
Published: (2025)
Watch Out! Simple Horizontal Class Backdoor Can Trivially Evade Defense
by: Ma, Hua, et al.
Published: (2023)
by: Ma, Hua, et al.
Published: (2023)
When Chain of Thought is Necessary, Language Models Struggle to Evade Monitors
by: Emmons, Scott, et al.
Published: (2025)
by: Emmons, Scott, et al.
Published: (2025)
Gradient Coding in Decentralized Learning for Evading Stragglers
by: Li, Chengxi, et al.
Published: (2024)
by: Li, Chengxi, et al.
Published: (2024)
CLOAK: Contrastive Guidance for Latent Diffusion-Based Data Obfuscation
by: Yang, Xin, et al.
Published: (2025)
by: Yang, Xin, et al.
Published: (2025)
Output Supervision Can Obfuscate the Chain of Thought
by: Drori, Jacob, et al.
Published: (2025)
by: Drori, Jacob, et al.
Published: (2025)
Evading Data Contamination Detection for Language Models is (too) Easy
by: Dekoninck, Jasper, et al.
Published: (2024)
by: Dekoninck, Jasper, et al.
Published: (2024)
Understanding the Staged Dynamics of Transformers in Learning Latent Structure
by: Saha, Rohan, et al.
Published: (2025)
by: Saha, Rohan, et al.
Published: (2025)
An Invariant Latent Space Perspective on Language Model Inversion
by: Ye, Wentao, et al.
Published: (2025)
by: Ye, Wentao, et al.
Published: (2025)
Inference Time Policy Optimization for Offline RL with Differentiable World Models
by: Deb, Rohan, et al.
Published: (2026)
by: Deb, Rohan, et al.
Published: (2026)
Generative Model for Small Molecules with Latent Space RL Fine-Tuning to Protein Targets
by: Sob, Ulrich A. Mbou, et al.
Published: (2024)
by: Sob, Ulrich A. Mbou, et al.
Published: (2024)
Command-line Obfuscation Detection using Small Language Models
by: Outrata, Vojtech, et al.
Published: (2024)
by: Outrata, Vojtech, et al.
Published: (2024)
Injecting Undetectable Backdoors in Obfuscated Neural Networks and Language Models
by: Kalavasis, Alkis, et al.
Published: (2024)
by: Kalavasis, Alkis, et al.
Published: (2024)
Neural Chameleons: Language Models Can Learn to Hide Their Thoughts from Unseen Activation Monitors
by: McGuinness, Max, et al.
Published: (2025)
by: McGuinness, Max, et al.
Published: (2025)
Latent Matters: Learning Deep State-Space Models
by: Klushyn, Alexej, et al.
Published: (2026)
by: Klushyn, Alexej, et al.
Published: (2026)
Steering Your Diffusion Policy with Latent Space Reinforcement Learning
by: Wagenmaker, Andrew, et al.
Published: (2025)
by: Wagenmaker, Andrew, et al.
Published: (2025)
Learning Dynamics in RL Post-Training for Language Models
by: Tomihari, Akiyoshi
Published: (2026)
by: Tomihari, Akiyoshi
Published: (2026)
Monitoring Latent World States in Language Models with Propositional Probes
by: Feng, Jiahai, et al.
Published: (2024)
by: Feng, Jiahai, et al.
Published: (2024)
LLMs Can Learn to Reason Via Off-Policy RL
by: Ritter, Daniel, et al.
Published: (2026)
by: Ritter, Daniel, et al.
Published: (2026)
A Deep Latent Space Model for Graph Representation Learning
by: Yang, Hanxuan, et al.
Published: (2021)
by: Yang, Hanxuan, et al.
Published: (2021)
Contrast Sets for Evaluating Language-Guided Robot Policies
by: Anwar, Abrar, et al.
Published: (2024)
by: Anwar, Abrar, et al.
Published: (2024)
AuthorMist: Evading AI Text Detectors with Reinforcement Learning
by: David, Isaac, et al.
Published: (2025)
by: David, Isaac, et al.
Published: (2025)
Evidence of Learned Look-Ahead in a Chess-Playing Neural Network
by: Jenner, Erik, et al.
Published: (2024)
by: Jenner, Erik, et al.
Published: (2024)
Inverse Optimization Latent Variable Models for Learning Costs Applied to Route Problems
by: Lahoud, Alan A., et al.
Published: (2025)
by: Lahoud, Alan A., et al.
Published: (2025)
SLAC: Simulation-Pretrained Latent Action Space for Whole-Body Real-World RL
by: Hu, Jiaheng, et al.
Published: (2025)
by: Hu, Jiaheng, et al.
Published: (2025)
Learning Abstract World Models with a Group-Structured Latent Space
by: Delliaux, Thomas, et al.
Published: (2025)
by: Delliaux, Thomas, et al.
Published: (2025)
When Your AIs Deceive You: Challenges of Partial Observability in Reinforcement Learning from Human Feedback
by: Lang, Leon, et al.
Published: (2024)
by: Lang, Leon, et al.
Published: (2024)
Evading Black-box Classifiers Without Breaking Eggs
by: Debenedetti, Edoardo, et al.
Published: (2023)
by: Debenedetti, Edoardo, et al.
Published: (2023)
Exploration Hacking: Can LLMs Learn to Resist RL Training?
by: Jang, Eyon, et al.
Published: (2026)
by: Jang, Eyon, et al.
Published: (2026)
LatentBreak: Jailbreaking Large Language Models through Latent Space Feedback
by: Mura, Raffaele, et al.
Published: (2025)
by: Mura, Raffaele, et al.
Published: (2025)
Geometry of Uncertainty: Learning Metric Spaces for Multimodal State Estimation in RL
by: Reichlin, Alfredo, et al.
Published: (2026)
by: Reichlin, Alfredo, et al.
Published: (2026)
Autoregressivity in the Latent Space of a GP-VAE Language Model: An Empirical Ablation Study
by: Ruffenach, Yves
Published: (2025)
by: Ruffenach, Yves
Published: (2025)
OSNIP: Breaking the Privacy-Utility-Efficiency Trilemma in LLM Inference via Obfuscated Semantic Null Space
by: Cao, Zhiyuan, et al.
Published: (2026)
by: Cao, Zhiyuan, et al.
Published: (2026)
Similar Items
-
Obfuscated Activations Bypass LLM Latent-Space Defenses
by: Bailey, Luke, et al.
Published: (2024) -
Frontier Models Can Take Actions at Low Probabilities
by: Serrano, Alex, et al.
Published: (2026) -
Evading Deep Learning-Based Malware Detectors via Obfuscation: A Deep Reinforcement Learning Approach
by: Etter, Brian, et al.
Published: (2024) -
Can Reasoning Models Obfuscate Reasoning? Stress-Testing Chain-of-Thought Monitorability
by: Zolkowski, Artur, et al.
Published: (2025) -
SALSA-RL: Stability Analysis in the Latent Space of Actions for Reinforcement Learning
by: Li, Xuyang, et al.
Published: (2025)