Neural Chameleons: Language Models Can Learn to Hide Their Thoughts from Unseen Activation Monitors
Fuente:
arXiv
Saved in:
| Main Authors: | McGuinness, Max, Serrano, Alex, Bailey, Luke, Emmons, Scott |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Path Integral Optimiser: Global Optimisation via Neural Schrödinger-Föllmer Diffusion
by: McGuinness, Max, et al.
Published: (2025)
by: McGuinness, Max, et al.
Published: (2025)
Skill Issues: An Analysis of CS:GO Skill Rating Systems
by: Bober-Irizar, Mikel, et al.
Published: (2024)
by: Bober-Irizar, Mikel, et al.
Published: (2024)
Obfuscated Activations Bypass LLM Latent-Space Defenses
by: Bailey, Luke, et al.
Published: (2024)
by: Bailey, Luke, et al.
Published: (2024)
Image Hijacks: Adversarial Images can Control Generative Models at Runtime
by: Bailey, Luke, et al.
Published: (2023)
by: Bailey, Luke, et al.
Published: (2023)
A Pragmatic Way to Measure Chain-of-Thought Monitorability
by: Emmons, Scott, et al.
Published: (2025)
by: Emmons, Scott, et al.
Published: (2025)
Gaming and Blockchain: Hype and Reality
by: McGuinness, Max
Published: (2024)
by: McGuinness, Max
Published: (2024)
ALMANACS: A Simulatability Benchmark for Language Model Explainability
by: Mills, Edmund, et al.
Published: (2023)
by: Mills, Edmund, et al.
Published: (2023)
Hide to Guide: Learning via Semantic Masking
by: Liu, Ruitao, et al.
Published: (2026)
by: Liu, Ruitao, et al.
Published: (2026)
Output Supervision Can Obfuscate the Chain of Thought
by: Drori, Jacob, et al.
Published: (2025)
by: Drori, Jacob, et al.
Published: (2025)
Dataset Clustering for Improved Offline Policy Learning
by: Wang, Qiang, et al.
Published: (2024)
by: Wang, Qiang, et al.
Published: (2024)
MetaExplainer: A Framework to Generate Multi-Type User-Centered Explanations for AI Systems
by: Chari, Shruthi, et al.
Published: (2025)
by: Chari, Shruthi, et al.
Published: (2025)
Evidence of Learned Look-Ahead in a Chess-Playing Neural Network
by: Jenner, Erik, et al.
Published: (2024)
by: Jenner, Erik, et al.
Published: (2024)
Short Rainbow Circuits in Regular Matroids
by: McGuinness, Sean
Published: (2026)
by: McGuinness, Sean
Published: (2026)
Cyclic Orderings of Paving Matroids
by: McGuinness, Sean
Published: (2023)
by: McGuinness, Sean
Published: (2023)
Skew circuits and circumference in a binary matroid
by: McGuinness, Sean
Published: (2024)
by: McGuinness, Sean
Published: (2024)
FedHide: Federated Learning by Hiding in the Neighbors
by: Park, Hyunsin, et al.
Published: (2024)
by: Park, Hyunsin, et al.
Published: (2024)
Can Watermarking Large Language Models Prevent Copyrighted Text Generation and Hide Training Data?
by: Panaitescu-Liess, Michael-Andrei, et al.
Published: (2024)
by: Panaitescu-Liess, Michael-Andrei, et al.
Published: (2024)
Exploration Hacking: Can LLMs Learn to Resist RL Training?
by: Jang, Eyon, et al.
Published: (2026)
by: Jang, Eyon, et al.
Published: (2026)
Genuine certifiable randomness from a black-box
by: McGuinness, Liam P.
Published: (2026)
by: McGuinness, Liam P.
Published: (2026)
Semantic Chameleon: Corpus-Dependent Poisoning Attacks and Defenses in RAG Systems
by: Thornton, Scott
Published: (2026)
by: Thornton, Scott
Published: (2026)
A Framework for Nonstationary Gaussian Processes with Neural Network Parameters
by: James, Zachary, et al.
Published: (2025)
by: James, Zachary, et al.
Published: (2025)
When Chain-of-Thought Fails, the Solution Hides in the Hidden States
by: Mehrafarin, Houman, et al.
Published: (2026)
by: Mehrafarin, Houman, et al.
Published: (2026)
Neural Language of Thought Models
by: Wu, Yi-Fu, et al.
Published: (2024)
by: Wu, Yi-Fu, et al.
Published: (2024)
RL-Obfuscation: Can Language Models Learn to Evade Latent-Space Monitors?
by: Gupta, Rohan, et al.
Published: (2025)
by: Gupta, Rohan, et al.
Published: (2025)
The Quantum Cramér-Rao lower bound (Why quantum computers won't work I)
by: McGuinness, Liam P.
Published: (2025)
by: McGuinness, Liam P.
Published: (2025)
Frontier Models Can Take Actions at Low Probabilities
by: Serrano, Alex, et al.
Published: (2026)
by: Serrano, Alex, et al.
Published: (2026)
Uncovering Latent Human Wellbeing in Language Model Embeddings
by: Freire, Pedro, et al.
Published: (2024)
by: Freire, Pedro, et al.
Published: (2024)
Can Language Models Learn Typologically Implausible Languages?
by: Xu, Tianyang, et al.
Published: (2025)
by: Xu, Tianyang, et al.
Published: (2025)
When Your AIs Deceive You: Challenges of Partial Observability in Reinforcement Learning from Human Feedback
by: Lang, Leon, et al.
Published: (2024)
by: Lang, Leon, et al.
Published: (2024)
Chameleon: A Flexible Data-mixing Framework for Language Model Pretraining and Finetuning
by: Xie, Wanyun, et al.
Published: (2025)
by: Xie, Wanyun, et al.
Published: (2025)
Reference-Specific Unlearning Metrics Can Hide the Truth: A Reality Check
by: Cho, Sungjun, et al.
Published: (2025)
by: Cho, Sungjun, et al.
Published: (2025)
Modeling Collapse of Steered Vine Robots Under Their Own Weight
by: McFarland, Ciera, et al.
Published: (2025)
by: McFarland, Ciera, et al.
Published: (2025)
Are Large Language Models Chameleons? An Attempt to Simulate Social Surveys
by: Geng, Mingmeng, et al.
Published: (2024)
by: Geng, Mingmeng, et al.
Published: (2024)
Can Interpretation Predict Behavior on Unseen Data?
by: Li, Victoria R., et al.
Published: (2025)
by: Li, Victoria R., et al.
Published: (2025)
Chameleon2++: An Efficient and Scalable Variant Of Chameleon Clustering
by: Singh, Priyanshu, et al.
Published: (2025)
by: Singh, Priyanshu, et al.
Published: (2025)
Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
by: Korbak, Tomek, et al.
Published: (2025)
by: Korbak, Tomek, et al.
Published: (2025)
Can AI Dream of Unseen Galaxies? Conditional Diffusion Model for Galaxy Morphology Augmentation
by: Ma, Chenrui, et al.
Published: (2025)
by: Ma, Chenrui, et al.
Published: (2025)
Beyond ReLU: How Activations Affect Neural Kernels and Random Wide Networks
by: Holzmüller, David, et al.
Published: (2025)
by: Holzmüller, David, et al.
Published: (2025)
State-Space Constraints Can Improve the Generalisation of the Differentiable Neural Computer to Input Sequences With Unseen Length
by: Ofner, Patrick, et al.
Published: (2021)
by: Ofner, Patrick, et al.
Published: (2021)
OOD-Chameleon: Is Algorithm Selection for OOD Generalization Learnable?
by: Jiang, Liangze, et al.
Published: (2024)
by: Jiang, Liangze, et al.
Published: (2024)
Similar Items
-
Path Integral Optimiser: Global Optimisation via Neural Schrödinger-Föllmer Diffusion
by: McGuinness, Max, et al.
Published: (2025) -
Skill Issues: An Analysis of CS:GO Skill Rating Systems
by: Bober-Irizar, Mikel, et al.
Published: (2024) -
Obfuscated Activations Bypass LLM Latent-Space Defenses
by: Bailey, Luke, et al.
Published: (2024) -
Image Hijacks: Adversarial Images can Control Generative Models at Runtime
by: Bailey, Luke, et al.
Published: (2023) -
A Pragmatic Way to Measure Chain-of-Thought Monitorability
by: Emmons, Scott, et al.
Published: (2025)