When Chain of Thought is Necessary, Language Models Struggle to Evade Monitors
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Emmons, Scott, Jenner, Erik, Elson, David K., Saurous, Rif A., Rajamanoharan, Senthooran, Chen, Heng, Shafkat, Irhum, Shah, Rohin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Pragmatic Way to Measure Chain-of-Thought Monitorability
von: Emmons, Scott, et al.
Veröffentlicht: (2025)
von: Emmons, Scott, et al.
Veröffentlicht: (2025)
Chain-of-Thought Reasoning In The Wild Is Not Always Faithful
von: Arcuschin, Iván, et al.
Veröffentlicht: (2025)
von: Arcuschin, Iván, et al.
Veröffentlicht: (2025)
Thought Branches: Interpreting LLM Reasoning Requires Resampling
von: Macar, Uzay, et al.
Veröffentlicht: (2025)
von: Macar, Uzay, et al.
Veröffentlicht: (2025)
Do I Know This Entity? Knowledge Awareness and Hallucinations in Language Models
von: Ferrando, Javier, et al.
Veröffentlicht: (2024)
von: Ferrando, Javier, et al.
Veröffentlicht: (2024)
RL-Obfuscation: Can Language Models Learn to Evade Latent-Space Monitors?
von: Gupta, Rohan, et al.
Veröffentlicht: (2025)
von: Gupta, Rohan, et al.
Veröffentlicht: (2025)
Emergent Misalignment is Easy, Narrow Misalignment is Hard
von: Soligo, Anna, et al.
Veröffentlicht: (2026)
von: Soligo, Anna, et al.
Veröffentlicht: (2026)
Can Confidence Estimates Decide When Chain-of-Thought Is Necessary for LLMs?
von: Lewis-Lim, Samuel, et al.
Veröffentlicht: (2025)
von: Lewis-Lim, Samuel, et al.
Veröffentlicht: (2025)
Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2
von: Lieberum, Tom, et al.
Veröffentlicht: (2024)
von: Lieberum, Tom, et al.
Veröffentlicht: (2024)
Aligned, Orthogonal or In-conflict: When can we safely optimize Chain-of-Thought?
von: Kaufmann, Max, et al.
Veröffentlicht: (2026)
von: Kaufmann, Max, et al.
Veröffentlicht: (2026)
Simple Mechanistic Explanations for Out-Of-Context Reasoning
von: Wang, Atticus, et al.
Veröffentlicht: (2025)
von: Wang, Atticus, et al.
Veröffentlicht: (2025)
All Random Features Representations are Equivalent
von: Sernau, Luke, et al.
Veröffentlicht: (2024)
von: Sernau, Luke, et al.
Veröffentlicht: (2024)
Steering Out-of-Distribution Generalization with Concept Ablation Fine-Tuning
von: Casademunt, Helena, et al.
Veröffentlicht: (2025)
von: Casademunt, Helena, et al.
Veröffentlicht: (2025)
Quantifying the Necessity of Chain of Thought through Opaque Serial Depth
von: Brown-Cohen, Jonah, et al.
Veröffentlicht: (2026)
von: Brown-Cohen, Jonah, et al.
Veröffentlicht: (2026)
Can Reasoning Models Obfuscate Reasoning? Stress-Testing Chain-of-Thought Monitorability
von: Zolkowski, Artur, et al.
Veröffentlicht: (2025)
von: Zolkowski, Artur, et al.
Veröffentlicht: (2025)
How Well Do Models Follow Their Constitutions?
von: Jakkli, Arya, et al.
Veröffentlicht: (2026)
von: Jakkli, Arya, et al.
Veröffentlicht: (2026)
Improving Dictionary Learning with Gated Sparse Autoencoders
von: Rajamanoharan, Senthooran, et al.
Veröffentlicht: (2024)
von: Rajamanoharan, Senthooran, et al.
Veröffentlicht: (2024)
Dense SAE Latents Are Features, Not Bugs
von: Sun, Xiaoqing, et al.
Veröffentlicht: (2025)
von: Sun, Xiaoqing, et al.
Veröffentlicht: (2025)
Uncovering Latent Chain of Thought Vectors in Language Models
von: Zhang, Jason, et al.
Veröffentlicht: (2024)
von: Zhang, Jason, et al.
Veröffentlicht: (2024)
When Your AIs Deceive You: Challenges of Partial Observability in Reinforcement Learning from Human Feedback
von: Lang, Leon, et al.
Veröffentlicht: (2024)
von: Lang, Leon, et al.
Veröffentlicht: (2024)
The Fragility of Chain-of-Thought Monitoring Across Typologically Diverse Languages
von: Onyame, Eric, et al.
Veröffentlicht: (2026)
von: Onyame, Eric, et al.
Veröffentlicht: (2026)
When Chain-of-Thought Backfires: Evaluating Prompt Sensitivity in Medical Language Models
von: Sadanandan, Binesh, et al.
Veröffentlicht: (2026)
von: Sadanandan, Binesh, et al.
Veröffentlicht: (2026)
ALMANACS: A Simulatability Benchmark for Language Model Explainability
von: Mills, Edmund, et al.
Veröffentlicht: (2023)
von: Mills, Edmund, et al.
Veröffentlicht: (2023)
All Code, No Thought: Current Language Models Struggle to Reason in Ciphered Language
von: Guo, Shiyuan, et al.
Veröffentlicht: (2025)
von: Guo, Shiyuan, et al.
Veröffentlicht: (2025)
Neural Chameleons: Language Models Can Learn to Hide Their Thoughts from Unseen Activation Monitors
von: McGuinness, Max, et al.
Veröffentlicht: (2025)
von: McGuinness, Max, et al.
Veröffentlicht: (2025)
Towards eliciting latent knowledge from LLMs with mechanistic interpretability
von: Cywiński, Bartosz, et al.
Veröffentlicht: (2025)
von: Cywiński, Bartosz, et al.
Veröffentlicht: (2025)
Convergent Linear Representations of Emergent Misalignment
von: Soligo, Anna, et al.
Veröffentlicht: (2025)
von: Soligo, Anna, et al.
Veröffentlicht: (2025)
Stop When Enough: Adaptive Early-Stopping for Chain-of-Thought Reasoning
von: Sun, Renliang, et al.
Veröffentlicht: (2025)
von: Sun, Renliang, et al.
Veröffentlicht: (2025)
Fine-Tuned Thoughts: Leveraging Chain-of-Thought Reasoning for Industrial Asset Health Monitoring
von: Lin, Shuxin, et al.
Veröffentlicht: (2025)
von: Lin, Shuxin, et al.
Veröffentlicht: (2025)
Measuring and Improving Chain-of-Thought Reasoning in Vision-Language Models
von: Chen, Yangyi, et al.
Veröffentlicht: (2023)
von: Chen, Yangyi, et al.
Veröffentlicht: (2023)
When LLMs Struggle: Reference-less Translation Evaluation for Low-resource Languages
von: Sindhujan, Archchana, et al.
Veröffentlicht: (2025)
von: Sindhujan, Archchana, et al.
Veröffentlicht: (2025)
Measuring Chain-of-Thought Monitorability Through Faithfulness and Verbosity
von: Meek, Austin, et al.
Veröffentlicht: (2025)
von: Meek, Austin, et al.
Veröffentlicht: (2025)
Steering When Necessary: Flexible Steering Large Language Models with Backtracking
von: Cheng, Zifeng, et al.
Veröffentlicht: (2025)
von: Cheng, Zifeng, et al.
Veröffentlicht: (2025)
Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
von: Korbak, Tomek, et al.
Veröffentlicht: (2025)
von: Korbak, Tomek, et al.
Veröffentlicht: (2025)
Beyond Chain-of-Thought, Effective Graph-of-Thought Reasoning in Language Models
von: Yao, Yao, et al.
Veröffentlicht: (2023)
von: Yao, Yao, et al.
Veröffentlicht: (2023)
Scheherazade: Evaluating Chain-of-Thought Math Reasoning in LLMs with Chain-of-Problems
von: Miner, Stephen, et al.
Veröffentlicht: (2024)
von: Miner, Stephen, et al.
Veröffentlicht: (2024)
When More is Less: Understanding Chain-of-Thought Length in LLMs
von: Wu, Yuyang, et al.
Veröffentlicht: (2025)
von: Wu, Yuyang, et al.
Veröffentlicht: (2025)
When Chain-of-Thought Fails, the Solution Hides in the Hidden States
von: Mehrafarin, Houman, et al.
Veröffentlicht: (2026)
von: Mehrafarin, Houman, et al.
Veröffentlicht: (2026)
Is Reference Necessary in the Evaluation of NLG Systems? When and Where?
von: Sheng, Shuqian, et al.
Veröffentlicht: (2024)
von: Sheng, Shuqian, et al.
Veröffentlicht: (2024)
AtP*: An efficient and scalable method for localizing LLM behaviour to components
von: Kramár, János, et al.
Veröffentlicht: (2024)
von: Kramár, János, et al.
Veröffentlicht: (2024)
On the Representational Capacity of Neural Language Models with Chain-of-Thought Reasoning
von: Nowak, Franz, et al.
Veröffentlicht: (2024)
von: Nowak, Franz, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
A Pragmatic Way to Measure Chain-of-Thought Monitorability
von: Emmons, Scott, et al.
Veröffentlicht: (2025) -
Chain-of-Thought Reasoning In The Wild Is Not Always Faithful
von: Arcuschin, Iván, et al.
Veröffentlicht: (2025) -
Thought Branches: Interpreting LLM Reasoning Requires Resampling
von: Macar, Uzay, et al.
Veröffentlicht: (2025) -
Do I Know This Entity? Knowledge Awareness and Hallucinations in Language Models
von: Ferrando, Javier, et al.
Veröffentlicht: (2024) -
RL-Obfuscation: Can Language Models Learn to Evade Latent-Space Monitors?
von: Gupta, Rohan, et al.
Veröffentlicht: (2025)