Training on Documents About Monitoring Leads to CoT Obfuscation
Fuente:
arXiv
Salvato in:
| Autori principali: | Haskins, Reilly, Chughtai, Bilal, Engels, Joshua |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Distilled Circuits: A Mechanistic Study of Internal Restructuring in Knowledge Distillation
di: Haskins, Reilly, et al.
Pubblicazione: (2025)
di: Haskins, Reilly, et al.
Pubblicazione: (2025)
KEA Explain: Explanations of Hallucinations using Graph Kernel Analysis
di: Haskins, Reilly, et al.
Pubblicazione: (2025)
di: Haskins, Reilly, et al.
Pubblicazione: (2025)
Building Production-Ready Probes For Gemini
di: Kramár, János, et al.
Pubblicazione: (2026)
di: Kramár, János, et al.
Pubblicazione: (2026)
Difficulties with Evaluating a Deception Detector for AIs
di: Smith, Lewis, et al.
Pubblicazione: (2025)
di: Smith, Lewis, et al.
Pubblicazione: (2025)
Can Language Models Explain Their Own Classification Behavior?
di: Sherburn, Dane, et al.
Pubblicazione: (2024)
di: Sherburn, Dane, et al.
Pubblicazione: (2024)
Summing Up the Facts: Additive Mechanisms Behind Factual Recall in LLMs
di: Chughtai, Bilal, et al.
Pubblicazione: (2024)
di: Chughtai, Bilal, et al.
Pubblicazione: (2024)
CoT Red-Handed: Stress Testing Chain-of-Thought Monitoring
di: Arnav, Benjamin, et al.
Pubblicazione: (2025)
di: Arnav, Benjamin, et al.
Pubblicazione: (2025)
Transformer Circuit Faithfulness Metrics are not Robust
di: Miller, Joseph, et al.
Pubblicazione: (2024)
di: Miller, Joseph, et al.
Pubblicazione: (2024)
From Explicit CoT to Implicit CoT: Learning to Internalize CoT Step by Step
di: Deng, Yuntian, et al.
Pubblicazione: (2024)
di: Deng, Yuntian, et al.
Pubblicazione: (2024)
Detecting Strategic Deception Using Linear Probes
di: Goldowsky-Dill, Nicholas, et al.
Pubblicazione: (2025)
di: Goldowsky-Dill, Nicholas, et al.
Pubblicazione: (2025)
To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning
di: Sprague, Zayne, et al.
Pubblicazione: (2024)
di: Sprague, Zayne, et al.
Pubblicazione: (2024)
Noticing the Watcher: LLM Agents Can Infer CoT Monitoring from Blocking Feedback
di: Jiralerspong, Thomas, et al.
Pubblicazione: (2026)
di: Jiralerspong, Thomas, et al.
Pubblicazione: (2026)
EPiC: Towards Lossless Speedup for Reasoning Training through Edge-Preserving CoT Condensation
di: Jia, Jinghan, et al.
Pubblicazione: (2025)
di: Jia, Jinghan, et al.
Pubblicazione: (2025)
Unveiling and Causalizing CoT: A Causal Pespective
di: Fu, Jiarun, et al.
Pubblicazione: (2025)
di: Fu, Jiarun, et al.
Pubblicazione: (2025)
Robust Filtering -- Novel Statistical Learning and Inference Algorithms with Applications
di: Chughtai, Aamir Hussain
Pubblicazione: (2025)
di: Chughtai, Aamir Hussain
Pubblicazione: (2025)
To Think or Not to Think: The Hidden Cost of Meta-Training with Excessive CoT Examples
di: Kothapalli, Vignesh, et al.
Pubblicazione: (2025)
di: Kothapalli, Vignesh, et al.
Pubblicazione: (2025)
Demonstrations, CoT, and Prompting: A Theoretical Analysis of ICL
di: Tong, Xuhan, et al.
Pubblicazione: (2026)
di: Tong, Xuhan, et al.
Pubblicazione: (2026)
Self-Verifying Reflection Helps Transformers with CoT Reasoning
di: Yu, Zhongwei, et al.
Pubblicazione: (2025)
di: Yu, Zhongwei, et al.
Pubblicazione: (2025)
RL-Obfuscation: Can Language Models Learn to Evade Latent-Space Monitors?
di: Gupta, Rohan, et al.
Pubblicazione: (2025)
di: Gupta, Rohan, et al.
Pubblicazione: (2025)
CDW-CoT: Clustered Distance-Weighted Chain-of-Thoughts Reasoning
di: Fang, Yuanheng, et al.
Pubblicazione: (2025)
di: Fang, Yuanheng, et al.
Pubblicazione: (2025)
Low-Rank Adapting Models for Sparse Autoencoders
di: Chen, Matthew, et al.
Pubblicazione: (2025)
di: Chen, Matthew, et al.
Pubblicazione: (2025)
Decomposing The Dark Matter of Sparse Autoencoders
di: Engels, Joshua, et al.
Pubblicazione: (2024)
di: Engels, Joshua, et al.
Pubblicazione: (2024)
Data Shifts Hurt CoT: A Theoretical Study
di: Yin, Lang, et al.
Pubblicazione: (2025)
di: Yin, Lang, et al.
Pubblicazione: (2025)
Visual CoT Makes VLMs Smarter but More Fragile
di: Xu, Chunxue, et al.
Pubblicazione: (2025)
di: Xu, Chunxue, et al.
Pubblicazione: (2025)
Divide-and-Conquer CoT: RL for Reducing Latency via Parallel Reasoning
di: Mahankali, Arvind, et al.
Pubblicazione: (2026)
di: Mahankali, Arvind, et al.
Pubblicazione: (2026)
What Characterizes Effective Reasoning? Revisiting Length, Review, and Structure of CoT
di: Feng, Yunzhen, et al.
Pubblicazione: (2025)
di: Feng, Yunzhen, et al.
Pubblicazione: (2025)
CoT Information: Improved Sample Complexity under Chain-of-Thought Supervision
di: Altabaa, Awni, et al.
Pubblicazione: (2025)
di: Altabaa, Awni, et al.
Pubblicazione: (2025)
Exploring the Limitations of Mamba in COPY and CoT Reasoning
di: Ren, Ruifeng, et al.
Pubblicazione: (2024)
di: Ren, Ruifeng, et al.
Pubblicazione: (2024)
Enhancing Confidence Estimation in Telco LLMs via Twin-Pass CoT-Ensembling
di: Saenko, Anton, et al.
Pubblicazione: (2026)
di: Saenko, Anton, et al.
Pubblicazione: (2026)
What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret
di: Yuan, Yufeng, et al.
Pubblicazione: (2025)
di: Yuan, Yufeng, et al.
Pubblicazione: (2025)
Distribution-Aligned Sequence Distillation for Superior Long-CoT Reasoning
di: Yan, Shaotian, et al.
Pubblicazione: (2026)
di: Yan, Shaotian, et al.
Pubblicazione: (2026)
Compositional Generalization from Learned Skills via CoT Training: A Theoretical and Structural Analysis for Reasoning
di: Yao, Xinhao, et al.
Pubblicazione: (2025)
di: Yao, Xinhao, et al.
Pubblicazione: (2025)
Is continuous CoT better suited for multi-lingual reasoning?
di: Bashir, Ali Hamza, et al.
Pubblicazione: (2026)
di: Bashir, Ali Hamza, et al.
Pubblicazione: (2026)
The Quest for Efficient Reasoning: A Data-Centric Benchmark to CoT Distillation
di: Zhang, Ruichen, et al.
Pubblicazione: (2025)
di: Zhang, Ruichen, et al.
Pubblicazione: (2025)
Amalgam: A Framework for Obfuscated Neural Network Training on the Cloud
di: Taki, Sifat Ut, et al.
Pubblicazione: (2024)
di: Taki, Sifat Ut, et al.
Pubblicazione: (2024)
How Likely Do LLMs with CoT Mimic Human Reasoning?
di: Bao, Guangsheng, et al.
Pubblicazione: (2024)
di: Bao, Guangsheng, et al.
Pubblicazione: (2024)
CoT-UQ: Improving Response-wise Uncertainty Quantification in LLMs with Chain-of-Thought
di: Zhang, Boxuan, et al.
Pubblicazione: (2025)
di: Zhang, Boxuan, et al.
Pubblicazione: (2025)
To CoT or To Loop? A Formal Comparison Between Chain-of-Thought and Looped Transformers
di: Xu, Kevin, et al.
Pubblicazione: (2025)
di: Xu, Kevin, et al.
Pubblicazione: (2025)
The Ends Justify the Thoughts: RL-Induced Motivated Reasoning in LLM CoTs
di: Howe, Nikolaus, et al.
Pubblicazione: (2025)
di: Howe, Nikolaus, et al.
Pubblicazione: (2025)
Audio Flamingo Sound-CoT Technical Report: Improving Chain-of-Thought Reasoning in Sound Understanding
di: Kong, Zhifeng, et al.
Pubblicazione: (2025)
di: Kong, Zhifeng, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Distilled Circuits: A Mechanistic Study of Internal Restructuring in Knowledge Distillation
di: Haskins, Reilly, et al.
Pubblicazione: (2025) -
KEA Explain: Explanations of Hallucinations using Graph Kernel Analysis
di: Haskins, Reilly, et al.
Pubblicazione: (2025) -
Building Production-Ready Probes For Gemini
di: Kramár, János, et al.
Pubblicazione: (2026) -
Difficulties with Evaluating a Deception Detector for AIs
di: Smith, Lewis, et al.
Pubblicazione: (2025) -
Can Language Models Explain Their Own Classification Behavior?
di: Sherburn, Dane, et al.
Pubblicazione: (2024)