Output Supervision Can Obfuscate the Chain of Thought
Fuente:
arXiv
Saved in:
| Main Authors: | Drori, Jacob, Marks, Luke, Woodworth, Bryce, Cloud, Alex, Turner, Alexander Matt |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Can Reasoning Models Obfuscate Reasoning? Stress-Testing Chain-of-Thought Monitorability
by: Zolkowski, Artur, et al.
Published: (2025)
by: Zolkowski, Artur, et al.
Published: (2025)
Recontextualization Mitigates Specification Gaming without Modifying the Specification
by: Azarbal, Ariana, et al.
Published: (2025)
by: Azarbal, Ariana, et al.
Published: (2025)
Adversaries Can Misuse Combinations of Safe Models
by: Jones, Erik, et al.
Published: (2024)
by: Jones, Erik, et al.
Published: (2024)
Evading Deep Learning-Based Malware Detectors via Obfuscation: A Deep Reinforcement Learning Approach
by: Etter, Brian, et al.
Published: (2024)
by: Etter, Brian, et al.
Published: (2024)
Fooling SHAP with Output Shuffling Attacks
by: Yuan, Jun, et al.
Published: (2024)
by: Yuan, Jun, et al.
Published: (2024)
SALT: Steering Activations towards Leakage-free Thinking in Chain of Thought
by: Batra, Shourya, et al.
Published: (2025)
by: Batra, Shourya, et al.
Published: (2025)
The Jailbreak Tax: How Useful are Your Jailbreak Outputs?
by: Nikolić, Kristina, et al.
Published: (2025)
by: Nikolić, Kristina, et al.
Published: (2025)
Differentially Private Synthetic Data Release for Topics API Outputs
by: Dick, Travis, et al.
Published: (2025)
by: Dick, Travis, et al.
Published: (2025)
Thought Purity: A Defense Framework For Chain-of-Thought Attack
by: Xue, Zihao, et al.
Published: (2025)
by: Xue, Zihao, et al.
Published: (2025)
Strategic Dishonesty Can Undermine AI Safety Evaluations of Frontier LLMs
by: Panfilov, Alexander, et al.
Published: (2025)
by: Panfilov, Alexander, et al.
Published: (2025)
ObfuscaTune: Obfuscated Offsite Fine-tuning and Inference of Proprietary LLMs on Private Datasets
by: Frikha, Ahmed, et al.
Published: (2024)
by: Frikha, Ahmed, et al.
Published: (2024)
Weakest Link in the Chain: Security Vulnerabilities in Advanced Reasoning Models
by: Krishna, Arjun, et al.
Published: (2025)
by: Krishna, Arjun, et al.
Published: (2025)
Reasoning Under Pressure: How do Training Incentives Influence Chain-of-Thought Monitorability?
by: MacDermott, Matt, et al.
Published: (2025)
by: MacDermott, Matt, et al.
Published: (2025)
PoLO: Proof-of-Learning and Proof-of-Ownership at Once with Chained Watermarking
by: Deng, Haiyu, et al.
Published: (2025)
by: Deng, Haiyu, et al.
Published: (2025)
LLMs Can Covertly Sandbag on Capability Evaluations Against Chain-of-Thought Monitoring
by: Li, Chloe, et al.
Published: (2025)
by: Li, Chloe, et al.
Published: (2025)
Your Agent Can Defend Itself against Backdoor Attacks
by: Changjiang, Li, et al.
Published: (2025)
by: Changjiang, Li, et al.
Published: (2025)
Can Neural Decompilation Assist Vulnerability Prediction on Binary Code?
by: Cotroneo, D., et al.
Published: (2024)
by: Cotroneo, D., et al.
Published: (2024)
RMSL: Weakly-Supervised Insider Threat Detection with Robust Multi-sphere Learning
by: Wang, Yang, et al.
Published: (2025)
by: Wang, Yang, et al.
Published: (2025)
A Study on Semi-Supervised Detection of DDoS Attacks under Class Imbalance
by: Hallaji, Ehsan, et al.
Published: (2025)
by: Hallaji, Ehsan, et al.
Published: (2025)
Semi-Supervised Learning for Anomaly Traffic Detection via Bidirectional Normalizing Flows
by: Dang, Zhangxuan, et al.
Published: (2024)
by: Dang, Zhangxuan, et al.
Published: (2024)
Enhancing Internet of Things Security throughSelf-Supervised Graph Neural Networks
by: Atitallah, Safa Ben, et al.
Published: (2024)
by: Atitallah, Safa Ben, et al.
Published: (2024)
Data-Chain Backdoor: Do You Trust Diffusion Models as Generative Data Supplier?
by: Lu, Junchi, et al.
Published: (2025)
by: Lu, Junchi, et al.
Published: (2025)
TrustChain: A Blockchain Framework for Auditing and Verifying Aggregators in Decentralized Federated Learning
by: Hallaji, Ehsan, et al.
Published: (2025)
by: Hallaji, Ehsan, et al.
Published: (2025)
Concept Drift Adaptation Using Self-Supervised and Reinforcement Learning In Android Malware Detection
by: Sabbah, Ahmed, et al.
Published: (2026)
by: Sabbah, Ahmed, et al.
Published: (2026)
SequentialBreak: Large Language Models Can be Fooled by Embedding Jailbreak Prompts into Sequential Prompt Chains
by: Saiem, Bijoy Ahmed, et al.
Published: (2024)
by: Saiem, Bijoy Ahmed, et al.
Published: (2024)
Can Differentially Private Fine-tuning LLMs Protect Against Privacy Attacks?
by: Du, Hao, et al.
Published: (2025)
by: Du, Hao, et al.
Published: (2025)
CARACAS: vehiCular ArchitectuRe for detAiled Can Attacks Simulation
by: Kirdi, Sadek Misto, et al.
Published: (2024)
by: Kirdi, Sadek Misto, et al.
Published: (2024)
ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?
by: Wang, Zhun, et al.
Published: (2026)
by: Wang, Zhun, et al.
Published: (2026)
Distillation Robustifies Unlearning
by: Lee, Bruce W., et al.
Published: (2025)
by: Lee, Bruce W., et al.
Published: (2025)
InfiCoEvalChain: A Blockchain-Based Decentralized Framework for Collaborative LLM Evaluation
by: Yang, Yifan, et al.
Published: (2026)
by: Yang, Yifan, et al.
Published: (2026)
Packet Inspection Transformer: A Self-Supervised Journey to Unseen Malware Detection with Few Samples
by: Stein, Kyle, et al.
Published: (2024)
by: Stein, Kyle, et al.
Published: (2024)
Recalling The Forgotten Class Memberships: Unlearned Models Can Be Noisy Labelers to Leak Privacy
by: Sui, Zhihao, et al.
Published: (2025)
by: Sui, Zhihao, et al.
Published: (2025)
No, of Course I Can! Deeper Fine-Tuning Attacks That Bypass Token-Level Safety Mechanisms
by: Kazdan, Joshua, et al.
Published: (2025)
by: Kazdan, Joshua, et al.
Published: (2025)
Backdoor defense, learnability and obfuscation
by: Christiano, Paul, et al.
Published: (2024)
by: Christiano, Paul, et al.
Published: (2024)
GraphIP-Bench: How Hard Is It to Steal a Graph Neural Network, and Can We Stop It?
by: Zhao, Kaixiang, et al.
Published: (2026)
by: Zhao, Kaixiang, et al.
Published: (2026)
A Comprehensive Study of Supervised Machine Learning Models for Zero-Day Attack Detection: Analyzing Performance on Imbalanced Data
by: Lotfi, Zahra, et al.
Published: (2025)
by: Lotfi, Zahra, et al.
Published: (2025)
IoT-based Android Malware Detection Using Graph Neural Network With Adversarial Defense
by: Yumlembam, Rahul, et al.
Published: (2025)
by: Yumlembam, Rahul, et al.
Published: (2025)
TeleSparse: Practical Privacy-Preserving Verification of Deep Neural Networks
by: Maheri, Mohammad M, et al.
Published: (2025)
by: Maheri, Mohammad M, et al.
Published: (2025)
SAGA: A Security Architecture for Governing AI Agentic Systems
by: Syros, Georgios, et al.
Published: (2025)
by: Syros, Georgios, et al.
Published: (2025)
Stop Reasoning! When Multimodal LLM with Chain-of-Thought Reasoning Meets Adversarial Image
by: Wang, Zefeng, et al.
Published: (2024)
by: Wang, Zefeng, et al.
Published: (2024)
Similar Items
-
Can Reasoning Models Obfuscate Reasoning? Stress-Testing Chain-of-Thought Monitorability
by: Zolkowski, Artur, et al.
Published: (2025) -
Recontextualization Mitigates Specification Gaming without Modifying the Specification
by: Azarbal, Ariana, et al.
Published: (2025) -
Adversaries Can Misuse Combinations of Safe Models
by: Jones, Erik, et al.
Published: (2024) -
Evading Deep Learning-Based Malware Detectors via Obfuscation: A Deep Reinforcement Learning Approach
by: Etter, Brian, et al.
Published: (2024) -
Fooling SHAP with Output Shuffling Attacks
by: Yuan, Jun, et al.
Published: (2024)