Reasoning Under Pressure: How do Training Incentives Influence Chain-of-Thought Monitorability?
Fuente:
arXiv
Saved in:
| Main Authors: | MacDermott, Matt, Wei, Qiyao, Djoneva, Rada, Ward, Francis Rhys |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Can Reasoning Models Obfuscate Reasoning? Stress-Testing Chain-of-Thought Monitorability
by: Zolkowski, Artur, et al.
Published: (2025)
by: Zolkowski, Artur, et al.
Published: (2025)
Preemptive Answer "Attacks" on Chain-of-Thought Reasoning
by: Xu, Rongwu, et al.
Published: (2024)
by: Xu, Rongwu, et al.
Published: (2024)
LLMs Can Covertly Sandbag on Capability Evaluations Against Chain-of-Thought Monitoring
by: Li, Chloe, et al.
Published: (2025)
by: Li, Chloe, et al.
Published: (2025)
Password-Activated Shutdown Protocols for Misaligned Frontier Agents
by: Williams, Kai, et al.
Published: (2025)
by: Williams, Kai, et al.
Published: (2025)
The Reasons that Agents Act: Intention and Instrumental Goals
by: Ward, Francis Rhys, et al.
Published: (2024)
by: Ward, Francis Rhys, et al.
Published: (2024)
Output Supervision Can Obfuscate the Chain of Thought
by: Drori, Jacob, et al.
Published: (2025)
by: Drori, Jacob, et al.
Published: (2025)
BadThink: Triggered Overthinking Attacks on Chain-of-Thought Reasoning in Large Language Models
by: Liu, Shuaitong, et al.
Published: (2025)
by: Liu, Shuaitong, et al.
Published: (2025)
ExplainableGuard: Interpretable Adversarial Defense for Large Language Models Using Chain-of-Thought Reasoning
by: Guan, Shaowei, et al.
Published: (2025)
by: Guan, Shaowei, et al.
Published: (2025)
Strengthening Human-Centric Chain-of-Thought Reasoning Integrity in LLMs via a Structured Prompt Framework
by: Zhou, Jiling, et al.
Published: (2026)
by: Zhou, Jiling, et al.
Published: (2026)
PPMI: Privacy-Preserving LLM Interaction with Socratic Chain-of-Thought Reasoning and Homomorphically Encrypted Vector Databases
by: Bae, Yubeen, et al.
Published: (2025)
by: Bae, Yubeen, et al.
Published: (2025)
SFCoT: Safer Chain-of-Thought via Active Safety Evaluation and Calibration
by: Pan, Yu, et al.
Published: (2026)
by: Pan, Yu, et al.
Published: (2026)
From Thinking to Output: Chain-of-Thought and Text Generation Characteristics in Reasoning Language Models
by: Liu, Junhao, et al.
Published: (2025)
by: Liu, Junhao, et al.
Published: (2025)
Hidden Thoughts Are Not Secret: Reasoning Trace Exposure in LLMs
by: Lu, Yu-An, et al.
Published: (2026)
by: Lu, Yu-An, et al.
Published: (2026)
Hey GPT-OSS, Looks Like You Got It -- Now Walk Me Through It! An Assessment of the Reasoning Language Models Chain of Thought Mechanism for Digital Forensics
by: Michelet, Gaëtan, et al.
Published: (2025)
by: Michelet, Gaëtan, et al.
Published: (2025)
CoTSRF: Utilize Chain of Thought as Stealthy and Robust Fingerprint of Large Language Models
by: Ren, Zhenzhen, et al.
Published: (2025)
by: Ren, Zhenzhen, et al.
Published: (2025)
Large Language Model-driven Security Assistant for Internet of Things via Chain-of-Thought
by: Zeng, Mingfei, et al.
Published: (2025)
by: Zeng, Mingfei, et al.
Published: (2025)
Imitation Game for Adversarial Disillusion with Chain-of-Thought Reasoning in Generative AI
by: Chang, Ching-Chun, et al.
Published: (2025)
by: Chang, Ching-Chun, et al.
Published: (2025)
Thought-Transfer: Indirect Targeted Poisoning Attacks on Chain-of-Thought Reasoning Models
by: Chaudhari, Harsh, et al.
Published: (2026)
by: Chaudhari, Harsh, et al.
Published: (2026)
Reasoning Under Threat: Symbolic and Neural Techniques for Cybersecurity Verification
by: Veronica, Sarah
Published: (2025)
by: Veronica, Sarah
Published: (2025)
Know Thy Enemy: Securing LLMs Against Prompt Injection via Diverse Data Synthesis and Instruction-Level Chain-of-Thought Learning
by: Chang, Zhiyuan, et al.
Published: (2026)
by: Chang, Zhiyuan, et al.
Published: (2026)
Leaky Thoughts: Large Reasoning Models Are Not Private Thinkers
by: Green, Tommaso, et al.
Published: (2025)
by: Green, Tommaso, et al.
Published: (2025)
Stop Reasoning! When Multimodal LLM with Chain-of-Thought Reasoning Meets Adversarial Image
by: Wang, Zefeng, et al.
Published: (2024)
by: Wang, Zefeng, et al.
Published: (2024)
Beyond Content Safety: Real-Time Monitoring for Reasoning Vulnerabilities in Large Language Models
by: Wang, Xunguang, et al.
Published: (2026)
by: Wang, Xunguang, et al.
Published: (2026)
R-CoT: A Reasoning-Layer Watermark via Redundant Chain-of-Thought in Large Language Models
by: Zhang, Ziming, et al.
Published: (2026)
by: Zhang, Ziming, et al.
Published: (2026)
NEST: Nascent Encoded Steganographic Thoughts
by: Karpov, Artem
Published: (2026)
by: Karpov, Artem
Published: (2026)
The Promptware Kill Chain: How Prompt Injections Gradually Evolved Into a Multistep Malware Delivery Mechanism
by: Brodt, Oleg, et al.
Published: (2026)
by: Brodt, Oleg, et al.
Published: (2026)
Refusal Falls off a Cliff: How Safety Alignment Fails in Reasoning?
by: Yin, Qingyu, et al.
Published: (2025)
by: Yin, Qingyu, et al.
Published: (2025)
How to Train your Antivirus: RL-based Hardening through the Problem-Space
by: Tsingenopoulos, Ilias, et al.
Published: (2024)
by: Tsingenopoulos, Ilias, et al.
Published: (2024)
Echoes within the Reasoning: Stealthy and Effective Watermarking via Chain of Thought
by: Lu, Jiacheng, et al.
Published: (2026)
by: Lu, Jiacheng, et al.
Published: (2026)
Examining Attacks on Consensus and Incentive Systems in Proof-of-Work Blockchains: A Systematic Literature Review
by: Wijewardhana, Dinitha, et al.
Published: (2024)
by: Wijewardhana, Dinitha, et al.
Published: (2024)
How To Think About End-To-End Encryption and AI: Training, Processing, Disclosure, and Consent
by: Knodel, Mallory, et al.
Published: (2024)
by: Knodel, Mallory, et al.
Published: (2024)
Adversarial Threat Vectors and Risk Mitigation for Retrieval-Augmented Generation Systems
by: Ward, Chris M., et al.
Published: (2025)
by: Ward, Chris M., et al.
Published: (2025)
Offensive Security for AI Systems: Concepts, Practices, and Applications
by: Harguess, Josh, et al.
Published: (2025)
by: Harguess, Josh, et al.
Published: (2025)
ChainMarks: Securing DNN Watermark with Cryptographic Chain
by: Choi, Brian, et al.
Published: (2025)
by: Choi, Brian, et al.
Published: (2025)
Retrieval Augmented Generation Based LLM Evaluation For Protocol State Machine Inference With Chain-of-Thought Reasoning
by: Maklad, Youssef, et al.
Published: (2025)
by: Maklad, Youssef, et al.
Published: (2025)
Monitoring Decomposition Attacks in LLMs with Lightweight Sequential Monitors
by: Yueh-Han, Chen, et al.
Published: (2025)
by: Yueh-Han, Chen, et al.
Published: (2025)
Reasoning-Style Poisoning of LLM Agents via Stealthy Style Transfer: Process-Level Attacks and Runtime Monitoring in RSV Space
by: Zhou, Xingfu, et al.
Published: (2025)
by: Zhou, Xingfu, et al.
Published: (2025)
MonitoringBench: Semi-Automated Red-Teaming for Agent Monitoring
by: Jotautaitė, Monika, et al.
Published: (2026)
by: Jotautaitė, Monika, et al.
Published: (2026)
Thought Purity: A Defense Framework For Chain-of-Thought Attack
by: Xue, Zihao, et al.
Published: (2025)
by: Xue, Zihao, et al.
Published: (2025)
Proactive DDoS Detection and Mitigation in Decentralized Software-Defined Networking via Port-Level Monitoring and Zero-Training Large Language Models
by: Swileh, Mohammed N., et al.
Published: (2025)
by: Swileh, Mohammed N., et al.
Published: (2025)
Similar Items
-
Can Reasoning Models Obfuscate Reasoning? Stress-Testing Chain-of-Thought Monitorability
by: Zolkowski, Artur, et al.
Published: (2025) -
Preemptive Answer "Attacks" on Chain-of-Thought Reasoning
by: Xu, Rongwu, et al.
Published: (2024) -
LLMs Can Covertly Sandbag on Capability Evaluations Against Chain-of-Thought Monitoring
by: Li, Chloe, et al.
Published: (2025) -
Password-Activated Shutdown Protocols for Misaligned Frontier Agents
by: Williams, Kai, et al.
Published: (2025) -
The Reasons that Agents Act: Intention and Instrumental Goals
by: Ward, Francis Rhys, et al.
Published: (2024)