Hidden Thoughts Are Not Secret: Reasoning Trace Exposure in LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Lu, Yu-An, Tsai, Ci-Yang, Tsai, Yu-Lin, Popa, Raluca Ada, Yu, Chia-Mu |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Opal: Private Memory for Personal AI
by: Kaviani, Darya, et al.
Published: (2026)
by: Kaviani, Darya, et al.
Published: (2026)
A Framework for Evaluating Emerging Cyberattack Capabilities of AI
by: Rodriguez, Mikel, et al.
Published: (2025)
by: Rodriguez, Mikel, et al.
Published: (2025)
Onyx: Cost-Efficient Disk-Oblivious ANN Search
by: Rathee, Deevashwer, et al.
Published: (2026)
by: Rathee, Deevashwer, et al.
Published: (2026)
Concept-Aware Privacy Mechanisms for Defending Embedding Inversion Attacks
by: Tsai, Yu-Che, et al.
Published: (2026)
by: Tsai, Yu-Che, et al.
Published: (2026)
MiniScope: A Least Privilege Framework for Authorizing Tool Calling Agents
by: Zhu, Jinhao, et al.
Published: (2025)
by: Zhu, Jinhao, et al.
Published: (2025)
Differentially Private Fine-Tuning of Diffusion Models
by: Tsai, Yu-Lin, et al.
Published: (2024)
by: Tsai, Yu-Lin, et al.
Published: (2024)
One Leak Away: How Pretrained Model Exposure Amplifies Jailbreak Risks in Finetuned LLMs
by: Tan, Yixin, et al.
Published: (2025)
by: Tan, Yixin, et al.
Published: (2025)
BadThink: Triggered Overthinking Attacks on Chain-of-Thought Reasoning in Large Language Models
by: Liu, Shuaitong, et al.
Published: (2025)
by: Liu, Shuaitong, et al.
Published: (2025)
SUDP: Secret-Use Delegation Protocol for Agentic Systems
by: Yu, Xiaohang, et al.
Published: (2026)
by: Yu, Xiaohang, et al.
Published: (2026)
ExplainableGuard: Interpretable Adversarial Defense for Large Language Models Using Chain-of-Thought Reasoning
by: Guan, Shaowei, et al.
Published: (2025)
by: Guan, Shaowei, et al.
Published: (2025)
Understanding Secret Leakage Risks in Code LLMs: A Tokenization Perspective
by: Chen, Meifang, et al.
Published: (2026)
by: Chen, Meifang, et al.
Published: (2026)
SFCoT: Safer Chain-of-Thought via Active Safety Evaluation and Calibration
by: Pan, Yu, et al.
Published: (2026)
by: Pan, Yu, et al.
Published: (2026)
MPC-Minimized Secure LLM Inference
by: Rathee, Deevashwer, et al.
Published: (2024)
by: Rathee, Deevashwer, et al.
Published: (2024)
Strengthening Human-Centric Chain-of-Thought Reasoning Integrity in LLMs via a Structured Prompt Framework
by: Zhou, Jiling, et al.
Published: (2026)
by: Zhou, Jiling, et al.
Published: (2026)
Auditing Black-Box LLM APIs with a Rank-Based Uniformity Test
by: Zhu, Xiaoyuan, et al.
Published: (2025)
by: Zhu, Xiaoyuan, et al.
Published: (2025)
Red Teaming Large Reasoning Models
by: Chen, Jiawei, et al.
Published: (2025)
by: Chen, Jiawei, et al.
Published: (2025)
Security and Resilience in Autonomous Vehicles: A Proactive Design Approach
by: Tsai, Chieh, et al.
Published: (2026)
by: Tsai, Chieh, et al.
Published: (2026)
Exposing Hidden Interfaces: LLM-Guided Type Inference for Reverse Engineering macOS Private Frameworks
by: Kharlamova, Arina, et al.
Published: (2026)
by: Kharlamova, Arina, et al.
Published: (2026)
Multi-Designated Detector Watermarking for Language Models
by: Huang, Zhengan, et al.
Published: (2024)
by: Huang, Zhengan, et al.
Published: (2024)
Can Reasoning Models Obfuscate Reasoning? Stress-Testing Chain-of-Thought Monitorability
by: Zolkowski, Artur, et al.
Published: (2025)
by: Zolkowski, Artur, et al.
Published: (2025)
MaXsive: High-Capacity and Robust Training-Free Generative Image Watermarking in Diffusion Models
by: Mao, Po-Yuan, et al.
Published: (2025)
by: Mao, Po-Yuan, et al.
Published: (2025)
Protect Your Secrets: Understanding and Measuring Data Exposure in VSCode Extensions
by: Liu, Yue, et al.
Published: (2024)
by: Liu, Yue, et al.
Published: (2024)
DiffuseTrace: A Transparent and Flexible Watermarking Scheme for Latent Diffusion Model
by: Lei, Liangqi, et al.
Published: (2024)
by: Lei, Liangqi, et al.
Published: (2024)
Web Agents Should Adopt the Plan-Then-Execute Paradigm
by: Piet, Julien, et al.
Published: (2026)
by: Piet, Julien, et al.
Published: (2026)
LLMs Can Covertly Sandbag on Capability Evaluations Against Chain-of-Thought Monitoring
by: Li, Chloe, et al.
Published: (2025)
by: Li, Chloe, et al.
Published: (2025)
Reasoning Under Pressure: How do Training Incentives Influence Chain-of-Thought Monitorability?
by: MacDermott, Matt, et al.
Published: (2025)
by: MacDermott, Matt, et al.
Published: (2025)
Rethinking Reasoning: A Survey on Reasoning-based Backdoors in LLMs
by: Hu, Man, et al.
Published: (2025)
by: Hu, Man, et al.
Published: (2025)
PROMPTFUZZ: Harnessing Fuzzing Techniques for Robust Testing of Prompt Injection in LLMs
by: Yu, Jiahao, et al.
Published: (2024)
by: Yu, Jiahao, et al.
Published: (2024)
Committed SAE-Feature Traces for Audited-Session Substitution Detection in Hosted LLMs
by: Liu, Ziyang
Published: (2026)
by: Liu, Ziyang
Published: (2026)
Token-Efficient Prompt Injection Attack: Provoking Cessation in LLM Reasoning via Adaptive Token Compression
by: Cui, Yu, et al.
Published: (2025)
by: Cui, Yu, et al.
Published: (2025)
DPrivBench: Benchmarking LLMs' Reasoning for Differential Privacy
by: Wang, Erchi, et al.
Published: (2026)
by: Wang, Erchi, et al.
Published: (2026)
Why Do Aligned LLMs Remain Jailbreakable: Refusal-Escape Directions, Operator-Level Sources, and Safety-Utility Trade-off
by: Chen, Yu, et al.
Published: (2026)
by: Chen, Yu, et al.
Published: (2026)
CachePrune: Teaching LLMs What Not to Follow via KV-Cache Editing
by: Wang, Rui, et al.
Published: (2025)
by: Wang, Rui, et al.
Published: (2025)
Engineering Robustness into Personal Agents with the AI Workflow Store
by: Geambasu, Roxana, et al.
Published: (2026)
by: Geambasu, Roxana, et al.
Published: (2026)
Enhancing Reasoning Capacity of SLM using Cognitive Enhancement
by: Pan, Jonathan, et al.
Published: (2024)
by: Pan, Jonathan, et al.
Published: (2024)
Shattering the Echo Chamber: Hidden Safeguards in Manuscripts Against the AI Takeover of Peer Review
by: Ma, Oubo, et al.
Published: (2026)
by: Ma, Oubo, et al.
Published: (2026)
PPMI: Privacy-Preserving LLM Interaction with Socratic Chain-of-Thought Reasoning and Homomorphically Encrypted Vector Databases
by: Bae, Yubeen, et al.
Published: (2025)
by: Bae, Yubeen, et al.
Published: (2025)
Tricking LLM-Based NPCs into Spilling Secrets
by: Shiomi, Kyohei, et al.
Published: (2025)
by: Shiomi, Kyohei, et al.
Published: (2025)
On the Feasibility of Using MultiModal LLMs to Execute AR Social Engineering Attacks
by: Bi, Ting, et al.
Published: (2025)
by: Bi, Ting, et al.
Published: (2025)
Confusion is the Final Barrier: Rethinking Jailbreak Evaluation and Investigating the Real Misuse Threat of LLMs
by: Yan, Yu, et al.
Published: (2025)
by: Yan, Yu, et al.
Published: (2025)
Similar Items
-
Opal: Private Memory for Personal AI
by: Kaviani, Darya, et al.
Published: (2026) -
A Framework for Evaluating Emerging Cyberattack Capabilities of AI
by: Rodriguez, Mikel, et al.
Published: (2025) -
Onyx: Cost-Efficient Disk-Oblivious ANN Search
by: Rathee, Deevashwer, et al.
Published: (2026) -
Concept-Aware Privacy Mechanisms for Defending Embedding Inversion Attacks
by: Tsai, Yu-Che, et al.
Published: (2026) -
MiniScope: A Least Privilege Framework for Authorizing Tool Calling Agents
by: Zhu, Jinhao, et al.
Published: (2025)