Hidden Thoughts Are Not Secret: Reasoning Trace Exposure in LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lu, Yu-An, Tsai, Ci-Yang, Tsai, Yu-Lin, Popa, Raluca Ada, Yu, Chia-Mu
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910276035543040
author Lu, Yu-An
Tsai, Ci-Yang
Tsai, Yu-Lin
Popa, Raluca Ada
Yu, Chia-Mu
author_facet Lu, Yu-An
Tsai, Ci-Yang
Tsai, Yu-Lin
Popa, Raluca Ada
Yu, Chia-Mu
contents Reasoning traces have become a valuable form of learning signals for improving and transferring the capabilities of large language models. In particular, detailed traces can help distill reasoning behavior from stronger teacher models into weaker student models. The value of capability transfer has motivated many deployed systems with reasoning models to hide raw internal traces and expose at most summaries and answers to users. As a result, we ask whether such interface-level trace hiding prevents users from obtaining useful reasoning supervision through prompting. We study this question with Reasoning Exposure Prompting (REP), a lightweight in-context elicitation method that uses shadow-model-generated demonstrations wrapped in auxiliary code-like formats to raise user-visible reasoning traces from a victim model. Across the common reasoning dataset, different victim models, and different student model distillation, REP substantially increases similarity between exposed and REP-conditioned internal traces while preserving useful reasoning signals.
format Preprint
id arxiv_https___arxiv_org_abs_2606_00642
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Hidden Thoughts Are Not Secret: Reasoning Trace Exposure in LLMs
Lu, Yu-An
Tsai, Ci-Yang
Tsai, Yu-Lin
Popa, Raluca Ada
Yu, Chia-Mu
Artificial Intelligence
Cryptography and Security
Reasoning traces have become a valuable form of learning signals for improving and transferring the capabilities of large language models. In particular, detailed traces can help distill reasoning behavior from stronger teacher models into weaker student models. The value of capability transfer has motivated many deployed systems with reasoning models to hide raw internal traces and expose at most summaries and answers to users. As a result, we ask whether such interface-level trace hiding prevents users from obtaining useful reasoning supervision through prompting. We study this question with Reasoning Exposure Prompting (REP), a lightweight in-context elicitation method that uses shadow-model-generated demonstrations wrapped in auxiliary code-like formats to raise user-visible reasoning traces from a victim model. Across the common reasoning dataset, different victim models, and different student model distillation, REP substantially increases similarity between exposed and REP-conditioned internal traces while preserving useful reasoning signals.
title Hidden Thoughts Are Not Secret: Reasoning Trace Exposure in LLMs
topic Artificial Intelligence
Cryptography and Security
url https://arxiv.org/abs/2606.00642