Latent Reasoning with Supervised Thinking States

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Amos, Ido, Caciularu, Avi, Geva, Mor, Globerson, Amir, Herzig, Jonathan, Shani, Lior, Szpektor, Idan
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917259553800192
author Amos, Ido
Caciularu, Avi
Geva, Mor
Globerson, Amir
Herzig, Jonathan
Shani, Lior
Szpektor, Idan
author_facet Amos, Ido
Caciularu, Avi
Geva, Mor
Globerson, Amir
Herzig, Jonathan
Shani, Lior
Szpektor, Idan
contents Reasoning with a chain-of-thought (CoT) enables Large Language Models (LLMs) to solve complex tasks but incurs significant inference costs due to the generation of long rationales. We propose Thinking States, a method that performs reasoning {\em while} the input is processing. Specifically, Thinking States generates sequences of thinking tokens every few input tokens, transforms the thoughts back into embedding space, and adds them to the following input tokens. This has two key advantages. First, it captures the recurrent nature of CoT, but where the thought tokens are generated as input is processing. Second, since the thoughts are represented as tokens, they can be learned from natural language supervision, and using teacher-forcing, which is parallelizable. Empirically, Thinking States outperforms other latent reasoning methods on multiple reasoning tasks, narrowing the gap to CoT on math problems, and matching its performance on 2-Hop QA with improved latency. On state-tracking tasks, we show Thinking States leads to stronger reasoning behavior than CoT, successfully extrapolating to longer sequences than seen during training.
format Preprint
id arxiv_https___arxiv_org_abs_2602_08332
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Latent Reasoning with Supervised Thinking States
Amos, Ido
Caciularu, Avi
Geva, Mor
Globerson, Amir
Herzig, Jonathan
Shani, Lior
Szpektor, Idan
Computation and Language
Artificial Intelligence
Reasoning with a chain-of-thought (CoT) enables Large Language Models (LLMs) to solve complex tasks but incurs significant inference costs due to the generation of long rationales. We propose Thinking States, a method that performs reasoning {\em while} the input is processing. Specifically, Thinking States generates sequences of thinking tokens every few input tokens, transforms the thoughts back into embedding space, and adds them to the following input tokens. This has two key advantages. First, it captures the recurrent nature of CoT, but where the thought tokens are generated as input is processing. Second, since the thoughts are represented as tokens, they can be learned from natural language supervision, and using teacher-forcing, which is parallelizable. Empirically, Thinking States outperforms other latent reasoning methods on multiple reasoning tasks, narrowing the gap to CoT on math problems, and matching its performance on 2-Hop QA with improved latency. On state-tracking tasks, we show Thinking States leads to stronger reasoning behavior than CoT, successfully extrapolating to longer sequences than seen during training.
title Latent Reasoning with Supervised Thinking States
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2602.08332