LoopRPT: Reinforcement Pre-Training for Looped Language Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Tang, Guo, Jiang, Shixin, Chang, Heng, Chen, Nuo, Li, Yuhan, Fan, Huiming, Li, Jia, Liu, Ming, Qin, Bing
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914410976509952
author Tang, Guo
Jiang, Shixin
Chang, Heng
Chen, Nuo
Li, Yuhan
Fan, Huiming
Li, Jia
Liu, Ming
Qin, Bing
author_facet Tang, Guo
Jiang, Shixin
Chang, Heng
Chen, Nuo
Li, Yuhan
Fan, Huiming
Li, Jia
Liu, Ming
Qin, Bing
contents Looped language models (LoopLMs) perform iterative latent computation to refine internal representations, offering a promising alternative to explicit chain-of-thought (CoT) reasoning. However, existing reinforcement learning (RL) paradigms primarily target output tokens, creating a structural mismatch with looped architectures whose reasoning unfolds implicitly. In this work, we propose LoopRPT, a reinforcement pre-training framework tailored for LoopLMs. By reframing next-token prediction as a next-token reasoning task, LoopRPT assigns reinforcement signals directly to latent steps using an EMA teacher reference and noisy latent rollouts. This formulation enables RL to directly shape intermediate representations, compressing effective reasoning into fewer iterations. We instantiate LoopRPT on the Ouro architecture across multiple model scales. Results demonstrate that LoopRPT consistently improves per-step representation quality, achieving Pareto dominance in accuracy-computation trade-offs. Notably, significant gains on hard tokens indicate that LoopRPT enhances early-stage reasoning rather than merely encouraging premature exits. Our findings highlight reinforcement pre-training as a principled paradigm for learning efficient latent reasoning in LoopLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2603_19714
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle LoopRPT: Reinforcement Pre-Training for Looped Language Models
Tang, Guo
Jiang, Shixin
Chang, Heng
Chen, Nuo
Li, Yuhan
Fan, Huiming
Li, Jia
Liu, Ming
Qin, Bing
Computation and Language
Looped language models (LoopLMs) perform iterative latent computation to refine internal representations, offering a promising alternative to explicit chain-of-thought (CoT) reasoning. However, existing reinforcement learning (RL) paradigms primarily target output tokens, creating a structural mismatch with looped architectures whose reasoning unfolds implicitly. In this work, we propose LoopRPT, a reinforcement pre-training framework tailored for LoopLMs. By reframing next-token prediction as a next-token reasoning task, LoopRPT assigns reinforcement signals directly to latent steps using an EMA teacher reference and noisy latent rollouts. This formulation enables RL to directly shape intermediate representations, compressing effective reasoning into fewer iterations. We instantiate LoopRPT on the Ouro architecture across multiple model scales. Results demonstrate that LoopRPT consistently improves per-step representation quality, achieving Pareto dominance in accuracy-computation trade-offs. Notably, significant gains on hard tokens indicate that LoopRPT enhances early-stage reasoning rather than merely encouraging premature exits. Our findings highlight reinforcement pre-training as a principled paradigm for learning efficient latent reasoning in LoopLMs.
title LoopRPT: Reinforcement Pre-Training for Looped Language Models
topic Computation and Language
url https://arxiv.org/abs/2603.19714