Scaling Latent Reasoning via Looped Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhu, Rui-Jie, Wang, Zixuan, Hua, Kai, Zhang, Tianyu, Li, Ziniu, Que, Haoran, Wei, Boyi, Wen, Zixin, Yin, Fan, Xing, He, Li, Lu, Shi, Jiajun, Ma, Kaijing, Li, Shanda, Kergan, Taylor, Smith, Andrew, Qu, Xingwei, Hui, Mude, Wu, Bohong, Min, Qiyang, Huang, Hongzhi, Zhou, Xun, Ye, Wei, Liu, Jiaheng, Yang, Jian, Shi, Yunfeng, Lin, Chenghua, Zhao, Enduo, Cai, Tianle, Zhang, Ge, Huang, Wenhao, Bengio, Yoshua, Eshraghian, Jason
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917086246207488
author Zhu, Rui-Jie
Wang, Zixuan
Hua, Kai
Zhang, Tianyu
Li, Ziniu
Que, Haoran
Wei, Boyi
Wen, Zixin
Yin, Fan
Xing, He
Li, Lu
Shi, Jiajun
Ma, Kaijing
Li, Shanda
Kergan, Taylor
Smith, Andrew
Qu, Xingwei
Hui, Mude
Wu, Bohong
Min, Qiyang
Huang, Hongzhi
Zhou, Xun
Ye, Wei
Liu, Jiaheng
Yang, Jian
Shi, Yunfeng
Lin, Chenghua
Zhao, Enduo
Cai, Tianle
Zhang, Ge
Huang, Wenhao
Bengio, Yoshua
Eshraghian, Jason
author_facet Zhu, Rui-Jie
Wang, Zixuan
Hua, Kai
Zhang, Tianyu
Li, Ziniu
Que, Haoran
Wei, Boyi
Wen, Zixin
Yin, Fan
Xing, He
Li, Lu
Shi, Jiajun
Ma, Kaijing
Li, Shanda
Kergan, Taylor
Smith, Andrew
Qu, Xingwei
Hui, Mude
Wu, Bohong
Min, Qiyang
Huang, Hongzhi
Zhou, Xun
Ye, Wei
Liu, Jiaheng
Yang, Jian
Shi, Yunfeng
Lin, Chenghua
Zhao, Enduo
Cai, Tianle
Zhang, Ge
Huang, Wenhao
Bengio, Yoshua
Eshraghian, Jason
contents Modern LLMs are trained to "think" primarily via explicit text generation, such as chain-of-thought (CoT), which defers reasoning to post-training and under-leverages pre-training data. We present and open-source Ouro, named after the recursive Ouroboros, a family of pre-trained Looped Language Models (LoopLM) that instead build reasoning into the pre-training phase through (i) iterative computation in latent space, (ii) an entropy-regularized objective for learned depth allocation, and (iii) scaling to 7.7T tokens. Ouro 1.4B and 2.6B models enjoy superior performance that match the results of up to 12B SOTA LLMs across a wide range of benchmarks. Through controlled experiments, we show this advantage stems not from increased knowledge capacity, but from superior knowledge manipulation capabilities. We also show that LoopLM yields reasoning traces more aligned with final outputs than explicit CoT. We hope our results show the potential of LoopLM as a novel scaling direction in the reasoning era. Our model is available here: http://ouro-llm.github.io.
format Preprint
id arxiv_https___arxiv_org_abs_2510_25741
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Scaling Latent Reasoning via Looped Language Models
Zhu, Rui-Jie
Wang, Zixuan
Hua, Kai
Zhang, Tianyu
Li, Ziniu
Que, Haoran
Wei, Boyi
Wen, Zixin
Yin, Fan
Xing, He
Li, Lu
Shi, Jiajun
Ma, Kaijing
Li, Shanda
Kergan, Taylor
Smith, Andrew
Qu, Xingwei
Hui, Mude
Wu, Bohong
Min, Qiyang
Huang, Hongzhi
Zhou, Xun
Ye, Wei
Liu, Jiaheng
Yang, Jian
Shi, Yunfeng
Lin, Chenghua
Zhao, Enduo
Cai, Tianle
Zhang, Ge
Huang, Wenhao
Bengio, Yoshua
Eshraghian, Jason
Computation and Language
Modern LLMs are trained to "think" primarily via explicit text generation, such as chain-of-thought (CoT), which defers reasoning to post-training and under-leverages pre-training data. We present and open-source Ouro, named after the recursive Ouroboros, a family of pre-trained Looped Language Models (LoopLM) that instead build reasoning into the pre-training phase through (i) iterative computation in latent space, (ii) an entropy-regularized objective for learned depth allocation, and (iii) scaling to 7.7T tokens. Ouro 1.4B and 2.6B models enjoy superior performance that match the results of up to 12B SOTA LLMs across a wide range of benchmarks. Through controlled experiments, we show this advantage stems not from increased knowledge capacity, but from superior knowledge manipulation capabilities. We also show that LoopLM yields reasoning traces more aligned with final outputs than explicit CoT. We hope our results show the potential of LoopLM as a novel scaling direction in the reasoning era. Our model is available here: http://ouro-llm.github.io.
title Scaling Latent Reasoning via Looped Language Models
topic Computation and Language
url https://arxiv.org/abs/2510.25741