Scaling Latent Reasoning via Looped Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917086246207488 |
|---|---|
| author | Zhu, Rui-Jie Wang, Zixuan Hua, Kai Zhang, Tianyu Li, Ziniu Que, Haoran Wei, Boyi Wen, Zixin Yin, Fan Xing, He Li, Lu Shi, Jiajun Ma, Kaijing Li, Shanda Kergan, Taylor Smith, Andrew Qu, Xingwei Hui, Mude Wu, Bohong Min, Qiyang Huang, Hongzhi Zhou, Xun Ye, Wei Liu, Jiaheng Yang, Jian Shi, Yunfeng Lin, Chenghua Zhao, Enduo Cai, Tianle Zhang, Ge Huang, Wenhao Bengio, Yoshua Eshraghian, Jason |
| author_facet | Zhu, Rui-Jie Wang, Zixuan Hua, Kai Zhang, Tianyu Li, Ziniu Que, Haoran Wei, Boyi Wen, Zixin Yin, Fan Xing, He Li, Lu Shi, Jiajun Ma, Kaijing Li, Shanda Kergan, Taylor Smith, Andrew Qu, Xingwei Hui, Mude Wu, Bohong Min, Qiyang Huang, Hongzhi Zhou, Xun Ye, Wei Liu, Jiaheng Yang, Jian Shi, Yunfeng Lin, Chenghua Zhao, Enduo Cai, Tianle Zhang, Ge Huang, Wenhao Bengio, Yoshua Eshraghian, Jason |
| contents | Modern LLMs are trained to "think" primarily via explicit text generation, such as chain-of-thought (CoT), which defers reasoning to post-training and under-leverages pre-training data. We present and open-source Ouro, named after the recursive Ouroboros, a family of pre-trained Looped Language Models (LoopLM) that instead build reasoning into the pre-training phase through (i) iterative computation in latent space, (ii) an entropy-regularized objective for learned depth allocation, and (iii) scaling to 7.7T tokens. Ouro 1.4B and 2.6B models enjoy superior performance that match the results of up to 12B SOTA LLMs across a wide range of benchmarks. Through controlled experiments, we show this advantage stems not from increased knowledge capacity, but from superior knowledge manipulation capabilities. We also show that LoopLM yields reasoning traces more aligned with final outputs than explicit CoT. We hope our results show the potential of LoopLM as a novel scaling direction in the reasoning era. Our model is available here: http://ouro-llm.github.io. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_25741 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Scaling Latent Reasoning via Looped Language Models Zhu, Rui-Jie Wang, Zixuan Hua, Kai Zhang, Tianyu Li, Ziniu Que, Haoran Wei, Boyi Wen, Zixin Yin, Fan Xing, He Li, Lu Shi, Jiajun Ma, Kaijing Li, Shanda Kergan, Taylor Smith, Andrew Qu, Xingwei Hui, Mude Wu, Bohong Min, Qiyang Huang, Hongzhi Zhou, Xun Ye, Wei Liu, Jiaheng Yang, Jian Shi, Yunfeng Lin, Chenghua Zhao, Enduo Cai, Tianle Zhang, Ge Huang, Wenhao Bengio, Yoshua Eshraghian, Jason Computation and Language Modern LLMs are trained to "think" primarily via explicit text generation, such as chain-of-thought (CoT), which defers reasoning to post-training and under-leverages pre-training data. We present and open-source Ouro, named after the recursive Ouroboros, a family of pre-trained Looped Language Models (LoopLM) that instead build reasoning into the pre-training phase through (i) iterative computation in latent space, (ii) an entropy-regularized objective for learned depth allocation, and (iii) scaling to 7.7T tokens. Ouro 1.4B and 2.6B models enjoy superior performance that match the results of up to 12B SOTA LLMs across a wide range of benchmarks. Through controlled experiments, we show this advantage stems not from increased knowledge capacity, but from superior knowledge manipulation capabilities. We also show that LoopLM yields reasoning traces more aligned with final outputs than explicit CoT. We hope our results show the potential of LoopLM as a novel scaling direction in the reasoning era. Our model is available here: http://ouro-llm.github.io. |
| title | Scaling Latent Reasoning via Looped Language Models |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2510.25741 |