h1: Bootstrapping LLMs to Reason over Longer Horizons via Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Motwani, Sumeet Ramesh, Ivanova, Alesia, Cai, Ziyang, Torr, Philip, Islam, Riashat, Shah, Shital, de Witt, Christian Schroeder, London, Charles
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915556036182016
author Motwani, Sumeet Ramesh
Ivanova, Alesia
Cai, Ziyang
Torr, Philip
Islam, Riashat
Shah, Shital
de Witt, Christian Schroeder
London, Charles
author_facet Motwani, Sumeet Ramesh
Ivanova, Alesia
Cai, Ziyang
Torr, Philip
Islam, Riashat
Shah, Shital
de Witt, Christian Schroeder
London, Charles
contents Large language models excel at short-horizon reasoning tasks, but performance drops as reasoning horizon lengths increase. Existing approaches to combat this rely on inference-time scaffolding or costly step-level supervision, neither of which scales easily. In this work, we introduce a scalable method to bootstrap long-horizon reasoning capabilities using only existing, abundant short-horizon data. Our approach synthetically composes simple problems into complex, multi-step dependency chains of arbitrary length. We train models on this data using outcome-only rewards under a curriculum that automatically increases in complexity, allowing RL training to be scaled much further without saturating. Empirically, our method generalizes remarkably well: curriculum training on composed 6th-grade level math problems (GSM8K) boosts accuracy on longer, competition-level benchmarks (GSM-Symbolic, MATH-500, AIME) by up to 2.06x. It also transfers significantly to diverse out-of-distribution ReasoningGym domains and long-context benchmarks, indicating broader generalization. Importantly, our long-horizon improvements are significantly higher than baselines even at high pass@k, showing that models can learn new reasoning paths under RL. Theoretically, we show that curriculum RL with outcome rewards achieves an exponential improvement in sample complexity over full-horizon training, providing training signal comparable to dense supervision. h1 therefore introduces an efficient path towards scaling RL for long-horizon problems using only existing data.
format Preprint
id arxiv_https___arxiv_org_abs_2510_07312
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle h1: Bootstrapping LLMs to Reason over Longer Horizons via Reinforcement Learning
Motwani, Sumeet Ramesh
Ivanova, Alesia
Cai, Ziyang
Torr, Philip
Islam, Riashat
Shah, Shital
de Witt, Christian Schroeder
London, Charles
Machine Learning
Artificial Intelligence
Large language models excel at short-horizon reasoning tasks, but performance drops as reasoning horizon lengths increase. Existing approaches to combat this rely on inference-time scaffolding or costly step-level supervision, neither of which scales easily. In this work, we introduce a scalable method to bootstrap long-horizon reasoning capabilities using only existing, abundant short-horizon data. Our approach synthetically composes simple problems into complex, multi-step dependency chains of arbitrary length. We train models on this data using outcome-only rewards under a curriculum that automatically increases in complexity, allowing RL training to be scaled much further without saturating. Empirically, our method generalizes remarkably well: curriculum training on composed 6th-grade level math problems (GSM8K) boosts accuracy on longer, competition-level benchmarks (GSM-Symbolic, MATH-500, AIME) by up to 2.06x. It also transfers significantly to diverse out-of-distribution ReasoningGym domains and long-context benchmarks, indicating broader generalization. Importantly, our long-horizon improvements are significantly higher than baselines even at high pass@k, showing that models can learn new reasoning paths under RL. Theoretically, we show that curriculum RL with outcome rewards achieves an exponential improvement in sample complexity over full-horizon training, providing training signal comparable to dense supervision. h1 therefore introduces an efficient path towards scaling RL for long-horizon problems using only existing data.
title h1: Bootstrapping LLMs to Reason over Longer Horizons via Reinforcement Learning
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2510.07312